* fix(lbug): re-open the read pool when analyze rebuilds the index under it
The MCP read pool's initLbug early-returned on an existing pool entry with no
freshness check, so after analyze rebuilt or mutated the on-disk index the
pool kept serving the old (POSIX: unlinked-but-open) inode until LRU/idle
eviction — a silent stale-read window of up to IDLE_TIMEOUT_MS (5 min).
Record the file identity {ino, mtimeMs, size} on each PoolEntry at open, and
re-stat in initLbug: unchanged → reuse; changed & idle → closeOne + reopen the
new file; changed while a query is in flight → serve the current handle (a
later idle initLbug reopens, since closing an in-use connection is a native
use-after-free). A stat failure (ENOENT during a full rebuild's unlink window)
is treated as unchanged so the reader keeps its valid open inode until the new
file appears. Mirrors the bridge cache's mtime-invalidation pattern.
Step 1 of docs/plans/2026-07-21-...-analyze-atomic-swap-invalidation. The
end-to-end reopen-on-swap path is exercised by the reader-during-rebuild
integration test in a later step.
* fix(analyze): publish a full rebuild via an atomic swap (POSIX)
The full-rebuild path wiped the live index (wipeLbugDbFiles(lbugPath)) and
rebuilt it in place, so a concurrent MCP reader that opened mid-build could
see an empty/half-loaded DB, and a crash between the wipe and the end-of-run
left the index destroyed (recoverable only by --force).
Build the fresh index at <lbugPath>.new and swap it over the live index in one
atomic rename at the end. All DB work flows through the singleton connection,
so only initLbug/wipeLbugDbFiles take the temp target; the close already
checkpoint-consolidates the build to a single file (verified: no residual
.wal/.shadow), so the rename publishes a complete index in one step. A reader
opening mid-build only ever sees the previous complete index; a reader holding
the old inode keeps a consistent stale snapshot until the pool re-opens onto
the new one (the pool staleness invalidation from the prior commit). On
failure the swap is skipped, leaving the previous index byte-for-byte intact.
POSIX only: the common CLI/serve-worker analyze paths skip the native close
(closeLbugBeforeExit, #2264) and leave the build handle open at swap time.
POSIX renames an open file cleanly; a same-process open handle blocks the
rename on Windows, so Windows keeps the current in-place behavior
(buildPath === lbugPath) until that is resolved. The Windows atomic swap and a
deterministic concurrent reader-during-rebuild test are deferred follow-ups.
Steps 2b + partial 3 of docs/plans/2026-07-21-...-analyze-atomic-swap-invalidation.
Integration test asserts the no-temp-leak + inode-swap invariants and the
crash-safety guarantee (a load failure leaves the live index untouched).
* test(analyze): end-to-end read-pool reopen after an atomic swap
Adds the deferred reader-during-rebuild / pool-reopen integration test:
analyze v1 -> read pool serves it -> rebuild with a renamed function (atomic
swap) -> the same repoId's initLbug detects the swapped inode and re-opens the
pool onto the new index. Asserts the pool sees the renamed function and NOT the
stale v1 name, exercising #1 (invalidation) and #2 (swap) together end to end.
* fix(lbug): bound pooled read queries with setQueryTimeout
The read pool relied only on a JS-side Promise.race (QUERY_TIMEOUT_MS) that
frees the waiter but leaves the native call running. Set the engine-level
setQueryTimeout on every pooled connection so a pathological query is bounded
at the source too.
* fix(lbug): name the held-open cause for WAL checkpoint failures (#2599)
A WAL-checkpoint IO error that also carries a busy/lock signal means another
handle (a gitnexus mcp server, or this process's own reader) holds the store
open, not a disk fault. Add isLbugCheckpointBusyError (reusing the tested
isDbBusyError keyword set) and, when the checkpoint driver exhausts its retry
budget on such an error, annotate the surfaced error with the actionable
held-open cause instead of a raw IO string.
Note: overlaps in-flight work on repro/issue-2599-windows-wal-checkpoint;
bundled here at the maintainer's request.
* feat(analyze): opt-in atomic incremental + best-effort Windows swap
Extends the atomic-swap publish (POSIX full rebuild) to two more cases:
- Windows: the swap now applies when a real close is safe to release the build
handle before the rename — i.e. non-pdg runs (windowsSwapOk excludes --pdg,
the #2264 destructor-crash case), forcing a real close on the swap path.
UNVERIFIED on Windows (no Windows runner here); --pdg and any failure fall
back to today's in-place behavior, so it can never corrupt.
- Incremental (opt-in, GITNEXUS_ATOMIC_INCREMENTAL=1): copies the live index
into the temp, applies the incremental delete/writeback to the copy, and
swaps at the end. Off by default because the whole-file copy negates
incremental's speed premise — kept behind a flag pending a benchmark. The
escalation valve also targets the temp so an escalated write stays atomic.
Integration test covers the opt-in incremental path end to end (no temp leak,
the incremental change is reflected after the swap).
* refactor(lbug): centralize the read-pool + bridge open-retry budgets
The lbug-config retry registry documented the open/handle-release/query-time
budgets but the read pool's LOCK_RETRY_* (pool-adapter) and the bridge's
LBUG_OPEN_RETRY_* (group/bridge-db) kept private copies that could drift. Move
both into the registry as exported constants (POOL_OPEN_LOCK_RETRY_*,
BRIDGE_OPEN_RETRY_*) and alias the local names to them — one tuning surface,
no behavior change.
* fix: address CI regressions from the bundled follow-ups
- setQueryTimeout: guard the call so test doubles that don't model the engine
method don't break connection creation.
- atomic swap: skip the rename when the build produced no DB at buildPath (an
empty repo / mocked pipeline) instead of throwing ENOENT.
- #2599: don't wrap the checkpoint error in the driver (it hid the IO signature
the CLI's --wal-checkpoint-threshold hint keys on); name the held-open cause
at the CLI instead, beside that hint, keeping the original error intact.
- retry consolidation: revert to documentation-only — moving the pool/bridge
budgets into lbug-config broke every explicit lbug-config test mock. The
registry now catalogues all budgets with their in-file locations.
- analyze-wal-checkpoint-failure test: block both lbug.wal.checkpoint and
lbug.new.wal.checkpoint, since a full rebuild now checkpoints the temp.
* fix(analyze): publish the swap before stamping meta; identity-gate the reader (#2614 F1)
Review found a HIGH regression: the full-rebuild wrote the freshness stamp
(saveMeta, indexedAt=T_new) BEFORE the atomic swap, so a concurrent MCP reader
that reinited in the saveMeta->swap window opened the OLD inode, recorded
observed=T_new, and then never reinited again (ensureInitialized returns early
on 'current') — serving the pre-rebuild graph indefinitely. The build-into-temp
change inverted the pre-PR invariant that 'meta shows T_new' implied 'lbugPath
holds T_new data'.
Two coordinated fixes:
- run-analyze: move the final saveMeta AFTER the swap, so meta.indexedAt only
becomes visible once lbugPath resolves to the new inode. Verified nothing in
the span reads on-disk meta and registerRepo writes only the registry.
Leaving the dirty flag set across the swap also improves crash-safety.
- local-backend: the reader staleness gate now also compares the lbug file
IDENTITY (ino/mtime/size), reiniting on an inode change even when
meta.indexedAt is unchanged. This closes the swap-window latch and covers the
in-place incremental case — and is what actually makes the pool's dbIdentity
net reachable for the MCP reader (the indexedAt gate otherwise bypassed it).
* fix(lbug/analyze): WAL-aware incremental, residual-sidecar reconcile, Windows opt-in, #2599 anchor (#2614 F2-F4)
Review remediations:
- F3: gate atomic incremental on a CLEAN live index (inspectLbugSidecars) — the
main-file-only copy would drop an orphan .wal's delta; fall back to in-place.
- F4: on the swap, MOVE a residual <buildPath>.wal/.shadow beside the published
index (not orphan it) so a swallowed final checkpoint's delta is replayed.
- F2: record identity on the shared read-only Database and warn when a cached
handle is reused after its on-disk index was rebuilt while another consumer
holds it (unreachable via MCP — one consumer per lbugPath; a complete fix
needs per-inode handles, documented).
- Windows swap: opt-in (GITNEXUS_ATOMIC_WINDOWS_SWAP=1), default off — the
forced real close re-bets an unproven #2264 assumption and can't be verified
without a Windows runner, so the default Windows path stays in-place.
- #2599: anchor isLbugCheckpointBusyError to real held-open wording instead of
isDbBusyError's bare .includes('lock') over a message that embeds the DB path
(a repo under blockchain-app misclassified a disk fault as held-open).
- Docs: corrected retry-catalogue budgets (linear, not exp) and the
checkedOut>0 bound comment (load-bounded, not IDLE_TIMEOUT_MS).
* test(analyze): cover the production close path in the atomic swap (#2614 F5)
Adds a full-rebuild swap test with skipNativeCloseOnExit:true — the close path
the CLI and serve-worker actually ship (build handle left open at swap time),
distinct from the default real-close the other swap tests exercise. Asserts the
POSIX swap still publishes a single consolidated lbug with no .new temp and no
orphan sidecar.
* test(analyze): give the follow-up git commits an inline identity (CI fix)
The end-to-end reopen and atomic-incremental tests' second commits used a bare
`git commit`, which fails on CI runners with no global git identity (empty
ident name). makeRepo's initial commit already passes -c user.name/-c
user.email inline; apply the same to the rename/change commits. No code change.
* fix(mcp): route reader reinit through initLbug's active-query guard (#2614 review)
Review found an active-query retirement race: LocalBackend.ensureInitialized
detected an identity/stamp change and called closeLbug(poolKey) DIRECTLY, but
closeOne closes the shared Database at refCount 0 regardless of checked-out
connections. So a reader detecting the new generation could close the Database
a concurrent query is still executing on — a native use-after-free. This
bypassed the checkedOut>0 guard that initLbug itself has.
Fix (delegate, not close directly): initLbug now returns whether it actually
rolled the pool over; ensureInitialized calls initLbug (which serves the
current handle while a query is in flight and reopens only when idle) instead
of closeLbug. The observed IDENTITY is advanced only when the pool actually
reopened — if a query was in flight, the identity stays divergent and the
reopen retries on a later idle check rather than latching on the old handle.
The observed STAMP advances regardless so a same-file stamp change can't loop.
Old generation now stays alive until its in-flight queries drain (lazy
rollover); new requests during the busy window share the old handle until the
pool goes idle, then reopen. No parallel open-both-generations, but no UAF and
no stale latch.
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
loadIgnoreRules is called once per repo, per language/contract
extractor during group sync -- an N-repo group fans out to 6+
extractors each calling it, turning an uncached execSync per call into
O(extractors x repos) blocking subprocess spawns for the exact
many-repos scenario #2606 describes.
Both getGitInfoExcludePath and getCoreExcludesFilePath resolve to the
same value for the same fromPath for the life of the process, so
memoize by fromPath in a process-lifetime Map. One-shot CLI runs are
unaffected by staleness; the long-lived MCP server would need explicit
invalidation if this becomes a real concern.
The first fully successful real workflow_dispatch run on the self-hosted
runner (https://github.com/abhigyanpatwari/GitNexus/actions/runs/29843028596)
still failed: 17/18 sessions hit error_kind plan-evidence-invalid with
"phase changed unauthorized workspace path(s): .claude/.cc-writes,
.claude/commands, .env, .env.development, ...".
Reproduced directly on the runner (SSM, matching the real sandbox settings
exactly, including enableWeakerNestedSandbox): a single trivial "say OK"
prompt -- no real task, no real API key even -- is enough to make Claude
Code create a synthetic package.json/lockfiles/node_modules, a full set
of .env variants, and .claude/agents, .claude/commands, .claude/.cc-writes
in the workspace on every single session. None of this is something the
model decided to write; it's Claude Code's own internal bootstrap for
running inside an already-sandboxed environment, and it happens
regardless of task or prompt.
enforce_phase_workspace (the planning-phase boundary check: verify the
plan session touched only its one plan doc) already excludes .git for
exactly this class of reason -- harness/tool noise, not substantive diff.
Extends the same exclusion to the empirically-observed bootstrap set.
workspace_snapshot has exactly one use (this check, confirmed via every
caller), so widening its exclusion list can't hide anything in some other
context that actually cares about these paths changing.
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Replace the custom ~/.gitnexus/ignore file with the same two sources
real git itself consults for exactly this purpose (gitignore(5)):
- core.excludesFile: git's own all-repos global ignore file (defaults
to $XDG_CONFIG_HOME/git/ignore when unconfigured)
- $GIT_COMMON_DIR/info/exclude: per-repo, untracked, so it works
without push/commit access to the repo
Precedence mirrors git exactly (lowest to highest): core.excludesFile,
then info/exclude, then .gitignore, then .gitnexusignore -- each later
source can negate an earlier one via a `!pattern` line, same
last-match-wins semantics git itself uses.
Adds getCoreExcludesFilePath and getGitInfoExcludePath to git.ts,
following the same execSync + git-common-dir pattern as
getCanonicalRepoRoot. GITNEXUS_NO_GLOBAL_IGNORE (or noGlobalIgnore)
still skips both global sources, mirroring GITNEXUS_NO_GITIGNORE.
IgnoreService only read per-repo .gitignore/.gitnexusignore, so an
exclusion meant to apply across every indexed repo had to be repeated
per repo or hand-patched into node_modules (wiped on every upgrade).
loadIgnoreRules now also reads a global ignore file at
$GITNEXUS_HOME/ignore (default ~/.gitnexus/ignore), reusing the
existing global directory that already holds registry.json and
config.json. It is added first, so per-repo .gitignore/.gitnexusignore
rules can still negate it, mirroring the .gitignore -> .gitnexusignore
precedence already in place. GITNEXUS_NO_GLOBAL_IGNORE (or
noGlobalIgnore) skips it, mirroring GITNEXUS_NO_GITIGNORE.
GitNexus review-agent finding: stripDynBound's documented Box<dyn Trait>,
Rc/Arc<dyn Trait>, and auto-trait/lifetime bound-list (dyn Trait + Send)
shapes had no test anywhere — only the bare &dyn Trait parameter case was
exercised end-to-end. Add direct unit coverage on normalizeRustTypeName and
(via interpretRustTypeBinding) normalizeRustReturnType for these shapes.
call-summary-schema-version.test.ts pins INCREMENTAL_SCHEMA_VERSION as a
literal per bump, documenting the reuse-gate boundary for each version.
Update the "current" expectation to 11 and add the v10 pre-current case,
matching the v7/v8/v9/v10 precedent already in the file.
RUST_SCOPE_QUERY gained a function_signature_item capture, shifting the
capture fingerprint for every bench fixture with a required trait method.
Verified: node --import tsx bench/scope-capture/measure.mjs --check now
passes across all 14 languages (rust scaling 1.036 < 1.5 budget).
Addresses gitnexus-review-agent findings on PR #2608:
- MED: on a partial apply (a file's write throws), drop that file's edits
from total_edits/graph_edits/text_search_edits/changes so the reported
result describes what actually reached disk, not what was attempted. The
comprehensive enumeration otherwise let a failing file contribute its
entire line count as phantom 'applied' edits. failed_files still names
every dropped file. Counts are now derived once from the reported set.
- MED: hoist the word-boundary regexes out of the per-line loop (one compile
each instead of one per line), reused by the apply loop.
- LOW: apply loop reuses escapedOldName instead of recomputing the escape
formula inline (removes a preview/apply drift risk).
- Soften the in-code comment: enumeration gives per-call preview/apply
consistency; the pre-existing two-read TOCTOU (external write between
preview and apply) is out of scope and noted, not newly introduced.
Tests: add a mixed graph-ref + text_search multi-file case (asserts per-file
confidence and the never-downgrade guard, via a stubbed rg), and a
partial-write-failure case (asserts only landed files are reported). Assert
concrete graph_edits/text_search_edits splits, not just their sum.
* fix(eval): bind the resolved node to a fresh sandbox path, not one under /usr
#2607 bound the resolved `node` to /usr/local/bin/node, but that path lives
inside the /usr tree that _runtime_mount_args already read-only-binds
wholesale. The second real workflow_dispatch run on the self-hosted runner
(https://github.com/abhigyanpatwari/GitNexus/actions/runs/29840270554)
failed immediately in the bubblewrap preflight: "bwrap: Can't create file
at /usr/local/bin/node: Read-only file system" -- bwrap can't create a new
mount-point file inside a tree it already bound read-only when the real
path doesn't already exist there on the host, which is exactly the
self-hosted case this bind exists to fix.
Introduces SANDBOX_NODE (/opt/claude/node), a fresh path outside every
tree _runtime_mount_args binds, following the same pattern SANDBOX_CLAUDE
and SANDBOX_PYTHON3 already use. Updates the two real call sites
(sanitized_graph.py, runner_sessions.py) to use the constant instead of
the hardcoded literal, so the fix can't drift out of sync with itself
again, and re-exports it from runner.py alongside the other SANDBOX_*
names for the real-bwrap tests that reference it directly.
Adds a real-bwrap test (gated behind GITNEXUS_REQUIRE_BWRAP_CANARY, same
as the existing ones) that copies a real node binary to a path outside
every bound tree and actually launches bwrap against it -- an
argv-construction test alone can't catch a bwrap-level "Read-only file
system" error, only a real invocation can, and that's exactly the gap
that let #2607's version of this fix through review looking correct.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(eval): don't let the new real-bwrap test's node-mock break bwrap's own resolution
CI caught this immediately: the new test_real_bubblewrap_runs_node_from_outside_the_bound_trees
monkeypatched shutil.which to return None for anything but "node", but
prepare_sandbox's own bwrap/claude resolution (_resolve_executable) goes
through shutil.which too -- so the test broke bwrap discovery before the
sandbox it's supposed to exercise could even be built ("SandboxError:
required executable is unavailable: bwrap").
Delegate to the real shutil.which for every other name instead of
blanket-returning None.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
RUST_SCOPE_QUERY gained a function_signature_item capture (previous commit)
so abstract trait methods can now dispatch a CALLS edge through a &dyn
Trait receiver. The incremental write set only covers changed files, so a
top-up against a pre-v11 index would keep silently missing these edges for
every unchanged Rust trait file — same contract as v7/v10; force a full
re-analyze instead.
Expected drift from the query.ts change: abstract trait methods now emit a
scope + declaration capture, shifting captureGroups/digest for every rust-*
fixture containing a trait with a required (bodyless) method.
New minimal fixture (single trait + impl + &dyn Trait call site, no other
same-named callers) proves the dyn-dispatch CALLS edge discriminates: fails
against the pre-fix source (0 edges) and passes against the two preceding
commits' fix (exactly 1 edge, verified via the CLI analyze pipeline against
a standalone repo).
The existing rust-abstract-dispatch fixture was NOT extended for this,
deliberately: it already has other callers referencing the same method
names (process()'s repo.find()/save()/count()), and an existing resolution
fallback picks those up via simple-name matching regardless of receiver
type — masking this specific defect in the in-process test-pipeline path.
A dedicated, single-caller fixture keeps the regression test load-bearing.
Surfaced by the first real workflow_dispatch run on the self-hosted runner
(https://github.com/abhigyanpatwari/GitNexus/actions/runs/29836411744):
every session failed with error_kind infra-error, error_detail "bwrap:
execvp /usr/local/bin/node: No such file or directory", tripping the
outage-streak breaker after 5 consecutive failures.
sanitized_graph.py and runner_sessions.py invoke the sandboxed graph CLI
at the fixed path /usr/local/bin/node. _runtime_mount_args only binds
/usr, /bin, /lib, /lib64 wholesale, so that path resolves correctly when
node happens to live under /usr/local/bin on the host -- true on
GitHub-hosted runner images, but not on a self-hosted runner, where
actions/setup-node installs into its own tool-cache directory instead
(outside all four bound trees, so invisible to the sandbox regardless of
what PATH says on the host).
Fix lives entirely in the mount construction: resolve `node` via
shutil.which (correctly picks up wherever actions/setup-node put it,
since its tool-cache dir is already on PATH by the time this runs) and
bind it read-only to the same fixed sandbox path the two call sites
already expect. Neither call site needed to change. Backward compatible
with GitHub-hosted runners, where this resolves to the same path and
binds a harmless no-op self-mount.
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
rename() reported total_edits from a partial enumeration (definition line
only, one-edit-per-graph-file then break, and text search that skipped any
file already covered by the graph) while the apply step does a whole-file
\boldName\b global replace on every touched file. When a private symbol's
definition and all its call sites live in one file, only the definition line
was reported (total_edits: 1) even though apply rewrote every occurrence, in
both dry-run and apply.
Rebuild changes/total_edits/graph_edits/text_search_edits from one file set:
classify each file to rewrite (definition + graph refs = graph confidence;
rg-only files = text_search, never downgrading a graph file), then enumerate
every matching line per file with apply's exact escaped global regex. The
reported edit list now equals what apply writes. Apply behavior is unchanged.
Adds a regression test reproducing the issue's single-file Rust case (def +
3 same-file call sites, empty graph): total_edits is 4 in both dry-run and
apply, and equals the replacements that land on disk.
fn foo(&self) -> T; (no body) parses as function_signature_item, a grammar
node distinct from function_item that RUST_SCOPE_QUERY never captured. An
abstract trait method therefore had no Function scope and no declaration,
so populateClassOwnedMembers never wired its ownerId to the trait's Class
scope — invisible to the CALLS-edge receiver-bound resolution pass even
after a receiver's type resolves to the trait correctly.
Together with the previous commit's dyn-stripping fix, a call through a
&dyn Trait parameter now emits a CALLS edge to the trait's method (#2604).
normalizeRustTypeName/normalizeRustReturnType stripped reference sigils,
pointer sigils, and smart-pointer wrappers but never the `dyn` keyword, so a
`&dyn Trait`-typed receiver normalized to the literal string "dyn Trait"
instead of "Trait" — an unmatchable name that silently broke every
downstream receiver-type lookup for trait-object dispatch.
Part of the #2604 fix (root cause has a second, independent half: abstract
trait methods are invisible to scope resolution until function_signature_item
is captured — next commit).
Two gitnexus-review-agent findings on PR #2602:
- MEDIUM: the bodied-constant MRO-to-host-enum path (a qualified call to an
inherited, non-overridden enum method) was claimed in a comment but never
tested. Add EnumConst.A.log() -> EnumConst.log#0, exercising E$N's
@reference.inherits MRO arm end to end.
- LOW: `bodiedName ?? hostEnum` conflated "body-less" with "name synthesis
failed on a bodied constant" (reachable only on malformed/error-recovery
trees), silently binding an overriding constant's receiver to the host
enum — a wrong edge instead of no edge. Switch to `isBodied ? bodiedName :
hostEnum` so a bodied constant binds ONLY to its E$N class, mirroring the
object_creation_expression branch's skip-on-synthesis-failure. Verified
output-neutral on the well-formed bench corpus.
Rebaseline the java scope-capture fingerprint (a822cef9 -> d04298a9): the
bench corpus IS test/fixtures/lang-resolution, so the new dispatchInherited
fixture method shifts it (+6 capture groups); the logic change contributes
nothing (confirmed by isolating the fixture-only fingerprint). java.test.ts
242 passed; measure.mjs --check PASS (14 languages); tsc/prettier/eslint clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(eval): move skill-evolution to a self-hosted runner and fix the sandbox's Python 3 trust gap
GitHub-hosted runners hard-cap job execution at 6 hours, which is too
short once a benchmark session actually invokes Skill/MCP tools for
real (the --bare fix in #2584 means sessions no longer no-op). Move the
job onto a self-hosted runner (5-day cap instead) and document the
activation step in the workflow's own checklist.
Validating the self-hosted run surfaced a real bug: gitnexus-plan
sessions inside the bwrap sandbox failed with "planning must create or
modify exactly one plan artifact; observed 0". Root cause:
evidence-provenance.mjs's atomic plan-writer only trusts a Python 3
binary owned by root or by the current process. Inside this
--unshare-user sandbox only the calling uid is mapped (root isn't), so
the real, root-owned /usr/bin/python3 surfaces as the kernel's overflow
uid and gets correctly refused as untrusted. Fix: provision a small,
self-owned wrapper script (same pattern already used for
shell-prefix) that execs the real interpreter, so the sandbox has a
Python 3 candidate the existing trust check can actually accept --
without touching that security-sensitive validation logic at all.
Also add visibility so this class of failure isn't quiet next time:
report.md now shows why each row failed (error_kinds), not just
resolved 0/1, and the benchmark now exits non-zero when an incumbent
arm -- the currently-shipped skill -- resolves zero across every task,
since that reads as a broken harness rather than a normal candidate
miss.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(eval): close the broken_incumbent_arms zero-valid-runs gap; document runner exposure tradeoff
Addresses the two MEDIUM findings from the gitnexus-review-agent on this PR
(https://github.com/abhigyanpatwari/GitNexus/pull/2600#issuecomment-5033363096).
broken_incumbent_arms required valid_runs > 0 before flagging an incumbent,
so an incumbent that fails every run with an excluded-but-non-systemic
error_kind (e.g. evidence-unverified, which the outage-streak breaker
explicitly resets on rather than accumulates) never accumulated a single
valid run and sailed through silently -- the exact quiet no-promotion
outcome this guard exists to catch, and arguably worse than the
some-runs-resolved-zero case since here nothing completed at all.
aggregate() never marks an excluded/unverifiable row resolved=True, so
dropping the valid_runs requirement and checking resolved == 0 alone
correctly covers both cases. Added a test for exactly this all-excluded
scenario, which none of the existing three did.
Updated the workflow's own activation checklist to reflect what's actually
true now (the gitnexus-evolution environment's branch policy and the
self-hosted runner are both live, codified in infra/gitnexus-evolution/ in
a companion PR) and documented the exposure-window tradeoff the review
flagged: the runner is stopped between runs but not destroyed/recreated per
run, so it isn't fully ephemeral. Stopping already bounds the exposure
window to the job's own runtime on one day out of seven; full per-job
ephemeral provisioning is a deliberate non-goal for a job that runs at
most weekly, revisit if that changes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* docs(eval): remove public infra/ pointers from the activation checklist
PR #2603 (the Terraform codification this checklist pointed to) got closed
-- publishing the exact IAM roles, security group rules, and self-hosted
runner topology for a real, live AWS account isn't safe to do in a public
repo, even with no literal secrets or resource IDs in the diff. The
underlying AWS/GitHub setup is unaffected and still documented privately;
this just removes the now-dangling references to a directory that won't
exist in this repo.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* ci(actionlint): register the gitnexus-evolution self-hosted runner label
actionlint rejected `runs-on: [self-hosted, linux, x64, gitnexus-evolution]`
in gitnexus-skill-evolution.yml because it can't discover custom runner
labels. Register it in .github/actionlint.yaml so the Workflow Lint check
passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
The enum-constant receiver-dispatch fix adds one @type-binding.* capture
per enum constant, so the java scope-capture fingerprint shifts
(85fc7af9 -> a822cef9). Pure capture-additive drift; no bench fixtures
added; scaling 1.024 < 1.5 budget. Verified `measure.mjs --check` passes
for all 14 languages.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Calling a method on an enum-constant receiver (E.CONST.method()) emitted
no CALLS edge. The receiver "E.CONST" is a two-segment compound receiver;
resolveCompoundReceiverClass walks each dotted segment via the owning
class scope's typeBindings map, but enum constants had no typeBinding, so
the constant segment dead-ended and no target was ever resolved.
#2555/#2558 gave bodied constants a first-class synthesized E$N class with
an MRO that includes the host enum; this is the receiver-side follow-up.
synthesizeJavaAnonymousClassDeclarations now emits a class-scope
typeBinding for every enum constant's simple name -> its E$N class (bodied)
or the host enum itself (body-less), reusing the exact mechanism a field
declaration uses. The generic compound-receiver chain walk then resolves
E.CONST.method() with no change to any shared scope-resolution code.
Bodied dispatch (EnumConst.A.hook() -> EnumConst$1.hook#0) and body-less
inherited dispatch (Plain.A.m() -> Plain.m#0) are covered by new tests in
the existing java-enum-constant-body fixture; both were verified to fail
against the pre-fix tree.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reverses the prior convention: gitnexus-plan/gitnexus-work plan documents
under docs/plans/ are working artifacts and no longer travel with the PR.
Drops the require-node-22.18 plan doc from tracking; the .gitignore now
ignores all of docs/. The workflow_bench snapshot features scan the
filesystem, not git-tracked status, so they are unaffected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
test_eval_ci_uses_locked_uv_and_blocking_native_containment_jobs pins the
eval-containment-linux job's setup-node version; move it in lockstep with
the ci-tests.yml pin bumped to the 22.18 floor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The module.registerHooks compat seam and the onnxruntime resolvers cited
the old '>=22.0.0' floor as the reason their sub-22.15 fallback was
reachable. With the floor now ^22.18.0 || >=24.11.0 (all >=22.15), every
supported runtime exposes the API; the fallback stays as defensive
handling for below-floor runtimes (engines is advisory, not
engine-strict). Comments only - no behavior change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With the supported minimum raised to Node 22.18, retarget every lane and
pinned runtime that sat at a lower version so nothing builds or runs the
package on an unsupported (EBADENGINE-warning) Node:
- ci-tests.yml: node-floor-compat 22.14 -> 22.18.0 (name, comment, pin,
version assertion) so the floor gate guards the new minimum; its #2372
registerHooks failure mode cannot recur above 22.15. Containment-canary
pin 22.16.0 -> 22.18.0.
- gitnexus-review-agent.yml + the pinned review/canary runtime: the
reproducible runtime is version-locked in lockstep across
.github/{gitnexus-review-runtime,claude-canary-runtime}/package.json and
their lockfiles (engines), the workflow's node-version, its two
'node --version = v22.18.0' assertions, the lockfile-engines guard, and
NODE_VERSION. Moved all of them 22.16.0 -> 22.18.0.
- gitnexus-skill-evolution.yml: pinned runtime 22.16.0 -> 22.18.0.
- CONTRIBUTING.md prerequisite floor updated.
- review-agent-workflow.test.ts, which enforces the runtime lock, updated
to expect 22.18.0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Babel 8 (devDep for the bench mutation oracle, pulled in by dependabot
previous floor (>=22.0.0), so every dev install on Node <22.18 emitted
nine EBADENGINE warnings. Rather than pin Babel back to 7, adopt Node
22.18+ as the supported minimum: set engines to ^22.18.0 || >=24.11.0,
matching Babel 8 exactly so the warnings resolve honestly with no
dependabot ignore needed.
@types/uuid@11 is a deprecated stub - uuid@14 ships its own types and no
tsconfig references it. Lockfile edited by hand (engines + @types/uuid
entry) to preserve the libc platform metadata a newer npm wrote;
verified consistent via npm ci (exit 0).
BREAKING CHANGE: the gitnexus package now requires Node ^22.18.0 || >=24.11.0
(previously >=22.0.0). Node 22.0-22.17 are no longer supported.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Anchor isBenignDropFtsIndexError to the START of the message
(startsWith, not includes) so a future genuine failure that merely
mentions "Binder exception" or "Catalog exception" mid-message can't
be misclassified as benign. New test proves the old substring match
would have swallowed such a message.
- incremental-fts-drop-ordering.test.ts: probe FTS availability once in
beforeAll and skip VISIBLY via ctx.skip() in beforeEach (matching the
withTestLbugDB/lbug-vector-extension convention) instead of a silent
console.warn+return inside the test body, which reported a false pass
with zero coverage of the ordering invariant when FTS was unavailable.
The post-first-run FTS-index-built check is now a hard assertion
instead of a second soft skip, since the beforeEach gate already
proved the extension loads.
CI caught this: adding the record_declaration capture legitimately
changes the pinned java capture fingerprint, same as every prior
capture-behavior change to this language (#2550, #2555). Rebaselined
following the established _rebaselined_* precedent; scaling ratio
1.059 stays well within the 1.5 budget.
Fixes#2589: incremental analyze intermittently crashed with "FTS
index 'file_fts' is inconsistent: term is missing during delete" after
markdown-only commits, and --repair-fts also failed in that state.
deleteNodesForFiles' batched DETACH DELETE ran against tables that
still carried the FTS index built at the end of the PREVIOUS analyze
run -- createSearchFTSIndexes only drops+rebuilds every index in Phase
3, well after that delete already ran. LadybugDB's FTS extension is
not proven to survive DML against an indexed table (its own docs never
demonstrate the sequence). Call the new dropSearchFTSIndexes() up front
in the non-escalated incremental branch, before deleteNodesForFiles --
Phase 3 still rebuilds every index from the final row set regardless.
New end-to-end test drives a real runFullAnalysis full+incremental
cycle and confirms it fails without this change (file_fts and 50
sibling indexes still present at delete time) and passes with it.
dropFTSIndex previously caught and discarded every DROP_FTS_INDEX error
unconditionally. Extract isBenignDropFtsIndexError, a pure classifier
for the two legitimate "nothing to drop" cases (Binder/Catalog
exceptions: index never created, or the FTS function isn't registered)
verified end-to-end against @ladybugdb/core 0.18.x's real conn.query()
error text. Anything else -- e.g. the Runtime exception "FTS index is
inconsistent" class from #2589 -- now rethrows instead of being masked,
so a corrupted index can no longer persist across analyze runs
undetected.
Pulls the existing per-index dropFTSIndex loop out into its own exported
function so the incremental writeback can drop FTS indexes up front,
before deleteNodesForFiles runs (#2589). No behavior change here —
createSearchFTSIndexes calls the new function and still rebuilds every
index afterward.
Review finding: the record_declaration container-node fix (894110bf)
makes previously-uncaptured Record nodes and HAS_METHOD edges appear
for the first time, but the incremental write set only covers changed
files. Without this bump, an existing index would silently keep
omitting the Record node and its HAS_METHOD edges for unchanged
record files after an ordinary incremental analyze.
Same contract as v7 (#2437/#2522) and the two closest precedents, v8
(#2550) and v9 (#2555), which bumped this constant for the identical
"model X as first-class node" class of change.
new Local().inner() bound the whole object_creation_expression as
@reference.receiver, so its raw source text ("new Local()") became the
receiver name. That text can never match a scope binding, so the call
silently fell through to name-only fallback resolution and could
resolve to an unrelated same-named method on a collision.
Normalize the receiver to the constructed type's simple name (reusing
javaBaseSimpleNameOf, already used for the anonymous-class inheritance
edge) so Case 2 (class-name / static receiver) in
receiver-bound-calls.ts resolves it via its normal MRO walk. Mirrors
the existing normalizePhpReceiver precedent in php/captures.ts - a
language-local capture rewrite, no shared-pipeline change.
JAVA_QUERIES had no @definition.record capture, unlike its
class_declaration/interface_declaration/enum_declaration siblings and
unlike CSHARP_QUERIES' own record_declaration pattern. A Java record's
container node was never created, so its HAS_METHOD edges were dropped
at persistence even though ownership resolution computed a valid
ownerId for its methods.
Downstream label mapping, the class-extractor config, the dispatch
table, and ownership reconciliation already treated 'Record' correctly
- this was purely a missing structure-phase capture.
Wrap the version-directory scan in try/catch so a transient FS error
(permission denial, an AV file lock on Windows, a directory vanishing
mid-scan) fails closed to null instead of throwing — matching the
original callers' contract, and safer across the Windows/macOS/Linux
CI matrix where these error modes differ.
Also drop the redundant USERPROFILE/HOME manual chain in
extension-binary-real.test.ts in favor of the repo's established
os.homedir() convention (already used ~15 other places here), which
Node resolves correctly per-OS and already honors env overrides.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
resolveInstalledFtsExtension (extension-binary-real.test.ts) and
resolveSeedExtension (fts-extension-e2e.test.ts) both hardcoded the
on-disk FTS extension path as .lbdb/extension/<lbug.VERSION>/..., but
LadybugDB's native INSTALL/LOAD resolves its own extension-ABI version
directory, which does not always track the npm package version. Bumping
@ladybugdb/core from 0.18.1 to 0.18.2 in this PR still installs into a
0.18.1 directory, so both hardcoded lookups came up empty and failed
hard under GITNEXUS_REQUIRE_FTS=1 in CI (all platforms, shard 3).
Add findInstalledFtsExtension() to discover the real installed file by
scanning every version subdirectory, and use it from both test files.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix: remove hardcoded 300-flows cap for large repositories
The dynamicMaxProcesses was capped at 300 via Math.min(300, ...),
causing large repositories (280K+ nodes) to lose execution flows.
Change: Remove the Math.min(300, ...) cap, keep dynamic calculation.
Effect: 280K-node repo: 300 → 1617 flows.
* test: add regression for dynamic maxProcesses sizing (#2198)
Verify that processProcesses honours maxProcesses > 300 without truncation.
Addresses the optional follow-up suggested by @koriyoshi2041.
* test: exercise computeDynamicMaxProcesses at the phase layer (#2198)
Extract from the inline
expression in so the regression test can exercise the
function that actually contained the removed cap.
The previous test called directly with
, which passes regardless of whether the phase-level
cap is present — never had the cap.
The new test suite covers:
- floor (20) for tiny repos
- linear scaling in the 0–3000 range
- growth past 300 for large repos (the actual regression)
- explicit assertion that reintroducing Math.min(300, …) would fail
Addresses review feedback from @azizur100389.
---------
Co-authored-by: Ubuntu <ubuntu@localhost.localdomain>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* feat(embeddings): support GITNEXUS_EMBEDDING_REQUEST_DIMS=omit
What: Honor GITNEXUS_EMBEDDING_REQUEST_DIMS=omit by suppressing the request-body
`dimensions` field sent to HTTP embedding backends.
Why: Strict OpenAI-compatible backends return vectors in the model's native size
but reject an unfamiliar `dimensions` field, breaking `analyze --embeddings`
against them. The var was parsed but never propagated, so `omit` was a no-op.
How: Add `requestDimensions` to HttpConfig, return it from readConfig, and forward
`config.requestDimensions` (not the validation-only `config.dimensions`) to
`httpEmbedBatch`. Local dimension checks still use `config.dimensions`.
Details: Coexists with the retry/pacing fields introduced upstream; both feature
sets are preserved. Default behavior unchanged when REQUEST_DIMS is unset.
Impact: gitnexus/src/core/embeddings/http-client.ts; README; unit tests.
* fix(embeddings): name GITNEXUS_EMBEDDING_REQUEST_DIMS in its own config error
Address the review findings on #2574.
What:
- A malformed GITNEXUS_EMBEDDING_REQUEST_DIMS now throws an error naming
GITNEXUS_EMBEDDING_REQUEST_DIMS, not the sibling GITNEXUS_EMBEDDING_DIMS.
- isHttpEmbeddingDimsError recognizes both leads, so the CLI still classifies
the REQUEST_DIMS config mistake as a clean config error, not a stack dump.
- Tests: numeric-override decoupling (DIMS=1024 validates the response while
REQUEST_DIMS=512 is sent in the body), the omit aliases (none/off/false/0),
and the malformed-value error path (which also pins the naming fix).
- README documents the full accepted values: omit-aliases and integer override.
Why: readConfig reused the DIMS error lead for the REQUEST_DIMS branch, so
REQUEST_DIMS=garbage misdirected the operator to edit the wrong variable. The
feature's actual decoupling and its non-omit inputs had no test coverage.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: wangxc <wangxc_a_bj@si-tech.com.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
--bare hard-disables the Skill tool and every mcp__* tool by Claude
Code design (confirmed against the pinned 2.1.214 binary; --allowedTools
cannot restore what --bare removes). Every workflow_bench arm except
baseline_nomcp needs Skill and/or GitNexus MCP tools, so every one of
those sessions has been silently unable to invoke gitnexus-plan/work/
review or the CE comparator skills -- the last skill-evolution run
(gen 0) scored 0/3 resolution on both arms across every task with
error_kind "skill-not-invoked", not because the candidate was bad but
because the harness could never invoke either arm's skill at all.
Only baseline_nomcp keeps --bare (it explicitly wants zero Skill/MCP
access anyway). The rest drop --bare and rely on ANTHROPIC_API_KEY
alone; the sandboxed HOME has no OAuth/keychain state to conflict
with it, and there's no committed .claude/settings.json in this repo
for dropping --bare to newly pick up.
Outside --bare the built-in toolset defaults to everything (WebFetch,
Task, subagents, ...), and --allowedTools only pre-approves within
whatever's available -- it doesn't narrow it. Added --tools for
non-bare sessions so the intended tool scope is still enforced instead
of silently widening.
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
The Idempotency beforeAll runs runSkillsCli (analyze --skills) twice, each
capped at 45s, under a 90s hook budget — exactly 2x the per-call timeout,
with no headroom for fixture creation and git init. On slow Windows CI
runners the two analyzes plus setup exceed 90s and the hook times out
('Hook timed out in 90000ms'), failing the shard before the test's own
status===null timeout tolerance can apply. Every other describe hook in
this file already uses 120s; align this one.
Co-authored-by: Claude <claude@anthropic.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The 64 MiB floor was too small: LadybugDB's bulk COPY needs working
buffer-pool memory that scales with the repo, so a 64 MiB pool fails with
"buffer pool is full and no memory could be freed" on any non-trivial
repo (empirically: the 6-file skills-e2e idempotency fixture needs
>=128 MiB; the ~1800-file GitNexus checkout needs >=256 MiB). Introduce a
distinct ADAPTIVE_POOL_FLOOR (256 MiB) for the hint clamp, kept separate
from BUFFER_POOL_FLOOR (64 MiB), which still guards defaultBufferPoolSize
on tiny-RAM machines; the hint is still clamped up to the machine default
so it can never over-commit.
This keeps the change as what it actually is — a large-repo optimization:
GitNexus full analyze is 51s (2 GiB) -> 35s (adaptive ~414 MiB). Small
repos now open COPY-safely at 256 MiB instead of the 2 GiB default (same
wall time on Linux, where commit is lazy; the eager commit is far cheaper
than 2 GiB on Windows). The reframed comments drop the earlier
unrepresentative "3-file repo / 64 MiB fast" claim.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The factor is validated by timing a full `analyze --force` of a large
repo, not a build-free bench (the pool is a native eager allocation).
Benchmark (GitNexus self, 101k graph elements): adaptive 414MiB pool =
35.3s vs forced 2GiB = 50.8s vs forced 64MiB = 26.5s — the adaptive pool
is 31% faster than the old 2GiB default even on a large repo (the eager
commit dominates), with no under-sizing thrash. 4KiB/element kept.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
runFullAnalysis now sets the buffer-pool size hint from the built graph's
node+relationship count (after the pipeline, before initLbug), and clears
it at the top of each run so a prior run's size can't leak into a
pre-pipeline open. A small repo opens with the fast 64MiB floor instead of
eagerly committing the full 2GiB pool.
Measured on a 3-file repo (this box): analyze drops 4.78s -> 1.9s for both
fresh and incremental, matching a forced 64MiB pool; the 2GiB path still
reproduces the old 4.78s. This roughly halves the skills-e2e Idempotency
hook (two analyze passes) that was timing out on Windows CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds the sizing lever without changing behavior yet: a module-scoped
buffer-pool size hint plus estimateBufferPool(graphElementCount), read by
resolveBufferManagerSize with precedence env-override > clamp(hint, 64MiB,
default) > default. The hint can only shrink the pool from the default
(clamped to [floor, default]), so the 2GiB/80%-RAM cap and the
GITNEXUS_LBUG_BUFFER_POOL_SIZE escape hatch (incl. 0) are preserved. With
no hint set, resolveBufferManagerSize returns exactly what it did before.
Motivation: LadybugDB eagerly commits the buffer pool at DB open, so the
fixed min(2GiB,80%RAM) pool adds a measured ~2.8s to every analyze even on
a 3-file repo (dominant on Windows). Sizing the pool to the graph lets
small repos use the fast 64MiB floor while large repos keep the cap.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
trusted_gitnexus_runtime_mounts and sanitized_graph both hand the ~290 MiB
graph index through TaskAssetSnapshot.materialize(), which reflinks into
each arm clone or falls back to a buffered copy under a 16 MiB budget.
Neither ext4 (CI runners) nor 9p (this container) support FICLONE, so
every real materialization fell to the buffered path and blew the budget
instantly (CI run 29750271566).
Raises MAX_BUFFERED_FALLBACK_BYTES to 512 MiB - well above the real index
size, still a full 4x below MAX_TASK_ASSET_BYTES so a genuinely oversized
declaration still fails closed.
resolve-invocation.ts requires hooks/claude/resolve-analyze-cmd.cjs at
module load time, reached whenever the analyze command loads. The
sandbox's curated mount list never exposed hooks/, so every
benchmark-arm session failed with MODULE_NOT_FOUND (CI run 29742191562).
Appends the new mount after the existing six so the function's
hardcoded mounts[0]/[1]/[2]/[5] validation reads stay correct. Extends
the real-bwrap canary to require analyze.js directly, since --version
alone never reaches the lazy import that broke.
Every benchmark-arm session failed with "sanitized graph snapshot
preparation failed: clone has more than 1024 references; refusing
incomplete sanitization" (confirmed via a real workflow_dispatch run,
29738099937, after the prior activation fixes let the proposer succeed
end-to-end for the first time).
make_worktree() creates each arm's throwaway clone with a plain `git
clone`, which inherits every tag and branch from the source. This repo's
history has grown to 1144 tags (a v1.6.9-rc.N release-candidate series)
out of 1650 total refs, exceeding oracle_assets.MAX_CLONE_REFS=1024 -- a
fail-closed guard in sanitize_clone_for_hidden_oracles() that refuses to
proceed unless it can enumerate and delete every ref before handing a
sanitized snapshot to a benchmark session (so an agent can never discover
oracle answers via a ref the sanitization missed).
`ref` at every call site (evolve.py, runner.py, sanitized_graph.py) is
always a bare SHA or the literal "HEAD", never a branch name, so
`--single-branch --branch <ref>` isn't viable (git clone's --branch
requires a name). Tags are never used by the checkout fallback or by
sanitization's own delete-everything behavior, so dropping them via
--no-tags removes the 1144-ref majority without touching branch-fetch
behavior or the existing ref/origin-ref checkout fallback, and without
weakening MAX_CLONE_REFS itself.
Verified against the real repository (not just the test fixture): cloning
/workspace (1650 refs, 1144 tags) via the fixed make_worktree() now
produces a clone with 237 total refs and 0 tags.
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The proposer's second real-run failure (after the ENV_SCRUB permission fix
landed) exited 1 with subtype "success" and an EMPTY stderr_tail -- opaque:
the downloaded CI artifact showed num_turns:1, cost_usd:0, tokens:0,
duration_s:0.1, meaning the session terminated before any real model turn
completed (consistent with an early, pre-flight-style failure), but nothing
in the persisted record said why.
The actual JSON event stream (permission_denials, tool_use/tool_result,
is_error) lives in stdout, which run_managed already captures as
proc.stdout_tail -- it just never made it into the session's error_detail.
Add it there, bounded and truncated the same way stderr_tail already is.
It flows through evolve.py's existing whole-record redaction before being
written to disk / the uploaded artifact, so this closes the diagnostic gap
without a new blind CI dispatch: the next failure of this shape is
readable directly from the artifact.
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The first real skill-evolution run got past task binding, then the proposer
session exited 1 with "Permission mode forced to default —
CLAUDE_CODE_SUBPROCESS_ENV_SCRUB is set (allowed_non_write_users hardening)".
On 2.1.214 the permission resolver unconditionally forces permission mode to
"default" whenever CLAUDE_CODE_SUBPROCESS_ENV_SCRUB is set — `--permission-mode
dontAsk`, settings `permissions.defaultMode`, and `autoAllowBashIfSandboxed`
are all ignored for the mode decision. The proposer runs headless `-p --bare`
where Bash is its only writable tool (it writes the candidate overlay); under
forced "default" Bash was no longer auto-approved, so the session blocked.
We cannot set ENV_SCRUB=0 (it scrubs the Anthropic auth token from the
sandboxed proposer's Bash subprocesses). Instead, align with the forced mode:
pre-approve the proposer's exact tool surface via settings `permissions.allow`
(["Read","Grep","Glob","Bash"]) — under "default" a tool runs without a prompt
iff it matches an allow rule — and stop requesting a non-default mode so no
warning fires. ENV_SCRUB and the full sandbox filesystem/network lockdown are
unchanged. The real-binary containment canary is updated to the new invocation
(no --permission-mode) so the CI job is the authoritative empirical gate, and a
fast unit assertion pins the new permissions.allow / absent defaultMode.
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(ci): install root and shared node_modules for the evolution benchmark
The first real workflow_dispatch of the skill-evolution loop failed at task
binding: capture_task_dependency_binding aborted with
SandboxError: sandbox_copy path is unavailable: node_modules: No such
file or directory
The benchmark tasks sandbox-copy node_modules from three locations
(tasks.scenarios.yaml) — the monorepo root, gitnexus-shared, and gitnexus —
mirroring a full dev checkout. The install step only ran `npm ci` in
gitnexus/, so the root and gitnexus-shared node_modules never existed and
the loop died before any agent ran. Install all three (root, then build
gitnexus-shared, then build gitnexus), matching the per-package install in
ci-tests.yml plus the root deps the tasks require.
A new contract test pins all three installs so this fails in CI rather than
on the next real run — the same guard the workflow's other two P1 fixes got.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): only add the missing root install; the subpackage steps already exist
The initial fix redundantly rebuilt gitnexus-shared and gitnexus inside the
gitnexus step — but the workflow already builds both in their own dedicated
steps. Only the monorepo root node_modules was missing. Add a single
"Install monorepo root dependencies" step and leave the two subpackage
build steps untouched, so the benchmark's root sandbox_copy resolves without
double-building.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
---------
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(ci): review agent on Sonnet 5 with structured, linked reviews
Bump the pinned review model from claude-sonnet-4-5-20250929 to
claude-sonnet-5 (verified against the pinned Claude Code 2.1.214 with
subscription auth and --json-schema structured output).
Restructure the published review body: verdict-first summary, findings
ordered by severity, fixed section order, and every file or symbol
reference as a GitHub permalink pinned to the analyzed head SHA (or the
merge-base SHA for deleted and rename-old paths) instead of bare
path:line text, so references are clickable and render inline previews.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* feat(ci): review agent runs as a coordinated reviewer swarm
Implement the review skill's expert-lens section in CI: the main agent
spawns four trusted lanes in parallel via the Task tool — correctness,
security, blast-radius, and coverage — each a purpose-built persona
restricted to Read/Glob/Grep plus the read-only graph MCP tools.
Personas live in the canonical skill tree (mirrored to all shipped
copies) and are installed into the reviewer's user-scope agents dir from
the exact control SHA, so a hostile PR head can never define a lane.
Lane reports are treated as unverified claims: the main agent re-anchors
findings before publishing, and the publisher's context-evidence gate
still requires the main conversation's own successful context call.
Bash and the newer Agent tool remain disallowed for every context; the
analyze timeout gets swarm headroom (45 -> 60 minutes). The workflow
contract test now pins the swarm posture.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): sidechain tool calls can no longer satisfy the evidence gate
The review agent's own review of this PR found that proveGraphReview()
walked the flat transcript without reading parent_tool_use_id, so a
spawned lane's context call could satisfy the publisher's graph-evidence
gate the prompt reserves for the orchestrator. Entries with a non-null
parent_tool_use_id are still strictly validated (malformed linkage fails
the transcript) but are excluded from both candidate context calls and
qualifying results; a new fixture proves sidechain-only evidence is
rejected while mainline evidence beside sidechain turns still passes.
Also gives the orchestrator turn headroom for the four dispatched lanes
(--max-turns 100 -> 150), addressing the review's LOW finding.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* refactor(skills): swarm-lane dispatch belongs to the review skill
Move the lane orchestration out of the workflow prompt and into the
gitnexus-review skill itself: a new "Swarm lanes" section names the four
ci-persona lanes, defines when and how to dispatch them (parallel, one
message, per-lane context and file slices), and owns the verification
contract (lane reports are unverified claims; re-anchor, dedup, drop
unanchored findings; lanes structure the work but never gate it). Any
runner of the skill — the CI workflow or a local harness — now triggers
the lanes from one canonical definition.
The workflow prompt keeps only its CI-specific deltas: the lanes'
trusted-control-SHA install provenance, the Task-tool dispatch surface,
and the publisher's orchestrator-only context-evidence gate. Mirrors
synced; 122 contract tests pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* feat(skills): add adversarial finder lane and critic gate to the swarm
ci-adversarial-lens joins the parallel finder wave: it assumes the change
is broken and constructs reachable failure scenarios — interleavings,
hostile inputs, state corruption, abuse of newly exposed surfaces — each
verified to a concrete entry point before it may be reported.
ci-critic-lens runs last as a gate on the orchestrator's finished draft:
it audits anchoring, concreteness, severity calibration, format
conformance, and honesty, returning PASS or a numbered defect list with
the smallest repair per item. The skill bounds it to two passes and the
critic hardens the review without ever blocking it; the workflow inherits
both lanes automatically through the wholesale ci-personas install.
Mirrors synced across all three shipped trees; 122 contract tests pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* refactor(ci): workflow defers the whole swarm contract to the skill
Now that the skill's Swarm lanes section owns dispatch, verification, the
critic gate, and the fallbacks, the workflow prompt stops restating any
of it. It contributes only what CI alone knows: the lanes' control-SHA
install provenance, the concrete environment mapping for lane inputs
(diff, manifest, head and merge-base checkouts, exact SHAs), and the one
CI override — the publisher's context-evidence gate remains
orchestrator-only. Analyze timeout gains headroom for the critic's
sequential rounds (60 -> 75 minutes).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): dispatch swarm lanes via the Agent tool, not the renamed Task alias
On the pinned Claude Code 2.1.214 the subagent-dispatch tool is `Agent`
(`Task` was renamed to `Agent` in 2.1.63 and is now a legacy alias), and
permission rules evaluate deny before allow. The workflow allowed `Task`
and denied `Agent`, so the orchestrator could never dispatch a lane and
every review silently fell back to the inline single-agent path while the
text-only tests certified the broken config.
Use `Agent` consistently: add it to --tools, allow it scoped to the six
ci-personas (`Agent(ci-correctness-lens,...,ci-critic-lens)`), remove it
from --disallowedTools, and update the prompt. Tests now match the scoped
allowlist on the raw string (commas inside Agent(...) break a split) and
assert Agent is no longer bare-denied.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): harden swarm permissions — allow Glob/Grep + merge-base reads, quarantine PR-head agents
Three permission-hygiene gaps around the swarm dispatch:
- Glob/Grep were in --tools but had no allow rule, so the lanes' declared
tools could manufacture denied-tool errors; allow them (read-only,
sandboxed by cwd + add-dir).
- The prompt hands lanes the merge-base source checkout for deleted /
rename-old symbols, but no Read rule covered it; add a scoped Read()
allow (which grants access without triggering --add-dir agent discovery).
- The --add-dir PR-head copy is scanned for spawnable agent definitions and
the pinned runtime has no suppression env, so a PR could ship its own
.claude/agents/*.md. Drop that subtree from the materialized copy after
checkout-index (skills left intact), so only the trusted control-SHA
personas can ever be dispatched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* test(ci): pin the Agent allowlist to the ci-personas; require a dispatch canary
A text-only assertion cannot prove the pinned CLI actually dispatches the
lanes (print mode silently ignores invalid settings and does not validate
Agent(type) content at parse time) — that is what let the original
Task/Agent inversion pass CI. Two mitigations for the class:
- A cross-consistency test asserts the six names in the Agent(...) allowlist
equal the six ci-personas filenames and each persona's frontmatter name,
so a rename or typo in any of the three fails without auth.
- The activation checklist now requires the post-merge canary to prove a
positive dispatch AND an unlisted-type refusal before enabling the trigger.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): bound swarm transcript volume with per-persona maxTurns
The six lanes stream into the single execution transcript the publisher
validates, but the personas carried no turn budget, so a large-PR swarm
run could overflow the (hard-throw) transcript caps and brick a valid
review. Bound each lane deterministically — finders maxTurns 12, the
critic maxTurns 6 — which keeps the worst case (~2×(150+5×12+2×6) ≈ 444
messages) under the unchanged 1_000 cap, so no cap needs raising. A new
test encodes that invariant: it fails if a persona's maxTurns is bumped
without revisiting the cap. Applied byte-identically across all four
shipped skill trees.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* test(ci): independently pin both sidechain evidence-gate guards
The sidechain-exclusion guards at candidate registration and result
acceptance were mutually redundant on realistic transcripts (a real
sidechain turn carries parent_tool_use_id on both its call and result),
so deleting either guard alone still passed the whole suite. Add two
asymmetric cross-wired fixtures — a mainline call with a sidechain result
(pins the acceptance guard) and a sidechain call with a mainline result
(pins the registration guard), both expecting missing_graph_evidence.
Mutation-verified: deleting either guard alone now reddens the suite.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* docs(skills): require own evidence before dispatch; document critic fail-open and swarm naming
Strengthen the gitnexus-review "Swarm lanes" contract (all four mirrors):
- The orchestrator must make its own graph context call on a changed
symbol before dispatching any lane, so a fully-delegated run cannot
leave the publisher's evidence gate unsatisfied (mirrored into the
workflow prompt, with a test pinning the ordering phrase).
- Document that the critic's fail-open is deliberate (bounded to two
passes, cannot deadlock, review still gated by evidence + schema),
and distinguish it from the hard lane-7 gate in the separate
gitnexus-pr-swarm-review skill.
- Give a concrete local-harness registration pointer for ci-personas.
- Add a reciprocal cross-reference in gitnexus-pr-swarm-review (single
path — that skill is not part of the mirrored family).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* docs: record the review-agent swarm capability (AGENTS.md, CLAUDE.md, reviewer-swarm README)
Reflect the shipped swarm in the standing docs: bump AGENTS.md to 1.14.0
and CLAUDE.md to 1.8.0 with changelog rows, extend the gitnexus-review
description to mention the ci-personas swarm lanes, and refresh the
reviewer-swarm README so its differentiator names the real distinction
(interactive on-demand swarm vs the CI review agent's in-workflow lanes)
now that both run swarms. No CHANGELOG.md edit (feature-PR rule).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): close pre-push review findings on the swarm permission change
Adversarial review of the fix diff caught two issues introduced by the
permission-hygiene commit:
- Bare Glob/Grep in --allowedTools are separate tools that the Read()-scoped
path denies (/proc, github.workspace, ...) do not cover, opening an
undenied read path to the raw checkouts and host paths via a prompt-
injected lane. Drop the bare allow — under dontAsk they stay denied by
omission; lanes read via the scoped Read() rules and the graph MCP.
- The agents quarantine removed only the add-dir root's .claude/agents; make
it recursive so a nested (e.g. monorepo subpackage) .claude/agents cannot
survive and be discovered.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* feat(ci): post an "in progress" marker while the review swarm runs
Swarm reviews can take up to 75 minutes, and until now the PR showed no
sign a review was running. Add a dedicated write-scoped `acknowledge` job
that, under the same authorization gate as analyze, upserts a per-PR
"🔄 GitNexus review in progress" sticky comment linking to the live run
(and reacts 👀 to the trigger comment); the publisher removes that marker
when the review — or a clean failure — posts.
The marker lives in its own job so the model-facing analyze job stays
secretless and read-only: it cannot post to the PR, so per-lane live
progress isn't exposed there — the marker is a binary "running" state with
a link to the Actions run where lane-by-lane progress is visible.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(eval): run the skill-evolution loop online
Add a scheduled + dispatch-gated workflow that runs the offline
propose -> benchmark -> gate loop (workflow_bench.evolve) in CI with the
pinned Claude canary runtime and bubblewrap containment, uploads the
benchmark evidence as an artifact, and on a gate-passed promotion opens
a human-reviewed PR via the release App token. The applied overlay is
bounded to the canonical skill tree and its shipped mirrors; any escape
fails the run instead of reaching a PR.
The scheduled lane ships disabled behind GITNEXUS_EVOLUTION_ENABLED and
requires the new GITNEXUS_BENCH_AUTH_TOKEN secret (benchmark sessions
bill real API usage), mirroring the review agent's staged rollout.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): restructure promotion-PR script so no lint suppression is needed
Replace the inline single-quoted credential helper with a GIT_ASKPASS
file written via a quoted heredoc (the App token still reaches git only
through step env at push time), and assemble the PR body from quoted
heredocs plus double-quoted printf instead of a backtick-laden
single-quoted template. Every run script in the workflow now passes
shellcheck with zero findings and zero disables; the body and askpass
rendering are smoke-tested.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): apply gate-passing overlays in the evolution loop
The loop invoked workflow_bench.evolve without --apply, so
apply_promoted_overlay (its only working-tree writer, gated by
`if args.apply:`) never ran. git status stayed clean, promoted=false was
emitted every run, and the App-token/PR-open steps were unreachable dead
code — a gate-passing run went green as "No promotion this run".
validate_promotion_for_apply already runs before the apply gate, so
adding --apply lets a passing candidate reach the tree without weakening
the deterministic gate; the boundary check then confirms it stayed in the
skill trees.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): provision ~/GitNexus so the benchmark repo resolves on CI
Every scenario in tasks.scenarios.yaml addresses the target repo as
~/GitNexus; runner_tasks.py resolves it with expanduser().resolve() then
`git -C <repo> rev-parse`, which raises when the path is missing. On a
hosted runner the checkout lands in $GITHUB_WORKSPACE and nothing created
~/GitNexus, so the first real run failed at task-binding.
Symlink ~/GitNexus -> $GITHUB_WORKSPACE before the loop. The checkout uses
fetch-depth: 0 (full history for the parentless clone), and the benchmark
only clones the repo copy-on-write and mounts deps read-only, so the
checkout is never mutated. GITNEXUS_BENCH_ORACLE_ROOT stays unset — it
defaults to the in-repo oracles dir and is staged by the harness.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): harden promotion summary output and PR branch recovery
Three fixes to the promotion-detection and PR-open steps:
- GITHUB_OUTPUT summary used a fixed `PROMOTION_EOF` heredoc delimiter; a
value containing that marker on its own line could close the block early
and inject output keys. Use a per-run random delimiter, matching the
pattern already in tree-sitter-upgrade-readiness.yml.
- The summary concatenated every generation's promotion.json (including
rejected ones), so the PR body could show a losing generation's
decisions. The loop returns on the first promotion, so emit only the
highest-numbered gen-N/bench/promotion.json — the decision that fired.
- The promotion branch name omitted the run attempt. GITHUB_RUN_ID is
stable across re-runs, so a re-run after push-succeeds/PR-create-fails
could never push. Include ${GITHUB_RUN_ATTEMPT} (the artifact name
already does).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): least-privilege the promotion App token and gate on an Environment
The Mint-App-Token step passed only app-id + private-key, so the minted
token inherited every permission the Release App installation holds
(including Workflows: write) — far more than "push a branch, open a PR".
Switch to `client-id` (as publish.yml does) and request only
permission-contents: write + permission-pull-requests: write.
Bind the job to a protected Environment (gitnexus-evolution) so promotion
runs can be gated server-side. workflow_dispatch runs the workflow and
in-tree evolve.py from the *dispatched ref*, so a code-side ref guard is
removable by the dispatched branch itself; an Environment deployment-branch
rule (main only) is the boundary that holds. The admin steps to create it
and scope the secrets are documented in the activation checklist.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(ci): correct upload-artifact pin comment and add shell strict-mode
- The upload-artifact SHA 043fb46d… is v7.0.1 (labeled so in the sibling
workflows that pin it); the comment mislabeled it # v6.0.0. Correct the
comment; the pin is unchanged.
- Add `set -euo pipefail` to the two build steps that lacked it, matching
every other run block in the file (GitHub's default shell already sets
-eo pipefail; this adds -u and consistency).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* docs(ci): complete the skill-evolution activation checklist
- Add RELEASE_APP_ID / RELEASE_APP_PRIVATE_KEY to the required-secrets
checklist (the Mint step hard-fails without them on a promotion) and the
App-install-scope verification.
- Document the protected Environment admin step and why it is the real
boundary for the workflow_dispatch ref-secret exposure.
- Note that workflow_dispatch runs the billing loop regardless of
GITNEXUS_EVOLUTION_ENABLED.
- Justify the weekly cron against the README's ~90-day guidance and note the
355-minute timeout ceiling.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* fix(eval): redact API tokens from diagnostic fields before artifact upload
results.jsonl (runner.py) and proposer-session.json (evolve.py) serialize
session records whose error_detail can carry a stderr_tail that echoed the
API key. Transcripts are redacted before persistence, but these two sinks
were not, and both land in the 14-day evolution artifact.
Run each record's serialized JSON through the existing redact_text with the
run's auth token before writing. Scoped to these diagnostic sinks only: the
promoted overlay and proposal.md are left untouched (the overlay is the
applied artifact and must stay byte-identical for apply and the
shipped-skills-sync guard).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* test(ci): add a contract test for the skill-evolution workflow
No test exercised this workflow's path, which is why both P1 blockers
(missing --apply, unresolvable ~/GitNexus task repo) reached production.
Parse the workflow YAML and assert the structural contract: --apply is
passed, the task repo is provisioned, the promotion branch carries the run
attempt, the App token is permission-scoped and the job is Environment-
gated, the output summary uses a random delimiter and a single generation,
the artifact pin is labelled correctly, and every multi-line shell step
sets strict mode. Follows the review-agent-workflow.test.ts precedent.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* feat(ci): run the proposer on its own (stronger) model
One `model` input drove both the benchmark arms and the proposer/diagnosis
session. Split them: `model` stays the benchmark arms (match the model your
skill users run, so a promotion is valid for them and the tasks aren't
ceiling-saturated), and a new `proposer_model` input runs the proposer —
the harder meta-reasoning task that writes the candidate skill, and only one
session per generation, so a stronger model is cheap here. evolve.py already
supports --proposer-model; the workflow just didn't expose it.
Defaults: arms = claude-sonnet-5, proposer = claude-opus-4-8 (both
overridable via workflow_dispatch). The weekly cadence bounds the added
spend. Contract test asserts the split stays wired.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Bump the pinned review model from claude-sonnet-4-5-20250929 to
claude-sonnet-5 (verified against the pinned Claude Code 2.1.214 with
subscription auth and --json-schema structured output).
Restructure the published review body: verdict-first summary, findings
ordered by severity, fixed section order, and every file or symbol
reference as a GitHub permalink pinned to the analyzed head SHA (or the
merge-base SHA for deleted and rename-old paths) instead of bare
path:line text, so references are clickable and render inline previews.
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(ci): unblock the review agent dispatch and publisher lanes
The first workflow_dispatch validation run surfaced two defects:
- setup-node rejects `cache: false` (the YAML boolean arrives as the
string 'false' and v6 fails with "Caching for 'false' is not
supported"), killing the analyze job before the isolation preflight.
Omitting the input is the supported way to disable caching.
- The publisher held only `issues: write`, but GITHUB_TOKEN needs
`pull-requests: write` to create issue comments on a pull request,
so even the safe-failure comment died with "Resource not accessible
by integration".
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
* test(ci): align the publisher permission contract with PR commenting
The workflow contract test pinned the publisher to pull-requests: read,
which is exactly the permission set that made comment publication fail.
Encode the corrected scope and assert the publisher still cannot write
repository contents.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Va5uu9Ar3e45QZ5xFsG4AZ
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): add ce-plan — GitNexus+PDG implementation-planning skill
Adds .claude/skills/ce-plan: a planning-only skill that builds
implementation-ready plans from GitNexus graph navigation (query/context/
impact/trace), bounded statement-level PDG slices (pdg_query, impact
mode:pdg, explain), and targeted source verification, with a context
ledger to prevent repeated reads and a machine-readable implementation
context pack (stable contract for a future ce-implement). Whitelisted in
.gitignore and registered in AGENTS.md and CLAUDE.md outside the
auto-managed gitnexus block.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(skills): apply ce-plan validation findings (tool contract, consistency, conventions)
Tool contract: impact mode:'pdg' shape now includes the schema-required
direction param; CDG branch sense documented as the result 'label' field
(reason is cypher/raw-edge only); explain caveats corrected to its real
false-negative classes (cross-function TAINT_PATH is modeled).
Consistency: PDG slice homed in working memory (ledger keeps one-liners);
depth knob defined and category-overrides-baseline ordering stated;
call_depth (consumed by nothing) and content-hash bookkeeping dropped;
Never section folded into Hard rules; Phase 3 deduplicated to a pointer;
allowed-repeat escalations defined; budget/discard accounting clarified;
verification-commands gathering added to Phase 4; open_questions added to
the context pack.
From scenario runs: plans now pin the verified-at HEAD commit and index
freshness in a header, tag claims [verified]/[graph]/[inferred]/[assumed],
quote load-bearing tool output, prefer pre-hook-carrying npm scripts, and
support an out:<path> destination override; output path defined as the
Phase 1 target repo root.
Conventions: AGENTS.md 1.9.0 / CLAUDE.md 1.4.0 changelog rows + metadata
bumps; future ce-implement qualified as future.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): rename ce-plan → gitnexus-plan; add cross-CLI (Codex) entrypoints
Renames the skill dir, frontmatter, output filename convention, plan H1
(GitNexus Engineering Plan), the future executor handle
(gitnexus-implement), the .gitignore whitelist entry, and all
AGENTS.md/CLAUDE.md references. Follows the pr-swarm-review cross-CLI
pattern: SKILL.md is the canonical CLI-neutral spec, AGENTS.md § Engineering
planning is the Codex/any-agent entrypoint, and the README documents the
optional user-level ~/.codex/prompts/gitnexus-plan.md slash command plus an
invocation matrix. Skill prose de-branded from Claude Code (agent-neutral
verification layer).
Also fixes two post-review README contradictions: the anti-reread claim now
names the ledger's allowed escalations, and 'read-only by contract' is now
'planning-only' (the skill writes exactly one repo file — the plan); the
scope-creep rule and template §12 now agree on where deferred follow-ups
land. Drops the stale plugin-collision limitation.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(skills): document Codex user-level install path for gitnexus-plan
Codex discovers SKILL.md skills from ~/.agents/skills (same path the other
gitnexus-* skills install to); README now documents the cp install plus the
optional ~/.codex/prompts slash-command file, with the prompt body preferring
the repo copy and falling back to the user-level install.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): gitnexus-plan freshness gate + active PDG-layer refresh
Freshness is now a Phase 1 gate, not advisory: under the default
freshness:strict, a stale index is refreshed once per planning session via
node .gitnexus/run.cjs analyze --index-only (appending --pdg when the task
will reach the PDG phase), then the context resource is re-read. A missing
PDG layer likewise triggers the one permitted --index-only --pdg refresh
and re-probe instead of a passive recommendation. freshness:accept (or a
failed/impractical refresh) preserves the old behavior: plan on the stale
graph, source-weighted, labelled in the plan header. --index-only is the
load-bearing flag choice — it suppresses all file generation, so the
planning-only contract holds (only the .gitnexus store changes). Ledger
gains an index_refresh record; plan header states fresh / refreshed /
refresh-skipped-with-reason.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): gitnexus-plan runner build check before freshness refresh
When the target repo builds the analyzer from its own source (bin → dist/
mapping, as gitnexus/ does), the Phase 1 freshness gate now verifies dist/
is current before running the analyze refresh — rebuilding via the
package's build script when any analyzer source file is newer than the
built entrypoint — and prefers that freshly built CLI. Otherwise a stale
dist re-indexes with outdated extraction logic and the 'fresh' index lies.
Rebuilds are recorded in the ledger's index_refresh; the PDG-phase refresh
inherits the same check.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): add gitnexus-work executor and gitnexus-lfg pipeline
gitnexus-work executes a gitnexus-plan as verified atomic commits: consumes
the §11 implementation_context pack, drift-checks the plan's evidence pin
against HEAD, re-verifies assumptions before relying on them, runs impact
before every symbol edit and detect_changes before every commit (repo
mandates), builds tests from the plan's scenarios, and routes structural
drift back to gitnexus-plan Deepen mode instead of coding around it.
gitnexus-lfg is a thin orchestrator: gitnexus-plan → blocking user gate
(deepen / proceed / stop, deepen loops allowed) → gitnexus-work → review
via the existing gitnexus-pr-review skill (open PR, else branch diff vs
default). One bounded fix cycle for review findings; never pushes or opens
a PR on its own.
gitnexus-plan gains a Deepen mode (re-run freshness gate, escalate to
depth:deep, re-verify graph/inferred/assumed claims toward verified,
rewrite the same file); its 'future gitnexus-implement' placeholder is
retired in favor of gitnexus-work. Registered via .gitignore whitelists,
AGENTS.md 1.10.0 (section renamed to Engineering planning & execution),
CLAUDE.md 1.5.0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(skills): apply cross-skill review findings to the gitnexus skill family
Two P1s: gitnexus-plan Deepen mode now re-anchors before re-pinning
(diffs the old evidence pin over every [verified]-claim file and re-reads
or downgrades before the header moves — moving the pin without this
laundered stale claims as verified); the index-refresh budget is stated
once in Phase 1 (one --index-only refresh plus at most one Phase 3 --pdg
upgrade per session, Deepen = its own session) with ledger and pdg-slice
deferring to it.
Contract fixes: gitnexus-work's drift check now covers every file the
pack cites (not just files_to_modify) and parses the full pack incl.
primary/related symbols and acceptance_criteria (walked in Phase 4
alongside §13); a pre-completed check skips §7 steps already landed and
Deepen gains a reconcile-execution-state step, closing the mid-execution
route-back loop; pack assumptions must name what to check and how.
lfg: Lane 4 passes the merge-base to detect_changes compare (two-dot
diff misattributes upstream commits when default advanced), branch-diff
is the stated normal case, oversized review findings route to the plan
gate instead of overflowing direct mode, the one-fix-cycle cap is
explicit on re-run, and headless runs end at the plan gate with the plan
as deliverable. work: blank mode narrowed to *gitnexus-plan*.md with a
re-execution guard, direct-mode discipline spelled out, branch
meaningfulness defined against the plan slug, and the plan document is
committed as the branch's docs commit (review diff includes it).
Planning-only contract now names the dist/ rebuild as the second
permitted state change; Phase 5.1 names the four claim tags; stale
AGENTS.md anchors fixed.
Known latent issue left untouched: gitnexus/gitnexus-pr-review pairs a
three-dot example with a two-dot detect_changes compare — that skill is
also shipped by the plugin, so fixing it here would drift the copies;
lfg compensates by passing the merge-base.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): ship the engineering skill family with the gitnexus package
npm i -g gitnexus users now get gitnexus-plan / gitnexus-work / gitnexus-lfg:
the three skills are added to gitnexus/skills/ in directory form (SKILL.md +
references/), which installSkillsTo already enumerates dynamically and copies
recursively to every editor target (~/.agents/skills for Codex, Cursor,
OpenCode, Qoder, ...) on gitnexus setup — uninstall enumerates the same root,
so removal stays clean. The Claude Code plugin channel
(gitnexus-claude-plugin/skills/) carries the same copies plus the standard
per-skill mcp.json.
Global-install support in the skill text: gitnexus-plan Phase 1 now resolves
the analyzer runner explicitly — node .gitnexus/run.cjs analyze when the
project has a runner, else gitnexus analyze (installed CLI), else
npx gitnexus analyze — and all analyze mentions route through it, satisfying
the skills-steering policy (#1939/#1945) which sweeps the plugin copies.
New drift guard test/unit/shipped-skills-sync.test.ts asserts the npm and
plugin copies stay byte-identical to the canonical .claude/skills/ family
(plugin = canonical + mcp.json), same discipline as run.cjs ↔
resolve-invocation.ts. skills-steering + shipped-skills-sync: 11/11 green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(eval): workflow_bench — measure the skill workflow's token savings
Benchmarks gitnexus-plan → gitnexus-work against a baseline agent
(--disallowedTools Skill) on identical tasks, in fresh detached worktrees,
using real headless Claude Code sessions; every number comes from the CLI's
--output-format json usage report (field names validated against a live
2.1.207 session). Reports per-arm medians (input/cache/output tokens, cost,
wall time, turns), a savings row, and resolve status from a per-task verify
command — savings on failed tasks are flagged, not celebrated. Per-task
setup hook prepares fresh worktrees (deps); --permission-mode
bypassPermissions (default) lets sessions run unattended in the throwaway
trees.
Free-model support: --base-url/--auth-token/--model route headless sessions
through any Anthropic-compatible endpoint; free-model.litellm.yaml is a
ready litellm-proxy template for OpenRouter :free variants or local Ollama,
so benchmarking burns no paid tokens (README documents rate limits and the
small-model skill-following caveat).
Harness validated end-to-end with a stub CLI (worktree lifecycle, both
arms, plan→work chaining, verify, aggregation, report) and 4 pytest units
for the pure aggregation/savings/report helpers. AGENTS.md 1.11.0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(eval): record first workflow_bench calibration run
Trivial-task calibration (add -V alias): both arms resolved; workflow arm
~4.3x baseline cost — the documented overhead-dominated regime, recorded so
the regime boundary is empirical rather than asserted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(eval): workflow_bench scenario matrix — arm variants, task classes, churn
Ground-base measurement across scenarios: tasks.scenarios.yaml spans four
labeled classes (trivial → investigation-bug → investigation-feature →
cross-module) with deterministic verifies (prescribed test files). New arms:
workflow_direct (gitnexus-work direct mode — the middle option that locates
the routing boundary lfg's gate and work's triage encode) and baseline_nomcp
(no skills AND no graph tools — separates workflow-discipline value from
GitNexus-tool value; off by default). Records now carry task class and diff
churn (files/+ins/−del vs the starting commit) as an over-engineering proxy;
the report renders a class column and per-arm savings rows vs baseline.
5 pytest units + stub-CLI e2e of the full three-arm matrix.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(eval): record workflow_bench ground base; fix churn measurement bias
Ground base (3 classes x 3 arms, n=1/cell): every arm resolved every task —
pass/fail quality saturates at this difficulty, making the comparison pure
cost. Full plan→work never amortized its ~$9-11 fixed cost on tasks a
baseline finishes in ≤35 turns (−211% to −333% cost); workflow_direct sits
near baseline (−15% to −55%, once faster wall) with more test coverage.
Routing implication recorded: direct mode/plain agent below this scale,
full workflow for cross-module / multi-session / plan-as-deliverable work.
The cross-module cell and multi-run variance are the next measurements.
Churn fix: git add --intent-to-add -A before diffing (arms that never
commit no longer undercount new files) and :(exclude)docs/plans (the
committed plan doc no longer inflates workflow churn); this run's churn
numbers predate the fix and are omitted from the recorded table.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* perf(skills): cost-optimize the workflow from measured ground base
Every optimization targets a measured fixed-cost component
(eval/workflow_bench ground base: workflow arm −211% to −333% vs baseline,
all tasks resolved):
- Plan form is category-priced: compact form (core sections w/ § anchors
preserved, ≤80 lines excl. pack, mini-pack subset of the context pack)
for narrow/default categories; the full 13 sections only for deep work
(refactor/security/performance/concurrency/architecture). A compact plan
outgrowing its cap reclassifies to full rather than overflowing.
- Freshness gate is category-priced: compact categories default to accept
(source-weighted, refresh only when a graph claim becomes load-bearing);
strict stays the default for full-plan categories — the rebuild+re-index
was the largest single fixed cost.
- Turn economy: per-category tool-call budgets (~10 to ~45; architecture
uncapped); budget exhaustion routes open questions to §12 instead of
more digging.
- gitnexus-work fast path: HEAD == evidence pin → skip all citation
re-reading (the pin's entire point); mini-pack fields tolerated.
- lfg Lane 1 boundary triage: tasks below the measured ~35-turn boundary
get offered gitnexus-work direct mode before the plan lane is spent.
Copies re-synced (npm skills/, plugin, ~/.agents); steering + sync guards
green. Re-measurement of the workflow arm follows to verify the numbers
actually improve.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(eval): record optimization re-measurement — inv-bug workflow cell −20% cost
Same task, same conditions, post-830a0459 skills: $14.56→$11.70 (−20%),
83→72 turns, cache_read −24%; verified in-transcript that the compact form,
turn budget, and skipped rebuild/re-index all fired. Wall +15% from a work-
session test-debugging tail (n=1 variance). Regime unchanged (~3.5x baseline
on this class) — routing rule stands.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(eval): per-arm clone isolation — worktree ref-namespace leak contaminated an arm
The cross-module workflow_direct cell reported an impossible 28-turn solve
with churn byte-identical to the workflow arm: git worktree add shares the
repo's ref namespace, so the workflow arm's slug branch (created by
gitnexus-work Phase 2) survived worktree removal and the direct arm found
and adopted the completed work. Arms now get isolated git clone --shared
copies (object store via alternates, refs clone-local — agent branches and
stashes die with the clone; origin/<ref> fallback for non-default refs).
Leaked branch deleted; baseline arm verified clean (0 branch references in
its transcript); cell marked invalidated pending re-run.
Records the valid cross-module cells: workflow $18.32 vs baseline $18.03
(premium −1.6%, vs −211%..−333% on smaller classes) — fixed costs amortize
at this scale, with a less destructive diff and a plan artifact as bonus;
resolve rate still tied. Churn fingerprinting is what caught the
contamination — noted in the README as an integrity check.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(eval): complete cross-module cell — direct mode wins 47% cost / 56% wall
Clean clone-isolated re-run: workflow_direct resolved the hardest class at
$9.53/52 turns/15m vs $18.03/98/34m baseline and $18.32/107/37m full
workflow. The measured story across all four classes: the execution
discipline (gitnexus-work) is the consistent sweet spot and delivers real
token savings on hard tasks; the planning pass buys its artifact, not
same-session savings. Resolve rate tied everywhere (n=1/cell caveat).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(eval): add trajectory-gated skill evolution (#2431)
- Pair prompt candidates with incumbent workflow arms
- Gate promotions on pinned-model quality and efficiency
- Expire router evidence and document its lifecycle
* fix(eval): allow pr-review skill candidates
* feat(skills): rename and generalize GitNexus review
* feat(eval): external-comparator and review arms for workflow_bench
- ce_workflow / ce_workflow_direct: compound-engineering ce-plan/ce-work
arms prompted with the same structure as the gitnexus arms
- review / ce_review: gitnexus-review vs ce-code-review on an identical
diff applied by the task's setup
- plan handoff is snapshot-based: committed example plans in docs/plans/
tie on clone mtimes and broke the name-glob pick (executed a stale plan)
- verify output tail is recorded per run and the final working-tree patch
is kept, so failed rows are diagnosable after the clone is destroyed
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(skills,eval): address #2431 review — data-safe rename migration, fail-closed bench evidence
- setup: never delete a legacy renamed skill dir — the installer cannot
prove ownership (users customize or hand-write skills under these
names); warn with the path instead, and the test now asserts survival
- workflow_bench: fail closed when a session's --output-format json
report is empty, malformed, or missing usage fields — an exit-0 shell
with no parseable usage no longer counts as measured evidence
(5 parametrized regression tests)
- workflow_bench: document the trust model prominently (task setup/verify
are shell-executed, sessions run bypassPermissions with the parent env,
candidate overlays are prompt injection surface) in README + docstring
- free-model.litellm.yaml: master_key from LITELLM_MASTER_KEY env instead
of a static token; loopback-binding warning
- ci: run the eval workflow_bench pytest suite on ubuntu (pytest+pyyaml
only — no full eval stack)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(eval): demand observed foreground verification in headless work-arm prompts
In a headless -p session there is no later turn: a work arm backgrounded
its slow test run, scheduled wakeups that can never fire, and reported
done while two of its tests failed. All four work-arm prompts (both
skill families, symmetric) now require verification output to be
observed inside the session.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): ask plan depth up front instead of offering deepen afterwards
gitnexus-plan Phase 0 now asks one blocking question in interactive
sessions — quick / standard / deep, mapped onto the existing depth/form/
freshness knobs — when the invocation carries no explicit depth signal.
Explicit knobs and headless runs skip the question (category posture
unchanged, so benchmarks and automation behave as before).
gitnexus-lfg's plan gate slims to proceed/stop: depth was already the
user's up-front choice, so deepening is no longer offered by default —
an explicit deepen request at the gate and executor route-backs still
run Deepen mode, which remains the mechanism for strengthening an
existing plan document.
All shipped copies resynced (npm skills/, Claude plugin); AGENTS.md
1.13.0 and CLAUDE.md 1.7.0 pointers updated, including the analyzer's
regenerated index-stats block at this branch's head.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): taint pass, expert lenses, and post-work index refresh
gitnexus-review gains a PDG-backed taint-and-dependence pass (explain +
pdg_query, --pdg folded into the stale refresh on trust-boundary diffs) and
an Expert lenses section: domain reviewers derived from the graph's
clusters plus four cross-cutting lenses (architectural fit, language
conformance per the repo's own contract, Definition of Done, simplicity),
dispatched once after the evidence-gathering steps and scaled to the diff.
gitnexus-work Phase 4 now refreshes the knowledge graph after the DoD walk
via the resolved-runner ladder with analyze --index-only, so the lfg review
lane and later sessions query the finished work without dirtying the tree.
lfg's threshold-governance paragraph moves to its README; eval citations
are tagged as measured in the GitNexus repo. All shipped copies re-synced.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(cli): remove legacy gitnexus-pr-review on uninstall; cover the rename migration
uninstall's removal set now includes LEGACY_SKILL_DIR_NAMES derived from
RENAMED_SKILL_DIRS, so a pre-rename install is cleaned up instead of
orphaned. The rename warning gains behavioral coverage (fires with a legacy
dir present, silent without), and shipped-skills-sync asserts legacy names
stay absent from every shipped tree.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(eval): metric provenance, error-kind rows, skill-invocation verification, gate noise floor
The promotion gate defaults to cost_usd (the only metric that includes
subagent spend); token metrics carry an explicit main-loop-only warning in
the report and promotion.json. Rows are classified by error_kind
(session-error / verify-failed / infra-error), excluded from efficiency
medians, and the gate requires equal valid-run counts. Each session's
transcript is scanned for the expected Skill invocation and fails closed on
a verified miss; a one-run resolution edge no longer promotes (noise
floor). Per-run timeouts and setup failures record an infra-error row
instead of aborting the sweep. Overlays touching skills no candidate arm
exercises are rejected up front.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: fix skill routing paths, version headers, and skill rosters
Routing tables point at the tracked direct skill paths (matching the
post-#2434 generator output), AGENTS.md/CLAUDE.md headers match their
latest changelog rows, the 1.12.0 row describes what the migration actually
does, package/cursor READMEs list the full shipped skill roster, and the
swarm READMEs describe /gitnexus-review's expert lenses instead of calling
it single-agent.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: drift-guard workflow for skill copies; pin eval pip deps; track docs/plans
ci.yml ignores '**.md', so an md-only skill edit would merge without the
shipped-skills-sync test running — skill-sync.yml triggers exactly on the
guarded trees. The eval job's pip install is version-pinned, and
docs/plans/ is unignored so gitnexus-plan output can be committed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ci): keep the runner-invocation literal in gitnexus-review; add concurrency block to skill-sync
skills-steering requires skills with a stale-index hint to carry the exact
'node .gitnexus/run.cjs analyze' form — restore it with the fallback ladder
as a parenthetical instead of replacing it. skill-sync.yml gains the
top-level concurrency block the workflow-convention check enforces.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(skills): token-economy guidance for expert lenses
Merge lenses that ground in the same material into one reviewer, and use
cheaper model/effort tiers for mechanical lenses where the harness offers
them, reserving the strongest engine for adversarial judgment.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(eval): isolate transcript home on Windows
Ensure workflow_bench transcript tests set USERPROFILE alongside HOME so Path.home() resolves to the temporary test home on Windows.
* docs(skills): fold PR #2522 execution learnings into review/work/plan
Eight incident-backed hardenings from running the full skill cycle
(review -> plan -> work, 28-finding fix series) on PR #2522:
gitnexus-review:
- Expert lenses execute the code under review on candidate failing shapes
(empirical probe outranks source reading — every HIGH the language
lenses found came from a probe, not a read).
- Step 7 re-runs the exact CI check for refreshed baselines/fingerprints
(a stale committed artifact is invisible in the diff; caught a red
benchmarks arm).
- Step 8 treats version/invalidation constants as review surface
(INCREMENTAL_SCHEMA_VERSION class recurred verbatim from #2494).
gitnexus-work:
- Step 4 proves regression tests discriminate against the pre-fix tree.
- Step 5 rebuilds executed build output before every verification run
(parse workers load dist/; a correct fix 'failed' until rebuilt).
- Step 6 makes stage -> detect_changes -> commit one unbroken sequence.
gitnexus-plan:
- Phase 0 seeded-evidence mode: plan FROM a completed review's verified
findings instead of re-running the graph ladder.
- Template §7: fingerprint/golden-guarded output rebaselines once, at the
series tip.
All distribution copies resynced; shipped-skills-sync + skills-steering
24/24 locally.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(eval): close the skill-evolution loop with an automated proposer driver
workflow_bench.evolve adds the three arrows the README described as manual:
a proposer session that turns loser trajectories (results.jsonl rows,
transcripts, patches, the learning queue) into ONE bounded candidate
overlay, a driver that iterates propose -> paired benchmark -> deterministic
gate up to --generations, and an --apply step that copies a promoted
overlay onto the canonical skills and shipped mirrors as a working-tree
diff. The trust boundary is unchanged: overlays re-validate through
candidate_overlay_files before any benchmark or apply consumes them, and
committing, CI, and the PR merge stay human.
learnings.jsonl is gitignored: it is machine-local evidence, like the
session transcripts it complements.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(skills): route live-task friction into the evolution learning queue
Each family skill gains a short 'Skill feedback' section: on friction with
the skill's own instructions, append one JSON line to
eval/workflow_bench/learnings.jsonl (GitNexus repo only) — never self-edit
the skill from a live task. The proposer in workflow_bench.evolve consumes
the queue as hints; a learning reaches a shipped skill only by beating the
incumbent on the paired benchmark. All shipped mirrors re-copied byte-
identical.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci(tests): run the evolve helper tests in the eval pytest job
test_evolve.py needs only pytest+pyyaml, same as the harness tests the job
already runs — without this line the new module had no CI coverage.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(ci): comment-triggered GitNexus review agent for PRs
'@gitnexus review' from a maintainer (OWNER/MEMBER/COLLABORATOR; the action
re-validates write access) runs the repo's gitnexus-review skill headlessly
against the PR and posts the review as a sticky comment — remote triggering
with no local setup. Read-only by construction: contents: read token,
Write/Edit and web tools disallowed, Bash allowlisted to git reads and the
gitnexus CLI; analyze parses PR code with tree-sitter, never executes it.
Requires the ANTHROPIC_API_KEY repository secret; activates once the file
is on the default branch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(ci): dispatch lane + existing OAuth secret for the review agent
Align with claude.yml: same action pin and the CLAUDE_CODE_OAUTH_TOKEN
secret the repo already carries — no new secret to configure. Add a
workflow_dispatch lane (PR number input) so the agent can be triggered from
the Actions UI and tested before the issue_comment trigger reaches the
default branch. Allowlist gh pr view/diff and gh api, which the review
skill uses to pin PR SHAs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ci): close a fork-PR RCE vector in the review agent's tool allowlist
A live headless run of the exact workflow session against PR #2431 (66
turns, full gitnexus-review pass) surfaced a real HIGH-severity confused
deputy: .gitnexus/ is gitignored, not blocked — a fork PR can commit its
own .gitnexus/run.cjs, issue_comment checks out PR-head content, and the
skill's runner ladder tries 'node .gitnexus/run.cjs analyze' first. That
would execute fork-controlled JS inside a job holding
CLAUDE_CODE_OAUTH_TOKEN and a write-scoped GITHUB_TOKEN — the opposite of
the 'PR code is read, never executed' claim in the workflow's own header.
Fix: drop the run.cjs allowlist entry so analyze always resolves through
npx gitnexus (npm registry, not the checked-out tree); the skill's
documented fallback mode covers the resulting graceful degradation. Also
drop 'gh api' (not read-only — accepts -X POST/PATCH/DELETE) and downgrade
pull-requests: write to read (comment posting only needs issues: write;
the prompt already forbids formal review submission).
Same session flagged a latent evolve.py bug: select_evidence's cost sort
used dict.get's missing-key default, which doesn't cover an explicit JSON
null in a foreign --seed-results row and crashes proposer setup with
TypeError. Guarded with 'or 0.0' and added a regression test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: harden PR review and evolution trust boundaries
* ci: follow workflow concurrency convention
* fix(eval): make terminating error paths explicit
* fix: unblock hardened review runtime checks
* test: make containment canaries deterministic
* test: expose Claude canary tool failures
* fix: adapt clean shell environment for Claude
* fix(eval): accept the runner's transcript source key in evidence preflight
The proposer evidence preflight required transcript-artifact metadata to be
exactly {path, sha256, bytes}, but the runner stamps a fourth provenance key
(source=parent-captured-stream-json). Any --seed-results or generation>=2 run
therefore aborted with SandboxError before proposing or promoting. Pin the
producer literal as PARENT_EVENT_STREAM_SOURCE and validate it in the metadata
check, and round-trip real producer output through sum_sessions into the
preflight so the schema can't drift again.
* fix(eval): treat an unmeasured session cost as unavailable, not $0
well_formed validated only the nested usage block, so an otherwise-successful
session missing total_cost_usd was recorded as cost_usd=0.0 — and cost_usd is
the default promotion metric (lower wins), so a cost-less session scored as
free and could win promotion it never earned. Extract cost via measured_cost()
(None on absent/garbage, a measured 0.0 preserved), propagate None through
sum_sessions/aggregate/savings/report, and have the gate refuse to rank on a
metric that was not measured on every run in both arms.
* fix(eval): warn when ranking on the main-loop-only num_turns metric
num_turns comes from the CLI's top-level usage (main-loop session only), like
output_tokens, but selecting it emitted no metric_warning — so a subagent-heavy
candidate could look artificially efficient. Add num_turns to
MAIN_LOOP_ONLY_METRICS and broaden the warning to cover turns.
* fix(eval): fail closed when an overlay adds a file with no committed base
An overlay adding a new .md under gitnexus-{plan,work} passes the structural
overlay checks but has no committed base for committed_destination_base_digests
to bind against, so it raised an uncaught ValueError that crashed the evolve
driver (and runner --candidate-overlay) mid-run. Catch it at both call sites:
evolve reports NOT PROMOTED and exits, runner routes it through parser.error.
* feat(eval): circuit-break the runner sweep on a systemic outage
A sustained upstream outage used to pay out every remaining --timeout window
one session at a time. Track consecutive session/infra/cleanup failures via a
pure systemic_outage_streak helper; after --outage-streak (default 5) in a row,
stop the sweep, still write report.md/promotion.json from partial evidence, and
exit non-zero so evolve.py halts instead of proposing from truncated evidence.
A task's own resolved=False never trips the breaker.
* fix(cli): report a dirty working tree as stale in gitnexus status
status --json (and the human output) computed up-to-date from commit + runner
identity + completeness only, so a repo with uncommitted source changes at a
matching HEAD was reported up-to-date while analyze would still re-index it.
A graph-backed agent gating on that JSON could skip re-analysis on a stale
graph. Extract analyze's dirty-tree check into a shared isWorkingTreeDirty()
in storage/git and fold it into the status freshness decision.
* fix(ci): use single-slash deny globs in the review agent's disallowedTools
github.workspace already expands to an absolute path, so Read(/${{ github.workspace }}/**)
and Read(//proc/**),(//sys/**),(//dev/**) produced double-slash patterns that a
normalizing matcher may not match — silently no-opping the deny layer. Not
exploitable (the allowlist is the primary control and never grants those
paths), but the globs should be well-formed. Update the pinned test strings.
* ci: install gitnexus-shared with npm ci from the committed lockfile
The gitnexus-shared build floated its deps via npm install in three workflows
(skill-sync, ci-tests, and — most importantly — the release publish.yml) while
every other install step uses npm ci. The lockfile is committed and in sync, so
switch all three to npm ci for reproducible, locked installs.
* test(cli): make the shipped-skills drift guard reject symlinks
listFilesRecursive walked with readdirSync and snapshotDir read with
readFileSync, both of which follow symlinks — so a mirror file symlinked to the
canonical tree passed the byte-compare (and a symlinked mirror dir would be
followed too). Reject a symlinked root via lstat and any symlinked entry via
Dirent.isSymbolicLink, with negative tests (skipped on Windows).
* test(eval): guard the candidate-skill vs mirror-root coverage invariant
MIRROR_SKILL_ROOTS omits the Cursor tree, safe only because no candidate skill
is cursor-shipped. Pin that invariant: every CANDIDATE_SKILLS entry must exist
under canonical + every mirror root and must not ship to Cursor, so adding a
cursor-shipped skill to the candidate set (the PR #2488 asymmetric-sync class)
fails loudly instead of syncing three of four trees.
* docs(ci): describe the review agent's staged post-merge rollout
The DoD asked for a dry-run or triggered run before merge, but an issue_comment
(or newly added workflow_dispatch) workflow only ever executes the default-branch
copy, so it cannot be exercised from the PR that introduces it. Reword the DoD
and the activation checklist to a staged rollout: merge registered-but-disabled,
validate same-repo and fork execution post-merge, then enable the variable.
* fix: pin plugin skill mcp.json to the release version via #2445 tooling
The ten plugin skill mcp.json launched `npx -y gitnexus@latest mcp` on every
skill connect — non-reproducible and a supply-chain surface, and (unlike the
persisted setup config) never pinned. Extend sync-plugin-manifests.mjs with an
mcp surface kind that stamps the gitnexus@<version> launch arg, pin all ten to
1.6.9 now, and keep them byte-identical so the drift guard stays green. The
release lifecycle + publish.yml --check now re-stamp them like the four manifest
surfaces; only READMEs stay on @latest as docs.
* test(eval): prove the proposer's built-in file tools are confined
The real-Claude canary only exercised Bash + MCP, so it proved process/MCP
containment but not that the proposer's built-in file tools stay inside their
mounts. Add a canary over the exact PROPOSER_ALLOWED_TOOLS surface and the same
read-only /evidence mount as run_proposer (allowlist extracted to a shared
constant so it can't drift): Read reaches /evidence, a Write into the read-only
evidence mount is denied, and a Write lands in the output tree.
* fix(eval): apply the candidate overlay after task setup for fair arms
The candidate overlay was applied before the task's untrusted setup ran, so
setup could observe candidate prose and the incumbent/candidate arms started
from different pre-overlay state. Reorder within the sandbox: capture the base
(pre-overlay) skill digest, run setup against the base skills, verify setup did
not tamper them, then apply the overlay and capture the post-overlay digest the
model must preserve. apply_candidate_overlay stages path-specific overlay files,
so setup's uncommitted changes stay out of the baseline and churn is unchanged.
Graph freshness for the review arm is handled by the status dirty-tree fix plus
the review skill's stale-triggered re-index, not by reordering the cached
per-task-sha graph materialization (which is mechanically blocked).
* test(eval): end-to-end containment proof of the autonomous proposer
Drives the real run_proposer through bubblewrap with a deterministic scripted
model (no paid API): it reads the read-only evidence bundle and writes a
candidate gitnexus-plan skill edit plus a rationale into the sandbox output
tree; run_proposer enforces the trust boundary and copies only the validated
overlay + proposal out. This exercises the autonomous-proposal stage of the
self-evolution loop end-to-end in the eval/containment CI job (the gate and
apply stages are covered by test_workflow_bench_evolution and
test_promotion_apply). Env-gated on GITNEXUS_REQUIRE_CLAUDE_CANARY, so it runs
only where the pinned Claude binary and user namespaces are available.
* fix(eval): let the proposer author its overlay via Bash
Running the end-to-end proposer canary in the containment CI job surfaced a real
bug: run_proposer starts the session with --bare, which hard-disables the
Write/Edit tools ("Write exists but is not enabled in this context"), yet
allowlisted Edit/Write and omitted Bash. The proposer therefore had no working
way to write its candidate overlay — the self-evolution loop could never produce
a candidate. The sandbox settings already pre-authorize Bash
(autoAllowBashIfSandboxed) and confine writes to workspace/tmp/home, so switch
PROPOSER_ALLOWED_TOOLS to Read/Grep/Glob/Bash and tell the proposer to author
files with Bash. The end-to-end test now drives the real run_proposer through
bubblewrap and asserts a validated overlay + proposal are produced (this also
replaces the earlier file-tool canary, whose Write/Edit premise was moot).
* test(eval): author the proposer overlay with newline-free Bash content
The nested shell-sandbox prefix mangles embedded newlines, so the multi-line
overlay content never landed. Use single-line content for the deterministic
proposer canary.
* test(eval): drop the unverifiable end-to-end proposer canary
The scripted proposer overlay never materialized in the containment job across
runs, and the model tool-result content is not visible in CI logs, so the test
cannot be finalized without an environment where the sandbox can actually run.
Keep the verified production fix (Bash-authoring in run_proposer); the proposer
sandbox/containment stays covered by the existing Bash+MCP and process-tree
canaries.
* test(cli): drop run-analyze.ts from the windowsHide spawn-family list
U7 moved run-analyze.ts's only child_process call (the git status --porcelain
dirty check) into storage/git.ts (already covered by this test, with
windowsHide). run-analyze.ts no longer imports a spawn-family function, so the
windowsHide-regression test's 'must have >=1 spawn call' invariant failed for
it. Remove it from SRC_FILES.
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Zander Raycraft <zanderjraycraft@gmail.com>
Co-authored-by: Azizur Rahman <azizur100389@gmail.com>
* feat(java): JLS 13.1 immediate-host naming for anonymous bodies + v9 schema window (#2555, step 1)
`synthesizeJavaAnonymousClassName` generalizes to both anonymous-body
shapes (`object_creation_expression` with a `class_body`; `enum_constant`
with a `body:` field) and switches from topmost-host naming to JLS 13.1
binary names: the `$`-joined chain of enclosing host types
(`EnumWrap$Mode$1`), numbered per IMMEDIATE host in source order across
both shapes (javac's shared counter). Every existing fixture's immediate
host is its top-level type, so existing names are unchanged — proven by
the 11 #2550 tests passing untouched, not assumed. The owner walk's
anonymous branch also fires on `enum_constant` now (the synthesis returns
undefined for body-less constants, so the walk continues to
`enum_declaration` as before).
Identity window: INCREMENTAL_SCHEMA_VERSION 8→9, parse-cache SCHEMA_BUMP
18→19, U-C5 pin extended with the v8-stamp rejection (enum-constant
methods re-key `E.hook`→`E$1.hook`; nested-host anons re-key
`EnumWrap$1`→`EnumWrap$Mode$1`).
Enum-constant Class-node emission and scope-side ownership land in the
next commits per
docs/plans/2026-07-18-gitnexus-plan-enum-constant-bodies.md (plan is
local — docs/ gitignored).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(java): model enum constant bodies as first-class instances (#2555, steps 2-4)
`enum E { A { void hook(){} } }` — javac's other anonymous-class shape —
joins the #2550 instance model:
- Structure: `(enum_constant body: (class_body)) @definition.class` in
JAVA_QUERIES; `enum_constant` in javaClassConfig.typeDeclarationNodes
with extractName synthesis. The shouldSkipClassCapture guard now also
covers enum_constant — without it, extract()'s name fallback would
fabricate a Class node from the constant's own identifier (`A`).
- Scope: `(enum_constant body: (class_body) @scope.class)` + synthesized
`@declaration.class`/`@declaration.name` anchored on the body, so the
constant's methods are owned (`ownerId`) and re-keyed
(`Method:...:EnumConst$1.hook#0`).
- Inheritance: a body-anchored `@reference.inherits` naming the HOST
ENUM (javac semantics: E$N extends E) — `mroFor(E$N) ∋ E`, so bare
calls from the body to enum helpers pass the ownership gate's MRO arm
while the same-file bare-call leak for constant-body method names is
closed (discrimination evidence: the #2549 review's archived S1b probe
showed the identical shape resolving `local-call` pre-fix).
- Nested-host JLS naming verified end-to-end: `EnumWrap$Mode$1` (not
`EnumWrap$1`).
- Bench: java scope-capture fingerprint rebaselined (new captures + two
fixtures), `measure.mjs --check` PASS across all 14 languages.
Verified: full java.test.ts 230/230 twice sequentially; TS 254 + JS/
Kotlin 289 (shared-file spot set); schema/scope/owner unit suites 90.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(java): exempt $-chain anonymous class defs from nested-class qualification (#2555 review)
Review lens probe caught a HIGH collapse: same-named methods across
sibling enum constant bodies attributed to the FIRST body's Method node
(`M3$1.hook -> M3.log` where the log() call lives in C's body; the
same-target sibling edge vanished entirely under dedup).
Root cause: `populateClassOwnedMembers`'s qualifier chains a
constant-body class def to `M3.M3$2` — its Class scope's parent is the
enum's Class scope, unlike OCE anons whose parent is a Function scope —
and its methods to `M3.M3$2.hook`. The structure-phase node id encodes
`M3$2.hook`, so the graph-bridge's qualified key misses and falls to
the file-wide simple-name lookup: first-write-wins.
Fix: `qualify()` now skips CLASS-LIKE defs whose name already carries a
`$` chain — a synthesized anonymous binary name is complete by
construction (JLS 13.1). Narrowly scoped: `$`-named MEMBERS (legal and
real in JS/TS) still qualify against their class, and named nested
classes (`Outer.Inner`, #1978) are untouched.
Discriminating regression test: same-name/distinct-target sibling
bodies must each own their edge, and the misattributed cross-edge must
not exist.
Verified: full java.test.ts 231/231; Python+Kotlin 459 (heaviest
populateClassOwnedMembers consumers) — zero assertion failures.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(ci): prettier formatting + java bench rebaseline at the final corpus (#2555)
Two CI reds from the review-fix commit landing AFTER the bench
rebaseline: (1) prettier reformat of the new java.test.ts describe;
(2) the java scope-capture fingerprint drifted again because the
review fix added the java-enum-constant-same-name fixture to the
corpus — rebaselined at the true final corpus (196 fixtures,
ce104a76…, scaling 1.05 < 1.5), local `measure.mjs --check` PASS
across all 14 languages. Lesson honored going forward: the bench
rebaseline is the LAST artifact step — any post-review fixture
addition reopens it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(java): strict JLS 13.1 chaining through anonymous enclosing types (#2555)
Per review discussion: anonymous enclosing types now chain into the
binary name instead of flattening to the nearest NAMED host — the
immediately enclosing type per JLS 13.1 may itself be anonymous:
- anon inside an anon: NestHost$1$1 (was NestHost$2)
- anon inside an enum constant: N$1$1 (was N$2)
- named nested hosts (unchanged): EnumWrap$Mode$1
`nearestJavaAnonHost` becomes `nearestJavaEnclosingType` (named hosts OR
anonymous bodies); an anonymous enclosing type's prefix is its own
synthesized name (memo-bounded recursion); numbering is per immediately
enclosing type in source order. Top-level-hosted names are untouched —
the full existing suite passes unchanged.
New coverage: anon-in-anon chain, anon-in-constant-body chain (with
ownership), and a bodied constant in a NESTED enum (EnumWrap2$Mode$1 —
the one host combination previously untested). Rides the unreleased v9
identity window (doc wording tightened); java bench fingerprint
rebaselined at the final corpus, `--check` PASS across 14 languages;
prettier clean.
Verified: full java.test.ts 234/234 (one worker-crash flake rerun green
in isolation).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(lbug): bound the LadybugDB buffer pool instead of the native 80%-of-RAM default (#2557)
createLbugDatabase passed bufferManagerSize=0, which the native runtime
sizes at 80% of physical RAM. A long-lived gitnexus mcp process (or a
large incremental analyze) could balloon to that ceiling — 19.5 GiB
observed against a 105 MiB on-disk index — and OOM-kill the host session.
Resolve the pool at call time: default min(2 GiB, max(64 MiB, 80% of
totalmem)), overridable via GITNEXUS_LBUG_BUFFER_POOL_SIZE (bytes);
0 deliberately restores the native unbounded default; invalid values
warn and fall back, mirroring GITNEXUS_WAL_CHECKPOINT_THRESHOLD.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(readme): document GITNEXUS_LBUG_BUFFER_POOL_SIZE and GITNEXUS_LBUG_MAX_DB_SIZE (#2557)
Both env tables gain the new buffer-pool ceiling variable and the
previously code-comment-only GITNEXUS_LBUG_MAX_DB_SIZE, with the
mmap-vs-memory distinction the issue had to discover from source.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(lbug): keep parseBufferPoolSize module-private
Review finding: the export had zero importers — parseWalCheckpointThreshold
earns its export via the CLI flag validation, but the buffer-pool CLI flag
was deliberately deferred. Re-export when a consumer exists.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
---------
Co-authored-by: Claude <claude@anthropic.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(scope-resolution): stop platform builtins resolving to unrelated same-file symbols (#2545)
An unqualified call to a platform/language builtin (e.g. TypeScript's
global fetch()) could resolve to an unrelated same-file declaration
sharing that name, most visibly a Cloudflare Worker's
`export default { async fetch(req) {...} }` handler. Two contributing
gaps, both fixed:
- Object literals had no scope boundary in the TS/JS grammar queries,
so a method's/property-arrow's name auto-hoisted past the literal
into whatever lexically enclosed it (scope-extractor.ts's auto-hoist
logic had nowhere to stop). Give object literals a Block scope, like
6 other languages already do for lexical blocks.
- Independently, finalize's per-file bindings bucket
(materializeBindings in gitnexus-shared) flattens every local
declaration in a file onto its module scope for cross-file import
resolution, regardless of true nesting -- so free-call-fallback's
scope-chain walk could still hit the leaked binding at module scope.
Guard free-call resolution: when a match for a known builtin name
(LanguageProvider.isBuiltInName, already populated for TS/JS but
never consulted by this pass) has no binding reachable via the true
lexical scope chain, leave the call unresolved instead of emitting a
false CALLS edge.
Verified against the full TS/JS resolver suites plus every other
language populating builtInNames (Python, Go, C/C++, C#, Dart, Kotlin,
PHP, Ruby, Rust, Swift, Vue) -- 2333 tests, no regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(scope-resolution): extend the #2545 scope-leak fix to Kotlin and Java
Anonymous object-expressions (Kotlin `object { ... }`) and anonymous
class bodies (Java `new Runnable() { ... }`) have the same missing
scope-boundary gap that caused #2545 in TypeScript/JavaScript: a method
declared inside has no scope of its own to stop the auto-hoist at, so
its name leaks past the container into the enclosing scope.
- Kotlin: `(object_literal) @scope.class` (distinct from the already-
scoped named `object_declaration`/`companion_object`). Kotlin already
populates `builtInNames`, so free-call-fallback's isBuiltInName guard
(added for #2545) fully closes the equivalent leak here too --
verified with a `println`-shadowing regression test.
- Java: `(object_creation_expression (class_body) @scope.class)`,
matching PHP's existing `anonymous_class` handling. Java has no
`builtInNames` list, so the isBuiltInName guard doesn't engage --
the scope-tree fix is still correct and necessary (the anonymous
class's own methods are now owned by the right scope), but an
unqualified call to an unrelated same-file method sharing the
anonymous class's method name can still resolve via finalize's
per-file module-scope bucket (materializeBindings, shared/
language-agnostic, intentionally not touched by this PR). Documented
in the test as a known residual gap, same as TS/JS/Kotlin's own
non-builtin-name collisions.
Audited every other language for the same shape (a value/container
node with no @scope.* capture hosting a would-be-auto-hoisted named
declaration): PHP and Vue already handle it correctly (PHP scopes
anonymous_class; Vue's <script> delegates to the now-fixed TS/JS
query). Ruby, Python, Dart, C#, Swift, Go, Rust, and C/C++ have no
query pattern that treats a literal/container value position as a
named declaration in the first place, so the bug shape can't occur
there.
Verified: full Kotlin + Java resolver suites, 468 tests, no
regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(scope-resolution): dedicated Object scope kind for object literals (#2545, #2551)
Review of the #2545 fix surfaced two defects, both fixed here:
1. The isBuiltInName guard suppressed genuine cross-file imports whose
name matches a builtin (`import { fetch } from './fetch-polyfill'`
silently stopped resolving -- verified regression vs. main). The
leak the guard targets is inherently same-file (finalize's flat
bucket is per-file), so the guard now also requires
`fnDef.filePath === parsed.filePath`. New regression test covers
the polyfill-import shape.
2. The sibling-property case of the reported bug was still broken and
masked by a tautological assertion (`c.reason` -- a property that
doesn't exist; the real path is `c.rel.reason` -- so the test
passed regardless of behavior). In
`export default { fetch() {...}, handler: () => fetch(...) }`,
`handler`'s bare `fetch()` still resolved to its sibling. Reusing
the `Block` scope kind was the root cause: correct for a real
lexical block (a nested closure legitimately sees a sibling
`let`/`const` from an enclosing `if`/`for`), wrong for object
literals, whose members are reachable only via property access --
never as bare identifiers, not even by sibling property bodies.
Fix: a dedicated `Object` ScopeKind (gitnexus-shared) -- a hoist
boundary whose own bindings scope-chain walkers never consult while
still traversing past it to the parent. TS/JS object literals now
emit `@scope.object`; the four chain walkers in
scope-resolution/scope/walkers.ts (walkScopeChain,
findAllCallableBindingsInScope, findCallableBindingsAndAdlBlocker,
findExportedDefByName) and free-call-fallback's
hasGenuineLexicalBinding skip Object scopes' bindings. Kotlin's
anonymous `object {}` keeps `@scope.class` -- unlike JS object
literals it has real implicit-this sibling dispatch.
Verified with the full resolver matrix run sequentially (TS 254, JS/
Kotlin/Java/Python/Go + TS variants 960, C/C++/C#/Dart/PHP/Ruby 1049,
Rust/Swift/Vue/Cobol + route/flow/unit suites 828, scope-extractor/
scope-tree units 51). Worker-pool crashes under parallel suite load
reproduced on unrelated files and pass in isolation (known flake, not
caused by this change).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
* feat(java): model anonymous class bodies as first-class Class nodes (#2550, step 1)
`new Runnable() { public void run() {} }` now emits a synthesized
javac-style `Class` node (`Worker$1`, `$N` = source order within the
top-level class) and owns its methods: the enclosing-owner walk
attributes `run` to `Worker$1` (re-keyed `Method:...:Worker$1.run#0`,
HAS_METHOD from the anonymous class) instead of the lexically enclosing
named class.
- `synthesizeJavaAnonymousClassName` (ast-helpers): single naming
authority for every layer that keys the anonymous class; returns
undefined for `object_creation_expression` without a `class_body`
child, which also keeps it a no-op for C#'s same-named node type.
- `findEnclosingClassInfo`: anonymous-body branch before the generic
container walk.
- JAVA_QUERIES: `(object_creation_expression (class_body))
@definition.class` (no @name); `getLabelFromCaptures` now lets a
nameless `definition.class` through — the parse-worker's existing
`!nameNode && !extractedClassSymbol` gate still drops any nameless
class the extractor cannot name, so other languages are unaffected.
- `javaClassConfig.extractName` synthesizes the name on the extractor
path (worker node emission).
- Node identities move on unchanged files: INCREMENTAL_SCHEMA_VERSION
7→8 and parse-cache SCHEMA_BUMP 17→18 (the v5 Route-identity
precedent) force full re-analyze / cache invalidation.
Verified: new #2550 identity tests + resolve-enclosing-owner and
has-method suites (53 tests) green.
Prep for step 2/3 (scope-side ownership + receiver typeBinding) and the
free-call instance-ownership gate per
docs/plans/2026-07-18-gitnexus-plan-java-instance-scoped-freecalls.md
(plan file is local — docs/ is gitignored by repo policy).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(java): instance-scoped free-call resolution for anonymous-class methods (#2550, steps 2-4)
Completes the #2550 instance model on top of the Worker$N identity
commit:
- Scope-side ownership (java/captures.ts): synthesize
`@declaration.class` + `@declaration.name` (`Worker$N`) anchored on
the anonymous `class_body` — same range as its `@scope.class`, so the
def lands in that Class scope's ownedDefs, `populateClassOwnedMembers`
stamps `ownerId` on the anonymous class's methods, and the name
auto-hoists exactly like a named class declaration.
- Receiver typeBinding (java/captures.ts + type-extractors/jvm.ts):
`Runnable handler = new Runnable() { ... }` binds `handler` to the
ANONYMOUS class (`Worker$1`), not the declared JDK interface — in both
the scope-side TypeRef channel (receiver-bound Case 4) and the worker
typeEnv. `handler.run()` now resolves through the receiver path
(reason 'global', target `Worker$1.run#0`) instead of depending on
the free-call finalize-bucket leak — which is why the prior gate
attempt broke it (the #2550 landmine, now explained and structurally
removed).
- Instance-ownership gate (free-call-fallback.ts + contract + run.ts +
java opt-in): with `ScopeResolver.freeCallsRequireInstanceOwnership`,
a free call may resolve to a `Method` only when the caller's
enclosing class chain (self + MRO via `scopes.methodDispatch.mroFor`)
contains the method's owner. Same-file matches only — the
`materializeBindings` leak is per-file; cross-file Method matches come
through genuine import channels (suppressing them broke the
arity-narrowing parity suite, verified). Suppressions recorded as
`'free-call-instance-ownership'` outcomes. Java opts in; every other
language is byte-identical (flag off).
Result on the #2545 fixture: `process()`'s bare `run()` emits NO edge
to the unrelated anonymous method (the #2550 bug, closed), while
`handler.run()`, same-class implicit-this dispatch, and bare inherited
calls (MRO arm) all keep resolving.
Verified: full java.test.ts 223/223 twice sequentially (landmine gate);
cross-language matrix (TS/JS/Kotlin/Python/Go/C/C++/C#/Dart/PHP/Ruby/
Rust/Swift/Vue/Cobol + callable-value-flow + java-class-impact + core
units) — zero assertion failures; worker-crash flakes re-verified green
in single-file isolation.
Known deferral (documented): EXTENDS/IMPLEMENTS edges from the
anonymous class to its constructed type are not yet emitted, so a
same-file inherited-but-not-overridden member called ON the anonymous
instance does not resolve through the anon MRO; tracked as the
follow-up in #2550.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
* fix(java): anonymous-class inheritance, host coverage, and phantom-node guard (#2550 review)
Self-review of the instance model (gitnexus-review with empirical lens
probes) surfaced three defects, all fixed:
1. HIGH — the ownership gate suppressed TRUE bare calls to inherited
methods inside an anonymous body extending a same-file class
(`new Base() { void extra() { work(); } }` lost `extra -> work`):
the anon class had no inheritance edge, so `mroFor(Worker$N)` was
empty and the MRO arm could never pass. The synthesis now emits an
`@reference.inherits` for the constructed type, anchored on the
`class_body` so the reference's enclosing class resolves to the
SYNTHESIZED def (anchoring on the type node would sit outside the
anonymous scope and attribute the edge to the wrong class). Anon
classes now get real EXTENDS/IMPLEMENTS edges and inherited bare
calls pass the gate.
2. MEDIUM — hostless anonymous bodies materialized a phantom Class
node named after the CONSTRUCTED type (`Class:...:Runnable`) via
extract()'s extractTypeNameFromNode fallback. New
`shouldSkipClassCapture` in javaClassConfig drops the capture when
no name can be synthesized.
3. MEDIUM — enum/interface/record-hosted anonymous bodies silently
fell back to the pre-#2550 model (mis-attribution + open leak).
The topmost-host walk now accepts all four host type declarations
(JAVA_ANON_HOST_TYPES), so `EnumHost$1` etc. are modeled; the
phantom-node shape disappears for those hosts as a side effect.
Also: per-parse-tree WeakMap memo for the `$N` numbering — the helper
is called from four independent layers per anonymous body and each call
re-scanned the host subtree (`descendantsOfType`), quadratic on
anon-heavy files (old-style listener-per-widget Java); and the
scope-capture bench fingerprints rebaselined for java/typescript/
javascript/kotlin (`measure.mjs --check` now passes all 14 languages —
it failed for every scope query this PR touched; drift notes added per
the file's convention).
Verified: full java.test.ts 225/225; all 11 #2550 tests including the
new anon-extends-base and enum-host scenarios; bench --check PASS.
Known remaining (documented, unchanged-old behavior): enum CONSTANT
bodies (`A { ... }`) stay unmodeled; nested-host naming is top-level-
anchored (`EnumWrap$1`, not javac's `EnumWrap$Mode$1`).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
* test(storage): update the INCREMENTAL_SCHEMA_VERSION pin to v8 (#2550)
The U-C5 reuse-gate test deliberately pins the exact schema version so
a bump cannot land without consciously extending the gate expectations.
Extend for v8 (Java anonymous-class node identities, #2550): a v7 stamp
now fails the strict-equality reuse gate — a pre-v8 index would strand
old `Worker.run`-keyed Method nodes alongside the re-keyed
`Worker$N.run` ones on unchanged files — and v8 passes.
Caught by CI (tests/ubuntu coverage shard 2/3 on PR #2549); the local
matrix had not included this unit file. All 7 schema-referencing unit
suites verified green (109 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(analyze): degrade FTS search instead of aborting analyze on index-build failure
createSearchFTSIndexes re-tokenizes every stored row on every analyze run
(full or incremental). A native LadybugDB tokenizer error on a single
pre-existing row (e.g. "Failed calling LOWER: Invalid UTF-8") previously
propagated uncaught out of run-analyze.ts's main FTS phase, discarding an
otherwise-successful run's graph/embeddings work every time analyze ran
thereafter.
Add buildSearchIndexesOrDegrade(), which catches build/verify failures and
lets analyze finish with keyword search degraded for that run instead —
mirroring the existing sibling degrade path for a missing FTS extension.
The dedicated --repair-fts path is untouched and still fails loudly.
Fixes#2544, #2546.
* fix(analyze): keep capabilities.fts/ftsSkipped honest when index build degrades
ftsSkipped and capabilities.fts.status were keyed only on ftsAvailable
(extension loaded), which the new degrade path leaves true even when the
index build itself failed. Track that outcome in ftsReady and use it for
both, and update run-analyze-fts-repair.test.ts's coverage of this path
from asserting the old throw to asserting the new degrade contract
(ftsSkipped, log message, meta.json capabilities.fts.status).
* docs(plans): add provider-hook value-refs plan (#2437)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(plans): deepen #2437 plan to USES + property-dispatch design
Design revised after prior-art research (Kythe ref vs ref/call, Joern
METHOD_REF, Feldthaus field-based call graphs, CodeQL impliedReceiverStep):
registration sites emit reference-class USES, invocation is recovered by a
field-based property-dispatch pass synthesizing CALLS at member-call sites.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(scope-resolution): model provider-hook value references (#2437)
Functions referenced as object-literal property values (provider hooks like
emitScopeCaptures: emitCppScopeCaptures) previously produced no edge at all,
so impact/context reported a false-safe 0 upstream dependents.
Two coordinated halves, per prior art (Kythe ref vs ref/call, Joern
METHOD_REF, Feldthaus ICSE'13 field-based call graphs, CodeQL
impliedReceiverStep):
- Registration -> USES: new ReferenceKind 'value-ref'; TS/JS queries capture
pair values and shorthand properties (with @reference.property-key);
emitted as a reference-class USES edge, reason 'scope-resolution:
value-ref'. Resolution is callable-gated so plain values emit nothing.
- Dispatch -> CALLS: new shared pass emitPropertyDispatchCalls synthesizes
CALLS (reason 'property-dispatch', confidence 0.7, per-key fan-out cap 32
calibrated on this repo's 16-provider hook tables) from member-call sites
to every function registered under the same property key.
Deviation from plan: the pass owns value-ref resolution entirely via the
post-finalize findCallableBindingInScope walker — the shared registries only
see pre-finalize local bindings, so imported hooks (the c-cpp.ts case) were
unresolvable through lookupForSite; Reference.propertyKey passthrough
dropped as unnecessary.
SCHEMA_BUMP 13 -> 14: ParsedFile gains value-ref sites + propertyKey.
Verified end-to-end: impact(emitCppScopeCaptures, upstream) now reports 8
impacted / HIGH with extractParsedFile (true dispatch caller) at d=1 via
property-dispatch and the c-cpp.ts registration via USES.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(scope-resolution): cover value-ref registration and property dispatch (#2437)
Integration: same-file/cross-file/aliased/shorthand registrations emit USES;
non-callable and destructuring values emit nothing; dispatch sites gain
property-dispatch CALLS (incl. JS twins and per-language partitioning);
fan-out-capped keys are dropped entirely; factory-call values unchanged.
Unit: capture-shape pins for @reference.value-ref + @reference.property-key.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(scope-resolution): surface dropped property-dispatch keys in stats (#2437)
Review finding: skippedKeys was returned but discarded — a hook table
larger than the fan-out cap silently reopened the #2437 gap for those
keys. Log dropped keys and fold value-ref USES + dispatch CALLS into
referenceEdgesEmitted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(plans): add callable reference-flow implementation plan
* fix(scope-resolution): close property-dispatch review gaps
* feat(scope-resolution): add callable flow facts
* feat(scope-resolution): resolve callable value flow
* feat(scope-resolution): resolve callable references across providers
* fix: harden callable reference flow resolution
* fix(scope-resolution): preserve callable binding semantics
* docs(plans): add pr-2522-review-fixes plan
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(storage): bump INCREMENTAL_SCHEMA_VERSION for callable-value-flow edges
Callable-value-flow CALLS/USES edges (#2437) can connect two files whose
content did not change, but the incremental write set only covers changed
files — a top-up against a pre-v7 index would silently omit the new edges
for every unchanged file pair, indefinitely. Force the one-time full
re-analyze (review finding 1, #2522).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(storage): sanitize callable-flow sites per-site at load, log drops
The load-time validator rejected the WHOLE ParsedFile when one site was
malformed or over-bound, with no logging — and C++ legitimately emits
empty-string parameterTypes entries ('' = unknown, the
ReferenceSite.argumentTypes convention) for cv-only/ERROR-recovered types,
so real repos fell into a permanent, silent warm-cache-miss reparse loop
through the #1983-sensitive main-thread path (review finding 7, #2522).
Now: '' entries are valid in type arrays; a malformed/over-bound site drops
only itself (counted, warned once per load); only non-array garbage —
evidence the serialization itself is untrustworthy — rejects the file.
Deviation from plan §6 wording: validator-side tolerance replaces emit-side
clamps — smaller diff, same asymmetry closed at the single chokepoint.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(scope-resolution): keep declarations in the union for reassigned callable cells
The binding-lookup suppression for fact-constrained cells was wholesale:
reassigning a declared function through its own name (greet = other;
greet()) deferred the call to the solver, which then refused the lexical
lookup that resolves the declaration — an unresolvable RHS yielded zero
CALLS for a call that resolved pre-flow (review finding 8, #2522).
Suppression now applies only to cells bound by FORMAL facts — its actual
purpose (a parameter whose grammar emits no declaration binding must not
adopt a same-named outer function). Copy/alias/store/load destinations keep
their declaration as an inclusion seed (Andersen-style union).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(scope-resolution): count forfeited deferred sites in the budget-bailout warning
On work-budget exhaustion the deferred invoke sites end the run with zero
CALLS — free-call fallback and reference emission already skipped them —
but the warning said 'ordinary graph emission remains untouched', which is
false for exactly those sites. The warning context now carries the
unresolved deferred-site count and the comment states the real cost
(review finding: budget-bailout honesty, #2522).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(scope-resolution): surface dropped property-dispatch keys in stats and warn payload
The over-cap warning carried only a count; the dropped key NAMES were
discarded and RunScopeResolutionStats had no field, so the PR-body claim
'includes them in resolver statistics' was unimplemented (review finding,
#2522; reviewer ask on the fan-out cap). The warn payload now names up to
20 dropped keys and the stats carry propertyDispatchSkippedKeys.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(scope-resolution): drop producer-less ownerQualifiedName from formal sites
No capture emitter anywhere produces @callable-flow.owner-qualified-name —
the solver branch consuming it was unreachable in production, yet the field
was typed, parsed, validated, and unit-tested with hand-built input (review
finding 16, #2522; YAGNI). Re-add with a real producer if C++ qualified
member declarators ever need it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(scope-resolution): drop dead callable-flow knobs
CallableFlowPassingMode 'callable-object' had no producer and no consumer
distinguishing it, and CallableFlowCaptureOptions.extractCallArguments had
no language providing it (unlike its live sibling extractCallCallee) —
review finding 17, #2522 (YAGNI). The invocation-kind 'callable-object'
is a different, live concept and stays.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ingestion): bind subscripted callable cells to the container, not the index
terminalIdentifier iterates children in reverse, so tbl[i] = handler seeded
the INDEX variable's cell (polluting a same-named formal) and tbl[i](7)
looked up the callee under i in a different scope — no join, no CALLS edge
for the classic function-pointer-array dispatch (review finding 12, #2522).
Subscript nodes now recurse into their container field only, in both
bindingIdentifier and terminalIdentifier, across the fielded grammars
(C/C++/JS/TS/Python/Go/Java).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ingestion): make cross-function file-scope callable bindings resolvable
Two stacked gaps killed the canonical C callback-registration pattern
(fp assigned in init(), called in run()) — the exact #2437 false-safe this
PR exists to fix (review finding H1, #2522):
1. isVisibleValueBinding only consulted assignment regions and formals, so
a call in a function OTHER than the assigning one emitted no invoke
fact. A declared callable-typed binding is now a value binding wherever
its declaration is visible (visibleCallableSignature).
2. The C scope query had no @declaration.variable pattern for function-
pointer declarators — void (*fp)(int); created no scope-tree binding,
so the seed (init) and invoke (run) cells canonicalized to different
keys and never joined. Both bare and initialized forms now bind.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(c): detect variadic parameters via the named variadic_parameter node
tree-sitter-c materializes '...' as a named variadic_parameter node; the
anonymous-token checks never matched, so variadic function-pointer
signatures were emitted with a wrong fixed arity and no '...' sentinel
(review finding, #2522). C++ is unaffected ('...' stays an anonymous token
there); the token checks remain for such grammars.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ingestion): emit invoke facts for field-stored callable member calls
The C ops-vtable pattern (o->run = handler; o->run(1)) captured the store
but never the call — the member path in emitCallFacts bailed for languages
without protocol methods, and the value-binding index recorded the member
store under the OBJECT's name ('o'), not the member's ('run') (review
finding 11/M3, #2522). Member destinations now also record their terminal
member name, and a member call whose name-cell has a visible store emits an
indirect invoke — gated on the store so plain accessor calls (map.get)
stay inert.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(cpp): disambiguate (obj->*ptr)() ERROR recovery by token order
tree-sitter-cpp groups the recovered '->*' two ways depending on
error-recovery cost (identifier lengths): [identifier, ERROR '->*m'] or
[ERROR 'obj->*', identifier]. The recovery assumed the first shape, so the
second silently swapped receiver/member and dropped the call site — the
committed test passed only by name luck (review finding H2, #2522). The
identifier's position relative to '->*' inside the ERROR now decides roles.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(cpp): class members are never file-local in hasFileLocalCallableLinkage
The name-keyed file-local set is populated from every static declaration,
so an in-class 'static void make();' (external linkage — in-class static
means no-instance) and any member sharing a name with a static free
function were over-marked, refusing legitimate cross-file
declaration/definition joins (review finding 13/M2, #2522). Method and
Constructor defs now bypass the name-set, per the hook's own linkage-only
contract.
Deviation from plan step 13: the regression is a unit-level contract pin
rather than an end-to-end join test — C++ merges out-of-line member
definitions onto the member node by qualified identity, so the graph shape
cannot discriminate the join refusal for members.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(cpp): classify parameter passing mode from the declarator chain only
A whole-subtree scan for reference_declarator inverted copy vs alias:
void reg(void (*cb)(int& out)) marked the by-value pointer cb as
'reference' because of the NESTED parameter's int&, making the solver
back-propagate formal targets into every caller's argument cell — alias
semantics for a copy (review finding 14/M5, #2522). The chain walk never
descends into nested parameter lists; a reference anywhere ON the chain
(int& x, void (*&cb)(int)) still aliases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ruby): bare identifiers are calls, not callable references
Ruby parses a receiver-less zero-arg method call identically to a variable
read, so 'action = process' — which CALLS process and stores its return —
seeded action with the callable and minted a wrong CALLS edge from any
dispatch through it, confirmed end-to-end (review finding 15/HIGH, #2522).
New provider knob bareNamesAreCalls: a bare name that is not a provably
local value binding and not an explicit reference form (method(:x),
lambda/proc) emits no flow fact, on both the assignment and argument paths.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(go): pair multi-value := positionally instead of cross-wiring
The shared field fallback took the FIRST LHS identifier and the LAST RHS
identifier of Go's expression_list pair, cross-wiring 'a, b := f, g' and
synthesizing a garbage comma-joined qualified name — the real relationships
were silently dropped (review finding 16, #2522). extractAssignment may now
return multiple pairs; Go pairs list entries positionally and emits nothing
for a length mismatch (multi-return call RHS).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(java): drop get/test from callableProtocolMethods
'get' and 'test' collide with ubiquitous non-functional-interface APIs
(Map/List/Optional/Future.get), so every ordinary container access emitted
a spurious callable-object invoke fact — high-volume misleading graph facts
with a cross-wiring risk on receiver-name reuse (review finding 17, #2522).
Supplier.get/Predicate.test dispatch is deliberately traded away until the
check can gate on the receiver's declared type.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(rust): pin the qualified-name no-degrade guard as a hard invariant
Rust's scoped_identifier callable-reference capture over-includes unit enum
variants and associated constants (Shape::Square seeds as if callable);
they stay edge-free only because resolveSeedCandidates refuses to degrade
an unresolved qualified name to a simple-name lookup (review finding 18,
#2522). Capture-side type filtering would false-negative on tuple-variant
constructors, so the guard IS the contract: documented as a hard invariant
(Go's mis-shaped multi-value forms also rely on it) and pinned end-to-end.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(php): remove nonexistent optional_parameter node type
tree-sitter-php has no 'optional_parameter' — defaults ride on
simple_parameter — so the entry was dead weight the #1920 literal gate
does not cover for capture-option Sets (review finding 19, #2522).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(cobol): detect procedure pointers on fixed-format sources
Two stacked defects made the feature a no-op on classic sequence-numbered
fixed format (review finding 20/H3, #2522):
1. parseDataItemClauses' USAGE alternation knew POINTER but not
PROCEDURE-POINTER/FUNCTION-POINTER, so the dataItems filter was dead.
2. The raw-line fallback scanned UNCLEANED text, where the sequence number
satisfied the leading digits and the LEVEL NUMBER got captured as the
pointer name. It now scans preprocessed lines and requires a letter-
initial name (COBOL data names must contain a letter).
161 COBOL preprocessor/copy-expander tests stay green; free-format matrix
case unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(cobol): skip comment lines in SET seed/copy scans
A commented-out SET (indicator-column '*'/'/' or free-format '*>')
produced a live seed and a false CALLS edge from dead code (review
finding 21/M1, #2522). The scan now skips indicator-column comment lines
and strips inline '*>' tails before matching.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(architecture): document callable-flow-only mode and skipped-key reporting
The Callable-value flow section omitted scopeResolutionEdgeMode:
'callable-flow-only' — a real emit-pipeline branch that suppresses all
ordinary emission for standalone providers (review finding 22, #2522) —
and predated the skipped-key names/stats surfacing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(scope-resolution): correct value-ref resolution attribution and stale pdg-gating comments
The value-ref contract comment claimed MethodRegistry resolution — the
mechanism is the post-finalize findCallableBindingInScope walker owned by
emitPropertyDispatchCalls (resolveReferenceSites skips these sites). Three
'only under --pdg' calleeIdSink comments were falsified by the #2437 gating
change (callee-id-sink.ts's header was updated; these copies were missed).
Review finding 23, #2522.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(ingestion): direct unit coverage for synthesizeCallableFlowCaptures
The 1,100-line shared synthesizer had no test naming it — only downstream
consumers were covered (review finding 24, #2522). Pins seed/invoke/
formal/argument emission, subscript container binding, store-gated member
invokes, produced-value guards, and the bareNamesAreCalls knob over a
minimal options object so assertions target the synthesizer's own
semantics.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(resolvers): deepen shallow-language coverage; fix Kotlin/Swift reassignment gaps it exposed
Adds the COBOL SET x TO y copy-branch scenario and conditional-assignment
scenarios for Kotlin, C#, Swift, and Dart (10 languages previously had one
generic case each — review finding 25, #2522). The new scenarios exposed
two real capture gaps, fixed here:
- tree-sitter-kotlin's 'assignment' node is fieldless, so nested
reassignments (chosen = ::target inside a block) produced no flow facts;
Kotlin's extractAssignment now decomposes it positionally.
- tree-sitter-swift fields its assignment as target:/result:, neither in
the shared fallback's field lists; both added.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(infra): literal-validation gate for callable-capture option Sets
The #1920 gate validates query literals and exported configs but not the
module-private *_CALLABLE_CAPTURE_OPTIONS Sets consumed by the shared
synthesizer — a typo'd node type silently captures nothing (PHP shipped a
dead 'optional_parameter'; review finding 26, #2522). Every <key>NodeTypes
Set literal is now validated against its language's grammar; name-carrying
sets (callableProtocolMethods, memberPointerOperators) are deliberately
outside the contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(storage): centralize corrupt-fixture casts into makeStoreEntry
The callable-flow store tests scattered 'as unknown as' double-casts per
fixture (review finding 27, #2522; standing no-as-any rule). One typed
helper now owns the single controlled escape hatch for building malformed
serialization-boundary payloads.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(bench): refresh capture fingerprints after review fixes
python-scope: the committed baseline (8d5c3699) never matched this
branch's code — CI's benchmarks arm was red on the PR head (review
finding 2/HIGH, #2522); regenerated (a99e69ab), scaling 1.04 in budget.
scope-capture: ruby/cpp/swift/java/kotlin drifted from the review-fix
commits (bare-name suppression, passing modes + ->* recovery, assignment
fields, protocol narrowing, positional assignment); all 14 languages
re-verified PASS with ratios <= 1.18 against the 1.5 budget.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(docs): untrack docs/plans working documents
docs/ is gitignored (local working docs); the plan files were force-added
past the ignore. Untracked from the index only — they stay on disk.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(golden): regenerate captures goldens after callable-flow review fixes
The per-language digest guards (csharp/go/php/python/ruby/rust/swift)
locked the pre-fix capture output; the review-fix series intentionally
changed it — store-gated member invokes, subscript container binding,
Ruby bare-name suppression, Swift assignment fields, positional pairing.
Regenerated with UPDATE_GOLDEN=1; clean verification run 59/59; all other
parity/golden guards (pipeline-graph, spring-route, python parity) pass
untouched at 33/33.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ingestion): prototypes are callees, not callable value cells
The cross-function visibility fix indexed EVERY signature-bearing
declaration as a value binding — including plain function/method
prototypes (void f(int);). Every call to a declared function then became
an indirect invoke, and with emitCanonicalInvokeReference (C/C++) minted a
free-call reference that resolved through the registry, bypassing the
precise passes' two-phase/ambiguity/subobject suppression — eight phantom
CALLS edges in the cpp resolver suite on CI.
Only declarations whose binding identifier sits under a pointer/
parenthesized declarator (callable-typed variables like void (*fp)(int);)
create value cells now. cpp resolver suite 331/331; callable-value-flow +
C/C++ suites 181/181 (the cross-function fp regression still passes); cpp
fingerprint rebaselined, both bench gates PASS.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(wiki): allow explicit HTTP LLM hosts
Keep wiki LLM HTTP endpoints fail-closed by default while adding a narrow exact-host opt-in for LAN/self-hosted models.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(wiki): simplify insecure LLM flag name
Rename the wiki HTTP opt-in flag to --allow-insecure-connection per review feedback.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(wiki): simplify insecure connection env
Rename the wiki HTTP allowlist environment variable and align validation errors with the CLI flag naming.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
The busiest Windows platform shard reached 14m57s against the 15 minute
watchdog on the rc.19 green run and has timed out once since. CI now
sets GITNEXUS_CROSS_PLATFORM_TIMEOUT_MINUTES=20 (the job timeout stays
25), the stale comfortably-under comment reflects reality, and the
runner always logs status, signal, spawn code and elapsed time so the
next status-null death is diagnosable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The RC path bumped only gitnexus/package.json, so every v1.6.10-rc tag
through rc.28 shipped the four plugin manifest surfaces frozen at 1.6.9
and failed its own unit suite. The npm version lifecycle script now
runs a fail-closed sync whenever npm version executes, in CI or on a
maintainer's laptop; publish.yml verifies the result and stages the
surfaces into the detached release commit, and the stable path refuses
to publish a tag whose manifests drifted. The sync is textual so a
release commit carries a one-line change per surface instead of
reformatting churn.
Design follows the proposal by @100yenadmin in #2445, moved onto the
standard npm version hook.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The skip-git ignore test asserted the error payload while relying on
exit 0; since the output() guard an error payload also exits 1, so the
test now captures the payload from the exec failure and pins both.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The zero-symbol File fallback from #2455 writes File embedding rows,
but the filePath-scoped delete sweeps joined through EMBEDDABLE_LABELS
only. Docs repos accumulated duplicate rows on re-analyze and deleted
files left orphans. Free for code repos: no File rows exist to match.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Moves the #2469 guard from cypherCommand into output() so all seven
tool commands that print backend results share the exit semantics.
Adds query and context regression cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An fs failure while resetting a chunk generation now warns and
continues like the neighboring durable-store paths instead of failing
the analyze. Workers recreate the directory on write.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Composed with the read-only and repository policies in the CallTool
handler: read-only assert, then budget resolution, then scoped dispatch
with the transport arg stripped.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The C#, Java and Go flag-off digests shift with the canonical community
projection from #2478 stacked on the walker sort from #2482. Regenerated
via vitest -u; mini-repo pipeline golden was already correct.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Merged with the read-only policy: both filters compose in server.ts.
Review hardening: assertResourceUri compares the opaque host case
insensitively (GITNEXUS://GROUP bypass) and gains the fail-closed
fallback for unparseable URIs; the backend proxy now also intercepts
queryClusters/queryProcesses/queryClusterDetail/queryProcessDetail;
scrubGroupDescription is shared from read-only-policy.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
crossDepth and subgroup are inert outside @group routing, but rejecting
them keeps the scrubbed schema and the dispatch contract in agreement.
Also documents that resource content scrubbing is cosmetic.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds the missing #2481 Rust regression test (importer before definer),
fails closed when a PHP namespace suffix matches directories under
different roots, and points the baselines note at #2481/#2482.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Non-loopback eval-server binds now require GITNEXUS_AUTH_TOKEN with an
exact Bearer header compared in constant time. Review fixes: loopback
binds defer unreadable env file errors to a warning, and help text
notes IPv4 hostname resolution.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Loopback binds now warn and start when .env or .env.local exists but
cannot be read; non-loopback binds keep the fail-closed error. Help
text notes that hostnames resolve to IPv4.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Initial plan
* fix(mcp): read fresh metadata from disk in context resource to fix stale staleness banner and stats (#2438)
The context resource (gitnexus://repo/{name}/context) was serving stale
lastCommit and stats after an out-of-process `analyze --index-only` refresh.
The LocalBackend's in-memory RepoHandle is cached and only refreshes on
registry misses — never on a disk update from an external process.
Fix: read fresh metadata via loadMeta(repo.storagePath) on every context
resource read, mirroring the ensureInitialized hot-swap pattern. Use the
fresh lastCommit for the staleness check and fresh stats for the stats
block; fall back to cached values when the disk read fails.
Tests:
- 5 new unit tests in resources.test.ts covering fresh meta, fallback behavior
- 3 new integration tests in context-resource-staleness.test.ts covering
the exact reproduce sequence from the issue report
* fix(test): use @ladybugdb/core mock to prevent native addon load in integration test
The context-resource-staleness integration test used importOriginal() in
the vi.mock factories for pool-adapter.js and mcp/core/lbug-adapter.js.
That caused the real pool-adapter.ts to load @ladybugdb/core, which
tries to dlopen lbugjs.node — a native addon that requires postinstall
(not run with --ignore-scripts in CI).
Fix:
- Add a vi.mock('@ladybugdb/core') at the package boundary, matching the
pattern in lbug-pool-fts-load.test.ts and pool-wal-recovery.test.ts.
- Replace the two importOriginal()-based lbug mocks with direct factories
that provide all necessary exports (initLbug, executeQuery,
executeParameterized, closeLbug, isLbugReady) without loading the real
module.
All 3 integration tests now pass in CI (no lbugjs.node required).
* test(integration): run context staleness flow end-to-end without mocks
* Address PR review feedback (#2439)
- Restore libc selectors for optional native packages
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* fix(web): use repo path identity in switcher
* keep repo URL project names stable
* fix server repo path resolution
* fix repo path miss resolution
* fix(server): guard clone-dir deletion with path ownership check
Deleting a registry entry derived its clone dir from the entry NAME with
no ownership check, so deleting a local repo that shares a display name
with a server-cloned sibling wiped the sibling's checkout. Gate the
removal on cloneDirBelongsToEntry (canonicalized path equality), the
same entry.path-driven rule the handler's step 2b already mandates.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(server): fail closed on relative repo params and rate-limit GET /api/repo
Relative separator-containing ?repo= values (org/name, ./repo) were
canonicalized against the server CWD — an attacker-influenced
realpathSync probe on an un-rate-limited GET — before failing anyway.
Reject them immediately without touching the filesystem, drop the
redundant path.sep clause, document the resolver's two-tier contract,
and wire createRouteLimiter on GET /api/repo like its DELETE sibling.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(server): lock repo resolver branches and register for Windows CI
Lock in the resolver's remaining branches: first-wins for ambiguous
bare names, Windows-shaped input as a fail-closed path claim, the
repos[0] default, and the case-insensitive name fallback. Register the
suite in cross-platform-tests.ts so windows-latest actually runs the
path-shape logic it exists to protect.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(web): single repoIdentity helper with repoPath normalized end-to-end
The identity fallback chain was copy-pasted in Header and RepoLanding
while backend-client already owns BackendRepo and the repoPath
normalization. Export one repoIdentity helper, normalize fetchRepos
like fetchRepoInfo, and emit repoPath from GET /api/repos so the
scheme no longer silently relies on /api/repo.repoPath equalling
/api/repos.path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(web): persist and restore repo path identity in the URL
The URL persisted only ?project=<display name>, so refreshing after
switching to a duplicate-name repo silently restored the first
same-named sibling. Persist ?repo=<server-resolved path> alongside the
readable ?project= at both write sites, prefer it on restore (legacy
project-only URLs still work), keep failed path restores fail-visible
(no name fallback), and strip stale identity params when deleting the
active or last repo.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(web): analyze completion connects by path identity
RepoAnalyzer's completion callback passed the display name, so
analyzing a repo whose basename collides with an existing one
reconnected the first same-named sibling. The SSE terminal payload now
carries the job's repoPath (both emit sites), the analyzer passes that
identity to onComplete while the done screen keeps showing the display
name, and old servers without repoPath degrade to today's behavior.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(web): scope code-reference file reads to the active repo identity
The code viewer passed the display name as the repo scope, so with
duplicate-name repos it rendered the wrong repo's file contents under
the right filename. Pass the active path identity (currentRepo) with
the display name as fallback, and collapse the two dead repo fields
that were already shadowed by the readFile spread.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(web): show display names instead of absolute paths in labels
The path-identity switch leaked raw filesystem paths into three
user-facing surfaces: the re-analyze progress label, the repo-switch
overlay, and the agent prompt's project name via loadGraphAnyway.
Resolve display names at render time (registry lookup, then basename
fallback) while state keeps holding the identity; loadGraphAnyway
passes the name explicitly because initializeAgent's empty-deps
closure would otherwise fall through to the literal 'project'.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(web): stop initializeAgent from clobbering repo identity with display names
initializeAgent fell back to writing overrideProjectName (a display
name) into the repo identity, so any future name-only caller — the
pre-PR idiom — would silently kill the Active badge and re-admit the
duplicate-name ambiguity through the agent path. Only opts.repo may
write the identity now.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(web): drop dead initializers flagged by CodeQL
pNameStr's and repoIdentity's initial values were never read: both are
assigned on the success path before any use and the catch returns
early. Bare declarations resolve CodeQL alerts 825/826.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* style(web): fix tailwind class order per root prettier plugin
The worktree pre-commit hook resolved prettier-plugin-tailwindcss
through symlinked node_modules and sorted scrollbar-thin differently
than CI's clean-room install. Re-formatted with the root lockfile
environment; no behavior change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(web): e2e coverage for every #2419 duplicate-name ambiguity
Provision two live repos with the same basename under different parents
via POST /api/analyze, then drive a real browser through each item of
the issue's "Actual behavior" list:
- duplicate rows render and the ACTIVE one is identifiable before and
after switching (active-state must not compare repo.name)
- switching between duplicates swaps the loaded graph, verified by
per-repo marker files (onSwitchRepo must not receive repo.name)
- re-analyze targets the clicked duplicate's exact path (POST body),
tracks progress on that row only, and the completion reconnect
requests that same path — never the same-named sibling
- delete requests target exactly the chosen duplicate's path; the
sibling stays registered and loaded
- backend ?repo= resolution is path-first: landing selection loads the
exact repo, ?repo= survives F5, and a stale path fails closed to the
repo picker instead of retargeting the sibling
Adds four data-testids to Header (switcher trigger/row/reanalyze/
delete, rows expose data-active) so the spec has stable selectors, and
broadens the post-analyze reconnect retry in App to any BackendError:
the server may still be reinitializing when the SSE complete event
fires, and that surfaces as transient 5xx/binder errors, not only 404.
The re-analyze and delete tests deliberately assert identity at the
request level and tolerate two pre-existing server races that are
unrelated to the #2419 identity contract (freshly-analyzed DB briefly
unreadable after SSE complete; registry validate-prune clobbering a
concurrent unregister) — see the in-test comments.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
* test(web): isolate repo-path-identity e2e onto a spec-owned backend
The spec is the only e2e file doing write operations (analyze,
re-analyze, delete). Running its force re-analysis against the shared
CI backend while parallel workers held connections took the whole
server down (run 29145679019: the jobId poll died with ECONNRESET and
every later test in every file failed to connect).
Spawn a dedicated `gitnexus serve` on port 4799 with an isolated
GITNEXUS_HOME in beforeAll instead: writes can no longer perturb the
other suites, a crash is contained to this spec (its output is captured
and printed, which CI otherwise loses), and the registry is hermetic by
construction — the previous leftover-purge and shared-registry cleanup
are gone. Every page is pointed at the spec backend through
useBackend's supported localStorage override, which covers both the
probe-driven landing flow and the ?server= auto-connect. Verified
self-sufficient (6/6 with no shared server running) and non-interfering
(full suite 39/39 with the shared server up).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* stabilize repo path identity e2e
* fix(server): don't report analyze complete before the index is settled
The analyze worker reports `complete` over IPC before its on-disk
finalization (LadybugDB checkpoint, native handle release, metadata
write) is visible at the storage path — observed up to ~6.5s behind the
IPC message. The launcher's "reinitialize backend BEFORE marking
complete" ordering was meant to make the repo queryable by the time the
client sees the SSE complete event, but it never verified that: clients
reconnecting on that event read a database still being written. Locally
that surfaces as "Binder exception: Table CodeRelation does not exist"
or a silently empty graph, and the open can quarantine the in-flight
WAL; on slow CI runners the native layer racing the rewrite has killed
the whole server (signal exit, no output — run 29146867959).
Gate the complete transition on the index actually settling: LadybugDB
file and metadata both rewritten by THIS job (mtime >= job start — bare
existence is not enough, a re-analysis leaves the previous index in
place while it works) and no transient WAL/shadow/checkpoint sidecars
remaining. Bounded (60s) and proceed-on-timeout, so a job whose
analysis legitimately rewrites nothing cannot wedge. Also evict the
server's cached DB handle before reinitializing — same invalidation
DELETE /api/repo performs — so post-completion reads cannot be served
from a pre-rewrite handle.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(web): assert re-analyze completion identity at the request level
The strict form (Ready + marker on the re-analyzed duplicate) still
trips a deeper pre-existing storage race that makes a freshly
re-analyzed database transiently unreadable to the reconnect even with
the settle gate in place — unrelated to the #2419 identity contract
this test covers. Keep the identity assertions (the reconnect targets
the exact duplicate's path and never the same-named sibling) and leave
a pointer to tighten once the storage race is fixed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(server): resolve the settle-gate path from the registry, not the request
CodeQL flagged the settle gate's stat/exists probes as js/path-injection:
the probed path derived from the user-provided analyze `path`. Resolve
it from the repo's registry entry instead — the user value is now only a
comparison key, and the probes run against the server-owned storagePath
record, which is also the authoritative path readers resolve through.
Re-resolved each poll round because the worker registers the repo as
part of the same finalization the gate is waiting out.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(tree-sitter): recover from embedded NUL bytes
Normalize embedded NUL bytes only in parser input so tree-sitter keeps recovering through the full source shape. Pass the file label through the worker for diagnostics and cover both direct-string and callback parse paths with regression tests.
* fix(review): align safe-parser contract count
* test(tree-sitter): cover worker NUL diagnostics (#2430)
* fix(hook): emit MCP query hint when server owns DB lock (#2396)
When the GitNexus MCP server holds the lbug write lock, the PreToolUse hook's CLI `augment` cannot run (LadybugDB is single-writer) and previously skipped silently — disabling graph augmentation in the most common deployment (server online). Since the same session already has the MCP `query` tool live, the owner branch now emits an additionalContext hint pointing the agent at mcp__gitnexus__query for that pattern, via the same sanctioned stdout channel the augment-success path uses (Codex-safe, #2369).
Rejected the alternative of having the hook query the server: it runs over stdio (no port/pipe from the separate hook process) and cross-process read-only access can't coexist with the write lock — both are large architecture changes. Applied to all three gated hook copies (claude .cjs, claude-plugin .js, antigravity .cjs); the cursor hook has no owner gate and is untouched. The stderr `augment skipped: MCP server owns DB` diagnostic stays GITNEXUS_DEBUG-gated (#1913). Owner-path tests flipped from stdout-empty to hint-present.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(hook): reword MCP-query hint to be conditionally truthful (#2396)
The #2396 owner branch emits the hint on every DB-owner path — a confirmed
`gitnexus mcp` owner, a `gitnexus serve` owner, and the fail-closed/timeout
paths (the probe collapses timeout and owned to one boolean). The old text
claimed "Knowledge graph is live via the MCP server" and named
mcp__gitnexus__query unconditionally, which is untrue on a fail-closed probe
where no server is confirmed and misdirecting for a serve-only owner
(review C2/C4).
Reword the hint (byte-identical across all three hook copies) to state that
local augment is unavailable and to condition the MCP call on the tools
actually being live ("if the GitNexus MCP tools are live in this session").
This is truthful on every owner path; the needles the assertions rely on
(mcp__gitnexus__query, query, search_query, the pattern) are preserved.
Fix the 10 stale owner/fail-closed unit tests that still asserted empty
stdout (review C1, the macOS platform-sensitive 2/3 blocker): flip them to
assert the hint via parseHookOutput, keep their stderr/GITNEXUS_DEBUG
expectations, and rename the two 'SILENTLY' titles. The GITNEXUS_DEBUG=''
owner-hint case is restored (the PR's new loop only covered '0'/'false').
Probe and its white-box tests untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(hook): de-orphan the JSDoc in the claude hook copy (#2396)
The #2396 change inserted buildMcpQueryHint between the pre-existing
"PreToolUse handler" JSDoc and handlePreToolUse, orphaning that doc onto the
helper and leaving handlePreToolUse undocumented (review C5). Move the helper
(with its own doc) above the handler doc so the "PreToolUse handler" comment
again precedes handlePreToolUse, matching the clean plugin copy. Pure move; no
behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(hook): throttle the MCP-owner hint to once per repo per window (#2396)
Previously the hint emitted on every qualifying search while a GitNexus process
owned the DB, so an owner-locked session (the common deploy) was nudged toward
the MCP query tool on every Grep/Glob/Bash — context bloat and ~2x query
amplification (review C3).
Add shouldEmitMcpHint(gitNexusDir) to all three hook copies: a per-repo
.gitnexus/.mcp-hint-shown mtime marker emits the hint at most once per window.
Window via GITNEXUS_MCP_HINT_THROTTLE_MS (default 10min; 0/invalid disables).
Best-effort — any fs error falls back to emitting, so the hint is never lost to
a marker failure. The stderr skip diagnostic still fires regardless (only the
hint is throttled).
Tests: hookEnv disables the throttle by default (gitNexusDir is shared across
the suite, so a marker would otherwise throttle sibling owner tests); a dedicated
macOS-lane test sets a real window and asserts emit-then-throttle with the marker
gating it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(hook): README reflects the MCP-owner query hint, not a silent skip (#2396)
The 'Hook augmentation/notifications are silently skipped' section still
described the MCP-server-owns-DB path as a silent augmentation skip (review
docs finding). That path now hands the agent a conditional MCP-query hint via
additionalContext (throttled per repo). Reword the section to describe the hint
and its GITNEXUS_MCP_HINT_THROTTLE_MS throttle, and keep the GITNEXUS_DEBUG
stderr-diagnostic guidance. No CHANGELOG edit (owned at release time).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(hook): guard hint-copy drift + pattern JSON-escaping (#2396)
Two gaps the review flagged (R7):
- Drift guard: buildMcpQueryHint and shouldEmitMcpHint are triplicated across
the three hook copies with no shared module. A source-level byte-identity
check (runs on every platform, unlike the macOS-only owner tests) fails if any
copy diverges — the institutional pattern the repo already uses for mirrored
hook metadata.
- Escaping: an adversarial Grep pattern (embedded quote + newline) must not
break the additionalContext JSON envelope. A macOS-lane owner test drives the
real hook with such a pattern and asserts parseHookOutput still yields valid
JSON containing the literal characters (JSON.stringify escapes them).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): shard platform-sensitive matrix + spawn built CLI to fix Windows cross-platform timeout
The `windows-latest (platform-sensitive)` job was hitting its 15-min internal
vitest watchdog in run-cross-platform.ts. It's cumulative slowness, not a hang:
the fixed 72-file suite is dominated by ~50 CLI/worker process spawns, and
Windows is ~5x slower than macOS at process startup (macOS ran the same set in
~3min of tests). Two complementary changes bring it back under the watchdog with
headroom, without touching any test assertion:
- Shard the platform-sensitive matrix (windows/macos × shard [1,2]) and forward
`--shard=i/2` through run-cross-platform.ts to vitest, which partitions the
fixed file list deterministically (sha1, equal file-count) — halving each
runner. macOS/Ubuntu were already under budget.
- New test/helpers/cli-entry.ts (`CLI_SPAWN_PREFIX`): spawn the built
`dist/cli/index.js` when `GITNEXUS_E2E_CLI=dist` (set on the cross-platform job,
which already builds) instead of `node --import tsx src/cli/index.ts`, which
re-transpiles the whole CLI on every spawn. Defaults to tsx-on-source so local
runs always reflect current source; `GITNEXUS_E2E_CLI=dist` on an unbuilt tree
throws an actionable "run npm run build" error. dist is opt-in only — never
inferred from a generic `CI` env — so an ambient `CI=1` can't silently run a
stale build. Converted 8 spawn-based e2e suites; added test/unit/cli-entry.test.ts.
The Ubuntu coverage job leaves `GITNEXUS_E2E_CLI` unset, so the tsx-on-source path
stays exercised in CI too (both entry points covered).
Measured on Linux: cli-limit-e2e 121.5s→91s, cli-e2e 289s→217s (~25%); larger on
Windows where the transpile is a bigger share of each spawn.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(ci): derive platform-sensitive shard count from one source (#2394)
The shard total was hardcoded in three coupled, unenforced places (matrix
length, job-name suffix, --shard denominator); editing one without the others
silently dropped a shard's tests with green CI. Add a checkout-free shard-plan
job whose single TOTAL generates both the shard index list (consumed via
fromJSON) and the /N denominator (job name + --shard arg), so they cannot
drift. Asserts TOTAL>=1 to rule out an empty-matrix silent skip. No behavior
change — still 2 shards per OS.
Addresses PR #2394 tri-review finding F2.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): 3 shards for real Windows headroom + honest sharding comments (#2394)
vitest shards by file COUNT, not runtime, so the heaviest spawn suites cluster
into one shard: live CI showed Windows shard 1/2 at 12m12s (~81% of the 15-min
watchdog) vs shard 2/2 at 3m0s. The old comments claimed "comfortable/generous
headroom", which the count-based split doesn't deliver at 2 shards. Bump TOTAL
to 3 (one line, single source) so even the busiest Windows shard clears the
watchdog, and reword the comments to describe count-based (not time-based)
sharding.
Addresses PR #2394 tri-review finding F1.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(ci): extract testable parseShardArg from run-cross-platform (#2394)
The --shard parse/forward glue had no unit test. Extract it into a pure
scripts/shard-arg.ts (mirroring the computeSpawnPrefix extraction precedent) so
the branch logic is lockable without the script's top-level execFileSync, and
add test/unit/shard-arg.test.ts (absent -> undefined, valid token -> passed
through, found amid other args). Behavior unchanged; U4 adds the malformed
fail-loud on top.
Addresses PR #2394 tri-review finding F3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): fail loud on a malformed --shard arg (#2394)
A shard-shaped-but-malformed arg (--shard=1, --shard, --shard=abc) was silently
ignored, dropping the shard flag so both legs ran the full unsharded ~50-spawn
suite — re-arming the Windows watchdog timeout with no signal. parseShardArg now
throws an actionable error on any --shard/--shard=… arg that fails the strict
regex (unrelated flags like --shardx= pass through), and the call site in
run-cross-platform.ts catches it into console.error + exit 1, kept outside the
execFileSync try so the message isn't swallowed by that catch's watchdog-only
branch.
Addresses PR #2394 tri-review finding F4.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(test): fail loud on an unknown GITNEXUS_E2E_CLI value (#2394)
computeSpawnPrefix silently degraded any unknown GITNEXUS_E2E_CLI value to
tsx-on-source, so a typo (e.g. `dsit`) would make CI believe it tests the dist
entry point while actually running src. Throw on any value other than
'dist'/'src'/unset (the safe tsx default is preserved for unset/''/'src', so it
still never selects dist without an explicit opt-in). Flip the unknown-mode unit
test to assert the throw and add the missing {mode:undefined, distExists:true}
case. Only ci-tests.yml sets the var (=dist), so no existing suite is affected.
Addresses PR #2394 tri-review findings minor-a/b.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(ci): run cli-entry.test.ts on the cross-platform matrix (#2394)
cli-entry.test.ts resolves CLI_SPAWN_PREFIX from a real path, and its last
assertion (cli[/\\]index) has a Windows backslash branch that only Ubuntu
exercised. Register it in PLATFORM_LOGIC so it runs on the Windows/macOS matrix
too. (shard-arg.test.ts stays out — pure string logic, OS-independent.) List
grows 73 -> 74; the generated shard matrix keeps coverage complete.
Addresses PR #2394 tri-review finding minor-c.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(test): share tsxLoaderUrl(), dedup the last tsx-loader boilerplate (#2394)
bridge-cache-reopen.test.ts carried its own copy of the tsx-loader-resolution
boilerplate (createRequire -> resolve('tsx/package.json') -> pathToFileURL) —
the one site the PR's CLI_SPAWN_PREFIX migration didn't cover (it spawns a seed
script, not the CLI). Export the existing tsxLoaderUrl() from cli-entry.ts and
reuse it here; the resolved loader URL is byte-identical.
Addresses PR #2394 tri-review finding minor-d.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(test): make skipUnlessFtsAvailable install FTS on miss so shards are self-sufficient (#2394)
Sharding the platform-sensitive suite into 3 exposed a latent test-isolation
bug: load-only FTS primitives (test/integration/lbug-core-adapter.test.ts) only
passed because a sibling installer test happened to co-locate in the same shard
and install FTS into the shared ~/.lbdb first. At 3 shards, lbug-core-adapter
landed in a shard with no installer sibling, so its load-only loadFTSExtension()
failed deterministically on macOS+Windows shard 2/3 under GITNEXUS_REQUIRE_FTS=1.
Make the gate self-sufficient: on a load-only miss under REQUIRE_FTS, install
FTS with `auto` (LOAD-first, then one bounded network INSTALL) before treating
it as a hard failure — mirroring withTestIndexedDB. A pre-installed extension
still costs no network (auto is LOAD-first); offline/local runs (no env var)
still skip gracefully. Verified: with a fresh HOME (no pre-installed FTS) +
REQUIRE_FTS=1, lbug-core-adapter now passes 15/15 (previously threw).
Addresses the 3-shard CI failure surfaced while validating PR #2394's F1 fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: warm-cache the LadybugDB FTS extension across platform shards (#2394)
Follow-up to the FTS self-install fix: cache ~/.lbdb/extension per OS + lockfile
so a warm run skips the network install entirely and the parallel shards share
one download across runs. Pure reliability/speed — on a cache miss the tests
still self-install FTS on demand (test/helpers/fts-availability.ts), so this is
never a correctness dependency, just a way to cut the network-install surface
that made the sharded FTS tests flaky. Keyed by lockfile hash (a LadybugDB
version bump re-installs); per-OS since the extension is a native binary.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): pass shard via env to clear zizmor template-injection (#2394)
Interpolating ${{ matrix.shard }} (now sourced from the shard-plan job output)
directly into the run: shell tripped zizmor's template-injection audit
(code-scanning alert #824, ci-tests.yml:147). Move the value into a SHARD env
var — assigned via ${{ }} but referenced as "$SHARD" in the shell, which is not
an injection sink — and set shell: bash so the expansion is uniform across the
windows + macOS matrix (the default run shell is pwsh on Windows, where $SHARD
would be empty and trip the new malformed-shard fail-loud). Verified locally
with zizmor: the :147 template-injection finding is gone.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: shard the ubuntu coverage job and merge blobs before the threshold gate (#2394)
The coverage job ran the full suite unsharded (~16 min). Shard it like the
cross-platform matrix, then merge the per-shard coverage before enforcing the
threshold gate:
- shard-plan now also single-sources the coverage shard count (cov_total /
cov_shards), so the coverage matrix + /N denominator can't drift.
- The `tests` job becomes a coverage shard matrix: each shard runs
`vitest run --shard --coverage --reporter=blob` with thresholds forced to 0
(a single shard's partial coverage can never meet the gate) and uploads its
blob. FTS self-installs per shard, so sharding the full suite is safe.
- New `coverage-merge` job (needs: tests) reduces the blobs with
`vitest --mergeReports`, enforcing the REAL config thresholds on the combined
('new') coverage — this is the gate. It also emits the merged test-results.json
and runs the unsharded web + docker suites, so the `test-reports` artifact
keeps the exact shape ci-report.yml consumes for its base-branch ('baseline')
vs new coverage delta.
The shard arg goes through a SHARD env var + shell: bash (no template-injection).
Validated locally: shard blobs write and merge into a coverage-summary.json +
merged test-results.json; the merge enforces thresholds on the union. CI Gate
still aggregates the coverage-merge result via the reusable-workflow call.
Note: the coverage check names change (ubuntu / coverage 1/3 … + merge) — update
any pinned branch-protection required checks.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): include hidden files when uploading the coverage blob (#2394)
The coverage shards write their blob to gitnexus/.vitest-reports/ (a dotdir).
actions/upload-artifact excludes hidden files by default, so the coverage-blob-*
artifacts uploaded empty — the merge job then downloaded 0 artifacts and
vitest --mergeReports failed with ENOENT scandir '.vitest-reports'. Set
include-hidden-files: true on the blob upload so the blobs actually ship.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): group shard-plan GITHUB_OUTPUT writes to satisfy shellcheck SC2129 (#2394)
Adding the coverage shard outputs (cov_shards/cov_total) made the shard-plan gen
step write four individual `>> "$GITHUB_OUTPUT"` redirects, which shellcheck
(run by the actionlint check) flags as SC2129. Group the echoes into a single
`{ …; } >> "$GITHUB_OUTPUT"` block.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(test): cost-balanced shard sequencer to cut CPU contention (#2394)
vitest's default --shard hashes file paths and splits by file COUNT, which
clustered the spawn-heavy suites onto one runner (Windows platform shard 1 ran
~4x the others). Add a custom sequence.sequencer that overrides only shard()
and balances by estimated WORK instead:
- specWeight() weights the fileParallelism:false spawn-heavy suites (cli-e2e,
lbug-db — already isolated to run sequentially) far above the parallel default
files, plus file size as a cheap finer signal. Deterministic per checkout.
- assignShards() does greedy longest-processing-time bin-packing (heaviest file
into the currently-lightest shard). The partition stays complete and disjoint
— verified: on the 74-file cross-platform set the three shards weigh
7611/7610/8064 (the sequential-heavy files spread ~7/7/8) with zero overlap and
no file dropped, vs the hash split's count-only balance.
sort() is left to the base sequencer so project groupOrder / duration-cache
ordering is untouched. Pure logic split into shard-balance.ts with a unit test
locking the disjoint+complete, balance, and determinism properties.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): install + cache FTS up front on the coverage (and cross-platform) shards (#2394)
coverage 3/3 failed on extension-binary-real.test.ts: it uses the file-path FTS
gate (requireFtsResourceOrSkip), which resolves ~/.lbdb/extension at MODULE LOAD
and cannot self-install the way the load-path gate (skipUnlessFtsAvailable, U8)
does. The coverage job had no FTS cache and relied on an installer test running
first in the shard — the balancing sequencer reshuffled the shards and dropped
extension-binary-real into a shard with no installer, so FTS was absent.
Remove the ordering dependency: add scripts/ensure-fts.ts (init a throwaway lbug
db, loadFTSExtension with policy:auto → LOAD-first, INSTALL on miss) and run it
up front on every coverage AND cross-platform shard, after restoring the per-OS
FTS cache. The coverage job now shares that same cache key (it previously had
none — this is the "share the cached FTS with coverage" the failure pointed at).
Cold cache installs once; warm cache is a no-network load. Verified locally:
ensure-fts installs FTS into a fresh HOME and is a no-op when already present.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(routes): add pure Python string-constant resolver (#2391 U1)
* feat(routes): extract Python module constants from tree (#2391 U2)
* feat(routes): capture non-literal FastAPI decorator args + per-file constants, bump parse-cache schema (#2391 U3)
* feat(routes): resolve composed decorator route constants in parse-impl + skip floor (#2391 U4)
* feat(routes): resolve composed FastAPI route constants in group HTTP-contract layer (#2391 U5)
* test(routes): multi-hop, ingestion↔group parity, and warm-cache regression locks (#2391 U6)
* docs(routes): mark the language-agnostic seam for cross-language const resolution (#2391)
* refactor(routes): extract language-agnostic constant-fold core; Python becomes a binding (#2391)
The fold, cycle guard, and depth cap now live in constant-resolver.ts and take a
pluggable ImportResolver. python-const-resolver.ts supplies the Python import
semantics + tree extractor and re-exports the same surface, so no call site
changes. A Spring/Kotlin/C# binding can now reuse the core with its own resolver
(proven by constant-resolver.test.ts driving it with a Java-style resolver).
* fix(routes): treat the constant-fold cycle guard as a recursion stack (#2391)
The `visited` set in `foldName` was added-to but never removed on unwind, so a
constant referenced more than once in a single fold — `A + A`, a reused
separator (`SLASH + PATH + SLASH`), or a diamond `X = P + Q` where P and Q share
a base — tripped the cycle guard on its second occurrence and the whole route
was silently dropped by the skip floor. Pop the guard in `finally` so it tracks
the ACTIVE resolution stack, not every name ever seen: a true cycle (a name
still on the stack) is still caught, but a name that already resolved and popped
folds again. Re-computation stays bounded by MAX_RESOLVE_DEPTH, so no blowup is
reintroduced.
Locked in constant-resolver.test.ts (A+A, reused separator, shared-base
diamond); the pre-existing real-cycle and depth-cap cases still return null.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(routes): make module-constant binding writes mutually exclusive (#2391)
`extractPythonModuleConstants` kept `literals`, `exprs`, and `imports` as three
independent maps: `setName` cleared literals+exprs but never `imports`, and an
import never cleared a prior literal/expr. Since `foldName` checks
literals > exprs > imports regardless of source order, a name that was both
imported and locally (re)assigned kept both bindings and the wrong one won —
`from .c import ROUTE; ROUTE = os.getenv(...)` resolved the STALE import instead
of dropping, a confidently wrong route path (the exact skip-floor invariant
this feature is meant to uphold).
Treat the three maps as one logical namespace: any write to one clears the
other two for that name (via `imports.delete` in `setName` and a `bindImport`
helper), so last-binding-in-source-order wins, matching Python. An import both
imported and dynamically rebound now drops. Folding `+=`/`+` onto an imported
base remains deferred (it drops safely, never a stale value).
Locked in python-const-resolver.test.ts: dynamic-rebind drops, literal-shadows-
import, import-shadows-literal, and `+=`-on-import drops.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(routes): widen the group cost-gate to catch literal-leading concats (#2391)
`NONLITERAL_ROUTE_DECORATOR_RE` required the first decorator argument to START
with an identifier, so a string-literal-leading concat like
`@router.get("/api" + SUFFIX)` never tripped `hasComposedRoute`. When such a
route was the ONLY composed shape in a repo, the group layer left `constantsByFile`
empty and dropped the route, while the ingestion side (which has no gate)
resolved `/api/users` and emitted a Route node — an R4 provider/graph parity break.
Widen the gate to also fire on a string-literal-leading `+`-concat, detected by a
`+` before the closing paren on the decorator line. Gating on the `+` (not merely
a leading quote) keeps a plain literal route `@router.get("/x")` OFF the gate, so a
literal-only repo still pays no parse pass.
Locked in fastapi-composed-provider.test.ts: a sole literal-leading concat now
resolves (parseCalls>0 + provider emitted), plus previously-uncovered
`@app.<verb>(CONST)` EXPR-branch resolution; the literal-only no-parse gate case
still passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(routes): correct package-init and over-deep relative import resolution (#2391)
Two edges in `resolvePythonImport`:
- `from . import X` (empty module after the dots) resolved to a sibling
`<dir>.py` instead of the package `<dir>/__init__.py`. Resolve the bare-package
case to `__init__.py`.
- An over-deep relative import (more extra dots than the importing file has
directory levels) silently clamped `dirOf('')` to `''` and could match an
unrelated root-level `<name>.py` — a wrong file. Guard with `walk > depth →
null` so an import that escapes above the repo root drops (skip floor).
Both preserve the exact-match / ambiguity→null behavior for ordinary relative and
absolute imports.
Locked in python-const-resolver.test.ts: `from . import` → `__init__.py` (and
null when absent), and an over-deep import returns null even when the clamped
target file exists.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(routes): bound parseConstOperands recursion depth (#2391)
`parseConstOperands` recursed on `binary_operator` children with no depth bound.
A stack overflow is not currently reachable (tree-sitter caps expression nesting
below the JS stack limit, so it throws on a deep `+`-chain before this runs), but
add a depth guard (cap 64, mirroring the fold engine's MAX_RESOLVE_DEPTH) as
defense-in-depth: a pathological chain now floors to null (skip) rather than
relying on tree-sitter's limit. The `depth` parameter defaults to 0, so all
existing callers are unaffected.
Locked in python-const-resolver.test.ts: a 100-term `+` chain yields no binding
(null) instead of throwing; ordinary short chains still fold.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(routes): read each .py once in buildPythonRepoContext (#2391)
The group repo-context builder read every `.py` file from disk twice: once in
the `include_router` cross-file pre-pass and again in the #2391 constant
cost-gate loop — an unconditional 2x read on every Python repo, on every group
extraction. Hoist a single read pass that populates one `pyContents` map (and
computes the composed-route cost gate); both the include_router pre-pass and the
constant-map pass now consume the cached content. Behavior-preserving — a
literal-only repo still does one read and zero parses.
Covered by the existing group unit + integration suites (R4 parity and
include_router prefix joins unchanged).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(routes): tidy constant-resolver docs and declaration order (#2391)
Three no-behavior nits from the PR #2393 review:
- Name `conditional_expression` (`x if c else y`) in the `parseConstOperands`
jsdoc list of shapes that deferred to null.
- Move `NONLITERAL_ROUTE_DECORATOR_RE` above `buildPythonRepoContext`, which
references it — it read as a forward reference before (runtime-safe, but
confusing).
- Correct the integration-test comment that called `/v2/api/v1/widgets/get`
"ingestion-only garnish": the group side emits it too (asserted separately);
the four paths in that block are the shared-parity set.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(routes): fold `X += "…"` onto an imported base constant (#2391)
Previously `from .c import BASE; BASE += "/v1"` dropped (the extractor could not
represent "the imported prior value" as an operand without self-referencing X and
tripping the cycle guard). Preserve the imported prior under a synthetic `$imp$N`
key — `$` can never appear in a Python identifier, so it cannot collide with a
real name — and reference it, so the augmented assignment folds to
`<imported BASE>/v1`. Extractor-only: no change to the `Operand` type, the fold
core, or the cache shape, so no SCHEMA_BUMP. An imported base that is itself
unresolvable still drops (skip floor preserved — never a wrong path).
Locked in python-const-resolver.test.ts: single and chained `+=` fold onto an
imported base; an unresolvable base still drops.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(routes): resolve bare decorator constants via the by-name entry (#2391)
The group `resolveExprArg` hand-built `[{ kind: 'ref', name }]` and called
`resolveOperands` for a bare-constant decorator argument — exactly what the
language-agnostic core's `resolveConstant(file, name, repo)` seam does. Call it
directly for the identifier case. This gives the previously test-only by-name
entry point a real production caller (it is the documented reuse seam for future
JVM/other bindings), drops the synthetic operand construction, and lets the now-
unused `Operand` type import go. Behavior-identical — the `+`-concat path still
parses to an operand list and folds via `resolveOperands`.
Guarded by the existing group provider suite (bare-constant and concat cases).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(routes): parse each .py once in buildPythonRepoContext (#2391)
The repo-context builder ran two parse loops — the include_router prefix pre-pass
and the #2391 constant-map pass — so an include_router file in a composed repo was
tree-sitter-parsed twice. Merge them into a single pass that parses each `.py` at
most once and feeds both extractions from the same tree; a file that needs neither
pass is still not parsed at all (cost gates unchanged). Complements the earlier
single-read-pass change (this is the single-parse counterpart).
Behavior-preserving (prefixes, R4 parity, and cost gates verified by the group +
integration suites). Locked with a parseCalls assertion: a file needing both
passes is parsed once, not twice.
Note: a cross-run (cross-process) constant-map cache — the other deferred perf
idea — remains out of scope; it needs disk persistence + invalidation and would
add hashing/IO cost on the common path, so it fails the minimal-change bar this
single-parse dedup meets.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(routes): bound constant-fold work and output to prevent OOM (#2391)
The `finally`-popped cycle guard (recursion-stack semantics) correctly folds
diamonds/repeated refs, but popping the guard removed the accidental work cap the
old seen-ever set provided: a wide shared-descendant DAG re-folds each child once
per reference, and a self-multiplying concat (`X = A + A; A = B + B; …`) builds a
genuinely exponential string. Reviewers reproduced ~16.8M folds escalating to
`RangeError: Invalid string length` and heap OOM — and neither fold call site is
wrapped in try/catch, so it crashed the whole phase rather than dropping the route.
Two complementary bounds, both flooring to null (skip), never a wrong value:
- a never-popped `memo` in `foldName` caps recomputation at O(nodes) (successes
only — a null may be transient on a cyclic branch);
- a `MAX_FOLD_LENGTH` (8192) cap in `foldExpr` drops a fold whose output grows
past any real route path, bounding the string size the depth cap does not.
Corrects the prior "≤ 2^8 folds" comment (output grows multiplicatively, not
additively). Locked with a 64^4-fanout construction that now drops in ~ms instead
of OOMing; diamonds/cycles/depth-cap behavior unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(routes): snapshot assignment RHS refs at the assignment line (#2391)
`ROUTE = BASE` was stored as a lazy `ref(BASE)`, resolved against BASE's FINAL
binding. So `ROUTE = BASE; BASE += "/v1"` (or `ROUTE = API; API = "/other"`)
resolved ROUTE to the MUTATED value — a confidently wrong path, since Python
assigns by value at the `ROUTE =` line. This was latent for local constants at
the base of this feature and the `+=`-on-import work extended it to imports.
Snapshot each assignment/`+=` RHS reference to a bound name into that name's
current frozen value at the assignment line (`freeze`/`snapshot`): a literal
value, a copy of the current expr (whose refs are already frozen), or an import
preserved under a `$imp$N` alias. Unbound refs (forward references) stay lazy.
A later rebind of the aliased name can no longer change the earlier binding.
`freeze` also unifies the previous `currentOps` + inline import-alias logic.
Locked in python-const-resolver.test.ts: aliased-import-then-`+=`,
aliased-local-then-`+=`, aliased-local-then-rebind all resolve to the pre-mutation
value; normal reference chains still fold.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(routes): fold group identifier args via resolveOperands for parity (#2391)
Resolving a bare-constant decorator arg through `resolveConstant` entered
`foldName` at depth 0, whereas the ingestion side folds `routePathOperands`
through `resolveOperands([{ref}])`, entering at depth 1. At the MAX_RESOLVE_DEPTH
boundary the group tolerated one more hop than ingestion, so a deep alias/re-export
chain resolved in the group provider set but dropped from the graph Route nodes —
an R4 parity break. Restore the operand-list path in the group so both subsystems
share identical fold-entry depth. (`resolveConstant` reverts to the documented
agnostic-core seam.)
Locked in constant-resolver.test.ts: a 4-hop chain that `resolveOperands([ref])`
drops but `resolveConstant` resolves, documenting why the group must use the
operand-list entry.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(routes): match multiline literal-leading concats in the cost gate (#2391)
`NONLITERAL_ROUTE_DECORATOR_RE` used `[^)\n]*` so it only saw a literal-leading
`+`-concat when the `+` was on the same line as the opening quote. A
Black-formatted `@router.get(\n "/api"\n + SUFFIX\n)` therefore failed the gate,
and when it was the only composed route in a repo the group dropped it while
ingestion (which parses the tree, not the raw line) resolved it — an R4 parity
break. Drop the `\n` exclusion: `[^)]*` spans the wrapped argument but stays
bounded by the decorator's own closing paren, so a plain literal route still
never trips the gate.
Locked in fastapi-composed-provider.test.ts with a multiline concat fixture.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(routes): bump SCHEMA_BUMP for changed extractor output + E2E snapshot lock (#2391)
`extractPythonModuleConstants` now emits DIFFERENT `moduleConstants` for the same
source (binding mutual-exclusivity clears stale imports; RHS refs are snapshotted;
`$imp$N` aliases). That output is cached verbatim in the parse cache, so a warm
shard built at the pre-fix version would replay stale — in one case actively
wrong — folded values, and the correctness fixes would silently no-op on upgrade.
Bump SCHEMA_BUMP 11→12 to force re-extraction (same warm-cache-replay class the
original 10→11 bump addressed for the field addition).
Also adds the first end-to-end coverage for the new behavior through the real
ingestion pipeline: app/snapshot.py aliases a constant (`SNAP = API_V1`) then
mutates the source (`API_V1 += "/mutated"`), and the test asserts the Route node
is `/api/v1`, never `/api/v1/mutated` — a case the pure-function unit tests
covered but the pipeline did not.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(lbug): classify FTS extension load errors with Windows missing-dependency guard (#2374)
Add classifyExtensionLoadError() — a pure-string, lbug-free four-way
classifier (missing_file / corrupt_file / missing_dependency / unknown).
The Windows catch-all guard keys missing_dependency strictly on the
error-126 signal, never LadybugDB's generic 'Failed to load library …
needed by extension' wrapper, so 127/5/1114 and truncated (193) files
route correctly.
* feat(fts): surface classified missing-dependency remedy in doctor, repair-fts, and degrade warnings (#2374)
Route the FTS load reason through classifyExtensionLoadError at all four
surfaces (doctor, --repair-fts error, analyze degrade log, ftsDegradedWarning).
For the Windows missing-dependency class, emit the runtime-install remedy
(VC++ redist, then OpenSSL) instead of the wrong reinstall-over-network
guidance; other classes keep their existing routing. Path redaction preserved
on the client-facing warning.
* test(fts): assert doctor surfaces the classified remedy end-to-end (#2374)
Extend the broken-file e2e: doctor now prints the corrupt-file re-download
remedy through the real CLI, and the Windows missing-dependency remedy
(VC++/OpenSSL) must not misfire on a corrupt file — the catch-all guard,
verified end-to-end. Also assert the repair path does not misfire.
* style(fts): apply prettier formatting to #2374 diagnosis files
* feat(fts): language-independent hedged fallback for Windows load failures (#2374)
The Windows OS-error tail is localized, so matching only en/zh 126 text left
other locales on the generic 'run doctor' remedy. lbug's 'Failed to load
library' wrapper is English on every platform and present for all load
failures, so use it as a fallback: when the localized tail matches no specific
class, emit a hedged remedy that points the user at their own OS error and
offers both branches (install runtime / --repair-fts) without prescribing the
wrong single fix. Precise en/zh 126 keeps its definite remedy.
* feat(fts): language-independent structural classifier via binary inspection (#2374)
Add diagnoseExtensionLoad: pull the extension's file path out of lbug's own
English wrapper and inspect the binary header (PE/ELF/Mach-O magic + arch)
directly, so corrupt-vs-valid is decided by the file itself, not the localized
OS-error tail. A valid binary that still failed to load ⇒ missing_dependency
(runtime dep), decided in any OS display language and on all three platforms.
Falls back to the string classifier (with its hedged fallback) when the file
can't be read. Wire all four surfaces to it. Event Viewer / GetLastError-via-FFI
were dead ends (lbug catches the failure — no crash event; no native FFI dep).
* test(fts): exercise the structural classifier on real binaries (#2374)
Add an integration suite that runs inspectExtensionBinary/diagnoseExtensionLoad
against genuine binaries — the running node executable, the real lbugjs.node
addon, and the installed FTS extension (valid); a truncated real binary and a
real text file (corrupt). Registered in cross-platform-tests PLATFORM_LOGIC so
it runs on the Windows + macOS matrix, proving the PE and Mach-O header parsing
on real PE/Mach-O files (ubuntu covers ELF).
* fix(fts): honor a corrupt_file verdict over a structurally-valid header (#2374)
The structural probe in diagnoseExtensionLoad inspects only the first 4 KB, so a
download truncated after its header reads 'valid' and was routed to the "install
VC++, reinstalling will NOT help" remedy — the exact loop #2374 exists to kill,
for the truncated-download case the module docstring claims it handles. Honor the
loader's own corruption report ("file too short" / Windows error 193
"not a valid Win32 application") before defaulting to the dependency remedy;
localized corrupt tails stay hedged missing_dependency, preserving
language-independence.
Addresses PR #2383 review finding F1.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(fts): return indeterminate for a PE header beyond the read window (#2374)
The structural probe reads only BINARY_HEADER_BYTES (4 KB). A valid PE with a
large DOS stub whose e_lfanew points past that window was wrongly called
'corrupt', routing a fine DLL to "re-download". A garbage e_lfanew from a truly
corrupt file is indistinguishable from here, so widen the header verdict with
'indeterminate' and return it in that case; the caller then defers to the
loader's own report instead of asserting a false verdict. Fat Mach-O stays valid
(LadybugDB ships thin per-arch binaries). Also covers the unmapped-arch and
garbage-PE-signature branches.
Addresses PR #2383 review finding F1-secondary.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(fts): drop contradictory reinstall guidance from the analyze degrade log (#2374)
For a missing runtime dependency the extension file is present, so appending
FTS_UNAVAILABLE_MESSAGE (which tells the user to install it "with network
access") to the remedy ("reinstalling will NOT help") produced self-contradictory
guidance on the main analyze surface. Lead the missing_dependency degrade log
with the class-neutral sentence (FTS_UNAVAILABLE_LEAD) and append only the
classified remedy; other classes keep FTS_UNAVAILABLE_MESSAGE unchanged.
Addresses PR #2383 review finding F2.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(fts): cache the load diagnosis so the degraded warning does no per-request I/O (#2374)
ftsDegradedWarning() runs on every degraded /api/search response and MCP query,
and it was calling diagnoseExtensionLoad — a synchronous openSync/readSync of the
extension file — on every call. Compute the diagnosis once at mark-unavailable
time (the single load-failure sink, run per Database not per request), cache it on
ExtensionCapability, and have the warning read the cached result (falling back to
the pure, no-I/O string classifier if it is absent). Loader capability-shape
assertions relax from toEqual to toMatchObject for the new optional field.
Addresses PR #2383 review finding F3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(fts): cover the missing_dependency remedy on the --repair-fts path (#2374)
The repair-fts error interpolates the classified remedy, but no test reached the
missing_dependency branch — only the corrupt/invalid-ELF path. Add a Windows
error-126 case asserting the thrown error carries the VC++ redistributable remedy
and omits the old "retry the network install" tail, and that no index is dropped.
Addresses PR #2383 review finding F6a (--repair-fts surface).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(fts): share the VC++ redistributable install hint (#2374)
The Microsoft Visual C++ redistributable name and aka.ms URL were duplicated
verbatim in WINDOWS_MISSING_DEPENDENCY_REMEDY and
STRUCTURAL_MISSING_DEPENDENCY_REMEDY. Factor a single VC_REDIST_INSTALL_HINT
constant so the pointer cannot drift between them; the composed remedy strings are
byte-identical (existing exact-text assertions unchanged). Also adds a test
covering the previously-unexercised structural remedy branch.
Addresses PR #2383 review finding F5a.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(fts): guard FILE_CORRUPTION_SIGNATURES parity with the installer script (#2374)
The corruption-signature list is deliberately duplicated between
extension-load-error.ts and scripts/install-duckdb-extension.mjs (the .mjs cannot
import the .ts), with nothing guarding against drift — a one-sided edit would
desync the FORCE-INSTALL verb from remedy classification. Export the array from
both and add a parity test that compares regex source + flags element-wise.
Addresses PR #2383 review finding F5b.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(test): run extension-binary-real in the sequential lbug-db vitest project (#2374)
extension-binary-real.test.ts imports @ladybugdb/core but ran in the parallel
`default` project, contrary to TESTING.md's rule that native-LadybugDB tests live
in the sequential `lbug-db` project. Add it to the lbug-db include list and the
default exclude list; it now runs under lbug-db and no longer under default.
Addresses PR #2383 review finding F6c.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(fts): fail loud, not silent-skip, on missing FTS artifacts under REQUIRE_FTS=1 (#2374)
The real-binary structural tests gated on raw .skipIf(!lbugNative) /
.skipIf(!installedFts), so under GITNEXUS_REQUIRE_FTS=1 a missing artifact would
silently vanish from a green CI run (the #2299 trap). These tests inspect the
extension file directly and need its path, not a loaded connection — so
skipUnlessFtsAvailable (which needs an initialized LadybugDB) does not fit. Add
requireFtsResourceOrSkip: skip gracefully offline, throw under REQUIRE_FTS=1. The
always-on process.execPath assertion still runs everywhere.
Addresses PR #2383 review finding F6d.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style(fts): apply prettier formatting to the #2383 fix files (#2374)
Line-wrapping only; the quality/format CI check flagged three files. No behavior
change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(lbug): store exact symbol content snippets
* fix(ingestion): emit 0-based line numbers for COBOL/JCL/scope/markdown nodes
COBOL/JCL processors, the scope-graph emitter, and the markdown Section
emitter stored 1-based startLine/endLine, unlike every tree-sitter node
(0-based). The exact-content slice (#2379) then dropped each symbol's
declaration line for those languages. Convert to 0-based at the graph-node
emission boundary via toZeroBasedLine — leaving parser-internal .line values,
L${line} node/edge IDs, and containment checks untouched.
Refs #2377, #2379
* refactor(lbug): single source of truth for symbol-content labels
Extract SYMBOL_NODE_LABELS so the exact-content label set can't drift the way
the inline copy did in #2379. csv-generator derives EXACT_SYMBOL_CONTENT_LABELS
from it; manifest-extractor's near-identical allowlist is left behavior-unchanged
(intentional subset, #2325-test-locked) with a documented cross-reference.
Refs #2379
* test(ingestion): cover 0-based emitter output and pin exact-content slicing
- csv-pipeline: replace the blank-buffer fixture (a +/-1 shift silently passed)
with directly-adjacent neighbors; add one-line-symbol and Section (+/-2 fallback)
cases.
- cobol resolver: assert COBOL Module and JCL job/step emit 0-based startLine.
- markdown CRLF: update Section startLine/endLine expectations to 0-based.
Refs #2377, #2379
* feat(mcp): present 1-based line numbers in context/query/impact tools
GraphNode startLine/endLine are stored 0-based (tree-sitter rows), which
surprised users querying them (they don't line up with editors/sed). Add
toDisplayLine and apply it at the context/query/impact response boundaries so
line numbers are editor/sed-aligned. Raw cypher stays 0-based (documented in the
schema resource); BasicBlock/PDG statement lines (already 1-based) and internal
join params are left untouched.
Refs #2377
* test(mcp): assert 1-based tool exposure with raw cypher staying 0-based
context() reports startLine+1 (editor/sed aligned); a raw cypher RETURN of the
same node keeps the stored 0-based value. Guards against double-conversion and
leaking the display shift into raw results.
Refs #2377
* fix(mcp): stop query() double-converting BM25 line numbers
bm25Search applied toDisplayLine to its result rows, and query()'s
aggregation loop applied it again, so BM25-matched symbols reported
lines shifted +2 (stored 0-based 41 read as 43, not 42) while
semantic-matched symbols were correct. bm25Search is called only from
query(); return raw 0-based rows and let the single aggregation-loop
conversion handle both retrievers.
Adds a query() BM25 regression test asserting stored 41 -> 42 (would
be 43 if double-converted), which the prior mcp-line-display test —
covering only context()+cypher — never exercised. (#2380, #2377)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): use ?? not || so first-line symbols keep their line number
`sym.startLine || sym[4]` treated a legitimate 0-based startLine of 0
as absent, so context()/query() dropped startLine/endLine for every
symbol on line 1 of its file — every COBOL Module (toZeroBasedLine(1)
= 0) and markdown h1. `??` only falls through to the positional
fallback on null/undefined, preserving a real 0. This also repairs the
rename definition-edit path, which consumes context()'s value.
Adds a context() first-line (startLine:0 -> 1) assertion. (#2380, #2377)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): make group/cross-repo trace line numbers 1-based consistently
A group/cross-repo trace presented 1-based endpoints (via
resolveSymbolForGroup) but 0-based hops (tagHops copies port.trace
output verbatim), so one response mixed bases. Wrap the trace port
adapter (traceForGroup) to convert hop lines to 1-based too, matching
the endpoints. Single-repo trace dispatches directly (not through this
port) and stays 0-based — full single-repo parity is a tracked
follow-up. core/group stays display-agnostic (no mcp import).
Extends the cross-trace e2e test to assert hops share the endpoints'
base (checkout 10 -> 11, getUsers 1 -> 2). (#2380)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): present explain/pdg_query anchor line 1-based
resolveBlockAnchor converted its ambiguous-candidate lines to 1-based
but left the resolved-target anchor raw 0-based, so the same tool
reported two bases depending on whether the target was ambiguous.
Convert the display anchor to 1-based via toDisplayLine. The BasicBlock
join param (symStart: sym.startLine + 1) is untouched — it targets the
1-based BasicBlock id space, not display.
Asserts the resolved anchor is 1-based (targetFn stored 10 -> 11). (#2380)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): bump schema + PDG result versions for the line-number change
The 0-based storage flip for COBOL/JCL/markdown/scope (#2377/#2379)
changed on-disk line semantics, and the PDG result startLine is now
1-based (#2380). Neither shipped a version bump, so an incremental
re-analyze would preserve old 1-based rows (mixed-base index rendered
one line too high) and PDG consumers got no signal.
- INCREMENTAL_SCHEMA_VERSION 5 -> 6 (forces a one-time full re-analyze)
- PDG_RESULT_VERSION 1 -> 2 (result-shape discriminator)
Updates the version-pinning tests, the pdgResultVersion result type,
and the tools.ts PDG output-contract doc. (#2380)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(group): guard manifest label list against SYMBOL_NODE_LABELS drift
manifest-extractor's CUSTOM_CONTRACT_RESOLVE_QUERY hand-lists the
contract-resolvable labels as a deliberate subset of the shared
SYMBOL_NODE_LABELS, guarded only by a comment — the same drift class
(#2379) the shared-set refactor eliminated elsewhere. Derive the
query's label set and assert it is a strict subset whose difference is
exactly {Namespace, Variable, Module}, so adding a symbol label without
a conscious manifest decision fails. Query string stays literal
(#2325-test-locked). (#2380)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(mcp): document which tools present 1-based vs 0-based line numbers
The schema-resource note listed only context/query/impact as 1-based.
After the trace/anchor fixes it now enumerates the full set —
context, query, impact, group/cross-repo trace, and explain/pdg_query
anchors are 1-based; raw Cypher and single-repo trace stay 0-based
(full single-repo-trace parity is a tracked follow-up); BasicBlock/PDG
statement lines are separately 1-based. (#2377, #2380)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(mcp): pin impact() line-value display (close the coverage gap)
The prior mcp-line-display test only asserted context() + raw cypher,
which is why the query() double-conversion (#2380) shipped green. Adds
an impact() line-value assertion via the ambiguous-candidate path (the
only impact response that surfaces a per-candidate line): two same-name
symbols force ambiguity and the candidate at stored 0-based 41 must
read 42. (#2380)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(mcp): fix stale rename #2283 mock after 1-based context display
rename resolves its symbol via context(), which now presents startLine
1-based (#2377), then subtracts 1 to recover the 0-based file index.
The #2283 mock stored startLine:1 but put `oldName` on the file's line
0, so after the 1-based shift the definition edit no longer matched and
the write-failure path never fired — the test read 'success' instead of
'partial'. Align the mock content to its stored line (oldName on
0-based line 1). Pre-existing failure surfaced once ubuntu/coverage
completed on this branch. (#2380)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(mcp): consolidate line-display tests into one shared DB block
The query()/BM25 case had spun up a second full LadybugDB + FTS setup;
fold it into the single existing block (adding FTS + the Zqxwvbm seed
there) so the file builds one DB, not two. Trims per-file setup cost —
relevant to the Windows platform-sensitive suite's under-load 15-minute
timeout. Same five assertions, all green. (#2380)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: kigland <shuaizhicheng336@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(setup): install Codex PreToolUse/PostToolUse hooks (#2328)
Codex CLI supports lifecycle hooks with Claude Code's exact
{hooks: {Event: [...]}} JSON schema, stdin payload, and
hookSpecificOutput response contract, registered in a dedicated
~/.codex/hooks.json (https://developers.openai.com/codex/hooks).
Parameterize installClaudeCodeHooks into installClaudeSchemaHooks
(claude | codex): both runtimes share the installer, the bundled
gitnexus-hook.cjs adapter, and its helpers. A codex HookTarget in
editor-targets.ts makes uninstall and the setup-uninstall round-trip
tripwire cover the new surface with no uninstall.ts changes.
SessionStart is deliberately not registered: Codex reads AGENTS.md
natively, which already carries the GitNexus context block.
Closes#2328. Closes#244 (Codex setup support is now complete:
MCP + skills + hooks).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(plugin): make the GitNexus plugin installable from Codex (#1131)
Codex's plugin system (https://developers.openai.com/codex/plugins/build)
reads a .codex-plugin/plugin.json manifest and a repo-root
.agents/plugins/marketplace.json registry. The existing
gitnexus-claude-plugin/ is already Codex-compatible as-is — Codex sets
CLAUDE_PLUGIN_ROOT for hook-command compatibility, loads the same
SKILL.md skills, hooks/hooks.json, and .mcp.json — so a second manifest
in the same folder replaces PR #1131's duplicated plugin tree with zero
copied skills or hooks. The .gitignore .agents/ scratch rule narrows to
re-include only the registry file.
Install: codex plugin marketplace add abhigyanpatwari/GitNexus
Supersedes #1131.
Co-authored-by: jublin <1799126+jublin@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: document Codex full support (MCP + skills + hooks + plugin)
Promote Codex to Full in both editor tables, document the
~/.codex/hooks.json hook install, and add the Codex plugin
marketplace install path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(release): extend the version-lockstep guard to the Codex manifests
The always-on drift guard asserted only the Claude plugin manifests
against gitnexus/package.json, so a release could ship stale versions in
.codex-plugin/plugin.json and .agents/plugins/marketplace.json without
CI noticing. Mirror the Claude lockstep test for the two Codex files and
extend the CONTRIBUTING §Releases lockstep list to match.
Verified guard semantics: a deliberate local version mutation of the
Codex marketplace entry turns the new test red.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(plugin): quote the hook command path for space-containing plugin roots
Both plugin hook commands ran `node ${CLAUDE_PLUGIN_ROOT}/hooks/...`
unquoted, which breaks whenever the substituted plugin root contains a
space — the common case on Windows user profiles. Both Claude Code and
Codex substitute the placeholder before shell execution, and Claude
Code's plugin docs mandate the double-quoted form in shell-form hooks.
No commandWindows entry: Codex source (codex-rs hooks engine) falls back
to `command` on Windows with identical placeholder substitution, so an
identical-content override would be pure duplication.
Verified: space-in-root smoke test (old form exits 1 MODULE_NOT_FOUND,
quoted form exits 0), `claude plugin validate` passes, and a local
`codex plugin marketplace add` parses the marketplace + plugin cleanly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(setup): pin fail-closed behavior for unreadable/corrupt Codex hooks.json
The non-ENOENT suite covered Claude settings.json (EACCES) and Codex
config.toml (EACCES) but not the new ~/.codex/hooks.json surface, and the
mergeHooksJsonc "is corrupt" branch had zero coverage for either editor.
A future refactor dropping the isEnoent rethrow or the parse gate could
silently rewrite a user's hooks.json gitnexus-only with no CI tripwire.
Two regression tests: EACCES leaves hooks.json byte-identical and reports
"Codex hooks: EACCES"; corrupt content is preserved and reported via
"Codex hooks: hooks.json is corrupt".
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(readme): add the Codex plugin-marketplace install path to the npm README
The root README documents the one-step plugin route but the package
README (what npmjs.com renders) only showed the setup-CLI path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(setup): rename claudeHook to hookCfg in installClaudeSchemaHooks
The local held a codex HookTarget on the codex branch since the installer
was parameterized, so the claude-specific name misled. Pure local rename,
no behavior change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(readme): document Codex SessionStart exclusion, /hooks trust gate, and install-route choice
Three behaviors were only recorded in code comments and the PR body:
SessionStart is deliberately not registered (Codex reads AGENTS.md
natively), setup-installed hooks need one-time /hooks approval in Codex,
and the setup CLI and plugin are alternative install routes whose hooks
load alongside each other if both are used.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(test): share one logLines helper across setup.test.ts describes
The corrupt-hooks.json test inlined the console.log-flattening
expression that the non-ENOENT describe already defined locally. Hoist a
single file-scope logLines so the two stay in sync.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(setup): add CodeBuddy and Qoder coding-agent integrations
Adds Tencent CodeBuddy and Alibaba Qoder to gitnexus setup/uninstall,
fitted to the editor-targets registry and --coding-agent selection.
- CodeBuddy: MCP entry written into the first existing file of its
documented priority chain (~/.codebuddy/.mcp.json recommended,
~/.codebuddy/mcp.json deprecated, ~/.codebuddy.json legacy) so a
populated deprecated config is never shadowed; skills to
~/.codebuddy/skills/ (https://www.codebuddy.ai/docs/cli/mcp)
- Qoder: MCP entry in ~/.qoder.json, skills to ~/.qoder/skills/
(https://docs.qoder.com/cli/using-cli, /extensions/skills)
- editor-targets gains optional legacyFiles; uninstall sweeps them
- roster strings updated (CLI help, i18n en/zh-CN, READMEs); en/zh-CN
setup descriptions were stale (missing Antigravity) and are refreshed
Supersedes and credits PR #1030 by @zykai0302, re-fitted to the
post-#2168 selective-agent architecture with documented config paths.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(cli): assert stable zh-CN setup-description fragment
* fix(setup): surface non-ENOENT config read/stat failures instead of clobbering
* fix(setup): report corrupt legacy MCP files informationally during uninstall
* test(setup): cover multi-candidate uninstall sweep combinations
* fix(setup): detect CodeBuddy/Qoder installs via existing MCP config files
* fix(setup): skip empty and non-file candidates in the MCP config chain
* docs: add CodeBuddy and Qoder manual MCP configuration sections
* test(ci): run the setup-uninstall round-trip in the cross-platform matrix
* fix(setup): never claim "not configured" when uninstall recorded errors
* refactor(cli): share the isEnoent predicate via editor-targets
* refactor(setup): share chain-file install detection between CodeBuddy and Qoder
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat: flat workspace index follows the checked-out branch (#2354)
A plain `gitnexus analyze` now always targets the flat workspace slot,
updating it incrementally across branch switches instead of auto-routing
non-owner branches into `branches/<slug>/` sub-indexes (disk bloat) or
nagging with the primary-inversion "run gitnexus clean" warning. No new
CLI flag or config key: the smart behavior is the default.
- Placement: only explicit `--branch` consults resolveBranchPlacement;
plain runs resolve to the flat slot, `meta.branch` becomes an
informational "last analyzed branch" label restamped each run.
- Fast path: a same-commit clean-tree branch flip restamps the label and
registry entry (adoptFlatBranchLabel, no-op for unregistered repos).
- Shadow cleanup: when the flat slot adopts a label that has a pinned
sub-index, the now-unreachable `branches/<slug>/` dir and its registry
summary are removed together.
- MCP: applyBranchScope always falls back to the on-disk flat meta before
throwing "not indexed", so long-lived servers resolve a freshly
restamped workspace branch.
- status: no more "current branch not indexed" dead end — falls through
to the workspace index with an informational line and the usual
commit-based staleness verdict.
- Deleted primaryInversionWarning; explicit `--branch` pinning, the
checkout-mismatch guard, detached-HEAD/CI behavior, and `clean
--branch` are unchanged.
Supersedes the flag-based approaches in #2358/#2359.
Closes#2354.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(storage): check registry before deleting shadowed sub-index (#2364 review F2)
adoptFlatBranchLabel ran the branches/<slug>/ rm before its own
unregistered-repo no-op check, so a repo in the #2264 half-finalized
state (up to date but unregistered) lost its pinned sub-index on a
same-commit branch flip while the run still failed. The registry
lookup now precedes the deletion, making the no-self-heal rule
(#2264/#1169) cover disk as well as registry state.
The 'never self-heals' unit test now materializes a sub-index dir and
asserts it survives; the run-analyze #2354 fast-path test registers
its repo under an isolated GITNEXUS_HOME (deletion is only legitimate
for registered repos) with a new unregistered variant pinning
dir survival.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(storage): keep branch summary when sub-index rm fails (#2364 review F4)
The shadow-cleanup fs.rm swallowed every error while the registry
summary was dropped unconditionally. On Windows an lbug held open by a
live MCP server fails the rm with EBUSY/EPERM, and once the summary is
gone 'clean --branch' can never target the leftover dir (it resolves
solely via the recorded summary) — stranding the exact un-cleanable
disk bloat adoptFlatBranchLabel exists to prevent.
The summary is now dropped only when the directory is verifiably gone
(post-rm existence check); on failure the summary is retained, a
warning names the path and errno, and the informational branch label
still restamps. Later adopts retry the rm.
New repo-manager-rm-failure.test.ts uses the delegating fs/promises
mock idiom (vi.spyOn cannot intercept ESM namespace exports).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(core): restamp fast path adopt-first and tolerate read-only storage (#2364 review F3)
The fast-path label sync stamped meta before adoptFlatBranchLabel, so
a crash or adopt failure between the two flipped the retry guard
(existingMeta.branch !== branchLabel) and locked in the partial state:
every subsequent same-commit run skipped the cleanup and branch-scoped
queries kept routing to the stale pinned sub-index. The block also sat
outside any try/catch, so a same-commit branch flip on a read-only
.gitnexus mount (the documented Docker :ro workflow, #1549) failed a
byte-for-byte-current analyze over a purely informational label sync.
Adopt now runs first and saveMeta last — any partial failure leaves
the guard true and the next run self-heals — and the whole sync is
best-effort: read-only errors warn citing #1549, anything else warns
and retries next run. Safe because the block only fires on a
same-commit clean tree, where the flat DB content is byte-valid for
both labels. isReadOnlyFilesystemError is now exported.
New run-analyze-adopt-failure.test.ts covers retry-after-partial-
failure, adopt-before-stamp ordering, and EROFS/EACCES/EPERM (gaps 4
and 7); a detached-HEAD fast-path pin lands in run-analyze.test.ts
(gap 6).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(mcp): make flat meta authoritative in applyBranchScope (#2364 review F1)
applyBranchScope trusted two pieces of cached state before its flat-
meta disk fallback, and the handle cache only refreshes on a resolve
miss — never on a hit. Post-#2354 that stale window is the routine
case: (i) the handle.branch early-return served the flat handle under
the OLD label after a workspace flip, silently returning the new
branch's content as the old branch (the pool staleness reinit hot-
swaps content without updating handle.branch); (ii) a stale cached
branches[] summary routed to a branches/<slug>/ dir that
adoptFlatBranchLabel had already deleted (raw 'LadybugDB not found' or
POSIX ghost reads with staleness detection blinded).
The on-disk flat meta is now read before any cached-state trust. A
branches[] summary is served only when its sub-index lbug actually
exists (the lbug is what the pool opens — serviceability truth); the
cached label is trusted only when no readable flat meta contradicts it
(#2106 R4 legacy shapes preserved). One refreshRepos() fires on
detected staleness so subsequent calls see fresh handles. Safe against
mid-analyze reads: dirty stamps spread the existing meta, preserving
the old label until the end-of-run atomic write.
Fixtures now materialize the pinned sub-index lbug; new regressions
cover the stale-old-label error, adopted-summary fall-through to flat,
and the dangling-summary partial-failure window (test gaps 1-2).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(core): make end-of-run branch-label sync best-effort (#2364 review F5)
The end-of-run adoptFlatBranchLabel sat inside the pipeline try whose
catch rethrows, so a registry write failure (ENOSPC, ~/.gitnexus
perms) after a successful multi-minute analyze failed the whole run —
even though the index was complete and registered, the neighbouring
parse-cache save is deliberately wrapped for exactly this reason, and
adopt retries unconditionally on the next plain analyze. It now warns
and continues, mirroring the parse-cache wrapper.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(mcp): correct branch-not-indexed guidance for workspace index (#2364 review F6)
The error told users to 'Run: gitnexus analyze --branch <X>', but
post-#2354 that command hard-errors unless X is checked out — and this
message is now the common goodbye for a formerly-indexed branch whose
sub-index the workspace slot adopted. The guidance now explains that
the workspace index follows the checked-out branch and leads with the
checkout; the '(primary only)' fallback becomes '(workspace only)'.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: align primary/workspace vocabulary with the #2354 inversion (#2364 review F7)
The review flagged pre-inversion 'primary/non-primary' wording that
now misleads readers about the placement model: the isPrimaryBranch
JSDoc (field name kept — public API surface), the two branches? JSDoc
comments in local-backend, the base_ref gate comment in cli/analyze,
and four branch-scope test names. Comment/JSDoc/test-name edits only;
'Registry-primary' and 'primary key' senses untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(cli): clarify workspace index status wording (#2364 review F8)
'gitnexus analyze follows this branch' was ambiguous about WHICH
branch analyze follows — the recorded one on the line or the current
checkout. Both locales now say a re-run follows the current branch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(storage): re-read registry after the shadow rm in adoptFlatBranchLabel
The F2 reorder moved the registry read to the top of the function, so
the whole-file writeRegistry at the bottom persisted a snapshot taken
BEFORE the recursive rm of an entire sub-index — widening the unlocked
read-modify-write window from microseconds to the duration of a multi-
hundred-MB delete. A concurrent registerRepo/removeBranchIndex writer
in that window was silently clobbered (the #2106 R9 lost-update class;
registerRepo re-reads before writing for exactly this reason).
The top read is now a cheap membership gate only (the F2 no-op
guarantee); the mutate re-reads its own fresh snapshot after the rm.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: treat only provably-absent errno as gone in the new existence probes
Both probes added by this series inverted the codebase's provably-
absent polarity (listRegisteredRepos validate prunes only on
ENOENT/ENOTDIR): adoptFlatBranchLabel's dirGone check read ANY
fs.access failure — including a transient EACCES/EIO on a surviving
dir — as 'verifiably gone' and dropped the summary, recreating exactly
the stranded-bloat bug F4 fixed; applyBranchScope's sub-index check
read the same transient errors on a healthy pinned lbug as 'adopted/
deleted', producing a false 'not indexed' error. A resolved force:true
rm now proves absence without a probe; on failure the probe treats
only ENOENT/ENOTDIR as gone, and a non-missing lbug serves the handle
so the pool open surfaces the real error.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(mcp): harden applyBranchScope stale-state coherence
Four residual gaps in the new arm structure, found by post-fix review:
- The stale-label error listed the just-contradicted cached label as
indexed ('not indexed: main. Indexed branches: main'). The message
now derives the flat label from the authoritative meta and excludes
the requested branch from the hint list.
- A branch pinned AFTER the server cached its handle never triggered a
refresh (resolve hits skip the miss-refresh), erroring until restart.
Every miss now fires exactly one best-effort refreshRepos() before
the error, so the next call resolves; a refresh-once guard keeps
doubly-stale resolutions to a single registry re-scan.
- A registry entry claiming the branch both as flat label and pinned
summary (the rm-failed adopt-degraded state) could serve the stale-
vintage pin under a label the flat slot owns; the summary arm now
requires handle.branch !== branch and the degraded state errors
honestly.
- The flat-meta match path returned the cached handle's pre-restamp
branch/commit/stats; the meta that decided routing now also supplies
the metadata.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(core): keep the real error visible in restamp warnings; correct the end-of-run retry claim
The fast-path catch replaced the actual error with 'storage is
read-only (#1549)' for any EACCES/EPERM — mislabeling ownership
problems and transient Windows locks and discarding the only
diagnostic signal. The warning now carries the real message with the
#1549 hint appended.
The end-of-run best-effort comment claimed adopt 'retries
unconditionally on the next plain analyze'; same-commit runs take the
fast path whose guard compares the already-stamped meta label, so the
retry actually lands on the next content-changing run. The comment now
states the true retry semantics and why the interim state is safe
(flat meta stamped first; applyBranchScope trusts it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: unique tmpdir for the branch-scope fixture; drop redundant dynamic imports
The branch-scope describe materialized its sub-index stub under a
FIXED os.tmpdir()/gnx-2106-multi path — concurrent vitest runs on one
host (the documented parallel-agents workflow) could rm each other's
stub between beforeEach and the resolve under test, flaking the
pinned-branch tests. The fixture root is now mkdtemp-unique per run
with afterAll cleanup.
run-analyze.test.ts dynamically imported repo-manager inside test
bodies despite the module being statically imported at the top of the
file (no vi.mock exists there to justify it); the three call sites now
use the static import.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* docs: fix skill reference drift
* docs: complete guide tool coverage and graph schema (#2356 items 5-6)
- add the 6 undocumented MCP tools to the guide's Tools Reference
(route_map, shape_check, api_impact, tool_map, group_list, group_sync)
- document the experimental @groupName cross-repo trace mode
- expand the Graph Schema section to the real node/edge type surface,
pointing at gitnexus://repo/{name}/schema as the authoritative list
- sync the packaged gitnexus/skills copy
Item 7 of #2356 (Codex host naming / duplicated filename) does not
reproduce on current main - no remaining copy contains it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* docs(readme): restructure for readability, fix stale facts
Reorganize the README so the visible page reads as a short narrative
(Quick Start -> Two Ways -> Why -> What Your Agent Gets -> Editor Setup
-> CLI -> How It Works -> Docker -> Enterprise) and move deep
operational detail into 13 collapsible <details> sections: env vars,
.gitnexusrc, Cosign/Kubernetes verification, manual MCP configs,
install troubleshooting, and extended tool examples.
Accuracy fixes verified against gitnexus/src:
- MCP tools: 17 (15 per-repo + 2 group), not 16/11+5; drop
group_contracts/group_query/group_status (CLI + resources now, not
tools); add check, trace, explain, pdg_query, route_map, tool_map,
shape_check, api_impact rows from src/mcp/tools.ts
- Agent skills: 6 installed (adds Guide + CLI), not 4
- Wiki default model: minimax/minimax-m2.5, not gpt-4o-mini
- Resources: add gitnexus://setup and gitnexus://group/{name}/...
- CLI: document group impact, doctor, and the direct terminal query
commands (query/context/impact/trace/cypher/detect-changes/check);
note the optional branch param on per-repo tools (#2106)
Structural cleanups: dedupe the two Codex config blocks, move Community
Integrations out of the MCP setup flow, move Star History to the
bottom. No content deleted - verbose material is collapsed, not cut.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: fact-check and fix the remaining READMEs
Reviewed all 9 tracked non-root READMEs against the source; fixed the
four with stale facts, left the already-accurate ones untouched
(pr-swarm-review, .claude reviewer-swarm adapter, and the three
bench/ methodology docs).
gitnexus/README.md (npm package page):
- MCP tools table: 17 tools (15 per-repo + 2 group), was 7 rows
- Resources: add gitnexus://setup and gitnexus://group/{name}/...
- Skills: 6 bundled (adds Guide + CLI) plus --skills generated ones
- Languages: add Dart (14 total) to the list and feature matrix
- Wiki default model: minimax/minimax-m2.5, not gpt-4o-mini
- Requirements: Node >= 22 (package.json engines), not >= 18
- Claude Code hooks: PreToolUse + PostToolUse
- CLI: add --skills/--skip-skills/--skip-git/--workers, doctor,
trace, check, group impact
- Optional grammars note: include Proto, mention
GITNEXUS_SKIP_OPTIONAL_GRAMMARS=1
gitnexus-cursor-integration/README.md:
- 17 MCP tools, was 16; skills list: all 9 bundled skills, was 5
eval/README.md:
- Model list matches configs/models/: Claude Haiku 4.5 (was
'3.5 Haiku'), adds MiniMax M2.5 and DeepSeek
- Node.js 22+ for GitNexus, was 18+
.devcontainer/README.md:
- Add a table of contents (364 lines, ~15 sections, no navigation)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: empty commit to retrigger CI
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat: add Spring DI resolver for @Autowired List<T> injection
Addresses all P0/P1 findings from tri-review (#2200):
- P0: Register INJECTS in RelationshipType union (compiles)
- P0: Rewrite execute() to emit consumer→implementation edges from graph data only
- P1: Register in VALID_RELATION_TYPES, single-pass O(N) indexes
- P1: Java-only gate with early exit on non-Java repos
- P1: Update FULL_ORDER golden test
- 8 unit tests covering all edge cases
* test: make VALID_RELATION_TYPES size assertion array-driven (no hardcoded count)
The security test hardcoded toBe(16) for the relation type count, but PR #2200
added INJECTS, bumping it to 17. Replace the magic number with an
EXPECTED_RELATION_TYPES array whose .length drives the size assertion,
so future additions only need to append to the list.
Fixes CI failure on PR #2200.
* fix(ingestion): thread raw generic field types onto Property nodes so Spring DI matching works (review 4616076037 P0)
Production declaredType is generics-stripped by design (extractSimpleTypeName:
List<Shape> -> "List"), so the spring-di phase's anchored regexes could never
match real extraction output — the phase was a silent no-op on every real Java
repository, while its unit tests passed against hand-built node shapes.
Add FieldInfo.rawDeclaredType captured verbatim from the field's type node
(.text, generics and qualifiers preserved — same precedent as the JVM method
extractor), thread it through both parse-worker Property sites, add it to the
shared NodeProperties contract, and match on rawDeclaredType ONLY (no
declaredType fallback: it can never match real data and would mask future
plumbing regressions as quiet no-ops).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ingestion): gate Spring DI on real injection annotations, honest edge reason (review 4616076037 P1)
Extract Java field annotations (shared extractAnnotations helper, moved
verbatim from the method extractor) onto Property nodes and require
@Autowired or @Inject before a collection field becomes an INJECTS
candidate. Previously every edge's reason string fabricated "@Autowired"
without any annotation ever being checked, and any plain collection field
would have fanned out false edges once matching worked.
@Resource is deliberately excluded: JSR-250 resolves by bean name first
(defaulting to the field name), injecting a single named collection bean —
the opposite of the collect-all-implementers fan-out INJECTS models. Pinned
by a test.
An annotated candidate missing rawDeclaredType now logs an isDev warning
(plumbing-contract breach signal) instead of vanishing silently.
SCHEMA_BUMP 9 -> 10: Property nodes gained rawDeclaredType + annotations;
warm parse caches must invalidate or the DI phase silently no-ops on
replayed pre-upgrade nodes (the #2038 trap).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(ingestion): framework-neutral di phase + language-scoped Spring matcher registry (review 4616076037 P1)
spring-di was the only pipeline phase naming a language in shared
core/ingestion code (DoD.md language rule; the maintainer's direction is a
generic DI solution). Split it:
- di-extractors/spring.ts: the Spring matcher (annotation gate, collection
type parse, @Resource exclusion rationale, framework-specific reason
payload) — language-scoped home, mirroring route-extractors/.
- di-extractors/index.ts: DI_MATCHERS, a single-valued
ReadonlyMap<SupportedLanguages, DiFieldMatcher> mirroring the
SCOPE_RESOLVERS registry shape sanctioned by AGENTS.md. Constructor
injection deliberately out of scope; widen to arrays only when a second
same-language framework lands.
- pipeline-phases/di.ts (renamed from spring-di.ts): framework-neutral —
routes Property nodes to registered matchers by node language via a typed
guard, then runs the unchanged reverse-index fan-out. Zero language or
framework names remain (grep-verified).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ingestion): language- and qualified-name-scoped interface resolution for DI fan-out (review 4616076037 P2)
The interface index was built from ALL Interface nodes regardless of
language, keyed by bare simple name with last-writer-wins overwrite —
a polyglot repo with a TS and a Java 'Shape' could fan Java INJECTS edges
into TypeScript classes, and two same-named Java interfaces in different
packages silently collapsed to whichever parsed last (documented GitNexus
bug class: #2054, PR #1956).
Resolution is now per-language with qualifiedName as the primary key
(Interface nodes already carry package-qualified qualifiedName); dotted
element types resolve via qualifiedName, bare names via a per-language
simple-name index that records ambiguity and fails CLOSED. Ambiguity skips
are observable: DIOutput.ambiguousSkipped + an aggregated isDev debug log,
so 'no DI fields' is distinguishable from 'all candidates ambiguous'.
Same-package tiebreaking is a pinned, documented follow-up.
Order-independence pinned by running collision tests in both insertion
orders.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(ingestion): depth-aware Spring collection-type parser for idiomatic generics (review 4616076037 P3)
The two anchored regexes silently skipped idiomatic Spring shapes:
Map<Pair<A,B>, IFoo> (nested-generic key broke the [^,]+ split),
List<? extends IFoo> / List<? super IFoo> (bounded wildcards),
java.util.List<IFoo> (qualified wrapper), and whitespace/multi-line
declarations.
Replace them with a small scanner: whitespace normalization, wrapper
matched by last dotted segment, depth-aware top-level-comma split, wildcard
bound stripping, and a final plain-dotted-type-name gate so anything else
(nested-generic elements, arrays, unbounded wildcards, embedded comments,
unbalanced brackets) fails closed. Every accept and reject is documented in
the module docstring and pinned by 27 table-driven cases.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(integration): prove Spring DI end-to-end through the real pipeline (review 4616076037 P1)
Both no-op incarnations of this feature shipped with a green unit suite
because every test hand-built the exact graph shape the phase expected —
no test ever ran real Java source through the actual extraction pipeline.
Add test/integration/spring-di-pipeline.test.ts: real .java fixtures via
runPipelineFromRepo, pinning (a) the extraction contract on the annotated
field's Property node (declaredType 'List', rawDeclaredType 'List<IFoo>',
annotations ['@Autowired']), (b) set-equality on ALL INJECTS edges
(exactly Consumer->FooA and Consumer->FooB; the non-annotated 'plain'
field of the same type contributes nothing; no self-edges), and (c) a
negative-control fixture with no injection annotations producing zero
INJECTS edges. Either historical regression fails at least one of these.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(incremental): register INJECTS across product surfaces + delete-before-writeback (review 4616076037 P2)
INJECTS was allowlisted in VALID_RELATION_TYPES but invisible or unhandled
everywhere else. Register it deliberately:
- REL_TYPES (gitnexus-shared schema-constants): web-side validRelType()
otherwise silently rejects INJECTS filters (CLI/web single source of truth).
- mcp/tools.ts cypher edge list (agent-facing schema discovery).
- isGraphWideRelType: INJECTS validity is a whole-program property — a
change to a THIRD file (the interface, or a new/removed implementer)
creates/invalidates edges between two untouched files (the TAINT_PATH /
#2084 M4 U6 class), so incremental extraction must always re-include the
full fresh set.
- deleteAllInjects (lbug-adapter): mirrors deleteAllInterprocTaintPaths —
COUNT-then-DELETE under withConnLock, benign missing-table carve-out,
re-throw otherwise (CodeRelation has no PK and there is no read-side
dedup; a fail-soft delete + re-add would silently duplicate rows).
- run-analyze.ts: the delete is UNCONDITIONAL, next to the Communities
delete — deliberately NOT inside the options.pdg block: the di phase runs
on every persisting analyze while the graph-wide re-include is
unconditional, so a pdg-gated delete would append without deleting on
every non-pdg incremental run (N runs = N copies).
- local-backend.ts comment: opt-in traversal by design (not in default
impact()/context() lists; no IMPACT_RELATION_CONFIDENCE entry per the
WRAPS/FETCHES precedent — edges carry their own 0.8).
- ARCHITECTURE.md: 14 -> 15 phases, DAG diagram, phase table, skip-list.
Note: the tools.ts edge list also predates WRAPS/QUERIES/USES — that drift
is pre-existing and left for a follow-up.
Idempotency pinned end-to-end: two successive incremental runs (real
runFullAnalysis + real LadybugDB, unrelated-file touches) leave the INJECTS
row count stable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: describe INJECTS' actual precondition; drop stale fixed-at-16 comments (review 4616076037 P3)
The shared-schema doc for INJECTS claimed an @Autowired precondition the
code (pre-fix) never checked, and hardwired Spring semantics into what is
now a framework-neutral edge type. Reword: precondition is an injection
annotation recognized by a per-language matcher in di-extractors/;
framework specifics live in the reason payload, not the type contract.
security.test.ts comments still said the allow-list size 'stays fixed at
16' (it is 17 and the assertion derives from EXPECTED_RELATION_TYPES).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor: simplify DI surfaces — narrow matcher contract, dedup delete-alls, derive tools edge list
Post-implementation simplification pass (4 review angles):
- DiFieldMatch/CandidateField carried collectionType + matchedAnnotation
that no consumer read (the matcher bakes both into reason) — narrowed to
{elementTypeName, reason}.
- parseElementTypeName had two guard branches fully subsumed by the final
plain-dotted-type-name gate — deleted, rationale folded into the regex
comment.
- The three byte-identical delete-all-by-rel-type functions in lbug-adapter
(TAINT_PATH / CALL_SUMMARY / INJECTS) are now one parameterized helper +
thin wrappers with identical names, signatures, and message text
(character-diff verified) — the missing-table regex and abort policy now
live in exactly one place.
- The cypher tool's hand-maintained edge-type list (already missing
WRAPS/QUERIES/USES) is now derived from the canonical REL_TYPES — the
drift class is gone rather than patched.
- di phase: interface indexes are built only for languages that actually
have candidates; test builder gained a rawDeclaredType opt-out replacing
a hand-rolled node.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: apply Tier-2 review findings — qualified-name fail-closed, honest cypher docs, pinned delete contract, hook isolation
- byQualifiedName was last-writer-wins on duplicate qualified names
(reproduced: order-dependent INJECTS edges with ambiguousSkipped 0 —
same package+interface duplicated across monorepo modules/source roots;
Java qualifiedName has no file-path component). Both indexes now share
the AMBIGUOUS fail-closed sentinel; order-flip test added.
- The REL_TYPES-derived cypher edge list advertised pdg-gated types with
no caveat (LLM queries on them silently return zero rows on default
indexes) — caveat appended, INJECTS example added, impact relationTypes
description now names the DI fan-out opt-in.
- The delete-all re-throw contract (only defense against duplicate
CodeRelation rows) was untested — error classification extracted to a
pure classifyDeleteAllError and pinned exhaustively.
- extractRawType/extractAnnotations hooks lacked the per-hook try/catch
the pipeline applies elsewhere (#2286 pattern): a throwing hook would
silently drop every remaining file in the language group. Hardened,
degradation tested.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix: resolve Java cast-wrapped and this.method() call edges
Two fixes for missing call edges in Java method resolution:
1. compound-receiver.ts — cast expression handling:
- Strip (Type) cast wrappers from receiver text, tracking the
outermost meaningful cast type
- Resolve directly to the cast type class (not the field's
declared type), since the cast narrows the receiver type
- Add this.field chain walker for field-access receivers
- Replace text → workingText throughout the function body
2. scope-resolver.ts:
- Enable resolveThisViaEnclosingClass: true for Java
(activates Case 0.5 in receiver-bound-calls.ts)
Verified on a large-scale Java codebase with no regressions.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(scope-resolution): format compound-receiver.ts with prettier (#2353 review F10)
Mechanical prettier --write from repo root — 6 brace-expansion sites and one
ternary re-join, zero logic changes. Clears the quality/format CI failure
that was blocking CI Gate on PR #2353.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test(scope-resolution): pin working Java cast-receiver shapes (#2353 review F3)
Fixture-backs the cast resolutions PR #2353 gets right — simple cast,
nested/CFR cast, cast over this.field, and the deliberate declared-type
fallback for a resolvable-shape cast to an unindexed type — each with a
same-named decoy method on the receiver's declared type so later refactors
cannot silently regress them. No resolver changes; tests are green as-is.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(scope-resolution): resolve nothing for unparseable cast types (#2353 review F1)
A receiver paren-group that is type-shaped but unparseable — generic
(List<String>), array (Foo[]), fully-qualified (com.example.Foo) — is a
cast whose type cannot be looked up. Stripping it and falling through
resolved the pre-cast expression's own declared type, emitting a
confident wrong CALLS edge. Classification is now three-way per peel:
simple identifier → capture (outermost wins), type-shaped-unparseable →
resolve nothing (pre-#2353 behavior; noise casts after a captured type
still win), anything else → not a cast, text left untouched. Cast
candidates require a non-empty trailing expression, so plain
parenthesized receivers never capture a cast type.
Red-first: all four shapes reproduced the wrong edge before the fix;
golden digest byte-stable after.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor(scope-resolution): delete duplicate this.field walker, seed literal-this chain heads (#2353 review F4/F5/F7)
A/B against the fixture corpus confirmed the generic per-segment walker
(head resolved via the synthesized this typeBinding) already covers every
method-body this.field chain — only initializer contexts (instance
initializer block, field initializer) were walker-dependent, since no
function scope exists there to carry a this binding. Deleting the
duplicate walker removes the naive chainRest.split('.') (F5) and the
widened fieldFallback use (F7) with it; the findEnclosingClassDef head
seed is the deliberate residue covering initializer contexts —
head-resolution only, the per-segment walk stays the single shared
implementation. Post-seed edge set is byte-identical to pre-deletion.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(scope-resolution): gate cast stripping behind opt-in stripReceiverCastExpressions (#2353 review F2)
Cast handling in resolveCompoundReceiverClass now runs only for
languages that opt in via the new ScopeResolver toggle (default off);
Java is the sole opt-in. The peel loop is extracted into the pure,
exported stripCastWrappers helper (placed with the file's other pure
string helpers) so it can be unit-tested directly. Non-opting languages
see receiver text untouched — pre-#2353 behavior by construction
(golden digest unchanged, TS/C++/C# suites green, 796/796). Shared-code
comments are language-neutral per AGENTS.md; the contract JSDoc carries
the classifier grammar, the second-language escalation rule, and the
Case 3b/Case 4 pass-through non-goal.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(scope-resolution): cap cast-peel iterations in stripCastWrappers (#2353 review F8)
MAX_CAST_PEEL = 16 (each cast level costs at most two peels, so this
covers 8-level nesting with headroom — real cast nesting, including
decompiler output, is a handful of levels). Each peel rescans the
working text for its matching close paren, so pathological nested-paren
input was O(N²); the cap bounds it at O(N·16). Exceeding the cap bails
all-or-nothing with the original text (not-a-cast outcome). Adds the
helper's first unit tests: 14 scenarios covering capture, unparseable
shapes, redundant-paren unwrap, captured-type precedence, rawName
no-op, over/under-cap, and unbalanced-paren termination.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(scope-resolution): revert Java resolveThisViaEnclosingClass, pin Case 4 bare-this dispatch (#2353 review F6/F9)
Remove resolveThisViaEnclosingClass from the Java scope resolver: the
toggle's own contract doc prescribes keeping it disabled where Case 4
(the synthesized this typeBinding) already handles this, and Case 0.5's
C++-authored semantics (hiddenByName arity-hiding, method-before-field)
provably bypass the interface-dispatch fan-out only Case 4 emits.
A/B gate (new java-this-dispatch pinning fixtures): flag-off 7/7 green;
flag-on 2/7 red (hiddenByName drops the this.greet overload site —
masked by a free-call-fallback 'local-call' edge — and the
interface-dispatch fan-out is missing). Corpus A/B over all 54 java-*
fixtures: 2 fixtures differ — java-this-dispatch (reason
'local-call'→'global' on the bare-this overload site; +2
interface-dispatch fan-out edges flag-off) and java-this-field-chain
(2 initializer-context bare-this ACCESSES reads emitted only by Case
0.5, which Case 4 cannot resolve — no synthesized this binding without
a Function scope; the corresponding CALLS edges are unaffected via the
F4 commit's literal-this head seed).
Also (F9): insert Case 0.5 into the I4 case-order listings (contract +
receiver-bound-calls header, now 8-case, marked gated) so the next flag
flip is visible at review time; the two 'sole C++ language' comments
are accurate again unedited.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(scope-resolution): restrict literal-this head seed to initializer contexts (#2353 review follow-up)
Final-review finding (two independent reviewer angles): the literal-this
chain-head seed landed ungated in shared code, so any language's
this-headed chain in a scope without a synthesized this typeBinding —
including contexts where the language DELIBERATELY leaves this unbound
(object-literal methods, nested plain functions) — would seed from the
lexically enclosing class. isInitializerContext now permits the seed
only when no Function scope sits between the site and its class, which
is precisely the field-initializer / instance-initializer shape the
seed exists for. Adds a TS guard fixture pinning that an
object-literal method's this.field.method() chain emits no fabricated
edge (mechanism did not empirically reproduce even ungated — the
restriction is conservative hardening, and the pin keeps it that way).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(scope-resolution): attach stripCastWrappers JSDoc, fast-path non-paren receivers (#2353 review nits)
Two final-review nits: a blank line detached the helper's 30-line
classification-contract JSDoc from the declaration (IDE hover showed
nothing at call sites); and the gate now skips the helper call plus
result allocation for the majority of receivers that cannot be casts
because they do not start with '(' — the helper's own check stays as
the safety net.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* bench(scope-capture): rebaseline Java fingerprint for new #2357 fixtures
The scope-capture correctness fingerprint hashes captures over the
java-* fixture corpus; the three fixture dirs added by this PR
(java-cast-receiver, java-this-field-chain, java-this-dispatch) extend
that corpus, so the fingerprint moves. Verified purely additive: with
the three new dirs parked, the fingerprint reproduces the prior
baseline byte-identically — no emit/capture behavior changed.
--check now passes for all 14 languages.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: ww <ww@wwdeMacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* docs: add Kilo Code MCP workflow
* Fixed Space
Fixed 1 deleted blank line
* Fixed Readme Link
Fixed guide link from docs to documentation folder fixing 404 error
* Added Image and linked to .md file
* Removed Trailing Spaces
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* fix: consolidate icon imports, fix stale refs and package name collision
three components were importing directly from lucide-react instead of
going through the centralized @/lib/lucide-icons module like the rest
of the codebase. added the missing Keyboard and BarChart2 exports to
the icons module and updated the imports.
also:
- removed duplicate mermaid init comment in ProcessFlowModal
- replaced placeholder issue #XXX with a descriptive note in git.ts
- updated stale KuzuDB reference to LadybugDB in ARCHITECTURE.md
- renamed gitnexus-web package.json name from "gitnexus" to "gitnexus-web"
to avoid collision with the CLI package
* fix: complete package rename in lockfile and cite #2054 in git.ts comment
Address tri-review findings: package-lock.json name fields (top-level and
packages[""]) now match the renamed gitnexus-web package, and the
getCanonicalRemote doc comment cites #2054 instead of dropping the issue
reference.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(autofix): apply prettier + eslint fixes via /autofix command
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(embeddings): use system-matched onnxruntime-node CUDA build so CUDA 13 hosts use the GPU
transformers.js exact-pins a CUDA-12 onnxruntime-node while gitnexus' own dep floats to a CUDA-13 build. npm/pnpm cannot dedupe an exact pin against a range, so npm i -g installs two copies and the gitnexus overrides block (root-only) is inert. On a CUDA-13-only host the nested CUDA-12 provider cannot load libcublasLt.so.12, the CUDA EP fails, and embeddings silently fall back to CPU (isCudaAvailable() also only probed .so.12).
Add onnxruntime-node-resolver.ts (module.registerHooks redirect to the host-matching build, no-op elsewhere) mirroring onnxruntime-common-resolver.ts; probe libcublasLt .so.12 OR .so.13 against the copy that actually loads; unit test with 12 cases.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(embeddings): wire CUDA-13 build-match resolver into MCP query embedder
The MCP query-time embedder (src/mcp/core/embedder.ts) has its own,
separate initEmbedder() used for semantic search — it only called
ensureOnnxRuntimeCommonResolvable() before importing transformers.js, so
the CUDA-13 build-matching redirect added for the analyze/CLI embedder
never applied here. A CUDA-13 host running MCP search with
--embedding-device cuda still loaded the mismatched default onnxruntime-node
build.
Wire ensureOnnxRuntimeNodeMatchesSystem() into the same call site, mirroring
the core embedder's ordering (registered after the common-resolver fallback,
before the dynamic transformers import).
* fix(embeddings): gate CUDA redirect decision on registerHooks availability
decide() computed the CUDA-major redirect independent of whether Node's
module.registerHooks API actually exists — only ensureOnnxRuntimeNodeMatchesSystem()
checked that. On Node 22.0-22.14 (allowed by this package's engines
floor; registerHooks needs >=22.15), isCudaAvailable() could therefore
report a redirect target that ensureOnnxRuntimeNodeMatchesSystem() then
silently failed to install, so transformers.js loaded the mismatched
default onnxruntime-node build while the embedder still requested
device:'cuda' against it — reintroducing the uncatchable native crash
this probe exists to prevent.
Move the registerHooks check to the top of decide() so the probe and the
loader can never disagree, and skip CUDA-major probing entirely in that
case (a redirect could never install anyway).
Also fixes a related test-helper bug found while writing this unit's
tests: loadResolver's destructuring default (`registerHooks = vi.fn()`)
silently substituted a real mock function even when a test passed
`registerHooks: undefined` to simulate Node < 22.15 — meaning the
existing 'no-ops... when registerHooks is unavailable' test never
actually exercised that path. Distinguish 'omitted' from 'explicitly
undefined' via an 'in' check.
* test(embeddings): drive decide() -> redirect:true and assert the resolve() closure
The PR's actual shipped behavior — the installed registerHooks resolve()
closure, and the full redirect-active decision path — had zero executed
test coverage. All 4 prior ensureOnnxRuntimeNodeMatchesSystem tests avoided
driving decide() into redirect:true because require.resolve/createRequire
were never mocked, so the two-distinct-directory comparison decide()
depends on always resolved against whatever's actually installed in this
test's real node_modules (a single real copy, not the PR's two-copy
scenario).
Extend loadResolver()'s existing node:module mock to also fake createRequire,
keyed by call origin, so resolveOurOrtNodeDir/resolveDefaultOrtNodeDir can be
driven to two distinct fake directories with distinct CUDA majors — reaching
redirect:true without adding any injection points to production code. Then
capture the installed resolve() closure (mirroring the sibling
onnxruntime-common-resolver.test.ts's captureResolve() pattern) and assert
its three branches directly: onnxruntime-node redirect, onnxruntime-common
redirect, and passthrough for any other specifier.
* fix(embeddings): distinguish ldd detection-failure from no-CUDA-provider
ortCudaMajor treated any execFileSync('ldd', ...) failure with no usable
stdout (missing ldd binary, permission-denied .so, sandboxed exec)
identically to 'CUDA provider genuinely absent'. The pre-PR detection
(hasOrtCudaProvider) only used existsSync, never ldd, so this is a
regression: a CUDA-12 host that worked fine before this PR can now
silently fall back to CPU if ldd itself can't run, even though the
provider .so and system CUDA libs are both genuinely present.
readSoNeeded now reports whether ldd produced any usable output at all,
distinct from 'ldd ran and just found no matching NEEDED entry' (the
existing, already-handled '=> not found' case). When detection genuinely
fails, log a warning so an operator can tell 'CPU fallback because
detection itself failed' apart from 'CPU fallback because no CUDA build
shipped' — the return value stays null either way (the type can't
distinguish a third state), but the two cases are now observably
different via the log.
* fix(embeddings): check ourDir independently of whether defaultDir resolved
decide()'s ourDir fallback lookup was nested inside
'if (systemMajor != null && defaultDir)', so a null defaultDir (transformers'
own onnxruntime-node resolution failing outright, e.g. a partial/broken
install) skipped checking ourDir entirely — getEffectiveOnnxRuntimeNodeDir()
returned null even when gitnexus' own matching CUDA-13 copy would have
resolved fine and worked.
defaultDir resolving is not a precondition for the comparison: an
unresolvable default already counts as 'the default doesn't match', so the
ourDir check now runs whenever systemMajor is known, regardless of whether
defaultDir resolved.
* fix(embeddings): prefer CUDA 13 globally across the env-var directory scan
detectSystemCudaMajor's CUDA_PATH/LD_LIBRARY_PATH scan returned on the
first CUDA-major match within a single dir/sub pair, so a stale .so.12
found early (e.g. a leftover CUDA_PATH entry from a prior install) shadowed
a genuine .so.13 found later in the search path, even though the scan's
own ordering (checking 13 before 12 within each pair) was clearly intended
to prefer 13 wherever possible.
Keep scanning the full search space once a 12 is found, only returning
early once a 13 is found (the best possible answer) or the space is
exhausted.
* fix(embeddings): have onnxruntime-common-resolver defer to the effective onnxruntime-node dir
onnxruntime-common-resolver.ts independently re-derived transformers'
default onnxruntime-node dir (its own copy of the 'resolve transformers'
main entry, then onnxruntime-node' walk) to compute which onnxruntime-common
to pair with — duplicating onnxruntime-node-resolver.ts's own walk, and
capable of disagreeing with it: when the CUDA-major redirect is active,
this hook would still pair onnxruntime-common with transformers' default
(unredirected) onnxruntime-node, not the redirected copy the other hook
just switched onnxruntime-node itself to.
Have it call the already-exported getEffectiveOnnxRuntimeNodeDir() instead
— the same decision the CUDA-major redirect hook uses — so both hooks
always agree on which onnxruntime-node they're pairing onnxruntime-common
against, and the duplicated resolve-walk is removed entirely rather than
merely factored out.
* fix(embeddings): cache the effective CUDA major to remove redundant subprocess spawns
isCudaAvailable() in embedder.ts re-invoked ortCudaMajor/detectSystemCudaMajor
directly even though decide() (via getEffectiveOnnxRuntimeNodeDir) had
already computed both to make its redirect decision — a second, wasted
ldconfig + up to 2 ldd spawns on every initEmbedder() call.
Add effectiveMajor to the memoized Decision, computed once inside decide()
alongside effectiveDir/systemMajor, and export a single
isEffectiveCudaAvailable() that reads straight from the cached decision.
embedder.ts's local isCudaAvailable() wrapper (and its now-unused
getEffectiveOnnxRuntimeNodeDir/ortCudaMajor/detectSystemCudaMajor imports)
is replaced by this one exported function.
* fix(embeddings): surface CUDA redirect state at info level and in doctor
A successful CUDA-build redirect logged only at logger.debug (filtered
by the default 'info' level), and gitnexus doctor's embeddings section
never mentioned the redirect at all — leaving no diagnostic path for
'why is my CUDA-13 host still on CPU' after this PR ships.
Log the successful-redirect line at info (no-redirect/failure paths stay
at debug, since those are the common, expected case). Add
cudaRedirectDoctorStatus(), a pure summary of decide()'s already-computed
decision mirroring doctor.ts's existing localEmbeddingDoctorStatus shape,
and print it as a new literal (non-i18n) 'CUDA:' line in doctor's
embeddings section alongside the existing 'Support:' line, matching that
line's established convention.
* test(embeddings): register onnxruntime-node-resolver.test.ts in the cross-platform subset
The new test file guards on process.platform (linux/darwin cases) but was
absent from cross-platform-tests.ts's PLATFORM_LOGIC list, which
TESTING.md says platform-sensitive tests should be added to — so it never
ran on the Windows/macOS CI matrix, only Ubuntu.
Note: the sibling onnxruntime-common-resolver.test.ts has the identical,
pre-existing gap (it predates this PR) — left as-is here, since fixing
unrelated pre-existing test-registration debt is out of scope for this
PR's own follow-up fixes.
* test(embeddings): strengthen weak assertions, add garbled-output and CUDA_PATH coverage
Three of the four ensureOnnxRuntimeNodeMatchesSystem tests only asserted
'doesn't throw' rather than a concrete outcome — including one literally
named 'idempotent' that never asserted a call count on its own spy.
Strengthen each to assert real outcomes (module stays functional after a
no-op; spy call counts; return-value shape), while keeping the true
install-once idempotency proof in the redirect-active test added earlier
(this file's no-redirect scenario can't exercise it, since registerHooks
is never called either way).
Add the missing edge cases flagged in review: a CUDA_PATH-only fallback
scan test (mirroring the existing LD_LIBRARY_PATH one), and garbled/
unrecognized ldconfig and ldd output cases for both detectSystemCudaMajor
and ortCudaMajor, confirming neither falsely matches a CUDA major on
unparseable input. Also parameterize the non-linux platform test across
both darwin and win32 rather than darwin alone.
Not changed: the process.env reassignment vs. Object.defineProperty
'inconsistency' flagged in review — process.env, unlike process.platform,
has no getter-only restriction, so plain reassignment is already correct
and switching it to Object.defineProperty would be unnecessary ceremony.
* docs(embeddings): note the npm link/symlinked dev-checkout resolution caveat
resolveOurOrtNodeDir/resolveDefaultOrtNodeDir anchor to this module's own
real (post-symlink) location via import.meta.url, so a linked local dev
checkout may resolve against its own node_modules rather than the
consuming app's. Narrow, dev-only blast radius (regular npm/pnpm installs
are unaffected) — document-only, no structural fix warranted.
* fix(test): point the windowsHide spawn-family registry at the file that actually spawns
hooks.test.ts's windowsHide regression check still listed
gitnexus/src/core/embeddings/embedder.ts as a child_process-spawning
file, but this PR itself already moved all execFileSync usage out of
embedder.ts and into the new onnxruntime-node-resolver.ts — without
updating this registry. The check was silently failing at the PR's own
head commit (confirmed: 0 spawn-family calls found in embedder.ts,
'expected 0 to be greater than 0'), a pre-existing gap this fix-pass
surfaced via a full-suite run rather than something introduced by any of
the preceding follow-up commits.
Swap the registry entry to onnxruntime-node-resolver.ts, which does
import execFileSync (ldd + ldconfig, both already correctly passing
windowsHide: true).
* fix(test): make onnxruntime-node-resolver.test.ts path comparisons OS-agnostic
Registering this file in cross-platform-tests.ts's PLATFORM_LOGIC (a
prior commit in this series) means it now runs on the Windows CI matrix,
not just Ubuntu — and several of the fakeDirs-based tests (redirect:true,
ourDir-independent, subprocess-count, doctor-status) compared the
resolver's real join()/dirname() output against hardcoded forward-slash
fixture strings via exact-match or .startsWith().
Node's module is bound to path.win32 (or path.posix) based on the
REAL host OS at process start — stubbing process.platform later, as these
tests already do for the resolver's own platform branching, has no effect
on it. So on a genuine Windows runner, join(effectiveDir, 'package.json')
backslash-normalizes even under a faked platform:'linux', silently
breaking every forward-slash comparison in this file: the createRequire
dispatch would route to the wrong fake require, throw MODULE_NOT_FOUND,
get swallowed by ensureOnnxRuntimeNodeMatchesSystem's outer try/catch, and
registerHooks would never fire — the redirect-active tests would fail
outright on Windows CI.
Normalize every comparison point (the createRequire dispatcher, and the
shared execFileSync/existsSync mocks) with a single toPosix() helper.
Added a forceWin32Path test option (using path.win32's real join/dirname
behavior) to prove this holds without needing an actual Windows runner —
confirmed by temporarily reverting the fix and observing the new test
fail with the exact predicted mismatch before restoring it.
* chore(autofix): apply prettier + eslint fixes via /autofix command
* fix(embeddings): keep CUDA auto-detect working on Node < 22.15 when the default build already matches
The registerHooks guard in decide() returned effectiveMajor: null
unconditionally, so on Node 22.0-22.14 / 23.0-23.4 (engines floor is
>=22.0.0) isEffectiveCudaAvailable() was always false and a CUDA-12 host
whose default onnxruntime-node build already matched — which needs no
hook at all to use the GPU — silently regressed from CUDA to CPU on the
auto device path (pre-PR isCudaAvailable() behavior).
Probe the system and the default copy regardless of registerHooks
availability; only the ourDir redirect branch stays gated on it, so the
probe still never reports a redirect target that cannot be installed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix: add --limit i18n, negative guard, correct property paths, and zh-CN translations
- Add i18n keys for context/impact/cypher/detect-changes --limit options
- Add zh-CN translations for all 4 --limit option descriptions
- Add Math.max(0, parseInt()) guard to prevent negative --limit
- Fix ALL property path mismatches discovered by audit:
- context: callers/callees → incoming.calls/outgoing.calls+accesses
- impact: upstream/downstream → affected_processes/affected_modules/byDepth
- cypher: rows → row_count cap (rows embedded in markdown string)
- detect-changes: affected_flows → affected_processes
- Change query command from required to optional positional arg with -q alias
- Update @ladybugdb/core from ^0.16.1 to ^0.17.1
- Update typescript from ^5.4.5 to ^5.9.3
* test: add E2E tests for --limit flag across all 5 CLI commands
Tests context, impact, cypher, detect-changes, and query with
--limit 1, baseline comparison, and --limit 0 (falsy/no-op).
detect-changes output is formatted text (not JSON), so those
tests count symbol lines matching 'Type name -> filePath' pattern.
14 tests, all passing. No regressions in 6455 existing tests.
* fix: address Copilot review feedback on --limit guards
- Add Math.max(0, ...) guard to queryCommand limit parsing
- Change if(limit) to if(limit !== undefined) in all 5 commands
(prevents --limit 0 from being treated as falsy/no-op)
- Make queryText parameter optional (Commander may pass undefined)
- Fix usage error strings: --search to -q, --query (en + zh-CN)
* chore(autofix): apply prettier + eslint fixes via /autofix command
* fix(cli): centralize --limit parsing, slice cypher markdown, fix usage text
Address PR review feedback on --limit handling:
- Add a shared parseLimit() helper (Number.isInteger(n) && n > 0), used by all
5 tool commands. Non-numeric / 0 / negative --limit now means "no limit"
instead of the `options.limit ? Math.max(0, parseInt(...)) : undefined` path,
where a string like "abc" is truthy and yields NaN -> slice(0, NaN) -> the
guardrail commands (impact/context/detect-changes) silently emptied results
with exit 0.
- cypher: slice the markdown table to --limit data rows so the reported
row_count matches what is actually printed (was capping row_count while
printing every row).
- Fix query usage string: [search_query] (optional positional) and
`--query <text>` invocation form, not the option-definition
`-q, --query <search_query>` syntax (en + zh-CN).
- Add an E2E regression test for non-numeric --limit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cli): escape newlines in cypher markdown cells
A multi-line cell value (e.g. a symbol's `content`) was rendered with raw
newlines via String(v), so one logical row spanned multiple physical lines.
That corrupts the markdown table and breaks `cypher --limit`'s line-based
slice (it kept the wrong number of rows, often zero, while row_count
over-claimed). Collapse newlines in formatCypherAsMarkdown so one physical
line == one row; the existing CLI slice is now correct and the pre-existing
un---limited corruption is fixed too. (#2310 review)
* test(cli): de-vacuum the --limit truncation tests
The truncation it()s used the repo-banned vacuous-pass pattern (early-return
on status===null, assertions guarded by if(Array.isArray), bounds-only
toBeLessThanOrEqual — DoD.md:82) against `validateInput`, which has only 1
caller, so context/impact/query --limit 1 compared 1>=1 and stayed green even
if the slice were deleted. Rewrite with unconditional, exact assertions and
target `logMessage` (2 callers, 4 processes) so the no-limit baseline truly
exceeds the limit; detect-changes now mutates two real function bodies (two
changed symbols). Adds a multi-line-cell cypher --limit regression. (#2310)
* test(ci): run cli-limit-e2e in the cross-platform matrix
The --limit E2E suite spawns the real CLI (child_process) but was not in
SPAWN_CLI, so it ran only on Ubuntu — the cross-platform check only fails on
listed-but-missing files, not the reverse (TESTING.md §Cross-platform). Register
it so the --limit regression guard also runs on Windows/macOS, where path
separators, CRLF and the formatted-output arrow differ. (#2310)
* fix(cli): document impact --limit affected-list cap, drop dead byDepth re-slice
`impact --limit` also caps affected_processes/modules, but the help only
mentioned the per-depth cap — so JSON consumers reading the affected lists got
a silently-truncated array. Update en + zh-CN + the command description to say
so. Also remove the client-side byDepth re-slice: the backend already
paginates byDepth to the same limit (paginationLimit = clamp(limit,1,10000),
offset applied backend-side), so the client slice was a guaranteed no-op. (#2310)
* fix(cli): reconcile detect-changes --limit summary, list, and overflow
formatDetectChangesResult computed the "... and N more" overflow from the
already---limit-sliced array length, so under `--limit` the header (true
summary total), the listed rows, and the marker disagreed — e.g. "2 symbols"
in the header but a list of 1 with no marker. Base the overflow on the true
summary.changed_count / affected_count instead, and add the same marker to the
affected-processes list, so header + list + marker stay consistent. (#2310)
* feat(cli): add -l shorthand to impact --limit
The PR added the -l alias to context/cypher/detect-changes but left impact on
the long --limit only, so `impact -l 5` errored while `context -l 5` worked.
Add -l for parity and update the help-i18n OPTION_DESCRIPTION_KEYS key to the
new `-l, --limit <n>` flag string so the description still resolves. (#2310)
* fix(cli): bound all context --limit array categories
context --limit sliced only incoming.calls / outgoing.calls / outgoing.accesses
/ processes, leaving the other relType buckets unbounded — notably
incoming.accesses (bounded on outgoing but not incoming) plus imports/extends/
uses/… and typed_properties. Replace the hardcoded slices with a generic loop
over every array-valued bucket under incoming/outgoing, plus typed_properties
and processes, so --limit caps the whole context payload. (#2310)
* refactor(cli): parse --offset with a parseLimit-style helper
impactCommand parsed --offset with the legacy parseInt/Number.isFinite idiom
while --limit had moved to parseLimit, leaving two parsing styles side by side.
Add a sibling parseOffset helper (non-negative — offset 0 is valid) and use it,
so both options share one idiom; as a bonus it now rejects negative/fractional
offsets instead of forwarding them to the backend. (#2310)
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(deps): bump @ladybugdb/core to 0.18.0
Pins the release containing LadybugDB/ladybug#605 (TransactionManager
lock-order-inversion deadlock fix). Checked for known post-release
regressions specific to 0.18.0 via the Ladybug issue tracker — none found.
* fix(lbug): re-validate version-coupled comments and regexes for 0.18.0
Extends the LADYBUGDB-CONTRACT re-validation to two spots the marker
convention doesn't catch (bridge-db.ts's LBUG_OPEN_RETRY_PATTERNS,
conn-lock.ts's serialization rationale). Confirms via upstream source
diff (v0.16.1..v0.18.0) that every matched error-text string is
unchanged; conn-lock.ts's rationale is unaffected by #612/#623 since
neither addresses concurrent queries on one connection. Adds a
stemmer-sweep test proving the bundled 0.18.0 FTS extension accepts
every entry in SUPPORTED_FTS_STEMMERS, not just the default porter.
A live-trigger test for isMissingShadowSidecarError was attempted but
abandoned after empirical probing showed it isn't reliably
reproducible (even a SIGKILL-simulated crash didn't reproduce the
error on reopen) — documented as inspection-verified instead of
overclaiming test coverage that doesn't exist.
* test(lbug): add concurrent multi-connection deadlock stress test (#2338)
Directly validates LadybugDB/ladybug#605 — the TransactionManager
lock-order-inversion deadlock between a commit()-triggered checkpoint
and a concurrent beginAutoTransaction() — under a shape close to
GitNexus's real concurrent-writer load, independent of conn-lock.ts's
app-level serialization.
Comparison run against 0.17.1 (pre-fix): 1 of 4 runs hung for the full
60s timeout, a direct reproduction of the deadlock. 9 consecutive runs
against 0.18.0 (post-fix) all passed cleanly. Production is unchanged —
conn-lock.ts still serializes every write; this test validates the
engine-level fix without shipping multi-writer as a default.
* fix(test): address code review findings in multiwriter deadlock test
- Reuse lbug-config.ts's createLbugDatabase (via GITNEXUS_WAL_CHECKPOINT_THRESHOLD)
instead of a hand-duplicated 9-arg raw constructor call whose stated
justification (needing to bypass createLbugDatabase for the threshold
override) was incorrect — the env var already provides it.
- Close every QueryResult via the existing closeQueryResults helper
(write loop, read loop, verify query, setup query) instead of leaking
native cursors, matching lbug-adapter.ts's established pattern.
- Move all cleanup (timers, connections, db close, env var restore) into
the outer finally block so it runs on every exit path, not just the
happy path — a timeout or a writer exhausting its retry budget no
longer leaves dangling timers/connections/abandoned query loops.
Verified: 8 consecutive runs after the refactor, all passing cleanly.
Found via 8-angle parallel code review (medium effort); the two other
findings (isDbBusyError not recognizing LadybugDB's 'Only one write
transaction' message, and shadow-file poll timing sensitivity) are
noted in the PR description as residual — the first is a production-code
change beyond this validation test's scope, the second is inherent to
observing a transient native sidecar file and not cleanly fixable
without overengineering.
* fix(test): apply ce-code-review autofix findings
Fixes from an 8-persona parallel review round (correctness/testing/
maintainability/project-standards/reliability/adversarial/agent-native/
learnings):
- Extract the duplicated skipUnlessFtsAvailable/FTS_UNAVAILABLE_NOTE
helper (previously copy-pasted between lbug-core-adapter.test.ts and
fts-stemmer-sweep.test.ts) into a shared test/helpers/fts-availability.ts.
- Fix a native connection leak: verifyConn in the deadlock test's final
verification block is now pushed into the readers array the outer
finally already closes, so it's cleaned up even if the count query
throws.
- Fix a latent TypeScript type error (tsconfig.test.json catches it,
tsconfig.json doesn't): conn.query() types as QueryResult |
QueryResult[]; narrow to the single-result case before calling
.getAll() rather than assuming the array branch never happens.
- Replace repeated inline InstanceType<typeof import(...)> expressions
with local LbugDatabase/LbugConnection type aliases.
Verified: 12 consecutive runs of the deadlock test all pass, full
lbug-db project (336 tests) green.
Cross-reviewer-confirmed but left as residual (design judgment calls,
not mechanical fixes) for the PR description: isDbBusyError doesn't
recognize LadybugDB's 'Only one write transaction' message (pre-existing
production gap, confirmed independently by 3 reviewers); the deadlock
test's timeout path doesn't cancel in-flight writer/reader loops before
closing connections; the reader loop has no bounded retry for transient
errors during the race window; pinning @ladybugdb/core with a caret
range trades automatic patch updates for less re-validation certainty.
* docs: trim task-referencing JSDoc artifacts, add operator notes
The U2 re-validation pass left verbose 'Re-validated on the
0.17.0->0.18.0 bump (#2338): ...' paragraphs stacked onto 5 production
files' docstrings, alongside the already-updated version numbers. That
narrative (SIGKILL-probe methodology, diff commands run, issue
cross-references) belongs in the PR description, not in code comments
that will accumulate a new paragraph on every future bump and confuse
readers who just want the current fact. Trimmed each to state only the
durable, current-state fact:
- lbug-config.ts, sidecar-recovery.ts, lbug-adapter.ts, bridge-db.ts:
dropped the bump-narrative paragraphs; kept only genuinely durable
notes (e.g., which matchers are inspection-verified vs live-tested,
what upstream wording changed).
- conn-lock.ts: compressed a 12-line, 3-issue-number enumeration into
2 lines stating the current conclusion (no upstream 0.18.0 fix
addresses the same-connection-concurrent-query risk this lock
guards against).
Also added operator-facing notes to GUARDRAILS.md and RUNBOOK.md's
existing 'LadybugDB lock' sections: an isDbBusyError gap found during
this validation (LadybugDB's 'Only one write transaction...' message
isn't recognized by our busy/lock retry matcher) means that specific
error can surface unretried. Documented so it's recognized as the same
single-writer conflict, not a new failure mode.
* refactor(test): use gitnexus-shared's withRetry in multiwriter deadlock test
Replaces the hand-rolled writeWithRetry/sleep loop with the existing
gitnexus-shared retry helper (already used by embeddings/hf-env.ts)
instead of duplicating the pattern.
* fix(test): guarantee non-zero retry delay in deadlock test's writer loop
withRetry's isRetryable previously returned {retry: bool} with no afterMs,
so computeBackoffMs's exponential-jitter formula gave a deterministic
zero-delay on the first retry (floor(random()*1) is always 0 at attempt=0).
This contradicted the file's own documented tuning, which specifically
needs a non-zero 1-3ms delay to avoid tripping a different native guard.
Return an explicit afterMs override on the retryable branch instead.
* docs(test): remove dangling doc references from deadlock test JSDoc
The JSDoc pointed to a local-session-only docs/plans/2026-07-01-001-...
path (docs/ is repo-gitignored, so this never existed for anyone but the
implementing session) and to "the PR description" as a source of truth
that stops being current once the PR merges. Replace both with
self-contained prose and durable references (issue/PR numbers, commit
SHAs, GUARDRAILS.md/RUNBOOK.md) that stay resolvable after merge.
* fix(search): harden SUPPORTED_FTS_STEMMERS against external mutation
Type as ReadonlySet<string> to match this codebase's established
convention for exported validation allowlists (EVAL_SERVER_TOOLS,
STRUCTURAL_LABELS). Type-only change — no behavior change; both the
internal .has() check and the sweep test's spread-iterate pattern
continue to work unchanged.
* docs(guardrails): fold Known-gap note into the LadybugDB Sign's Why label
GUARDRAILS.md's own convention is strictly Trigger/Do/Why per Sign
entry (stated in the file's header, followed by all 5 other entries).
The new isDbBusyError gap note introduced a 4th label; fold it into
Why instead, which is what it's actually explaining.
* fix(test): run the multi-writer deadlock test on Windows too
itLbugMultiwriter mirrored lbug-core-adapter.test.ts's win32 skip, but
that pattern exists for a close-then-reopen-same-path lock lingering
bug (kuzudb/kuzu#3872). This test never reopens the database — it
holds connections open for the whole run — so the skip excluded the
one test validating issue #2338's deadlock fix from the platform
conn-lock.ts actually ships native bindings for.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(indexing): keep full text file content searchable
* fix(indexing): flush CSV chunks by byte size
* fix(fts): flatten newlines/tabs in indexed content so multiline files are searchable (#2317)
The end-to-end FTS test (review follow-up F2) exposed that removing the 10KB
cap alone does NOT fix#2317: Ladybug's FTS tokenizer splits ONLY on the space
character — \n, \r, and \t are not delimiters. So multiline file/symbol content
indexes as a few giant cross-line tokens that no word query matches; full
content is stored but stays unsearchable. (Verified: identical 8KB content is
fully searchable when space-separated and entirely unsearchable when
newline-separated.) The existing fts-description-search test never caught this
because all its seed content is single-line.
Collapse \r\n\t -> single space in the FTS-indexed text (extractContent's File
and snippet content, plus the description column) via normalizeFtsText. This
rewrites the stored column too, so File content returned via the graph API is
space-flattened — an accepted trade for making file/symbol text searchable.
Add the real end-to-end guard test/integration/fts-fullfile-search.test.ts:
write a >16KB file, load it through the real streamAllCSVsToDisk -> COPY ->
createSearchFTSIndexes path, and assert searchFTSFromLbug returns a needle past
10KB (plus a short-content no-regression and a stored-cell-not-truncated
guard). It drives the COPY path a Cypher-seed test would bypass, reusing
withTestLbugDB's FTS-availability gating via a new before-FTS load hook.
* docs(lbug): note the deliberate File-unbounded / snippet-capped asymmetry
The File branch returns full content (whitespace-normalized for FTS, bounded
upstream by the walker cap) while the symbol snippet path 11 lines down stays
MAX_SNIPPET-capped. Comment the intent so the uncapped File branch doesn't read
as a forgotten guard. No behavior change.
* test(lbug): update #2203 overlap round-trip for FTS whitespace normalization
The newline/tab→space normalization (a170915a, #2317) flattens stored File
content, so the #2203 overlap test's "File content == original multiline
source" assertion no longer holds. The test's actual invariant — overlap path
== serial path, byte-for-byte — is unchanged and still asserted; BasicBlock
text (not FTS-indexed) still round-trips raw. Update only the File-content
expectation to the whitespace-flattened form and document why.
* fix(lbug): collapse CSV flush to a single byte threshold
BufferedCSVWriter flushed on row-count (FLUSH_EVERY=500) OR byte-count
(FLUSH_BYTES=8MB) — two independent triggers for one job. Byte count is
the only one tied to the actual risk (an unbounded buffer.join('\n')
string), so drop FLUSH_EVERY and make shouldFlushCSVBuffer single-arg.
Rather than tune FLUSH_BYTES by guesswork or expose it as an env knob,
derive its safety margin from constants the codebase already hard-enforces:
a single row is capped at TREE_SITTER_MAX_BUFFER (32MB, clamped regardless
of GITNEXUS_MAX_FILE_SIZE) and at most doubled by escapeCSVField's
quote-escaping, so the worst-case joined chunk (FLUSH_BYTES + 2 *
TREE_SITTER_MAX_BUFFER ≈ 72MB) sits >7x under Node's MAX_STRING_LENGTH
(~512MB) — the ceiling that throws RangeError: Invalid string length.
A new test pins that margin numerically so it can't erode unnoticed, which
covers the "configurable" alternative better than a knob would: there's no
evidence any deployment needs a different value, and an unbounded env var
would let an operator silently walk the margin back into the danger zone.
Also updates the two tests tied to the removed row-count path: the
FLUSH_EVERY-boundary integration test now crosses FLUSH_BYTES with real
oversized File content instead of relying on row count, and the
shouldFlushCSVBuffer unit test drops to the new single-arg signature.
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* fix(group): use labels(n) IN allowlist instead of LadybugDB-incompatible multi-label Cypher (#2325)
manifest-extractor and http-route-extractor built Cypher with the openCypher
label disjunction `MATCH (n:A|B|C)`, which LadybugDB's parser rejects. The
error was swallowed by try/catch, so manifest contracts silently fell back to
synthetic UIDs with empty filePath and http-route cross-file handler
resolution silently returned null.
Replace all 7 queries with `MATCH (n) WHERE labels(n) IN [...]`. LadybugDB
returns labels(n) as a single string, so this is an exact allowlist — a 1:1
behavior-preserving syntax translation (validated against LadybugDB 0.17.1).
Export the two http-route query constants so integration tests can run the
exact production strings against a real DB, and add per-branch real-DB
regression coverage (the bug shipped because no test exercised these queries).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): import CypherExecutor from contract-extractor in #2325 test
The new manifest regression test imported `CypherExecutor` from
`group/types.js`, which does not export it — the type is defined only in
`group/contract-extractor.js` (as all production extractors import it).
This was a real TS2305 under `tsc -p tsconfig.test.json`, masked from CI
because the default tsconfig excludes `test/` and `import type` is erased
at runtime. Split the import so the type resolves from its real module.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(group): run #2325 native-LadybugDB tests in the lbug-db project
Per TESTING.md, every test that opens a real `@ladybugdb/core` handle must
be registered in the sequential `lbug-db` Vitest project (and excluded from
`default`) to avoid native-mmap file-lock conflicts across parallel forks on
Windows. The two new group integration tests use `withTestLbugDB`/pool-adapter
but were in neither list, so they ran under the parallel `default` project.
Add both to `lbug-db.include` and `default.exclude`, matching every sibling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(group): export custom-contract resolve query for #2325 test
The #2325 integration test hand-copied the 21-label `custom`-branch
resolve query into a local `LABELS_CUSTOM_QUERY` constant, so editing the
production allowlist would silently desync the canary. Promote the query to
an exported `CUSTOM_CONTRACT_RESOLVE_QUERY` (mirroring http-route-extractor's
exported query strings) and import it in the test, so the canary always runs
the exact production query. Behavior unchanged — same query string.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(group): de-brittle the #2325 custom-query label assertion
The unit test asserted a fixed 7-label ordered substring of the 21-label
custom-branch allowlist, coupling it to label order and no-space formatting —
a harmless reorder would have broken it. Replace with order/spacing-tolerant
membership checks for a spread of individual labels, keeping the unconditional
`not.toContain('Function|Method')` guard as the real regression check.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(group): correct #2325 http-route docstring + add real-trigger canary
The http-route test claimed `MATCH (n:Function|Method|CodeElement)` "which
LadybugDB rejects" — but that 3-label disjunction actually PARSES. Verified
against the real parser, the genuine #2325 trigger is a *reserved-keyword*
label in the disjunction: `Macro` and `Union` both are, and only the manifest
custom branch (21-label list) and the lib branch (missing `Package` table)
actually threw. The http-route conversion to `labels(n) IN [...]` was a
consistency change, not a parser fix.
Correct the misleading docstring and add a rejection canary pinned to the real
cause (`MATCH (n:Function|Macro|Union)` rejects), so a future query that
reintroduces a reserved-keyword disjunction is caught.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(group): cover the thrift package-strip path against a real LadybugDB
The thrift-only branch of resolveSymbol strips a `package.` prefix from the
service name (`com.example.AuthService` -> `AuthService`) before the
Class/Interface lookup — previously exercised only with a mocked executor.
Add a service-contract integration case (no method, so it takes the
package-strip path, not the grpc-identical method path) that resolves the real
`cls:AuthService`. Without the strip the lookup matches nothing and falls back
to a synthetic uid, so this is a non-vacuous guard for the strip.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): drop vestigial 'Package' label from lib contract lookup
The `lib` branch allowlisted `labels(n) IN ['Package','Module']`, but there is
no `Package` node table (see NODE_TABLES) — the entry only ever matched
nothing. Restrict to `['Module']`, the label libraries actually resolve to.
Behavior-neutral: the lib integration case still resolves its Module symbol.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(group): update PIPELINE label-scoped queries to labels(n) IN form
The resolveSymbol label-scoping bullets still showed the banned
`MATCH (n:A|B)` disjunction; a contributor copying them would reintroduce
#2325. Rewrite them in the actual `labels(n) IN [...]` form, note the real
trigger (LadybugDB rejects a disjunction naming a reserved keyword such as
`Macro`/`Union`), and reflect the lib allowlist as `['Module']` after dropping
the vestigial `Package` label.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(group): correct #2325 root-cause comments in the extractors
The production comments claimed LadybugDB rejects the `MATCH (n:A|B)`
disjunction "outright". Verified against the real parser, it rejects only when
a label is a reserved keyword (`Macro`, `Union`) or names a missing node
table. So only the manifest `custom` branch (reserved keywords in its 21-label
list) and the `lib` branch (missing `Package` table) actually threw; the
http-route/grpc/thrift/topic disjunctions parse fine and were converted to
`labels(n) IN [...]` for consistency and future-proofing, not because they were
broken. Rewrite the comments to say so accurately. No behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(group): make #2325 test prose name the real reserved-keyword trigger
The manifest test docstring/title and the unit-test comment said LadybugDB
rejects the `MATCH (n:A|B)` disjunction generally. It rejects only when a label
is a reserved keyword (`Macro`/`Union`) or a missing table. Reword the docstring
(custom + lib branches threw; others parsed), retitle the rejection canary to
"its list names reserved keywords Macro/Union", and correct the unit-test
comment. The rejection canary still passes — the custom 21-label list does
contain Macro/Union. No behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): cache read-only bridge handle to fix Windows @group reopen (#2274)
A long-lived MCP server opened bridge.lbug read-only, queried, and closed
it on every @group trace/impact call. On Windows the in-process reopen of the
same file fails (the OS handle is not fully released before the next open races
in), so repeated @group calls broke. #2269 fixed Linux/macOS by skipping
CHECKPOINT on read-only handles; Windows stayed broken.
Instead of fighting LadybugDB's Windows close/reopen timing: cache one
read-only handle per groupDir and reuse it across calls (open-once-per-process
already works on Windows). getCachedBridgeReadOnly:
- reuses a single handle keyed by resolved groupDir,
- invalidates on mtime change (external writer / re-sync),
- invalidates explicitly before same-process writes (writeBridge),
- guards concurrent first-open with an in-flight promise (no handle leak),
- closes all handles on process exit.
closeBridgeDb now no-ops for the cached handle (cache owns its lifetime);
uncached/writable handles are unaffected. ensureBridgeReady uses the cache.
The in-process write->read reopen of the same bridge.lbug file remains a known
LadybugDB Windows limitation, so the existing reopen tests stay win32-skipped.
A new cache-aware itCacheReopen gate applies to the 3 new tests whose setup
requires write-then-read in the same process (same class as itLbugReopen). The
cache itself exercises read->read reuse and is unaffected.
* fix(group): harden bridge RO-handle cache for concurrency, lifetime & Windows (#2313 review)
Addresses the tri-review + Copilot findings on the read-only bridge-handle cache:
- P1 (F2): serialize queryBridge per cached handle via a per-handle FIFO lock
(the conn-lock.ts chain mechanic, keyed per cache entry, not the global lock).
Two concurrent @group callers sharing one lbug.Connection can no longer
dispatch two queries at once (the heap-corruption hazard). Uncached/writable
handles skip the lock at zero cost.
- P1 (F3): refcount lease — getCachedBridgeReadOnly acquires, closeBridgeDb
releases (no caller change). The native close is deferred until in-flight
readers drain (refs===0) and runs exactly once (closeStarted guard).
invalidateBridgeCache and the mtime-evict path share one evict/close path.
- Windows: bounded drain in evictBridgeEntry — a concurrent group_sync waits
(<= WINDOWS_DRAIN_TIMEOUT_MS) for readers to release before the atomic rename
on win32 so it stays clean; POSIX remains fully non-blocking; single-threaded
sync still closes-before-rename on all platforms.
- P0 (F1/F6): gate the mtime cache test with itCacheReopen (win32-skipped) and
drop the manual invalidate so writeBridge self-invalidation is under test;
add an external-writer (fsp.utimes) reopen case.
- Windows coverage (F9): new cross-process integration test seeds bridge.lbug
in a separate tsx process, so read->read handle reuse is proven on win32 CI
(not skipped). Plus concurrent cold-open dedupe coverage.
- P2/P3: scope the Windows NOTE to read->read (F4); JSDoc the closeBridgeDb
release/close contract (F5); drop the if-branch in the B2 probe (F7); revert
incidental Prettier churn in cross-impact.ts (F14); fix the stale describe
header (F15); document the beforeExit/signal and ENOENT-mtime behavior
(F11/F13).
tsc clean; group unit + integration suites green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(group): run the B2 rename-clash probe on win32 via cross-process seed (#2313 review)
Moves the B2 "external rename while a cached RO handle is held" probe out of the
unit suite (where it was win32-skipped, because its in-process writeBridge->RO-open
is the unfixed Windows reopen) into the cross-process integration test, where a
separate-process seed makes the RO open clean. The probe now RUNS ON WIN32 CI and
empirically answers whether an open RO handle blocks an external atomic rename over
bridge.lbug — the assumption under writeBridge's invalidate-before-rename and the
win32 drain.
Hardened (per adversarial review) so a win32 RED is the real steady-state share-mode
signal, not an artifact:
- use production retryRename (not bare fsp.rename) so transient EBUSY/EPERM from the
Windows AV/indexer scanning the fresh temp file is absorbed; a RED then means the
rename is blocked even after retries (FILE_SHARE_DELETE absent -> invalidate-before-
rename is load-bearing).
- stage the byte-identical replacement BEFORE opening the RO handle, so no second OS
handle touches bridge.lbug while LadybugDB holds it (avoids a FILE_SHARE_READ red for
the wrong question).
- drop the post-rename query (handle survival is covered by the reuse test); the probe's
sole verdict is whether the rename is blocked.
Removes the old win32-skipped unit B2 (a strict subset of the new probe).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): stabilize api_impact response shape for same-URL multi-verb routes
After #2302 made Route identity method-aware, a same URL exposes one Route
node per HTTP verb, so a bare-URL api_impact lookup could silently flip from a
direct route object to the wrapped { routes, total } envelope. Surface each
route's `method` (via the shared fetch) so multi-verb results are
distinguishable, and add an optional `method` selector that narrows a
multi-verb URL/file to one verb and forces the singular shape. A verb that
matches no route returns a clear error. Document the match-count contract in
the tool schema.
Refs #2308
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(mcp): cover same-URL multi-verb api_impact contract
Regression coverage for #2308: bare-URL and bare-file lookups of a same-URL
GET+POST pair return the wrapped form with distinct per-route methods; the
method selector collapses to the singular shape (case-insensitively); an
unmatched verb returns a verb-not-found error; and verbless routes surface a
null method.
Refs #2308
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(review): apply autofix feedback
- tools.ts: correct api_impact contract docs — `method` narrows to one verb
but the singular shape only holds when exactly one route remains after
filtering (substring route/file matches can still wrap); cover file lookups;
enumerate verbs.
- local-backend.ts: surface `method` in route_map and shape_check output (the
shared fetch already returns it; agents discover verbs there before
api_impact).
- local-backend.ts: compute routeCountByHandler from the unfiltered match so a
method-scoped api_impact still flags a multi-verb handler's partial middleware.
- tests: add file+method and verbless-exclusion cases; assert unconditionally
via toMatchObject; lowercase the verb-not-found input to exercise error
uppercasing.
Refs #2308
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): treat wildcard '*' routes as matching any api_impact method selector (#2308)
Method-agnostic routes (Django function views) persist with Route method
'*', not null. The api_impact method selector used exact verb equality, so
'*' routes were excluded and api_impact({route, method:'POST'}) falsely
reported 'No routes found' for a route that handles every verb. Treat '*'
as matching any requested verb, and correct the comment + tool-description
strings that wrongly grouped Django wildcards with null/verbless routes.
* fix(mcp): harden api_impact method input against non-string and empty values (#2308)
The MCP envelope is not schema-validated, so a non-string `method` reached
`.toUpperCase()` and threw a TypeError. Widen the param to `unknown` and guard
it with a typeof check that returns a structured error (mirroring the
resolveAliasString pattern from #2175), and collapse empty/whitespace verbs to
no selector.
* fix(mcp): distinguish url-not-found from verb-not-found in api_impact error (#2308)
The verb-not-found error appended 'with method "X"' even when the URL/file
itself did not exist, implying the URL exists with other verbs. Gate the verb
clause on matched.length > 0 so a non-existent URL/file gets the plain message.
* fix(mcp): clarify api_impact middlewareNote wording for verbless siblings (#2308)
The partial-middleware note claimed 'other methods in this handler' even when
the co-located sibling is a verbless (null) route rather than another HTTP
verb. Refer to 'other route exports' instead, which covers both cases.
* docs(mcp): document and test the method field on route_map and shape_check (#2308)
The shared fetchRoutesWithConsumers change surfaced a method key on route_map
and shape_check responses too, but their tool descriptions never mentioned it
and no test covered it. Document the field on both descriptions and add unit
tests asserting it (shape_check rows carry responseKeys + a consumer so they
survive shape_check's keys-and-consumers filter).
* test(mcp): cover middlewareDetection 'partial' survival under a method filter (#2308)
The diff's core behavioral line counts verbs-per-handler from the unfiltered
match set so a method-scoped query still flags a multi-verb handler's partial
middleware, but no test exercised it (every verbRow hardcoded middleware:null).
Add a middleware param to verbRow and a test that fails if the count is taken
from the post-filter set instead. Verified via mutation: matched->routes fails it.
* test(mcp): add live-LadybugDB integration coverage for route method round-trip (#2308)
The new n.method query column was only unit-mocked. Add a self-contained
integration suite that seeds GET+POST /api/orders and a method-agnostic '*'
Django route, then asserts api_impact surfaces method, narrows by verb, and
matches the '*' route end-to-end (the U1 fix), plus route_map surfacing.
Own seed + no FTS so it neither perturbs api-impact-e2e nor silently skips.
* refactor(mcp): type the api_impact response shape instead of Promise<any> (#2308)
Replace apiImpact's Promise<any> with an explicit ApiImpactResult union
(single route | wrapped { routes, total } | { error }) and a typed
ApiImpactRoute. The results.map is annotated so the response builder is
checked against the declared shape. Behavior unchanged; sibling MCP methods
keep their Promise<any> convention.
* fix(mcp): express the route-or-file requirement in the api_impact schema (#2308)
The inputSchema left route/file as bare optionals, so the 'at least one of
route/file' rule the handler enforces was invisible to clients. Add an optional
anyOf to ToolDefinition (forwarded verbatim by the ListTools handler) and an
anyOf:[{required:[route]},{required:[file]}] on api_impact. Matches runtime
(both allowed, route wins); 'at least one' not 'exactly one'.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(ingestion/routes): give Route nodes a (method, url) identity (#2289)
Route node identity was URL-only, so a same-URL multi-verb pair
(GET /x + POST /x) collapsed into a single node and the second verb's
handler and execution flow were silently lost. Route identity is now
(method, url) via routeNodeKey(method, url): a known, specific verb keys
as "METHOD url", while a method-less route (filesystem routes — Next.js /
Expo / PHP — and Laravel resource/apiResource) or a wildcard "*" route
(e.g. Django function views) falls back to URL-only. The fallback is
byte-identical to the previous URL-only ids, so only genuine
declaration-style multi-verb routes split into separate nodes.
The identity key is shared across the three phases that must agree on the
Route node id:
- routes phase: registry key + node id + handler-symbol lookup; the Route
node still carries the bare URL as its display name.
- call-processor: resolveRouteHandlerSymbols re-keyed by identity so each
verb resolves its own handler; a verb-less fetch() consumer matches by
URL and connects to every Route node at that URL (one per verb).
- processes phase: ENTRY_POINT_OF targets the identity-keyed node id.
Bumps INCREMENTAL_SCHEMA_VERSION 4 -> 5: persisted pre-v5 Route nodes use
the old url-only ids, so an incremental top-up would strand them alongside
new composite-keyed nodes — force a full re-analyze instead.
Part of #2280.
* fix(ingestion/routes): address PR #2302 review (P1/P2/P3)
P1 — Schema v5 fast-path bypass (run-analyze.ts):
Adds a schemaVersion-mismatch guard above the alreadyUpToDate early-return,
mirroring the pdgModeMismatch slot. Without it, a same-commit re-analyze on
a pre-v5 stamp returned alreadyUpToDate without ever reaching the
isIncremental gate, defeating the v5 schema bump's migration intent.
Regression test covers: analyze (stamps v5) → meta downgrade to v4 → same
commit re-analyze must NOT early-return and meta restamps to v5.
P2 — ENTRY_POINT_OF handler-aware linking (processes.ts):
Pre-fix routesByFile fanned every same-file Route to every same-file
process, cross-wiring same-file GET/POST handlers. Now reads handlerSymbolId
off the Route graph node (the source of truth routes.ts stamps) into
routesByHandlerId, with a routesWithoutHandlerByFile fallback — mirrors the
Tool linking precedent 10 lines below. Two regression tests: weak form
(only one handler has a process; sibling verb does not get spuriously
attached) and strong form (both handlers form distinct processes; each
Route links to exactly its own entryPoint, 2 edges not pre-fix 4).
P2 — Roundtrip composite-id (route-{method,handler-symbol}-roundtrip):
Both tests now seed the Route node with
generateId('Route', routeNodeKey('POST', '/api/orders')) and run the
Cypher MATCH against the composite id, exercising the literal-space-in-id
through CSV→COPY→HANDLES_ROUTE_QUERY. A space-in-id escape regression
would surface here instead of being silently swallowed by the extractor's
catch.
P3 — doc-drift + test if:
- route-path.ts:4 — header updated to "(method, url) via routeNodeKey"
- java.ts:684 — drop "Route nodes are URL-keyed"; #2289 closes that gap
- manifest-extractor.ts:196 — explicit that Route node *id* is composite
while route.name remains the bare URL
- multi-verb-route-identity.test.ts:88 — forEachRelationship+if rewritten
as a .filter().map() chain (no test-level conditional). New
route-process-linking tests are also if-free.
Validation: tsc clean, prettier clean, 9 touched suites / 43 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(ingestion/routes): drop routes.ts re-export, fix CI fast-path tests
Two follow-ups on PR #2302's CHANGES_REQUESTED review:
1. Drop `routes.ts` re-export of `normalizeExtractedRoutePath` /
`normalizeRouteMethod` / `routeNodeKey` (per @magyargergo's inline
comment at routes.ts:153 — the symbols already live in
`route-extractors/route-path.ts` and consumers should import them
from the source, not via a routes-phase indirection that was kept
only as a compat shim during the #2289 refactor). Updated the two
remaining callers (blade-template-routes / spring-route-extractor-
parity tests) to import directly from `route-extractors/route-path.js`.
`call-processor.ts` and `processes.ts` already import from the source.
2. Fix two `run-analyze.test.ts` fast-path tests that started failing
on CI after the schema-version mismatch guard landed (
"creates .gitnexus/.gitignore on the already-up-to-date fast path"
and "reports isPrimaryBranch false for an up-to-date non-primary
branch"). The test fixtures hand-built a RepoMeta with NO
schemaVersion field; with the guard now checking
`existingMeta.schemaVersion !== INCREMENTAL_SCHEMA_VERSION`, that
pre-versioning shape was treated as a mismatch and forced a rebuild,
short-circuiting the fast path the tests exercise. Stamp the current
schemaVersion on those fixtures so they reflect the post-#2289 meta
shape production actually writes (`runFullAnalysis` always stamps
the field on git repos — see meta save site).
Validation: tsc clean, prettier clean, 11 touched suites / 80 tests pass.
---------
Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* fix(search): index description column for FTS so doc comments are keyword-searchable
Closes#2299. descriptionExtractor (#2286) populates the `description`
column for every symbol table, but FTS only indexed name+content on 5
tables, so doc-comment keywords (Javadoc/KDoc/godoc/Rust ///) were
invisible to BM25 keyword search.
- Add `description` to the Function/Class/Method/Interface FTS indexes
(File has no description column, left as name+content).
- Add FTS indexes for the remaining EMBEDDABLE_LABELS symbol tables
(Struct, Enum, Trait, Impl, Macro, Namespace, Constructor, TypeAlias,
Typedef, Const, Property, Record, Union, Static, Variable).
- createSearchFTSIndexes now drops-then-creates each index so the schema
change reaches existing DBs on incremental re-analyze and --repair-fts
(createFTSIndex is idempotent-by-name and would otherwise skip stale
indexes).
Tests: fts-schema column-subset + coverage guards; drop-before-create
order; e2e doc-comment keyword search (Java class + Rust struct found by
description-only terms). bm25-search assertions derive from FTS_INDEXES.
* fix(review): apply autofix feedback
- Guard the --repair-fts path on FTS-extension availability before
createSearchFTSIndexes drops-then-creates indexes (P1 regression:
without the gate, an unavailable extension could drop existing indexes
then fail to recreate them, leaving the DB index-less). Mirrors the
analyze path's ftsAvailable gate and fails loudly first.
- Add a re-analyze upgrade integration test: seed an old name+content-only
DB (no Struct index), run the real createSearchFTSIndexes(), and assert
description keyword search + the previously un-indexed Struct now resolve.
Proves drop-then-create upgrades a live stale index end-to-end.
* fix(ci): add loadFTSExtension to --repair-fts test mocks
The R3 review fix added a loadFTSExtension availability gate to the
--repair-fts path, but run-analyze-fts-repair.test.ts mocked the lbug
adapter without that export, so both repair tests threw `No
"loadFTSExtension" export`. Add loadFTSExtension to the two mocks
(returning true to preserve their original intent) and add a dedicated
test proving the guard fails loudly — and does NOT drop any index —
when the extension is unavailable.
* test(fts): run fts-description-search in the sequential lbug-db project
It was the only FTS-index-creating integration test left in the parallel
`default` vitest project; every other ftsIndexes-using test (search-core,
search-pool, augmentation, …) runs in the `lbug-db` project, which forces
fileParallelism: false to avoid LadybugDB native mmap file-lock conflicts
in parallel forks (Windows). Add it to the lbug-db include list and the
default exclude list to match the convention and remove the flake risk.
* test(ci): fail loudly when FTS extension is unavailable, never silently skip
FTS-dependent lbug integration suites (search-core, search-pool,
augmentation, fts-description-search, …) self-skip via ctx.skip() when the
LadybugDB FTS extension can't load, emitting only a console.warn while the
job stays green. That means a broken/missing FTS extension in CI would make
these integration tests silently vanish with no signal — false confidence.
withTestLbugDB now honors GITNEXUS_REQUIRE_FTS=1: when set and the extension
is unavailable, setup() throws instead of skipping, so the suite fails
loudly. The CI test jobs (ubuntu coverage + windows/macOS cross-platform)
set the flag; local/offline runs leave it unset and keep skipping
gracefully. (Verified the extension currently loads on all three runners,
so this is a guard against regression, not a behavior change today.)
* test(ci): run fts-description-search on macOS/Windows cross-platform jobs
The new FTS description-search suite was registered in the sequential
lbug-db vitest project (ubuntu/coverage) but absent from LBUG_NATIVE, so
the macOS/Windows platform-sensitive jobs (which run only the explicit
ALL_CROSS_PLATFORM allowlist via run-cross-platform.ts) never executed it.
The GITNEXUS_REQUIRE_FTS=1 hardening on those jobs guarded the old FTS
fixtures but not the new 20-index/description path. Add the suite to
LBUG_NATIVE so the new path is validated cross-platform too.
Refs #2299.
* fix(search): verify FTS indexes cover description, not just queryability
verifySearchFTSIndexes probed each index with QUERY_FTS_INDEX and treated
'queryable' as 'present'. A stale name+content-only index left on a
pre-#2299 DB stays queryable yet silently misses the description column, so
verification would pass green while doc-comment search stayed broken.
Switch to a single CALL SHOW_INDEXES() that exposes property_names per
index, and report an index as missing when it is absent OR does not cover
its configured columns. Return contract (string[] of table.indexName) is
unchanged, so both run-analyze.ts call sites are untouched. The per-index
string interpolation is gone, so the now-dead safeIdentifier helper is
removed.
The real caller of the live function in tests is bm25-search.test.ts (the
repair test mocks verifySearchFTSIndexes wholesale); its two probe-shaped
cases are rewritten to feed SHOW_INDEXES rows and now assert column
coverage, plus an absent-index case.
Refs #2299.
* test(search): assert description search via the public query surface
The #2299 integration suite only exercised the searchFTSFromLbug helper.
Add a third block that drives the public LocalBackend.callTool('query')
path — which resolves the repo via the registry and routes BM25 through the
pool adapter (a different connection context than the core-adapter helper) —
and asserts a description-only keyword returns the seeded class. Reuses the
existing description-only SEED and production FTS_INDEXES; partial-mocks
repo-manager so listRegisteredRepos points at the test DB while
cleanupOldKuzuFiles and the rest stay real.
Refs #2299.
* test(search): make lbug-core-adapter FTS gate honor GITNEXUS_REQUIRE_FTS
lbug-core-adapter.test.ts has its own per-test FTS gate (skipUnlessFtsAvailable)
that called ctx.skip() whenever the extension could not load — bypassing the
GITNEXUS_REQUIRE_FTS=1 hardening that withTestLbugDB already honors. Since this
file is in LBUG_NATIVE it runs on the ubuntu/macOS/windows jobs that all set
GITNEXUS_REQUIRE_FTS=1, so an FTS regression on a runner would have let these
FTS-primitive tests silently vanish from a green run — the exact gap #2299's
test-infra hardening set out to close.
Make the helper mirror withTestLbugDB: when GITNEXUS_REQUIRE_FTS=1 and the
extension is unavailable, throw (hard fail) instead of skipping. Offline/local
runs (no env var) still skip gracefully.
Refs #2299.
* feat: ✨ resolve Nuxt/Nitro auto-imports in TypeScript scope resolver
* fix: 🐛 skip self-referential edges in Nuxt auto-import emission
* fix: 🐛 address Sourcery review -- gate Nitro scan on imports.d.ts and pre-index explicit imports
* fix: scope Nuxt auto-import resolution
* fix: address Nuxt auto-import review follow-ups
* fix(ingestion): capture only LHS binding names in Nitro server-util exports
The Nuxt server-util export scanner ran a declarator regex over the whole
`export const …` right-hand side, so it registered RHS tokens as auto-import
names: arrow-function parameters (`export const f = (event) => …` → `event`),
object-literal keys (`export const c = { onError } ` → `onError`), and bare
operands. It also dropped generic-typed declarators
(`export const x: Map<a, b> = …`) because the type-annotation skip broke at the
comma inside the generic. Both produced wrong/missing auto-import CALLS edges.
Capture only the leading binding name of each top-level declarator via a
depth-aware comma splitter (tracks (), [], {}, <>), skipping destructuring
patterns. Nitro auto-imports only surface top-level binding names, so the RHS
is never parsed. Adds unit coverage for the param/object-key/operand/generic
and multi-declarator forms.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU
* fix(ingestion): stop Nitro server callers resolving client composables
`getNuxtAutoImportEntry` fell back to the client composable map when a
`server/api|routes|middleware` caller's name had no `server/utils` entry. But
Nitro only auto-imports `server/utils/**` into the server context — app
`composables/` are Vue-app-only — so that fallback minted CALLS/IMPORTS edges
Nitro never creates (e.g. a server route "calling" a composable it cannot see
without an explicit import).
Server callers now resolve the server map only. Restructure the barrel-directory
integration test to use a client caller (which legitimately auto-imports the
composable) so `index.*` resolution stays covered, and add a negative assertion
that `server/api/route.ts` emits no edge to `composables/*` while its real
`server/utils` call still resolves. Unit test locks that a server caller does
not fall back to a client-only name.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU
* fix(ingestion): let unresolved explicit imports shadow Nuxt auto-imports
The explicit-import suppression index only recorded import local names whose
edge resolved to a file (`edge.targetFile !== null`). An explicit import from an
unresolved external package — `import { useAuto } from '@vueuse/core'; useAuto()`
— therefore escaped suppression, and the post-resolution hook emitted a spurious
Nuxt auto-import CALLS edge for a name the file already imports explicitly.
Record the local name regardless of whether the import resolved: an explicit
import is authoritative shadowing intent. Adds an integration fixture importing
from an external package and a (non-vacuous) assertion that it emits no nuxt
edge.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU
* fix(ingestion): let type-annotated params shadow Nuxt auto-imports
hasLocalBindingInScopeChain only consulted scope.bindings, but type-annotated
function parameters live in scope.typeBindings (the TS scope query records them
as `@type-binding.parameter`, not `@declaration`). A parameter named like a
composable therefore failed to suppress the auto-import, leaking a spurious
CALLS edge.
Also check scope.typeBindings for the name (same-file scopes only). typeBindings
holds value-space binders' type facts (parameter annotations, `self`, variable
annotations) and never a pure type that belongs to callable space, so this
cannot over-suppress a real auto-import.
Documents the residual: function-typed params (`p: () => void`), untyped params,
destructured locals, and catch-clause vars are captured by neither map and still
leak — closing that needs shared scope-query changes beyond this feature, left
as a follow-up. Also adds a no-vacuous-pass guard to the shadowing/noise test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU
* fix(ingestion): treat server/plugins and server/tasks as Nitro runtime
isNitroServerRuntimeFile only matched server/api, server/routes, and
server/middleware. Nitro also auto-imports server/utils into server/plugins
and (since Nitro 2.6) server/tasks, so callers there were misrouted to the
client composable map. Extend the prefix set (now a named constant) to cover
them.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU
* fix(ingestion): merge duplicate JSDoc on collectImportsDts
Two consecutive JSDoc blocks preceded collectImportsDts; tooling (IDEs,
TypeDoc) attaches only the last one, silently dropping the descriptive block.
Fold the "returns true when read" line into the descriptive block as a
`@returns` tag.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU
* fix(ingestion): contain .nuxt/imports.d.ts source resolution to the repo
A crafted `.nuxt/imports.d.ts` source such as `from '../../../../etc/passwd'`
passes the project-local relative-path check but resolves outside the analyzed
repo, causing fs.stat probes against arbitrary host paths. Skip any source that
resolves outside repoRoot before touching the filesystem.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4W25WLfYD1JNy8icxeLPU
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(ingestion): add shared leading-doc-comment description extractor (#2270)
Add `extractLeadingDocComment` plus a language-neutral
`createLeadingDocDescriptionExtractor` factory and a shared
`DOC_BEARING_LABELS` set to `utils/ast-helpers.ts`. The helper pulls the
normalized text of a leading doc comment off a definition node's preceding
named sibling, covering both block doc comments (Javadoc/KDoc/JSDoc/PHPDoc/
Doxygen, opened by double-star or bang) and runs of line doc comments
(triple-slash, bang-slash, or caller-supplied prefixes such as Go's
double-slash or Ruby's hash). Grammar-agnostic by prefix match; widens
`getDefinitionNodeFromCaptures` to accept the optional-valued capture map.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eh2UmA6f2p25F3ow75Hjzx
* feat(languages): surface leading doc comments as description for all languages (#2270)
Register the leading-doc `descriptionExtractor` on every documentable
provider so Javadoc/KDoc/JSDoc/Doxygen/godoc/RDoc/`///` doc text lands in the
`description` column and reaches the embedding metadata header — making
methods/types semantically searchable by doc-only terms, matching the
behavior Python (docstring) and PHP (Eloquent) already had.
- Java, Kotlin, TypeScript, JavaScript, C, C++, C#, Dart, Rust, Swift: default
config (block + triple-slash/bang-slash doc comments).
- Go: godoc double-slash leading comments.
- Ruby: leading hash (RDoc/YARD) comments.
- PHP: existing Eloquent metadata takes precedence, else PHPDoc docblock.
Field/property/variable/const docs are intentionally out of scope.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eh2UmA6f2p25F3ow75Hjzx
* fix(review): apply autofix feedback (#2270)
Code-review autofix pass on the leading-doc-comment extractor:
- Enforce start-row adjacency in the line-comment run so a doc run stops at a
blank line (godoc/RDoc/rustdoc semantics). Prevents a Go license/earlier `//`
block or a Ruby shebang + `# frozen_string_literal:` magic comment, separated
by a blank line, from being absorbed into the first declaration's
description. Adjacency uses startPosition.row (reliable across grammars).
- Fix the degenerate empty comment `/**/` producing a spurious `/` description.
- PHP: compose createLeadingDocDescriptionExtractor() as the docblock fallback
instead of duplicating its body, and widen the param to CaptureMap to match
the LanguageProvider hook contract.
- Drop the factory's unused `labels` option (no consumer overrides it).
- Add tests: degenerate `/**/`, multi-line `///` run, `//!` inner doc, `/*!`
Doxygen block, Go/Ruby blank-line non-attachment + two-block adjacency, and
PHP Eloquent-metadata-wins-over-docblock ordering.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Eh2UmA6f2p25F3ow75Hjzx
* fix(ingestion): resolve exported TS/JS JSDoc via export_statement wrapper
Exported TS/JS declarations dropped their JSDoc: the TS query captures the
inner function_declaration/class_declaration, whose previousNamedSibling is
null because the JSDoc precedes the wrapping export_statement (PR #2286 review,
reproduced). Add a wrapperNodeTypes option to extractLeadingDocComment (folded
into a LeadingDocCommentOptions object threaded through the factory); when the
captured node yields no doc and its parent type is a configured wrapper, retry
from the parent. TS/JS providers pass ['export_statement']. Language config
stays at the call site (RFC #909). Mirrors the existing walk-up in
languages/javascript/captures.ts for JSDoc params.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ingestion): bound DOC_BEARING_LABELS to embeddable labels
Module/Delegate/Annotation were doc-bearing but absent from EMBEDDABLE_LABELS,
so their descriptions were extracted and written to the DB yet never embedded
or searchable (PR #2286 review) — wasted work, and the factory JSDoc overstated
"becomes semantically searchable". Remove those three labels so DOC_BEARING_LABELS
is a subset of EMBEDDABLE_LABELS, narrow the JSDoc, and add a subset-invariant
unit test to guard against drift. Making those labels (and C++ `Template`)
searchable needs an embedding-pipeline/schema change and is left as a follow-up.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ingestion): skip file-top license/header blocks as descriptions
A file-top /** … */ license/copyright/overview block has no package/import
sibling to shield it from the first declaration, so it was absorbed as that
symbol's description and polluted the embedding text (PR #2286 review). The
block-comment branch already cannot use a strict row-adjacency check (grammars
fold the trailing newline into the comment node), so match header markers
instead — SPDX-License-Identifier, @license/@file/@fileoverview, "Licensed
under", and copyright-with-(c)/year. Markers are specific enough not to fire on
an ordinary doc that merely mentions the word "copyright" (over-fire guard test).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ingestion): ignore Go/Ruby directive & magic comments in doc runs
Go build/tool directives (//go:build, //go:generate, // +build, //nolint, //line)
and Ruby magic comments / shebang (# frozen_string_literal:, # encoding:, # -*-,
#!, …) sitting directly above a symbol were folded into its description and
polluted the embedding text (PR #2286 review). Add a lineDirectivePrefixes option;
a matching line is skipped in the doc run (skip-and-continue, so a real doc above
an interleaved directive is still collected — godoc/RDoc semantics). Go and Ruby
providers supply their own directive prefixes (RFC #909 — config at the call site).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ingestion): guard descriptionExtractor call against throws
A throw inside any provider's descriptionExtractor escaped processFileGroup to
the language-group catch, which treats any throw as "parser unavailable" and
silently drops every remaining file in the group (PR #2286 review). Wrap the
call in try/catch + reportWarning, mirroring the adjacent extractTemplateConstraints
guard. Defensive parity — no behavior change on the success path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ingestion): treat Rust //! and /*! as inner docs
Rust //! and /*! are INNER doc comments (they document the enclosing item/module),
not the following item, but the shared helper attached them to the next definition
(PR #2286 review; a test even enshrined the wrong behavior). Add a blockDocPrefixes
option (default ['/**','/*!']); the Rust provider opts out of both inner-doc markers
(lineCommentPrefixes ['///'], blockDocPrefixes ['/**']). Doxygen //! and /*! keep
working for C/C++ via the defaults. Flip the Rust //! test to a negative assertion
and add a Rust /*! negative case.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ingestion): strip bidi/zero-width controls from doc descriptions
Doc-comment text is attacker-influenceable (any indexed repo) and is returned
verbatim to MCP clients, so a description could smuggle Trojan-Source-style bidi
overrides or zero-width characters (PR #2286 review). Strip U+202A–202E,
U+2066–2069, U+200B–200D and U+FEFF in the doc-comment normalization path
(block + line). Scoped to the description path only — global sanitizeUTF8 is
deliberately left alone (pre-existing, affects all fields). Implemented with a
code-point predicate so no literal invisible bytes live in the source.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(ingestion): doc-comment helper maintainability cleanups
PR #2286 review nits (no behavior change): drop the unused `export` on
DEFAULT_LINE_DOC_PREFIXES (no importer outside ast-helpers.ts); widen
getLabelFromCaptures' captureMap param to `Record<string, SyntaxNode | undefined>`
to match getDefinitionNodeFromCaptures (all accesses are truthiness-guarded); and
merge the split ast-helpers import statements in dart/ruby/rust into one each.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(architecture): document descriptionExtractor LanguageProvider hook
descriptionExtractor is now a near-universal LanguageProvider field (issue #2270)
but was missing from the architecture "Key fields" table (PR #2286 review). Add a
row describing it and the shared createLeadingDocDescriptionExtractor factory.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(ingestion): end-to-end description searchability for exported symbols
The unit tests stop at the descriptionExtractor hook; nothing proved a doc
comment survives the full parse pipeline into node.properties.description (the
field the embedding metadata header reads) — the exact gap that hid the exported
TS/JS regression (PR #2286 review). Add an integration test running the real
worker pipeline over an exported, JSDoc'd TS function and asserting its node
description carries the doc text. Verified locally against a built worker
(20s); runs in CI via pretest:integration build.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style(ingestion): prettier-wrap a long line in the doc-comment test
Formatting-only follow-up to the U3/U7 test additions so `quality / format` is
green. No behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(routes): extract Spring method-level array-form routes in ingestion + extractor parity test (#2138 follow-up)
ingestion's `extractSpringRoutes` (route-extractors/spring.ts) matched only a
single string literal on `@(Get|...)Mapping`, so the array form
`@GetMapping({"/a","/b"})` produced no graph Route node — while the group-layer
`java.ts` scan did match it. That divergence was the root of the #2265 array-form
parse-skip gap.
- spring.ts: add the array-form alternation
`[(string_literal) @value (element_value_array_initializer (string_literal) @value)]`
to the two method-declaration query branches (positional + `path=`/`value=`),
mirroring the group query. A multi-element array yields one match per element,
so the Phase 2 loop emits one route per path with no other change. Class-level
`@RequestMapping` array prefixes remain single-literal (rare; left to a
follow-up).
- test: spring-route-parity runs one shared Java fixture through BOTH extractors
(ingestion `extractSpringRoutes` + group `JAVA_HTTP_PLUGIN.scan`) and asserts
identical provider {method,path} sets — the parity guard the maintainer asked
for in #2078, so the two Spring extractors can't silently drift again
(verified: reverting the array branch turns the parity test red).
* fix(ingestion/routes): suppress wrong unprefixed route under class-array @RequestMapping; cover named-array + class-array parity
Addresses PR review on #2281:
- P2 class-array wrong-route: class branches now match the array form only to detect it; a method-level array route under a class-level array-form @RequestMapping is suppressed rather than emitted with a dropped prefix, so ingestion stays a strict subset of the group scan. Scalar method paths under an array class prefix are unchanged (pre-existing). Full class-array cross-product support tracked in a follow-up.
- P2 named-array coverage: added value={...}/path={...} parity cases, a consumes/produces array false-positive case, and a dedicated empty-provider-set assertion.
- P3 stale comments: updated the routeCoverage comment in java.ts and the route-parse-skip test note; narrowed the parity test drift claim.
routeCoverage stays 'partial'.
---------
Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* fix(mcp): tolerate adapter-materialized line:0 in impact callgraph mode (#2279)
Some MCP client/agent adapters serialize an omitted optional numeric
field as `0` rather than dropping it, so callgraph `impact` calls arrive
carrying a spurious `line: 0`. `line` is a PDG-only statement anchor and
is meaningless on the callgraph path, so the backend rejected the call
("'line' is only supported with mode:'pdg'") and strict clients rejected
it client-side against the advertised `minimum: 1`.
Treat a literal `line: 0` as omitted in `_impactImpl` when mode !== 'pdg'
and let the normal symbol→symbol BFS run. The coercion is deliberately
narrow: only the literal 0, only on the callgraph path. A genuine
positive `line` on callgraph still errors (real mode mistake), negative/
fractional values still error, and pdg mode is untouched — `line: 0`
there is still rejected (there is no 1-based source line 0 to anchor on).
Regression tests pin the full matrix: callgraph + line:0 runs the BFS and
is byte-identical to omitting line; pdg + line:0 still errors; positive
line on callgraph still errors.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): log swallowed best-effort query degradations at warn, not error
`logQueryError` is the shared handler for query failures that every caller
catches and degrades past with a safe fallback (the operation still returns
a result). It logged all of them at `logger.error` (level 50) — the same
severity as fatal failures — so a gracefully-handled degradation raised a
false alarm and drowned genuine errors. This surfaced as an ERROR-level log
firing during a passing unit test that intentionally injects a slice-callees
query failure to verify the degrade path.
Make the severity match reality:
- benign missing optional table/label/column (a repo analyzed without
processes/communities, or a pre-v3 PDG index lacking the `calleeIds`
column — a query that fails on every pdg-downstream impact for such an
index) → debug, the normal-configuration case.
- any other swallowed failure → warn (handled degradation, still observable).
- error is reserved for failures that actually abort an operation, which
log directly rather than through this helper.
Also fix the sibling bm25/FTS fallback, which logged its swallowed
"FTS indexes may not exist" degradation at error while its own import-failure
fallback already used warn.
The slice-callees degradation test now captures the log and asserts it lands
at warn (40), not error (50), pinning the severity against regression.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): relax impact `line` schema minimum to 0 for adapter compatibility (#2279)
Strict MCP clients/agents validate against the advertised input schema and
reject a request before sending it. With `line` declaring `minimum: 1`, a
client that materializes the omitted optional `line` as `0` rejects a
perfectly valid callgraph impact call client-side — so the backend tolerance
added in the previous commit never gets a chance to run.
Lower the advertised `line.minimum` to 0 and document that 0 (or omission)
means "no statement anchor" while mode:'pdg' still requires a positive line.
The advertised schema is advisory (the backend self-validates and is the real
gate), so this cannot loosen any enforced contract — it only stops strict
clients from pre-rejecting `line: 0`. Negative lines are still rejected at the
client boundary.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(review): apply autofix feedback
Code-review autofix pass on the #2279 branch:
- Replace a newly-introduced `mode as any` cast in the #2279 it.each with the
narrow `mode as 'callgraph' | undefined` (strict-typing-no-any).
- Add a degradation test for the new logQueryError benign-missing-table → debug
branch (asserts no warn/error record surfaces, i.e. it routed to debug).
- Pin the bm25/FTS error→warn severity change with a _captureLogger assertion
in the existing #1489 test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): make swallowed-failure callers surface degradation; narrow benign-error match (#2283)
Tri-review (#2283) found the `error → {debug|warn}` rework reduced telemetry
for `logQueryError` callers that do NOT degrade safely, while the docstring
over-claimed "every caller degrades to a safe fallback". Address the substance
rather than only the log level:
- rename apply-edit: track failed writes and return status:'partial' with
`failed_files` instead of reporting `status:'success'` when a write was
swallowed. A partial rename is no longer indistinguishable from a clean one.
- detect_changes: a swallowed symbol/process query failure now sets
`partial:true` (rendered by the existing eval-server partial path) so the
pre-commit safety gate can't return a false-clean `risk_level:'low'` no-op.
- isBenignMissingTableError: scope the `not (defined|found)` arm to a schema
object (table/label/rel/column/property), mirroring lbug-adapter's
isMissingColumnError. An unscoped "not found" matched operation failures like
`rg: not found` / `Symbol not found` and silently demoted them to debug.
- logQueryError docstring: state the contract honestly — level reflects
telemetry severity, and mutating/safety-critical callers MUST also surface a
result-level degradation signal; `warn` alone is not a substitute.
- pdg dispatch: pass the normalized `effectiveLine` (not raw params.line) so
the validation gate and engine share one source of truth (identity today).
Tests:
- _captureLogger(level?) lets tests capture below info; the benign-missing-table
test now asserts the record IS emitted at debug (20), not merely absent —
no longer a vacuous pass if the call were deleted.
- new: a non-schema "not found" failure logs at warn (regex-narrowing guard);
rename write-failure degrades to status:'partial'+failed_files; line:-1 on
the callgraph path still errors (line:0 coercion is narrow); typed the
it.each tuple to drop a `mode as` cast.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(mcp): fix impact `line` description contradiction for whole-symbol pdg (#2283)
The new `line` schema description said "mode:'pdg' requires a positive line",
which contradicted the top-level impact description ("Without 'line', pdg
returns whole-symbol inter-procedural reach plus local whole-symbol PDG
diagnostics"). A pdg call without a line is a valid (degraded whole-symbol)
call, not an error — the old wording could push an agent to avoid valid no-line
pdg calls or synthesize line:0 (which then hard-errors).
Reword to: omit line for whole-symbol pdg; a positive line anchors a statement
slice; literal 0 is tolerated only as an omitted-line compatibility sentinel on
the callgraph path and is rejected for mode:'pdg'. Update the schema test to
pin the new, non-contradictory wording and assert "requires a positive line" is
gone.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(group): resolve Go inline provider handlers via line containment (#2276)
Widen the Go HandleFunc + framework-route handler capture to match
func literals and emit name:null + call-site line for them, so an
inline handler resolves to its containing/closure symbol instead of
file-level. Named identifier handlers keep resolving by name.
* feat(group): resolve Laravel closure provider handlers via line containment (#2276)
Capture the Laravel route handler argument; a closure (anonymous
function or arrow fn) now emits name:null + the registration line so it
resolves to its containing symbol (service-provider boot, controller
method) by containment. Named-controller routes keep the 'route' label.
File-scope closures stay file-level (PHP closures not yet indexed).
* feat(group): wire call-site line on FastAPI provider emits (#2276)
Set line on the FastAPI @app/@router provider detections (already
name:null) so the source-scan fallback resolves the decorated handler
by line-span containment. Best-effort: FastAPI routes are graph-backed
and the function span starts at def, so this lands the single-decorator
case. Flask add_url_rule already carried line.
* feat(group): wire call-site line on Kotlin/Java Spring provider emits (#2276)
Add line to the Kotlin and Java Spring @*Mapping provider detections
for parity with the consumer emits and a future inline DSL. Inert for
current resolution: a named Spring controller method resolves by name
and never falls through to line-span containment.
* fix(review): apply autofix feedback
Pin two documented limitations with tests: a file-scope Laravel closure
and a multi-decorator FastAPI handler both degrade to file-level rather
than mis-attributing (#2276 ce-code-review autofix).
* test(group): lock named gin framework-route resolves by name not registrar (#2276)
Reviewer verified named Go handlers still resolve by name across the
widened queries; the HandleFunc path was already pinned, this adds the
framework-route (gin/echo) path with a DB + enclosing registrar whose
span covers the registration line, proving the emitted line never
diverts a named provider to its registrar via containment.
* test(group): end-to-end inline Go provider resolution against real LadybugDB (#2276)
Closes the validation gap that all prior coverage mocked CONTAINING_QUERY:
runs the real pipeline over a Go file with an inline http.HandleFunc
func-literal handler, persists into a real LadybugDB, and runs the
production HttpRouteExtractor against the real executor — proving the
emitted call-site line lands inside main()'s real 0-based span and yields
source_scan_resolved, not the file-level fallback.
* fix(test): use fs.mkdtemp to satisfy CodeQL insecure-temporary-file gate (#2276)
The new integration test created its temp base via a predictable
os.tmpdir()+name join, which CodeQL flags as js/insecure-temporary-file
(1 high). Switch to fs.mkdtemp for an atomic, randomly-named base dir.
* fix(group): anchor Go provider @handler to the trailing argument (#2276)
The widened framework-route and HandleFunc handler captures
(`[(identifier) (func_literal)] @handler`) were unanchored, so a variadic
middleware route `r.GET("/x", mw, func(){})` produced two provider
detections — one for the middleware identifier and one for the closure.
The contractId-only merge then kept the middleware detection and
mis-attributed the route to it (and the pre-existing `mw, namedHandler`
shape had the same defect), silently neutralizing the inline-handler
containment resolution from #2276.
Add a trailing tree-sitter anchor (`@handler .`) so the handler binds the
LAST argument of the call, leaving middleware args before it unconstrained.
Verified against tree-sitter-go: the multi-arg shapes now yield exactly one
detection (the real handler) while every 2-arg case is unchanged. Adds two
regression tests pinning that a middleware + inline closure resolves to its
containing function and a middleware + named handler resolves by name.
* test(group): cover FastAPI @router inline-handler containment (#2276)
The @router/APIRouter provider emit gained a call-site `line` in #2276 but
only the @app path was tested; the existing @router tests call
`extract(null, …)` so the resolver/containment path never ran for @router.
Add two tests mirroring the @app cases: a single-decorator @router handler
resolves to its function via source_scan_resolved (which fails if `line` is
dropped), and a multi-decorator one degrades to file-level.
* fix(group): treat synthetic 'route' label as anonymous in cross-trace (#2276)
After #2276 an unresolved file-scope Laravel closure emits name:null, so its
persisted symbolName falls back to 'handler' — which providerLabel already
anonymizes to '<contractId handler>'. But an unresolved named-controller
route still carries the synthetic 'route' placeholder, which the sentinel
did NOT cover, so group_trace/group_cross_impact rendered it as the literal
'route' while equivalent closures showed '<... handler>'.
'route' is only ever the synthetic Laravel placeholder (php.ts), never a
resolved handler name, so add it to the unresolved-generic sentinel set
alongside 'handler'/'fetch'. The resolved branch is untouched, so a real
symbol genuinely named 'route' still displays its name. Adds a cross-trace
test pinning the anonymized label.
* fix(group): gate Spring provider line on a present method name (#2276)
The Java/Kotlin Spring @*Mapping provider emits set `line` unconditionally
while the method name is typed string|null. The 'a named provider never
reaches containment' guarantee held only because the grammar always captures
a method name — the type did not enforce it. A (grammar-impossible) null name
would emit name:null + line and resolve by containment to the enclosing class
body instead of staying file-level.
Emit `line` only when the method name is truthy, so a nameless provider
degrades to file-level (the safe no-mis-attribution outcome). Behavior is
unchanged for every real Spring route (name is always present), but the
inertness is now enforced rather than incidental.
* feat(group): resolve cross-file named HTTP handlers via unique repo-wide lookup
U1 of #2275. When a provider's named handler is defined in a file other than its
route registration (e.g. router.get('/x', listUsers) with listUsers imported),
the registration file's symbols don't contain it, so resolution fell back to the
file-level boundary. Add a repo-wide name query (RESOLVE_BY_NAME_QUERY, the
label-union pattern from manifest-extractor) consulted only after the file-scoped
lookup misses, and honored ONLY when exactly one Function/Method/CodeElement
carries that name (zero/many → keep the file fallback, no wrong-symbol
attribution). Provider-only, cached by name. 4 unit tests; 743 group tests pass.
* test(bench): cross-file named handler scenario (end-to-end proof of #2275)
U2 of #2275. Adds a fifth bench scenario: a backend route whose handler
(listUsers) is imported from another file than its registration, with a frontend
consumer. Asserts the provider resolves to the handler via the repo-wide unique
name lookup (sym=listUsers, uid set) and that the cross-repo trace is symbol-
precise (no file-level fallback). verify.mjs now 12/12 on the real pipeline.
* fix(review): apply autofix feedback
ce-code-review (autofix) — no correctness/security findings; applied test-coverage
+ robustness fixes: repo-wide query throw -> empty (no exception); by-name lookup
cache fires once across same-named handlers; consumers never consult the repo-wide
lookup; same-file-wins now asserts the global path is bypassed; bench provider find
scoped by contractId; clarified the uniqueness-guard comment. 167 extractor tests.
* fix(group): tri-review fixes for cross-file handler resolution
Two-engine PR tri-review (Claude swarm+ce, Codex gpt-5.5 swarm+ce+adversarial)
on #2277. Correctness/security clean (injection refuted, bind-param). Fixes:
- Named-provider wrapper-attach (Codex swarm P1 + Claude ce-adversarial,
cross-engine): a named handler that fails both name lookups no longer falls
through to line-span containment, which attached the route to the enclosing
registrar (e.g. a setupRoutes() wrapper) instead of leaving it empty.
Containment now applies only to consumers and inline-arrow providers.
- CodeElement/ORM empty-file nodes (Claude ce-adversarial reproduced +
ce-maintainability): RESOLVE_BY_NAME_QUERY gains 'AND n.filePath <> ""' so a
handler name colliding with a synthetic ORM model node (orm.ts emits
filePath:'') neither resolves to an edge-less node nor inflates the uniqueness
count and masks the real handler; + a defensive empty-filePath guard in
resolveSymbolByNameUnique. Added LIMIT 2 (Codex swarm P3 + ce-maintainability)
to bound homonym materialization (count guard stays exact).
- Documented the aliased-import limitation (Codex adversarial): the route-site
identifier is the local alias, fix deferred to #2275 import narrowing.
- README expected verdict 9/9 -> 12/12 (Codex swarm+ce P3).
Tests: +3 (wrapper-no-attach, empty-filePath reject, empty-registration-file
resolves) covering the cross-engine gaps. 170 extractor / 748 group+integration
pass; bench 12/12 end-to-end.
* feat(group): import-pinned handler resolution (fixes deferred alias case)
Resolves the tri-review's deferred item: cross-file named handlers are now pinned
to their import's target module instead of resolved by name alone, so aliases and
names that collide with a local symbol resolve correctly.
- node.ts builds a local-binding -> {declared name, module} map from the file's
named imports; the express handler emits the DECLARED name + a handlerImport
{name, module} (HttpDetection gains the optional field).
- resolveDetectionSymbol gains an imported-handler rung: resolveImportedSymbol
pins to the import's target file via RESOLVE_IN_MODULE_QUERY
(n.name= AND filePath STARTS WITH the resolved module path), unique-match
only. An imported handler never uses file-scoped lookup (it is defined
elsewhere); on a module miss it falls back to a unique repo-wide name match on
the DECLARED name, then null. Relative imports only; bare/non-relative imports
keep the repo-wide fallback. Cached by (module-prefix, name).
- Closes the Codex-adversarial alias finding: import { listUsers as handleUsers }
+ an unrelated handleUsers no longer mis-resolves — the route resolves to the
imported listUsers in its module, and the alias is never looked up.
- Shared toResolvedSymbol helper (dedups the row->symbol + empty-filePath guard).
Tests: alias-resolves-to-declared-name + module-pin-resolves-ambiguous-name unit
tests; same-file-wins reworked to a genuinely LOCAL handler. Bench scenario 6
(aliased import with a decoy) proves it end-to-end. 172 extractor / 751
group+integration pass; bench 14/14.
* feat(group): import-pinned resolution for Python aliased handlers
Extends the JS/TS import-pinning to Python. The Python analog of express
router.get(path, handler) is Flask's imperative add_url_rule(view_func=...),
whose view is often an imported (aliased) symbol.
- New Flask add_url_rule provider pattern (path + view_func handler + methods;
default GET, methods=[...] honored). High Flask-specificity keeps false
positives low — unlike bare path()/Route(), which the plugin deliberately
leaves to graph Route nodes.
- buildPythonImportMap resolves 'from .mod import name as alias' (and plain
'from mod import name') to the declared name + raw module spec.
- resolveModuleBase generalized to two relative-import dialects: path-style
(JS './h/users') and dotted (Python '.handlers.users', '..pkg.users' — leading
dots are package levels). Bare/absolute imports keep the repo-wide fallback.
- Django stays graph-resolved (handlerSymbolId); FastAPI/Flask decorators stay
same-file (decorated function). This only adds the imperative imported-view
case Python lacked.
Tests: Flask aliased add_url_rule unit test (relative dotted module pinned, alias
never queried) + bench scenario 7 (end-to-end, 16/16). 173 extractor / 752
group+integration pass.
* fix(lang-kotlin): support `fun interface` extraction via tree-sitter-kotlin re-vendor
Vendored tree-sitter-kotlin@0.3.8 (fwcd) parsed `fun interface Foo` as an
ERROR node and dropped the declaration plus its abstract method, so functional
(SAM) interfaces were never extracted. The fix landed upstream in
fwcd/tree-sitter-kotlin#169 (closes#87), merged to main 2025-04-25, but is not
in any npm release (latest tag 0.3.8; main is the unreleased 0.4.0).
Re-vendor the grammar from the unreleased fwcd main commit c8ac3d26:
- refresh src/{parser.c,scanner.c,node-types.json,tree_sitter/*.h} and
bindings/node/index.js; bump the vendor version 0.3.8 -> 0.4.0; record the
pinned SHA + rationale in _vendoredBy and the vendor README.
- switch the prebuild workflow's kotlin registry kind 'npm' -> 'vendored' (the
fix is unreleased on npm, so prebuilds must build from the vendored C source,
like swift/dart/proto).
- add a hold to .github/vendored-grammars.json so the weekly auto-update
monitor does not strict-inequality-revert the pin to the broken npm 0.3.8
(isNewer compares 0.3.8 != 0.4.0).
- add 3 regression tests + a fixture asserting fun interfaces extract as
Interface nodes with their abstract methods, and that plain-interface
heritage still resolves.
Existing KOTLIN_QUERIES need no change: the new grammar models `fun interface`
as a class_declaration with an "interface" keyword child (plus an extra "fun"
modifier child), which the existing interface rule already matches. Full Kotlin
suite green against the new grammar (300 unit/cfg/resolver + 233 integration).
NOTE: prebuilds/ are intentionally not in this commit. The version bump
auto-triggers .github/workflows/build-tree-sitter-prebuilds.yml, which
regenerates all 6 platform binaries from the vendored source in a separate PR.
Until that lands, CI loads the committed 0.3.8 prebuild, so the new kotlin
tests are red and the grammar change is inert at runtime. Merge the prebuild PR
first or together.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(ci): count kotlin's vendored hold as a 0.25-readiness blocker
The kotlin `hold` added in the previous commit makes the tree-sitter
upgrade-readiness report count it as a blocker — the report treats every
held vendored grammar as frozen below a runtime upgrade (same as the
intentionally-pinned tree-sitter-cpp and the ABI-held tree-sitter-c),
"in-range ABI or not". So the report's blocker count goes 2 -> 3.
Update the hardcoded count in
test_issue_update_summary_regex_matches_current_report (and the
_render_report docstring) accordingly — exactly as that test instructs:
"if a grammar is added/removed or a pin/hold changes, update the expected
counts". kotlin's ABI (14) is in range; the hold is what flags it, with the
reason recorded in .github/vendored-grammars.json.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(ci): refresh kotlin baselines for the grammar bump
Two committed baselines pinned the pre-bump kotlin state and broke when the
grammar was re-vendored (0.3.8 -> 0.4.0):
- cli-commands.test.ts pinned the vendored kotlin package version at 0.3.8 ->
update to 0.4.0.
- bench/scope-capture/baselines.json: the new kotlin-fun-interface fixture joins
the lang-resolution/kotlin-* corpus AND the new grammar parses `fun interface`
as a class_declaration (not an ERROR node), so the capture fingerprint drifts.
Rebaselined to the NEW grammar's fingerprint (verified by building the vendored
parser.c against tree-sitter@0.21.1 and running measure.mjs --check); scaling
~0.83 (linear).
Like the fun-interface integration tests, the scope-capture --check passes only
once the regenerated prebuilds land; until then CI loads the committed 0.3.8
binary, so it stays red.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci(prebuilds): rebuild + commit grammar prebuilds into the PR on vendored-source change
build-tree-sitter-prebuilds.yml previously rebuilt a grammar's native prebuilds
only when its package.json VERSION bumped, and delivered them via a separate bot
PR. Now any change to the vendored grammar source re-cuts the prebuilds and they
ride into the same PR.
- Trigger on any build-affecting change under gitnexus/vendor/tree-sitter-*/**
(parser.c, grammar.js, binding.gyp, scanner, bindings), not just version bumps.
The prebuilds/ subtree is negated in the paths filter AND excluded from the
guard's source diff, so the bot's own prebuild commit can never retrigger the
workflow (no build -> commit -> build loop).
- The guard builds a grammar when its recorded version changed OR its vendored
source changed vs the PR base.
- Same-repo PRs get the rebuilt prebuilds committed straight onto their own head
branch (included in the SAME PR) via a non-force push that only adds a commit
on top of head. Manual dispatch still opens a fresh chore/ PR; fork PRs stay
artifacts-only (a bot cannot push into a fork branch).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci(prebuilds): deliver rebuilt prebuilds to fork PRs via a trusted workflow_run stage
A fork PR's producer run has a read-only token and no secrets, so it can build and
validate the prebuilds but can't commit them. Add the safe two-stage handoff that
mirrors the pr-autofix producer/publish split.
- build-tree-sitter-prebuilds.yml (untrusted producer): on a fork PR, upload a
pr-meta artifact (schema, pr_number, head_sha, head_ref, head_repo, base_repo)
alongside the prebuild artifacts. Values flow through env + jq, never
interpolated into a shell.
- commit-fork-prebuilds.yml (trusted, workflow_run): downloads ONLY the artifacts
(never executes fork code — it checks out the pinned HEAD SHA solely to add
files), allowlist-validates every metadata field, cross-checks identity against
the workflow_run authority (head_sha / head_repo / pr_number, via
commits/{sha}/pulls for forks), then pushes the prebuilds onto the fork head
branch with --force-with-lease + http.extraheader auth. No PAT: this works when
the contributor left "Allow edits by maintainers" on; on push failure it posts a
sticky comment telling them to enable it or commit the downloaded artifacts.
zizmor: allowlist commit-fork-prebuilds.yml's workflow_run dangerous-trigger with
the documented mitigation, matching the existing ci-report / pr-autofix entries.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(vendor): rebuild tree-sitter-kotlin prebuilds for the re-vendored fun-interface grammar
The fun-interface re-vendor changed vendor/tree-sitter-kotlin source but left main's
old (0.3.8) prebuilds in place, so all 6 platform binaries were stale relative to the
new parser. Replace them with the freshly cross-built + ABI-validated binaries from
build-tree-sitter-prebuilds run 28010841458 — each .node was require()-loaded and
parsed a snippet on its target platform-arch before upload.
This is the manual equivalent of the commit-fork-prebuilds.yml delivery, which can't
run for this fork PR until it lands on main.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(lang-kotlin): read extension-function receiverType from the re-vendored grammar's `receiver` field
The fun-interface re-vendor changed the kotlin AST: an extension function's
receiver is now a `receiver_type` exposed via a named `receiver` field, where the
old grammar emitted a bare user_type before the name. extractReceiverType only
matched the old shape, so receiverType came back null
(method-extraction.test.ts > Kotlin MethodExtractor > extracts receiverType).
Prefer the `receiver` field (unwrapping it), and keep the old child-scan — now
also recognizing `receiver_type` — as a fallback.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* refactor(group): extract shared resolveBridgeNeighbors from cross-impact
Lift the uid-filtered consumer<->provider ContractLink join (direction +
queryBridge + row normalization + confidence sort) out of runGroupImpact's
inline Phase-2 block into an exported resolveBridgeNeighbors helper. Behavior
is unchanged for impact; the helper becomes the single shared bridge join so
the upcoming cross-repo trace path never forks its own copy of the neighbor
Cypher. Empty uid sets short-circuit without a DB round-trip.
Adds direct coverage (real bridge via writeBridge/openBridgeDbReadOnly) for
both directions plus the empty-set and unknown-uid edges.
* feat(group): cross-repo trace stitching (groupTrace + runGroupTrace)
Add GroupService.groupTrace and the pure runGroupTrace engine that stitches
per-repo CALLS/HAS_METHOD trace segments across one ContractLink boundary in
the group bridge:
from --(local trace)--> consumer --(ContractLink)--> provider --(local trace)--> to
- Resolves from/to across all members (symbol node id == bridge symbolUid);
same-repo endpoints delegate to a single local trace with no crossing.
- Single boundary crossing (MAX_SUPPORTED_CROSS_DEPTH); deeper crossDepth is
clamped with a note, mirroring cross-impact.
- Discriminated GroupTraceResult union (ok|not_found|ambiguous|error) with
per-hop repo tags, a typed crossings[] entry, and centralized degraded-state
note constants (TRACE_NOTES). No .
- Trace-specific pair query (keeps BOTH crossing endpoints) lives in this
module; the uid-filtered neighbor join (resolveBridgeNeighbors) is reused
where it fits. ensureBridgeReady exported for reuse.
- New GroupToolPort methods (trace/resolveSymbol/pdgFlows) are optional so
existing port mocks keep type-checking; runGroupTrace guards on presence.
PDG enrichment is wired as an opt-in hook (enrichSegment) — the port method is
stubbed until U4. Covered by unit tests over a real bridge + mocked port.
* feat(group): route trace tool to groupTrace on @group syntax
Wire the cross-repo trace through the existing @group dispatch:
- callTool routes trace with an @-prefixed repo to callToolAtGroupRepo, which
forwards from/to/uid/file/maxDepth/includeTests plus the experimental
pdg/crossDepth flags to GroupService.groupTrace. Member path in @group/path
is advisory for trace (resolution is whole-group).
- Port gains trace/resolveSymbol/pdgFlows adapters. resolveSymbolForGroup wraps
the shared resolveSymbolCandidates so groupTrace can locate the member repo
and recover each endpoint node id (== bridge symbolUid). pdgFlowsForGroup is
a degraded stub here (call-level only); U4 implements the REACHING_DEF walk.
- trace tool schema documents the @group entry point, pdg, and crossDepth.
Single-repo trace is untouched. Covered by dispatch-routing tests (@group ->
groupTrace, non-group stays local) and tool-schema assertions.
* feat(group): opt-in PDG data-flow enrichment for cross-repo trace
Implement _pdgFlowsForGroupImpl: the real REACHING_DEF anchor walk that backs
the port pdgFlows adapter (replacing the U3 call-level stub). When pdg:true and
the segment repo has a flows PDG layer, the boundary-adjacent segments carry
their intra-procedural def->use hops:
- Anchors by the boundary symbol UID (precise; avoids the by-name ambiguity the
resolveBlockAnchor path can hit), then reuses the same span-anchored,
bind-param-only flows query as pdg_query (BasicBlock id-prefix + [start+1,
end+1] line window; no rel-property index, so the anchor IS the bound).
- Stays intra-procedural: data flow never crosses the repo boundary.
- pdgStampForMode probe: false -> available:false (degrade with note); the
trace stays ok. Any query failure is swallowed (enrichment is auxiliary).
Covered by runGroupTrace enrichment tests: dataFlow attached on opt-in,
degraded note when no layer, and no pdgFlows call when pdg is omitted.
* test(group): evaluation-first cross-repo trace e2e (two real indexes)
End-to-end gate for the cross-repo trace: stands up two real LadybugDB indexes
(consumer 'frontend' + provider 'backend'), a real ContractLink bridge, and a
real LocalBackend with both repos registered, then drives the public
callTool('trace', { repo: '@grp', pdg: true }) and asserts:
- the stitched checkout -> callUsers -(CONTRACT_LINK)-> handleUsers -> getUsers
path, each hop tagged with its member repo
- real REACHING_DEF data-flow enrichment of the consumer segment (userId)
- a degraded 'No PDG layer in app/backend' note (provider has no PDG layer)
- single-repo trace against one member is unchanged (no crossings)
Hand-persists the minimal real graph (deterministic; a full two-repo analyze is
heavier than this gate needs) and exercises real Cypher across
resolveSymbolCandidates, _traceImpl, the bridge pair query, and
_pdgFlowsForGroupImpl. Windows-skipped (describeReopen) and registered in the
cross-platform native-lbug set.
Scoped to a single @group call: opening bridge.lbug read-only a SECOND time in
one process currently fails (shared bridge open/close lifecycle, also affects
impact @group) — the pdg-omitted/clamp variants are unit-covered.
* docs(group): document cross-repo trace + PDG enrichment
ARCHITECTURE.md: trace is now group-aware; describe the @group cross-repo
stitch over a single ContractLink boundary (CONTRACT_LINK hop, crossings[],
crossDepth clamp), the opt-in experimental PDG REACHING_DEF enrichment of
boundary-adjacent segments, the symbolUid-grain join between the two stores,
and the deferred full cross-program (SDG-like) data flow. PIPELINE.md: add the
cross-trace consumer of the bridge with its pair-query rationale.
Does not touch gitnexus/CHANGELOG.md (release-owned).
* fix(review): apply autofix feedback
Apply safe_auto findings from ce-code-review (run 20260622-094243):
- local-backend.ts: drop (r: any) in _pdgFlowsForGroupImpl row map; coerce
hop line via Number() so a nullish LadybugDB cell can't surface NaN.
- tools.ts: advertise the forwarded param in the trace schema and add
crossDepth maximum:10 (schema now matches what groupTrace reads).
- cross-trace.ts: parallelize per-member resolveSymbol/resolveRepo with
order-preserving Promise.all (matches groupContext/groupQuery); add a note
when pdg:true is passed to a same-repo trace (PDG only enriches at a
cross-repo boundary).
- tests: remove / tighten (no-any rule).
Residual gated_auto/manual findings (unbounded crossing query + loop,
whole-file PDG widening on absent span, error-vs-no_path masking, top-level
try/catch parity, helper dedupe, branch-coverage gaps) are recorded in the run
artifact for the PR body.
* fix(group): skip CHECKPOINT on read-only bridge close so it can reopen
Root cause of the in-process bridge.lbug reopen failure (which broke repeated
@group impact/trace calls in a long-lived MCP server): closeBridgeDb issued
CHECKPOINT on EVERY handle, including read-only ones. A CHECKPOINT on a
read-only connection has nothing to flush but leaves a WAL/shadow lock artifact
that makes the next read-only open of the same path fail (openBridgeDbReadOnly
returns null -> 'Could not open bridge.lbug read-only'). Reproduced: open ->
query -> closeBridgeDb -> open again returned null only when the close ran
CHECKPOINT; a non-checkpoint close reopened fine, and the raw native
open/close cycle was never the problem.
Fix: tag read-only handles (BridgeHandle._readOnly, set by openBridgeDbReadOnly)
and skip CHECKPOINT for them in closeBridgeDb. Writable handles are unchanged
(they still flush before close). This is the shared bridge-db close path, so
impact @group benefits identically.
- Regression test in bridge-db.test.ts: open/query/close/open/query/open in one
process now succeeds.
- Re-enabled the second @group call in cross-trace-e2e.test.ts (was scoped to a
single call for this very limitation).
* fix(group): bring bridge-db close to parity with the core adapter safeClose
The bridge open/close cycle was less robust than the main graph DB's: closeBridgeDb
closed the connection/database but skipped the post-close steps the core adapter's
safeClose performs, so a rapid in-process reopen could race the OS handle release
(Windows) or an orphaned WAL sidecar. That gap is why the close-then-reopen tests
had to skip Windows.
closeBridgeDb now mirrors safeClose after closing the handle:
- waitForWindowsHandleRelease(dbPath): probe the file (+ .wal) until the residual
Windows lock clears, so the next open does not race (warns if the budget is
exhausted, matching the core adapter).
- finalizeLbugSidecarsAfterClose(dbPath): quarantine an orphaned WAL (shadow
missing) so the next open replays a consistent file.
Both helpers are the same ones safeClose uses (Windows-proven via the core adapter
CI), and the bridge read open already retries transient locks. Combined with the
read-only CHECKPOINT skip, the bridge reopen is now robust on every platform, so
the close-then-reopen tests run on all platforms (Windows CI exercises them via the
cross-platform subset). No write-path behavior change; Linux/macOS unaffected.
* fix(group): bound cross-repo crossing fan-out (LIMIT + segment memoization)
Address the top review residual: the bridge crossing query was unbounded and the
crossing-selection loop could run an O(2*N) sequential trace-BFS over every
ContractLink between a repo pair.
- CY_CROSSINGS_BETWEEN now ORDERs BY confidence DESC and LIMITs to
MAX_CROSSINGS_TO_TRY + 1; listCrossingsBetween slices to the cap and reports
truncation. Exceeding the cap surfaces a note (no silent truncation), keeping
the highest-confidence crossings. Aligns with the repo's anchored+LIMIT-bounded
query discipline (LadybugDB has no rel-property index).
- The home-repo segment (from -> consumer) depends only on the consumer uid and
the target-repo segment (provider -> to) only on the provider uid, so each is
memoized by that uid. Many crossings sharing a consumer/provider (one client
call linked to several providers) now cost one trace per distinct endpoint
instead of one per crossing. A consumer whose segment already failed is skipped
for every later crossing that shares it.
Test: two links sharing a consumer (first provider unreachable, second reachable)
assert the from->consumer segment is traced exactly once and the second crossing
wins.
* fix(group): restore Windows skip for bridge reopen tests; drop ineffective close-side probe
The previous commit flipped the bridge close-then-reopen tests to run on Windows,
betting that a close-side waitForWindowsHandleRelease + finalizeLbugSidecarsAfterClose
probe (mirroring the core adapter safeClose) would make the in-process reopen work
there. Windows CI proved otherwise: 4 writeBridge->openBridgeDbReadOnly tests fail
('expected null not to be null' — the read open returns null). The writable-close ->
read-open handoff plus writeBridge's atomic sidecar rename does not release the OS
file handle before the read open races, and the existing open-side LBUG_OPEN_RETRY
only retries lock-pattern errors, not the post-rename sidecar database-id mismatch.
macOS passes; the core adapter's own reopen also passes — this is bridge+Windows
specific.
- Revert itLbugReopen to the Windows skip (the pre-existing, correct state).
- Remove the close-side probe + finalize from closeBridgeDb: it did NOT close the
Windows gap, and reviewers flagged it for hot-path latency (finalize ran on every
close, all platforms) and safeClose duplication.
- KEEP the load-bearing fix — skipping CHECKPOINT on read-only handles — which fixed
the reproduced Linux/macOS in-process reopen artifact (the real bug).
Net: Linux/macOS repeated @group impact/trace works in-process; Windows in-process
bridge reopen remains a documented limitation (unchanged from before this PR).
* fix(group): surface degraded members + cap truncation; honest crossDepth schema
Address the cross-engine-corroborated tri-review findings (Codex + Claude):
- resolveAcrossMembers / runGroupTrace now track member repos that could NOT be
queried (resolveRepo or resolveSymbol threw) and, when the result is not_found,
attach a degraded-member note. A transient/corrupt member DB is no longer
silently reported as a clean 'symbol absent' not_found. (Codex B1+B3 + ce-reliability.)
- The cross-repo not_found now carries a programmatic truncated:true flag (and a
clearer suggestion) when the MAX_CROSSINGS_TO_TRY cap was hit, so a consumer can
distinguish 'no path' from 'cap may have hidden a connecting ContractLink'.
(Codex B3 + ce-adversarial + ce-api-contract.)
- trace tool schema: crossDepth maximum 10 -> 1 to match the implementation's
single-hop clamp (the schema previously advertised an unsupported 2-10 range).
(ce-api-contract, conf 100.)
Test: a member whose resolveSymbol throws yields not_found WITH a degraded note
naming the unreachable repo (if-free responder map).
* docs(group): clarify trace @group/memberPath is advisory (resolves all members)
Tri-review (Codex ce, conf 100) caught a doc/impl inconsistency: ARCHITECTURE.md
lumped trace with query/context/impact as honoring @group/memberPath member
scoping, but cross-repo trace resolves from/to across ALL members (the member
path is advisory). Clarify the behavior and point to from_uid/to_uid for
disambiguating same-named symbols across members.
* feat(group): file-level boundary fallback so cross-repo trace works on HTTP contracts
Benchmark (bench/cross-repo-trace/) running the REAL pipeline (runFullAnalysis
--pdg -> real syncGroup -> trace @group) found that cross-repo trace returned
not_found for real HTTP links even though sync built the correct ContractLinks:
HTTP (and other source-scan) contracts hardcode symbolUid:'' (http-route-extractor),
and both cross-trace AND cross-impact join crossings by Contract.symbolUid, which
never matches an empty uid. (Pre-existing — impact @group has the same gap.)
Fix: when a crossing's symbolUid is empty, fall back to the contract's FILE — if
the user's from/to resolves into the contract file, that endpoint anchors the
boundary. CY_CROSSINGS_BETWEEN now returns consumer/provider filePath; a crossing
is kept if it can be anchored by uid OR file on each side; a fileBoundaryFallback
note flags that the boundary is file-level, not symbol-precise. This makes the
common 'trace from=<calling fn> to=<handler fn>' case work end-to-end (verified:
fetchUsers -> listUsers stitches with a CONTRACT_LINK hop + PDG enrichment, 2/2).
Limits (documented in the bench README + the note): anonymous handlers have no
named target; when several contracts share files the file fallback may attach the
wrong contractId to a correct path. The proper upstream fix is to populate
symbolUid in the HTTP extraction (benefits impact too) — the bench is its gate.
Adds a unit test pinning the empty-symbolUid file-fallback stitch.
* fix(group): resolve HTTP contract symbolUid by containment (fixes cross-repo trace + impact)
Addresses the root cause behind the cross-repo trace file-fallback: HTTP
contracts hardcoded symbolUid:'' (http-route-extractor), so both cross-trace and
cross-impact — which join crossings on Contract.symbolUid — could not traverse
HTTP links. (Also found: the pre-existing graph-assisted resolution queried the
wrong edge, CONTAINS instead of DEFINES, so it never resolved a uid either.)
Now the extractor resolves each detection to a real symbol:
- HttpDetection carries the call-site line (node.ts sets it on every express/
fetch/axios/jquery/nest detection; express also captures the handler arg).
- resolveDetectionSymbol resolves the named handler first, else the innermost
Function/Method whose line span encloses the call (consumer = the function
containing the fetch; provider = the named/inline handler), over the correct
File-[DEFINES]->symbol edge. Base-tolerant (0- vs 1-based startLine).
- Wired into both source-scan and graph-assisted provider/consumer paths.
Verified end-to-end (bench/cross-repo-trace): all 4 contracts now carry real
uids, trace is symbol-precise (GET pair -> http::GET, POST -> http::POST, no
file-fallback note), and impact @group fans out (cross_repo_hits 0 -> 1). The
cross-trace file-level fallback remains as the secondary path for truly
anonymous handlers. Adds 2 containment unit tests; 738 group/integration pass.
Languages other than JS/TS still resolve providers by handler name; their
consumers fall through to the file fallback until their plugins set the line.
* fix(group): extend HTTP symbolUid containment to all languages + nested methods
Completes the symbolUid resolution across every bundled HTTP plugin: Python, Go,
PHP, Kotlin and Java now set the call-site line on their consumer (and Feign/
named) detections, so their HTTP contracts resolve to the containing function
the same way Node/TS already did.
Also generalizes the containment query: it now matches Function/Method/CodeElement
by filePath (UNION ALL) instead of File-[DEFINES]->symbol. The DEFINES edge only
reaches a file's TOP-LEVEL symbols, so methods nested in classes (Java/Kotlin —
File defines the class, the class defines the method) were invisible; matching by
filePath reaches them. Verified against a real index (LadybugDB supports the
UNION); JS/TS still fully symbol-precise (bench 2/2), 709 group tests pass.
Residual is now only the inherent case — a fully anonymous handler with no named
callee — which keeps the cross-trace file-level fallback.
* feat(group): destination trace — follow a consumer to an anonymous handler
Handles the one inherent residual: an anonymous route handler
(`router.get('/x', (req,res) => …)`) has no symbol node at all (the file holds
only a Const + PDG BasicBlocks), so it can never be named as a trace `to`.
Adds a DESTINATION TRACE: omit to/to_uid/to_file on an @group trace and
`trace from=<consumer>` follows the consumer's outgoing HTTP call across the
bridge and reports where it lands — by route + file:line, with a notes[] entry
flagging the handler as anonymous. Implemented as a new branch in runGroupTrace
(p.destination) backed by CY_CROSSINGS_FROM (all ContractLinks leaving the
consumer repo) + stitchToDestination; the provider endpoint is labelled
'<METHOD /path handler>' when its symbolName is a generic token/file basename.
The MCP routing already omitted an absent `to`, so only the schema docs changed.
parseTraceParams now treats a missing `to` as a destination trace instead of an
error. Verified end-to-end: anonymous fixture reports
'app/frontend:fetchUsers -> app/backend:<http::GET::/api/users handler>'; named
fixture lands at the real function. Adds 2 unit tests; 915 group tests pass.
* fix(group): tri-review fixes for cross-repo trace + symbolUid resolution
Two-engine tri-review (Claude swarm+ce + Codex GPT-5.5 swarm+ce+adversarial)
surfaced these; cross-engine-corroborated unless noted.
Correctness (P1, all four lanes): destination trace reported the WRONG endpoint
— an empty-uid consumer made trace(from->from) trivially succeed, so the highest-
confidence same-file crossing won regardless of which call `from` makes.
stitchToDestination now collects ALL connecting crossings, prefers symbol-precise
hits, and returns `ambiguous` (with candidates) when it cannot disambiguate.
Correctness (P1, Codex): resolveDetectionSymbol early-returned null when
d.line==null, blocking NAME resolution for named providers that set no line
(Spring/Go/etc.). Name resolution now runs first; only containment needs a line.
Correctness (P2): resolveContainingSymbol OR-ed `line` and `line-1`, which could
mis-pick a one-line sibling. It now probes the base-correct `line-1` first and
falls back to `line` only if nothing matches.
Correctness (Codex): anonymous Express handlers emitted name:'handler' and could
attach to an unrelated fn literally named `handler`. node.ts now emits name:null
for non-identifier handlers (containment-only).
Robustness: drop the first-symbol-in-file pickSymbolUid guess from the graph
consumer/provider paths (a wrong uid would win the contractId merge); remove the
dead CONTAINS_QUERY fallback (CONTAINS is File->Folder, never a symbol) + the now
-unused pickSymbolUid/handlerName; seed destination notes with degraded-member
notes so a successful trace still surfaces them; providerLabel takes providerUid
so a resolved fn named `handler` is not mislabeled anonymous, and only true file
basenames (known extensions) — not any dotted name — count as anonymous.
API contract: a single-repo trace with no `to` now returns an actionable error
(destination trace is @group-only) instead of "symbol 'undefined' not found".
Maintainability/tests: narrow asLocalTrace per-field (drop as-unknown-as); fix the
PR's lone as-any (vi.mocked); if-free e2e teardown; qualify the bench README.
Adds ambiguous-destination, anonymous-handler-no-false-name, and single-repo-no-to
tests; redirects graph mocks CONTAINS->UNION ALL. 918 group/integration pass.
* fix(group): carry degraded-member notes through SUCCESSFUL group traces
A reviewer (koriyoshi2041, PR #2269) correctly flagged that degraded-member
resolution was surfaced only on not_found, not on a successful ok result. Group
trace resolves names across ALL members, so an ok is 'unique among the members
we could query' — if a member that threw during resolveSymbol also holds from/to,
the real answer could be ambiguous. The destination path already seeded the note
(prior commit); this extends it to the same-repo and cross-repo success paths by
seeding the dispatch notes with degradedNotes([...fromRes.degraded, ...toRes.degraded]).
Adds a regression test: reg-be throws while a same-repo trace succeeds in reg-fe;
the ok result now carries the 'could not be queried' degraded note (app/backend).
* test(bench): cover all implemented cross-repo trace cases in one runner
Replace the single named-handler script with a self-contained verify.mjs that
generates each fixture inline and exercises every implemented end-to-end case
against the real analyze -> sync -> trace/impact pipeline, asserting PASS/FAIL
(exit non-zero on failure). 10 checks across 4 scenarios:
- named handlers: 4/4 symbolUid resolved; symbol-precise GET vs POST crossing
selection; destination trace lands at the named handler.
- anonymous handler: empty symbolUid; destination trace reports it by route with
the anonymous note.
- impact @group fan-out (cross_repo_hits >= 1).
- multi-language (Python Flask + requests): link built, cross-repo trace stitches,
and the file-level boundary fallback is exercised when the provider has no uid.
Ambiguous-destination and degraded-member paths need synthetic inputs the real
analyzer cannot produce, so they stay in the unit suite (documented in the README
+ script header). Removes verify-named.mjs + fixtures-named/ (folded inline).
* test(group): pin destination degraded-success + precise-tier ambiguity
Adds the two regression guards koriyoshi2041 requested on PR #2269 after the
degraded-on-success fix:
- destination trace success with a degraded member: reg-fe resolves from and
follows the link to an anonymous handler while reg-be throws; the ok result
carries the anonymous endpoint AND the 'could not be queried' degraded note, so
the no-to path stays aligned with explicit to traces.
- multiple PRECISE destination hits: one from reaches two consumers with resolved
uids linked to different routes; the result is ambiguous (role: to) with both
route candidates. Distinct from the existing file-level ambiguous test, this
pins the stronger precise tier against a future change silently picking the
highest-confidence destination.
Both already pass against current behavior; 716 group tests pass.
* feat(routes): resolve + persist handler symbol on Route nodes (#2138 part 2, WIP)
Part 2 groundwork for #2138: give the graph-assisted HTTP provider path the
handler symbol directly, so it no longer re-parses source to recover the
handler name. (The remaining parse-skip in extract() + a call-count benchmark
land in a follow-up commit.)
- ExtractedDecoratorRoute gains `handlerName`; the Spring extractor captures
the decorated method's name (the method_declaration node is in hand).
- New `resolveRouteHandlerSymbols` (call-processor) resolves each route's
handler to a real symbol UID, keyed by normalized route URL — Laravel
framework routes (controller + method) and decorator routes (Spring/FastAPI)
both reduce to `(filePath, name) -> nodeId`. Threaded through the parse phase
onto `ParseOutput.routeHandlerSymbols`.
- routes phase stamps `Route.handlerSymbolId`; persisted end-to-end (schema +
Route CSV row + getCopyQuery COPY columns), mirroring Part 1's `method`.
- HttpRouteExtractor: `HANDLES_ROUTE_QUERY` returns `handlerSymbolId`;
`extractProvidersGraph` uses it as the authoritative symbol and SKIPS
`getDetections()` for resolved rows (CONTAINS is a cheap graph lookup for the
display name only — no tree-sitter parse). Fully backward compatible: an
unresolved/old-index route with no `handlerSymbolId` keeps the source-scan
fallback.
- Extracted `normalizeExtractedRoutePath` to `route-extractors/route-path.ts`
(shared by routes phase + resolver without an import cycle).
- SCHEMA_BUMP 6->7 (ParseWorkerResult gained `handlerName`); regenerated the
emit-persistence byte-identity baseline (route.csv header gained two columns).
- Tests: Spring pipeline asserts the Route node carries a handlerSymbolId
resolving to the handler method; extractor fast-path test proves the handler
resolves with zero source detections.
Refs #2138
* perf(group/http): skip source parse for graph-covered route files (#2138 Part 2)
Builds on the persisted Route.handlerSymbolId (U0–U3a). When a file's
HANDLES_ROUTE rows all resolve a handler symbol AND its language plugin
declares routeCoverage: 'complete' (Java/Python/PHP), the graph is
authoritative for that file's providers, so the source scan + tree-sitter
parse can be skipped — the scan would only re-discover routes the graph
already has. This is the measurable parse reduction #2167 could not show.
Consumer safety: routeCoverage: 'complete' asserts *provider* Route-node
completeness only. The scan() of those same languages also emits consumer
detections (RestTemplate/WebClient/OkHttp/Feign, Guzzle/Http::,
requests/httpx), and ingestion's FETCHES edges are JS/TS-only — so the
graph cannot back up server-side consumers. A provider-covered controller
that also calls out would otherwise lose its consumer contract. Guarded by
a cheap, parse-free text gate.
- types: HttpLanguagePlugin gains
- routeCoverage?: 'complete' | 'partial' (default 'partial')
- hasConsumerSignals?(content): false only when the raw source provably
has no outbound-HTTP call this plugin detects (conservative).
- java/python/php: mark routeCoverage 'complete' + implement
hasConsumerSignals with a token regex over their consumer idioms.
- http-route-extractor: run the graph provider pass first to build a
coveredFiles set; then keep a file covered only when
hasConsumerSignals(content) === false (read via readSafe, no parse).
scanFiles = files not covered → drives collectProjectDetections + both
source scans. Fail-open per file: any unresolved row, a 'partial'
language, a positive consumer signal, a missing hook, or an unreadable
file leaves the file in the scan set. The orchestrator names no
languages — token knowledge stays in the plugins.
Net: pure-provider controllers skip the parse (the win); controllers that
also call out are still parsed (no consumer loss); partial-coverage
languages and graph-less runs are unchanged.
- test: route-parse-skip integration test spies the real parseSourceSafe to
COUNT parses over a temp repo of Spring controllers with a mock DB —
baseline (every file parsed), fully-covered (0 parses), mixed (unresolved
file falls back, resolved stays skipped), and provider+consumer (a covered
controller that also calls restTemplate is parsed; its consumer contract
survives).
* fix(group/http): cover Spring HTTP Interface @*Exchange in Java consumer-signal gate
#2254 (merged) added Spring 6 HTTP Interface `@(Get|...)Exchange` /
`@HttpExchange` as a new Java *consumer* idiom. The #2138 parse-skip
consumer-safety gate must recognize it, or a provider-covered file carrying
an `@GetExchange` could be parse-skipped and lose that consumer contract.
Add `Exchange` to JAVA_HTTP_PLUGIN.hasConsumerSignals (conservative; also
matches `restTemplate.exchange(`).
* style(group/http): prettier formatting for #2138 Part 2 files
* style(ingestion): prettier formatting for call-processor.ts (#2138 Part 2)
* fix(group/http): P1 (Java over-claim) + P2 (handler mis-attribution) on top of #2268 (#2138 Part 2)
Re-applied on the maintainer's #2268 (expanded Java/Kotlin consumer
extraction) base.
P1 — `routeCoverage: 'complete'` over-claimed for Java: the graph provider
set is a strict subset of the group scan (array-form `@GetMapping({...})`,
interface-inherited routes, same-URL multi-verb have no graph Route node),
so parse-skip could drop those group-only providers.
- java/python → default 'partial' (always source-scanned). Java flips to
'complete' only once ingestion provider extraction matches the group scan
(a separate follow-up). Python was a no-op anyway (no handlerName resolved);
'complete' was a latent trap. PHP stays 'complete' (Laravel ingestion ⊇ the
group scan, the one language the skip engages for).
- python hasConsumerSignals widened to a true superset of scan() (uri=/url=
wrapper, aiohttp, urllib). Java's gate already covers #2268's consumer set
(same receivers; the @*Exchange token is present).
P2 — resolveRouteHandlerSymbols: reserve the URL slot on first encounter even
when unresolved (mirrors addRoute first-writer-wins, so a later same-URL route
can't stamp the node-winner's slot); refuse to guess on an ambiguous same-name
lookup (exactly one match → use it; zero/many → fail-open, never a wrong
handler). The cross-source case (filesystem route winning a URL a framework
route also normalizes to) is unchanged — the resolver never receives
filesystem routes — and stays fail-open.
Tests:
- route-parse-skip rewritten: the parse-skip win is proven on PHP (fully
covered → 0 parses; mixed fallback; consumer-covered file still parsed), plus
three Java P1 regression guards (array-form / interface-inherited / multi-verb)
asserting the group-only routes survive — verified they go red if Java is
flipped back to 'complete'.
- resolve-route-handler-symbols: direct unit tests (the fn had none) — unique
resolve, ambiguous/unknown fail-open, same-URL reservation, first-writer-wins.
- http-consumer-signals: each plugin's hasConsumerSignals is a superset of its
scan() consumer idioms; pure providers return false.
- route-handler-symbol-roundtrip: real-LadybugDB CSV→COPY→query for
Route.handlerSymbolId.
---------
Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* feat(group): extract RestTemplate URI.create(...) static paths (Java)
Widen the RestTemplate query @path captures to (_) and resolve
URI.create("/x") arguments via a new extractStaticPathExpression
helper. Variable-bound paths stay unresolved (no consumer).
* feat(group): resolve RestTemplate UriComponentsBuilder chains (Java)
Add appendPath + recursive extractUriComponentsBuilderPath (fromPath/
fromUriString/fromHttpUrl seeds, path/pathSegment append, build/query
passthrough). Host-bearing seeds are normalized downstream. Non-literal
segments stay unresolved.
* feat(group): infer OkHttp request verb from builder chain (Java)
Walk up from the matched .url(...) call to the sibling verb helper
(.post()/.method("X")), defaulting to GET. Variable-bound verbs stay
GET. Re-document the Kotlin OkHttp GET-default pin as an accepted
Java/Kotlin asymmetry (Kotlin verb-walk is a tracked follow-up).
* feat(group): Java HttpClient HEAD, .method("X"), and default-GET
Add HEAD to the verb-helper regex and two dedicated pattern families:
.method("VERB", body) (covers PATCH) and a bare .build() defaulting to
GET. The three families terminate at distinct calls, so each chain
matches exactly one (no double-emit); variable-bound verbs stay unresolved.
* feat(group): extract fully-qualified Java route annotations
Widen JAVA_ROUTE_ANNOTATION_PATTERNS @ann to match scoped_identifier
(predicate-free node-type change), normalizing to the trailing segment
via simpleName in the scan loop. Snapshot-verified: only previously
unmatched FQN annotations gain routes; existing contracts unchanged.
Brings Java to FQN parity with the Kotlin plugin (closes#2254 limitation).
* refactor(group): reuse simpleName helper in hasAnnotation
Drop the inline split('.').pop() (which shadowed the new module-level
simpleName) and call the helper. Note appendPath's deliberate divergence
from the shared joinPath so it is not accidentally unified.
* docs(test): fix reversed Java/Kotlin OkHttp asymmetry comment
The comment stated .kt->POST / .java->GET; it is the opposite — Java
infers the verb (inferOkHttpMethod) so .java emits POST, while Kotlin
still defaults so .kt emits GET. Aligns the prose with the assertion
below it.
* docs(group): correct stale kotlin.ts OkHttp parity comment
The comment claimed Kotlin's GET-default 'mirrors java.ts:OK_HTTP_PATTERNS'
and is 'the same trade-off Java has accepted'. The Java plugin now infers
the verb (inferOkHttpMethod), so this is now a documented Java/Kotlin
asymmetry; the Kotlin verb-walk is the tracked follow-up.
* test(group): pin OkHttp default-GET branch on a distinct path
The bare `.build()` (no verb call) now uses /api/bare-build and is
asserted individually, so the default-GET branch of inferOkHttpMethod
is no longer masked by the explicit .get() case collapsing into the
same {param} slot.
* fix(group): skip OkHttp emission for variable-bound .method(verb)
inferOkHttpMethod now returns string|null: an explicit .method(verb, …)
with a non-literal verb returns null and the loop skips it, instead of
asserting a wrong GET contract. A bare .url().build() with no verb call
still defaults to GET (OkHttp's real default). Matches WebClient
long-form, which also skips variable-bound verbs.
* fix(group): strip query from UriComponentsBuilder seed literal
A query baked into the seed (fromUriString("/base?x=1")) was returned
verbatim, so a later .path("/sub") glued onto it (/base?x=1/sub) and
normalizeHttpPath truncated the tail at ? to /base. Strip ?query at the
seed so .path() appends to a clean base → /base/sub. Host prefixes are
preserved and stripped downstream by normalizeConsumerPath.
* fix(group): add recursion depth guard to extractUriComponentsBuilderPath
The recursive builder-chain walk was unbounded; a pathological or
machine-generated chain could overflow the stack. Cap recursion at
MAX_BUILDER_DEPTH (100) and return null past it — consistent with the
project's other AST-depth guards.
* docs(group): document accepted FQN simple-name collision trade-off
The route discriminator matches on the trailing annotation segment, so a
non-Spring annotation sharing a route name (@com.evil.GetMapping) is
treated as a route — the same trade-off hasAnnotation makes and the
intended Kotlin parity. Note why package-origin gating is deliberately
not added.
* refactor(group): extract static-path helpers to java-static-path.ts
Move the URI.create / UriComponentsBuilder resolution helpers
(methodInvocation*, firstLiteralArgument, appendPath, extractUri*,
extractStaticPathExpression) out of java.ts (back under ~1000 lines).
java.ts imports the four it consumes; inferOkHttpMethod stays. Pure
move, behavior-preserving — full group suite unchanged.
* fix(group): walk builder chain for Java HttpClient verb (#2268)
Replace the three rigid JAVA_HTTP_CLIENT_* pattern families with one
.uri()-anchored query plus inferHttpClientMethod, which walks up the
fluent chain for the verb (mirroring inferOkHttpMethod). The walk is
transparent to intervening .header()/.timeout()/.version() calls, so a
header/timeout hop before the terminal no longer silently drops the
consumer contract.
Relocate both verb-walks onto a shared inferBuilderVerb in
java-static-path.ts and de-export the now-internal methodInvocation*
primitives; java.ts drops 1015 -> 910 lines.
* fix(group): append UriComponentsBuilder .path() verbatim (#2268)
Spring's UriComponentsBuilder.path(p) appends p as-is without inserting
a slash (then collapses duplicate slashes), unlike .pathSegment() which
slash-joins. The resolver used the always-one-slash appendPath for both,
so fromPath("/api").path("users") resolved to /api/users instead of
Spring's /apiusers. Switch the .path() branch to verbatim append plus a
colon-aware duplicate-slash collapse (preserving a host seed's ://);
.pathSegment() keeps appendPath.
* fix(group): skip empty-string verb literal in builder verb-walk (#2268)
`.method("", body)` produced a malformed `http::::/path` consumer:
unquoteLiteral('""') returns "" (not null), so the `=== null` guard
let an empty method through. Treat a falsy literal verb as unresolvable
(return null from the shared inferBuilderVerb) and switch the OkHttp and
HttpClient emission guards to falsiness, so an empty verb skips like a
variable-bound one.
* test(group): harden Java HTTP consumer coverage (#2268)
Add coverage beyond the tri-review findings: a count guard on the
UriComponentsBuilder query-seed test (so a double-emit can't slip past
the two find assertions), an exchange()+UriComponentsBuilder end-to-end
case (the widened (_) @path exchange capture was only covered with
URI.create), and an HttpClient .method("REPORT") custom-verb
pass-through pin.
* docs(group): document pre-path builder rigidity + fix stale refs (#2268)
Document the OkHttp pre-.url() limitation (a builder call before .url()
is missed) at OK_HTTP_PATTERNS, cross-referencing the Java-HttpClient
pre-.uri() dual the verb-walk rewrite leaves in place — so neither
comment overclaims that the chain is walked before the path call. Update
the now-stale 'inferOkHttpMethod in java.ts' references in kotlin.ts and
the test to point at java-static-path.ts after the relocation.
* feat(group): match Java HTTP consumer chains with a pre-path builder call (#2268)
The OkHttp .url() and HttpClient .uri() queries required the path call to
sit directly on the construction, so a builder call BEFORE it —
new Request.Builder().addHeader(...).url(...) or
HttpRequest.newBuilder().version(v).uri(...) — silently dropped the
consumer contract. Match the path call on any receiver and re-impose the
framework anchor in JS (okHttpUrlRootsAtBuilder / httpClientUriRootsAtNewBuilder:
the chain must root at new Request.Builder() / HttpRequest.newBuilder()), so a
preceding call is captured while an unrelated .url()/.uri() is rejected. The
verb-walk now scans the whole chain, so a verb set before the path call also
resolves. Also extract the HttpRequest.newBuilder(URI.create(...)) constructor-arg
form (skipped when a later .uri() overrides it). Resolves the deferred
follow-ups from the round-2 tri-review.
* feat(group): Kotlin OkHttp verb-walk parity with Java (#2268)
The Kotlin OkHttp consumer always emitted GET while the Java side walks
the builder chain to recover the verb — a documented Java/Kotlin
asymmetry. Mirror the verb-walk into kotlin.ts, adapted to the
tree-sitter-kotlin call_expression/navigation_expression grammar: match
.url("literal") on any receiver, gate to chains rooting at
Request.Builder() (kotlinUrlRootsAtRequestBuilder), and scan the whole
chain for the verb (inferKotlinOkHttpMethod — last-wins, null-skip for a
variable/empty .method(verb), resolves a named-argument .method(method="X")).
This brings .kt to full parity with .java — verb inference, a builder
call before .url(), and verb-before-url — pinned by two new Java<->Kotlin
parity-harness rows. Flips the former GET-default asymmetry test.
* [+] Add django route discovery to create cross-link for multi-repo
* [+] Update ingestion
* [~] Fix bugs and abstraction violation
* feat(python-http): add keyword url= and variable propagation for consumer detection
- Add REQUESTS_KEYWORD_URL_PATTERNS for requests.get(url='...') keyword args
- Add WRAPPER_URI_PATTERNS for generic wrapper.fetch(uri='...') calls
- Add WRAPPER_URI_VAR_PATTERNS + buildLocalStringMap for uri=variable propagation
- Add LOCAL_STRING_ASSIGNMENTS to track uri='...' assignments
- Wire both direct-string and variable-propagation loops in scan()
- Add normalizeConsumerPath() helper
Note: Automatic cross-link detection remains limited for runtime-computed URLs
(URLs built via .format(), string concat, or module constants). Manual
manifest links needed for known cross-repo contracts.
* [+] add extract uri and url keywork pattern for request http
* feat(python-http): add variable propagation for uri=/url= consumer patterns
Re-add LOCAL_STRING_ASSIGNMENTS, WRAPPER_URI_VAR_PATTERNS,
buildLocalStringMap(), and normalizeConsumerPath() lost during
cherry-pick merge of upstream keyword-URL commit.
Together with the upstream WRAPPER_URI_PATTERNS and
REQUESTS_KEYWORD_URL_PATTERNS, we now detect:
- requests.get(url='literal') keyword args
- wrapper.fetch(uri='literal') keyword args
- wrapper.fetch(uri=variable) where variable was assigned a string literal
* fix(group): discover Django roots relative to manage.py dir + multi-project (#1836 R1)
A Django project not at the repo root (e.g. backend/manage.py) discovered
zero routes: the settings module path was resolved repo-root-relative only,
so backend/myproj/settings.py was never found and discovery returned null.
Resolve settings, star-imported base settings, ROOT_URLCONF, and the root
urls.py against the manage.py's own directory first, then the repo root
(resolvedSettingsPath is now project-dir-aware so relative imports anchor
correctly). Iterate every manage.py so a monorepo with several Django
projects yields each project's root — the provider hook becomes plural
(discoverRootRouteFiles → string[]) and the main-thread pass loops over all
roots (inner-scoped continues, parser hoisted once per language).
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(group): remove dead code in Django root discovery (#1836 R9)
- Collapse the identical if/else in extractStarImports to one push.
- Drop the unreachable baseModule.startsWith('.') branch (baseModule is
always a resolved slash-path or a bare absolute module — never dot-prefixed).
- Import DjangoFileReader from django.ts instead of re-declaring the type.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): walk Django includes once per prefix, not per file (#1836 R2)
The include() recursion guard was keyed on file path alone and shared across
the whole walk, so a urlconf included under two prefixes (a "diamond" — the
same app mounted at /v1/ and /v2/) emitted routes for only the first mount.
Key the guard on (resolvedFilePath, accumulatedPrefix) at all three sites
(function entry, path()-wrapped include, bare include) so a file reached
under two distinct prefixes is walked once per prefix while a genuine cycle
(same file + same prefix) still terminates — null/'' prefixes collapse to one
key so a no-prefix re-entry is treated as a cycle. MAX_INCLUDE_DEPTH remains
the backstop. Adds diamond + self-include-cycle tests.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): extract Django routes from non-list urlpatterns (#1836 R3)
findUrlpatternsLists only accepted a list-literal RHS, so common shapes
yielded zero routes: concatenation (urlpatterns = a + b), wrapper calls
(format_suffix_patterns([...]), i18n_patterns, staticfiles_urlpatterns), and
tuples.
Add collectUrlpatternContainers to descend binary_operator operands, known
wrapper-call list arguments, and tuples. Inherently-dynamic forms (DRF
router.urls, comprehensions, bare names) still yield nothing but now emit a
debug log so the silent-zero case is observable rather than mysterious.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(group): thread Django parser explicitly, drop module singleton (#1836 R4)
extractDjangoRoutes relied on a module-level _djangoParser set via
setDjangoParser before each call — hidden state that would break if a second
language ever used the include re-parse path, and an easy-to-forget contract.
Pass the tree-sitter parser as an explicit parameter of extractDjangoRoutes
(the extractRoutes provider hook already receives it) and delete the global
plus its setter. The Python provider wires it directly; tests pass the parser
in place of the removed setDjangoParser() call.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ingestion): isolate a throwing extractRoutes in the cross-file route pass (#1836 R5)
The main-thread cross-file route pass called provider.extractRoutes without a
guard, so a throw (e.g. a future grammar edge case in the include() walk) would
propagate out of the parse phase and abort the entire analyze — unlike the
worker, which isolates per-file failures.
Wrap the per-root extractRoutes call in try/catch that logs a warning and
continues to the next root. Export extractCrossFileRoutes and add a unit test
driving a stub provider whose extractRoutes throws, asserting the pass returns
[] and does not propagate.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(ingestion): bucket only route-capable languages in cross-file pass (#1836 R6)
extractCrossFileRoutes runs in the deferred band on every analyze (incl. warm
all-cache-hit runs). It now derives the set of languages whose provider exposes
the cross-file route hooks once, returns early if none do, and buckets only
those languages' paths — so a non-framework repo no longer pays to bucket the
languages it doesn't use here.
Route results are intentionally not persisted across runs, so a Django repo
still re-derives its routes each analyze; documented inline that cross-run
route caching is a deliberate follow-up rather than implemented here.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style(group): prettier-format http-patterns/python.ts (#1836 R7)
The file was not formatted to the root .prettierrc (the consumer-path
normalizer used single-line try/catch and method chains), so the CI
quality/format check (`prettier --check .`) failed. Reflow only — no logic
change (`git diff -w` confines the change to normalizeConsumerPath's layout).
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): dedup Python URI detections by byte offset, not line arithmetic (#1836 R8)
The wrapper-URI dedup key was lineNum*1000+methodRow, which can collide for
distinct calls in files over 1000 lines (carry into the row term) and can
fail to dedup a genuine duplicate when a node straddles a line boundary.
Key on node byte offsets (`${pathNode.startIndex}:${methodNode.startIndex}`),
matching the sibling seenVarDetections dedup a few lines below.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(ingestion): end-to-end Django cross-file route extraction (#1836 R10)
Adds an integration test that runs runPipelineFromRepo against a Django
fixture whose project lives under backend/, asserting the resulting Route
graph nodes (/health, /api/items, /api/items/<int:pk>). This exercises the
previously-untested main-thread orchestration glue (discovery → parse →
extractRoutes → allExtractedRoutes → Route nodes) and, because the project is
in a subdirectory, regresses the subdir-discovery fix (R1) — a repo-root-only
resolver would discover nothing and emit zero Route nodes.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): anchor Django include() resolution at the project root (#1836 review F1)
resolveIncludedFile tried the bare repo-root candidate (app/urls.py) before the
project-relative one, so in a monorepo with both a repo-root app/ and a
backend/ Django project that also has an app/, include('app.urls') from the
backend project resolved to the WRONG service's routes.
Probe up-tree from the root urls.py for the nearest manage.py (the Django
project root / sys.path entry) and try that-anchored candidate first. Absolute
module paths like `app.urls` now resolve to <projectRoot>/app/urls.py
unambiguously. When no manage.py is reachable (e.g. unit tests with a urls-only
reader) the prior strategy order is preserved. Adds a monorepo wrong-app test.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): drop bogus Django provider source-scan, use graph routes (#1836 review F2)
The DJANGO_PATH_PATTERNS / DJANGO_URL_PATTERNS source scan emitted an HTTP
provider contract for every path()/re_path()/url() string literal, without
checking it was inside urlpatterns, without skipping include() mount points,
and without composing the include() prefix across files. For
`path('api/', include('app.urls'))` + child `path('items/', view)` it emitted
providers for `/api` (a mount, not a route) and `/items` (un-prefixed) — which
survived the exact-contract-ID dedup alongside the correct graph route
`/api/items`, polluting cross-repo matching with false providers.
Remove the Django provider patterns and their scan blocks. Django provider
contracts come from the graph Route nodes, which the ingestion route extractor
builds with includes already composed (and now correctly, per the other fixes).
Python HTTP *consumer* patterns (requests/wrapper) are unaffected.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): match method-agnostic Django providers to any-method consumers (#1836 review F3)
Django function views are method-agnostic, so extractDjangoRoutes emits
httpMethod '*'. That '*' was dropped by normalizeRouteMethod and then defaulted
to GET by the contract extractor, while the matcher only expanded wildcard
*consumers* — so a `POST /api/items` consumer never matched the Django
provider that was silently narrowed to GET.
- routes.ts: preserve '*' as a method-agnostic marker on the Route node, so the
contract layer emits a wildcard provider (http::*::path) instead of GET.
- matching.ts: make findMatchingKeys symmetric — a specific-method consumer
now matches an exact-method provider OR a wildcard (http::*::) provider on the
same path, mirroring the existing wildcard-consumer expansion.
Co-Authored-By: HuyNguyenDinh <61400397+HuyNguyenDinh@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Dinh Huy <huynd86@fpt.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(group): Kotlin Spring HTTP consumer extraction + provider parity with Java
Brings the Kotlin group/contract HTTP extractor up to parity with Java for
inter-service contract detection, and unifies the language-agnostic consumer
logic so it is not duplicated.
Consumers (new for Kotlin):
- @FeignClient interface @(Get|...)Mapping methods are emitted as OpenFeign
consumers (a remote call), not providers — previously mis-classified because
tree-sitter-kotlin models an interface as a class_declaration.
- Spring 6 HTTP Interface @(Get|...)Exchange (with optional class-level
@HttpExchange(url) prefix) — added for BOTH Java and Kotlin.
- Native OpenFeign @RequestLine, gated to interfaces (Feign proxies are
interfaces only), mirroring java.ts's findEnclosingInterface check.
Providers (Kotlin parity with java.ts scanSpringProject):
- A @(Get|...)Mapping on a non-Feign interface is a route *contract*, not a
served route; it is skipped in scan() so the implementing controller is the
sole provider (Java drops these implicitly via interface_declaration).
- scanProject inherits interface routes onto the implementing class, gated on
the class being a @RestController/@Controller (kotlinClassIsController handles
both the attached `modifiers` shape and the detached leading-arg-form
prefix_expression shape) so non-controller implementers don't emit phantom
providers.
Shared module:
- New spring-consumer-shared.ts holds the language-agnostic primitives
(REST_TEMPLATE_/WEB_CLIENT_/EXCHANGE verb maps, joinPath, parseRequestLine,
framework + confidence constants); java.ts and kotlin.ts both import it.
Array-of-paths (both languages):
- Route/Feign/Exchange annotation paths are `String[]`; a multi-element array
registers the route under EVERY element. The class/Feign/HttpExchange prefix
maps now accumulate all elements (were last-write-wins) and emission
cross-products prefixes × method paths, so `@RequestMapping(["/a","/b"])` +
`@GetMapping(["/x","/y"])` yields all four contract IDs. Array form is matched
via a predicate-free alternation over Kotlin `collection_literal` / Java
`element_value_array_initializer`.
Tests: comprehensive Java + Kotlin cases incl. consumer-vs-provider
classification, @*Exchange, @RequestLine (interface-only + plain-interface),
interface-based controller inheritance, non-controller negative case, detached
@RestController, single- and multi-element array paths (method-level and
class-prefix cross-product). 109 http-route + group tests pass; tsc/eslint/
prettier clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(group): apply @RequestMapping prefix to Kotlin @RequestLine consumers (#2254 P2)
A @RequestLine method on an interface with a class-level @RequestMapping
prefix but no @FeignClient(path) dropped the prefix in Kotlin while Java
applied it (java.ts merges the fallback into feignPrefixByInterfaceId).
Mirror the feignPrefixByClassId ?? prefixByClassId ?? [''] chain already
used by the @GetMapping-in-Feign path. Adds Kotlin twins for the
@RequestMapping-prefix and @FeignClient(path)-wins cases.
* fix(group): accept named-arg Kotlin @RequestLine(value=...) (#2254 P2)
The positional pattern's '.' anchor only matched @RequestLine("VERB /x"),
silently dropping the named @RequestLine(value = "VERB /x") form that
java.ts accepts. Add a dedicated named pattern constrained to #eq? @key
"value" (Java parity: non-value keys stay dropped). Adds Kotlin twins
for the named-value and non-value-key cases.
* fix(group): resolve Kotlin FQN annotations/supertypes by trailing segment (#2254)
A fully-qualified @org…RestController / supertype : a.b.Api parses to a
user_type with one type_identifier per dotted segment; kotlinAnnotationName
and collectKotlinSupertypes took the FIRST ("org"/"a"), so FQN controllers
were not recognised and FQN supertypes never matched their interface. Take
the trailing segment. Adds FQN controller + FQN supertype inheritance twins.
* refactor(group): remove dead prefix_expression branch in kotlinClassIsController (#2254)
AST probe (bare and realistic package+constructor forms) confirms the
arg-form @RestController("bean") attaches under the class `modifiers` as an
annotation/constructor_invocation, caught by the modifiers loop — the
prefix_expression sibling branch was unreachable. Remove it and correct the
false grammar comments (source + the arg-form test). The existing arg-form
test stays green via the modifiers branch, confirming no behavior change.
* feat(group): support Kotlin arrayOf(...) annotation arrays (#2254 P3)
arrayOf("/a","/b") (the explicit String[] form) parses to a call_expression,
not a string_literal/collection_literal, so it was missed across all five
annotation-array families. Add dedicated arrayOf query patterns (positional +
named) per family via a shared arrayOfArg fragment — kept out of the existing
[(string_literal) (collection_literal …)] alternation to avoid the
tree-sitter 0.21.x predicate-bucket hazard. Verified one match per element
(multi-element accumulates) with buildPath/produces/empty anti-overreach.
* feat(group): detect WebClient long-form in Java for Kotlin parity (#2254 P3)
Java deliberately deferred webClient.method(HttpMethod.X).uri(...); the
Kotlin plugin proves a single structural query suffices (same field-access
shape as REST_TEMPLATE_EXCHANGE). Add WEB_CLIENT_LONG_FORM_PATTERNS + scan
loop so .java and .kt detect it identically. Move WEB_CLIENT_LONG_VERB_RE to
the shared module (single source for both). Flip the now-obsolete java :1741
negative test to positive (verbs + no-double-emit) and add a Java var-verb
anti-overreach twin.
* refactor(group): share pushPrefix between java.ts and kotlin.ts (#2254)
The de-duping prefix accumulator was duplicated as kotlin.ts pushKotlinPrefix
and a java.ts closure. Hoist a single export pushPrefix into
spring-consumer-shared.ts; both plugins import it. No behavior change.
* test(group): add Kotlin interface-inheritance boundary twins (#2254)
Twins for the Java inheritance-boundary cases that had no Kotlin counterpart:
shared-leading-segment combine, prefix-less method overlap, ambiguous
duplicate-interface-name suppression, plus a positive multi-interface
implementer. These pin Kotlin's scanProject behavior before U8 extracts the
shared inheritance algorithm.
* refactor(group): share the Spring interface-inheritance scanProject algorithm (#2254)
scanKotlinProject and scanSpringProject were ~80-line near-duplicates over
structurally identical type records. Extract scanSpringInheritanceProject +
SharedSpringType into spring-consumer-shared.ts; collapse KotlinTypeInfo and
SpringTypeInfo into the shared type; both plugins' scanProject become thin
collect-and-delegate wrappers. The ownerPrefix-carrying intermediate is owned
by the shared function. Behavior-preserving — Java and Kotlin inheritance
suites (incl. the new Kotlin boundary twins) byte-identical; tsc clean.
* test(group): close Kotlin↔Java consumer test-parity gaps + assert confidence (#2254)
Add Kotlin twins for Java-tested consumer scenarios with no Kotlin coverage:
@RequestLine query-strip, mixed @RequestLine+@GetMapping, malformed-value
rejection, and @FeignClient(path)-wins-when-@RequestMapping-first. Add the
Java dual-role twin (interface as consumer + implementing controller as
provider). Add two-sided provider confidence (0.8) assertions on the
canonical Java and Kotlin interface-inheritance tests.
* docs(group): document Java FQN route-annotation limitation + pin it (#2254)
Per KTD6, the Java FQN route-annotation gap is documentation-first: the gap is
route-string-only (FQN controllers are already recognised via hasAnnotation)
and FQN-written annotations are vanishingly rare. Document the asymmetry with
Kotlin in JAVA_ROUTE_ANNOTATION_PATTERNS and pin current behavior with an
anti-overreach test. The scoped_identifier query change is deferred to avoid
re-keying existing contracts via the predicate-bucket hazard.
* test(group): add Java↔Kotlin contract set-equality parity harness (#2254)
Independent per-side twins can both pass while the emitted contract SETS
differ. Add a table-driven harness over the parity-critical families
(@RequestLine prefix-fallback, named @RequestLine, @FeignClient(path)+@GetMapping,
@HttpExchange+@GetExchange, WebClient long-form, interface inheritance) that
runs matched .java/.kt fixtures through both plugins and asserts the full
projected contract set (role+contractId+framework+confidence) is equal across
languages AND equal to the expected set — the durable guard for the
byte-identical goal. Gated on kotlinConsumerAvailable.
* style(group): apply prettier formatting to #2254 changes
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* fix(lbug): serialize singleton connection to stop --pdg analyze double-free
LadybugDB is single-writer and its Connection is NOT safe for concurrent
query execution. The WAL-checkpoint driver (5s setInterval) issued
`conn.query('CHECKPOINT')` on the same module-singleton `conn` the analyze
pipeline used for COPY. With --pdg the extra BasicBlock / REACHING_DEF / CDG /
POST_DOMINATE / TAINTED / CALL_SUMMARY / TAINT_PATH table COPYs outlast the
5s tick, so a checkpoint executed concurrently with an in-flight COPY on one
connection -> two libuv workers mutate shared native state -> heap corruption
("double free or corruption (out)" / SIGABRT, detected at the final
"Saving metadata..." free).
Fix: add conn-lock.ts (`withConnLock`, a promise-chain mutex) and run every
singleton-`conn` helper's full query + result-drain inside it: queryAndDrain
(when targetConn === conn), executePrepared, executeWithReusedStatement,
flushWAL, tryFlushWAL, getLbugStats, deleteAllInterprocTaintPaths,
deleteAllCallSummaries. Add an `if (inflight) return` reentrancy guard to the
driver tick so overdue ticks don't stack checkpoints. streamQuery is
intentionally NOT wrapped (read path, re-entrant per-row callback).
Reproduced the crash with concurrent queries on one raw Connection (serial =
stable); verified the fix drives the same overlap through the locked adapter
without crashing.
Tests: conn-lock serialization (no overlap / FIFO / throw-releases) and
driver reentrancy guard.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): lock deleteAllCommunitiesAndProcesses against the WAL driver (#2264)
The count + DETACH DELETE ran raw conn.query on the singleton connection during
incremental --pdg writeback while the WAL-checkpoint driver was live — the same
concurrent CHECKPOINT-vs-write double-free this branch fixes elsewhere. Wrap the
body in withConnLock, mirroring the already-wrapped deleteAllInterprocTaintPaths.
Adds test/integration/lbug-conn-serialization.test.ts (call-through withConnLock
spy) asserting the helper now acquires the lock, wired into the lbug-db vitest
project (and excluded from the default project so it doesn't run twice).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): lock queryImporters against the WAL driver (#2264)
queryImporters issued a raw conn.query on the singleton connection inside the
importer-BFS loop of incremental --pdg writeback, while the WAL-checkpoint driver
could fire a concurrent CHECKPOINT — the same double-free class. Wrap the read
(query + getAll + drain) in withConnLock.
Extends lbug-conn-serialization.test.ts with a routing assertion.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): lock deleteNodesForFile count query on the singleton path (#2264)
The per-table count read used a raw targetConn.query while the sibling DETACH
DELETE already routed through the locked queryAndDrain — an asymmetry that left
the count racing the WAL-checkpoint driver during incremental --pdg writeback.
Gate the count through withConnLock when targetConn === conn (the singleton),
matching queryAndDrain; per-query/temp connections stay lock-free.
Test asserts the count loop takes the lock once per filePath-bearing node table.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): drain DELETE results in the deleteAll* helpers (#2264)
deleteAllInterprocTaintPaths, deleteAllCallSummaries, and
deleteAllCommunitiesAndProcesses awaited conn.query(...DELETE...) but dropped the
returned QueryResult (only the count result was closed), leaking a native result
handle and violating the helpers' own "query + drain inside the lock" contract.
Close each delete result via closeQueryResults, matching the count handling.
Adds a seeded drain test (closeQueryResults fires for the DELETE result).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): make the conn-lock non-reentrancy invariant enforced, not just documented (#2264)
A future withConnLock-wrapped helper calling another wrapped helper would await
its own holder's tail and hang silently. Add an AsyncLocalStorage-based re-entry
guard: withConnLock throws a clear error when invoked from within a holding fn's
async context. A boolean flag can't do this — a legitimately-queued top-level
caller also runs while the lock is held; only AsyncLocalStorage distinguishes a
true nested call from normal contention.
Tests: re-entry throws (not deadlocks); sequential and concurrent top-level calls
do NOT false-fire; the lock releases after a re-entry throw.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* refactor(lbug): rename __resetConnLockForTests to _resetConnLockForTests (#2264)
Match the repo's single-underscore test-seam convention (_initLockPathForTest).
Pure rename of the @internal export and its sole importer.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* test(lbug): fix stale CHECKPOINT guard regex after the c.query refactor (#2264)
lbug-checkpoint.test.ts asserted exactly two CHECKPOINT sites by grepping the
literal `conn.query('CHECKPOINT')`. The connection-serialization refactor changed
flushWAL/tryFlushWAL to capture `const c = conn` and call `c.query('CHECKPOINT')`
inside withConnLock, so the literal grep found 0 and the test failed (expected 2).
Make the regex receiver-agnostic (`.query('CHECKPOINT')`) — preserves the guard's
intent (exactly two authorized CHECKPOINT sites; a third is a regression) while
tolerating the captured-receiver form.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): skip native close on CLI exit to dodge LadybugDB destructor double-free (#2264)
THE actual fix for the `analyze --pdg` crash. gdb shows the abort is a double-free
inside LadybugDB's own destructor during conn.close():
"double free or corruption (out)" -> abort
lbug::main::ClientContext::~ClientContext()
lbug::main::Connection::~Connection()
NodeConnection::Close(...) <- conn.close() from safeClose
It reproduces with the WAL driver OFF and with serial load, so it is NOT the
checkpoint/COPY concurrency the rest of this branch serialized — it's a native
LadybugDB engine bug (@ladybugdb/core 0.17.1, latest stable) triggered by the
larger --pdg write set, firing during teardown AFTER a fully-written, checkpointed
index.
Fix: closeLbug({ skipNativeClose }) CHECKPOINTs for durability (flushWAL) then
skips conn.close()/db.close(), leaving the handles referenced so no GC finalizer
re-runs the destructor. The CLI analyze command (success, error, and SIGINT paths
all process.exit) opts in via skipNativeCloseOnExit; long-lived callers (MCP
server, tests) keep the real close. Mirrors the pool adapter's fire-and-forget
native close and the ONNX native-cleanup philosophy.
Validated end-to-end: `analyze --pdg --force` now exits 0 with a 193,876-node
index; re-opening it (no --force) reads clean and reports up-to-date, proving the
CHECKPOINT-only persistence is durable without db.close().
Workaround for an upstream LadybugDB bug (ClientContext destructor double-free).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): keep conn.close()/db.close() literals out of the closeLbug comment (#2264 review P1-1)
The skipNativeClose comment in closeLbug contained the literal `conn.close()`/
`db.close()`, which the structural guard test (lbug-checkpoint.test.ts:52-53 —
"closeLbug must not inline conn.close()/db.close()") greps for and fails on.
Reword the comment to describe the native close without the literal tokens; the
code already delegates close exclusively to safeClose, so the guard's intent holds.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): real close on the analyze error path to avoid a hang under skipNativeClose (#2264 review P1-2)
The CLI error handler soft-returns (process.exitCode = 1) instead of forcing
exit, relying on the released native handles to let Node terminate. The earlier
commit made runFullAnalysis's error-path closeLbug skip the native close, leaving
live LadybugDB handles that keep the event loop alive forever — a post-init
analyze failure would hang. Only the SUCCESS path (which guarantees a following
process.exit) skips the native close; the error path now always closes for real.
A late-error close could still abort in the destructor, but that terminates the
process — it does not hang.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(server): skip native close in the analyze worker to avoid the LadybugDB destructor crash (#2264 review P2-3)
The forked server analyze worker runs runFullAnalysis then force-exits
(process.exit(0)). With a real native close inside runFullAnalysis, the LadybugDB
ClientContext destructor can double-free after --pdg writes and abort the worker
BEFORE it sends 'complete', failing the parent's analyze. Pass
skipNativeCloseOnExit: true so the worker checkpoints for durability and lets its
process.exit reclaim the handles — same about-to-exit contract as the CLI.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(cli): force exit on a soft error-return when LadybugDB handles are open (#2264 review P1)
The full-analysis success path skip-closes LadybugDB (handles left open, reclaimed
by process.exit). If a post-finalize step (assertAnalysisFinalized) then throws,
the outer catch soft-returns (process.exitCode = 1) — and with native handles
open the event loop never drains, so the process HANGS instead of exiting 1.
Guard once at the analyzeCommand wrapper, after the try/finally: if isLbugReady()
(handles still open) the analyze actually ran and we must force the exit. The
success path never reaches here (analyzeCommandImpl process.exit(0)s itself);
early-validation errors and unit tests that mock runFullAnalysis never open the DB
(isLbugReady() false), so the soft return is preserved.
Adds analyze-finalize-failure-exits.test.ts (force-exits when handles open; does
NOT when they aren't). The analyze-*.test.ts that mock lbug-adapter now also mock
isLbugReady (vitest throws on accessing an undefined export of a mocked module).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): skip the native close on the analyze error path too (#2264 review P2)
A real conn.close() on the error path after large --pdg writes can itself hit the
LadybugDB ClientContext destructor double-free → SIGABRT, degrading an actionable
exit-1 error into a raw native abort. Switch the error-path close to
skipNativeClose (mirroring the success path). Safe now that the CLI catch
force-exits when isLbugReady() (the prior commit): handles left open are reclaimed
by that guaranteed process.exit, so the process terminates without the abort and
without hanging. flushWAL keeps the partial index durable.
Depends on the prior commit (CLI force-exit guard).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): run loadCachedEmbeddings reads under withConnLock (#2264 review P2)
loadCachedEmbeddings issued raw conn.query reads on the singleton connection
outside withConnLock — safe today only because it runs before the WAL-checkpoint
driver starts, an ordering invariant not enforced by code. Wrap the whole read in
withConnLock so a future reorder can't race a CHECKPOINT on the connection. Leaf
read; no nested wrapped helpers.
Adds a routing assertion to lbug-conn-serialization.test.ts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* test(lbug): de-brittle the close/CHECKPOINT structural guard (#2264 review P2)
lbug-checkpoint.test.ts grepped the adapter SOURCE (comments included) for
conn.close()/db.close()/.query('CHECKPOINT') literals, coupling a passing test to
comment wording — a prior commit had to reword a comment just to keep it green.
Strip comments from the read source before the structural assertions so they
reflect code only; the invariant (exactly two CHECKPOINT sites; close calls only
in safeClose) is preserved and no longer breaks on a comment edit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* refactor(lbug): rel COPY uses the captured writeConn, matching node COPY (#2264 review P3)
The relationship COPY passed the module-level `conn` to copyCsvWithRetry while the
node COPY uses the captured `writeConn`. Use `writeConn` for both — one captured
reference for the whole bulk load, removing the latent identity dependency. Same
object during analyze (`conn` is only reassigned at open/close under the session
lock), so the queryAndDrain `targetConn === conn` lock gate still engages.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(server): make the analyze worker's IPC send() failure-safe (#2264 review P3)
The worker's send() used `process.send?.(msg)` — the `?.` guards an undefined
channel but not a throw from an already-closed one (ERR_IPC_CHANNEL_CLOSED). A
throw in the catch-branch send() would escape the message handler and skip the
scheduled `setTimeout(process.exit(0))`, stranding the worker (with skip-close
leaving native handles open, #2264). Wrap process.send in try/catch so the exit
always fires; a vanished child is a failure to the parent regardless.
Not unit-tested: send() is module-private and importing the worker registers
process signal handlers; the change is a defensive try/catch around one call.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(cli): bound the SIGINT cleanup CHECKPOINT so Ctrl-C stays responsive (#2264 review P3)
The SIGINT handler calls closeLbug({skipNativeClose:true}), whose flushWAL
CHECKPOINT queues behind the connection lock held by an in-flight COPY — so a
single Ctrl-C during a long --pdg COPY appeared hung until the COPY released.
Race the cleanup against a 2s timeout before process.exit(130); the WAL replays
on the next analyze. The double-Ctrl-C escape hatch is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(server): report analyze-worker errors over IPC, never swallow (#2264 P3)
The worker's send() swallowed IPC failures (and a prior pass logged them to
stderr). Per review, all worker errors must be reported back to the parent over
the existing IPC channel (send({ type: 'error' })) and nothing silently dropped.
- send() no longer catches: a dead channel (ERR_IPC_CHANNEL_CLOSED) throws
instead of being swallowed.
- Every handler (uncaughtException, unhandledRejection, SIGTERM, the analysis
message handler) reports its error via send() in try and schedules process.exit
in finally, so a throw from send() can no longer skip the exit and wedge the
worker — the P3 'schedule the exit so it always fires' fix, without a swallow.
- SIGTERM cleanup failures are now reported to the parent instead of an empty
catch {}.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(workers): report every caught parse-worker error over IPC (#2264)
The parse worker swallowed or only-locally-logged several caught errors, so they
never reached the pool: the per-language-group catch was an empty catch {} that
silently dropped the whole group on any throw (not just an unavailable grammar),
the per-file parse/query-execution catches only logger.warn'd (worker-thread
local), and the C++ template-constraint catch swallowed silently.
Route all work-path catches through a new reportWarning() helper that posts
{ type: 'warning', message } to the pool (which logs it on the main thread AND
resets the worker idle timer, so a worker grinding through failing files isn't
falsely idle-evicted), with a logger.warn fallback for the non-worker path. The
existing inline warning sites (query-compilation, the extractParsedFile callback,
CFG build) are migrated to the same helper.
The 4 optional-grammar module-load guards (Swift/Dart/Kotlin/C) stay silent: they
run before the 'ready' handshake and their absence is already surfaced via
result.skippedLanguages + the isLanguageAvailable gate. Fatal/group-aborting
errors continue to flow through the message handler's { type: 'error', errorStack }.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* test(cli): harden finalize-failure test against forked-worker death (#2264)
analyzeCommand calls installFatalHandlers(), which registers global
unhandledRejection/uncaughtException handlers that call the REAL process.exit(1).
Across this file's vi.resetModules() reimports they accumulate on `process`, and
under CI timing a stray async rejection fired one while no process.exit spy was
active — killing the forked vitest worker ("Worker exited unexpectedly"), which
only surfaced once the full test lanes finished (they were pending at review time).
Keep process.exit spied for the whole file (beforeAll/afterAll) so a fatal handler
can never really exit mid-run, strip the handlers installFatalHandlers added in
afterAll (preserving vitest's own, snapshotted up front) before restoring the real
process.exit, and reset process.exitCode so the worker exits clean. Passes in
isolation and grouped; behavior under test is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(analyze): don't take the up-to-date fast path for an unregistered repo (#2264)
A prior 'analyze --name X' that hit a registry name collision writes meta.json
(meta-save runs before registerRepo) but fails before registering — leaving the
index up-to-date but UNREGISTERED. A later 'analyze --name X --allow-duplicate-name'
then matched the up-to-date gate and early-returned WITHOUT registering, so the
repo stayed invisible to list_repos/MCP and the CLI's assertAnalysisFinalized
rejected it. --allow-duplicate-name could never heal it.
This was latent on main, masked by the very close-hang this PR fixes: the lingering
process pushed the cli-e2e #829 step-3 analyze past its 60s spawn timeout
(status===null → the test's vacuous early-return). With the hang gone the analyze
exits promptly, exit 1 surfaces, and the bug becomes deterministic on all platforms.
Fix: the up-to-date fast path now short-circuits only when the repo is actually
registered (new isRepoRegistered helper, sharing assertAnalysisFinalized's exact
canonical/case-folded membership check). An indexed-but-unregistered repo falls
through to the pipeline, which registers it honoring allowDuplicateName. Already
registered repos keep the fast path unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* chore: trigger CI re-run
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(analyze): gate up-to-date self-heal on --allow-duplicate-name (#2264)
The prior commit healed every up-to-date-but-unregistered repo by falling through
to register it — which broke the #1169 guard: a plain `analyze` of an up-to-date
repo whose registry entry is missing MUST fail loudly ("Analysis did not finalize")
rather than silently register a possibly half-finalized index.
Distinguish the two causes of "unregistered":
- collision-rejected + user re-runs with --allow-duplicate-name → explicit intent
to register, so fall through to the pipeline and register it (#829).
- plain analyze, registry missing/wiped → keep the #1169 fail-loud behavior.
So self-heal is gated on options.allowDuplicateName; isRepoRegistered is only read
on that opt-in branch, so the common fast path keeps its single-stat cost. Both
cli-e2e guards (#1169 fail-loud, #829 heal) now pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(server): skip the native close on analyze-worker SIGTERM cancellation (#2264 P2)
cancelJob() / the 30-min timeout (analyze-job.ts) send SIGTERM to the forked
analyze worker, but its SIGTERM handler still did a full `await closeLbug()`
(native conn/db teardown) — even though normal completion now skips it via
skipNativeCloseOnExit. A cancelled or timed-out --pdg server analyze could
therefore still hit the LadybugDB ClientContext destructor double-free, or block
behind the in-flight COPY's connection lock before exiting.
Mirror the CLI SIGINT path: a best-effort CHECKPOINT with
closeLbug({ skipNativeClose: true }) bounded by a 2s Promise.race timeout, then
process.exit(0) (which reclaims the handles). A CHECKPOINT failure is reported to
the parent over IPC rather than swallowed; the exit is in .finally so it always
fires.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* test(cli): import analyze once in the finalize-failure test (#2264 CI)
The test failed deterministically only on the ubuntu coverage lane (2/2 runs)
while passing locally and in isolation, incl. with --coverage. Cause: the
vi.resetModules() + per-test `await import('analyze.js')` re-instrumented the
ENTIRE analyze module graph on every test; under --coverage on the
memory-constrained CI runner that OOM/crashed the forked worker ("Worker exited
unexpectedly" → the assertion never ran).
Import analyzeCommand ONCE and drive the mocks per-test via mockReturnValue
(resetModules wasn't needed — the hoisted mocks are controllable per-test). Keeps
the whole-file process.exit spy + afterAll fatal-handler strip from the prior pass.
Behavior under test is unchanged; passes in isolation and with --coverage.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* test(cli): pre-set NODE_OPTIONS heap cap so ensureHeap can't re-exec (#2264 CI)
Root cause of the ubuntu-coverage-only failure (3/3 CI runs, passing locally):
analyzeCommand calls ensureHeap() (analyze.ts:715), which RE-EXECS the process —
spawning `node <heap-flags> <argv>` with vitest's argv — unless NODE_OPTIONS
already carries --max-old-space-size (analyze.ts:498). That re-exec killed the
forked vitest worker ("Worker exited unexpectedly" → the assertion never ran).
It only reproduced on the memory-constrained CI runner because locally a high V8
heap-size-limit also short-circuits ensureHeap (analyze.ts:501).
Reproduced locally with NODE_OPTIONS="--no-warnings" (no heap cap) → same failure;
fixed by pre-setting --max-old-space-size in beforeAll (restored in afterAll), the
same workaround cli-e2e uses. Verified: passes under the repro condition, normally,
and with --coverage.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(server): worker asserts finalization before reporting complete (#2264 P2)
The forked analyze worker reported {type:'complete'} straight after runFullAnalysis,
so a server/web analyze of a half-finalized repo (meta.json written but the global
registry entry missing — a prior collision-aborted run, or a wiped registry) was
reported successful while the repo stayed unregistered/invisible to list_repos. The
CLI already guards this with assertAnalysisFinalized; the worker did not.
Extract the run -> finalize -> report contract into a side-effect-free
analyze-worker-core seam (the entry module's top-level process.on handlers make it
untestable directly) and call assertAnalysisFinalized before sending complete — a
failure is reported as {type:'error'} instead of a false success. The seam is
dependency-injected and unit-tested with fakes; the entry module wires the real deps
and keeps owning process.exit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(server): coordinate worker SIGTERM cancellation with completion (#2264 P3)
The worker SIGTERM handler unconditionally sent {type:'error','Analysis cancelled'}
and didn't coordinate with the message handler that sends complete, so a cancel near
the finish line could report a cancelled job complete, or a late SIGTERM could flip
an already-complete job to failed.
Add a single terminal-outcome claim (createTerminalClaim) shared by the message
handler and the SIGTERM handler: whoever claims it first reports its terminal
message; the other skips its terminal send. Single-threaded JS makes the
check-and-set atomic. The cleanup + process.exit still run regardless.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(server): make a job's terminal outcome immutable on the parent side (#2264 P3)
Defense-in-depth complement to the worker terminal-claim: the launcher's message
handler and the job manager's updateJob both lacked a terminal-state guard, so a
late worker IPC message (a SIGTERM-driven 'error' after 'complete', or vice versa)
could re-release the repo lock and flip the reported status. (Touches parent-side
files outside the original PR diff — deliberate, clearly-scoped.)
- analyze-job.ts updateJob: drop any update once the job is already terminal (the
transition INTO terminal still applies, since status isn't terminal yet then).
- analyze-launch.ts message handler: return early when the job is already terminal,
mirroring its sibling exit handler.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* test(lbug): assert deleteAllInterprocTaintPaths + deleteAllCallSummaries route through withConnLock (#2264)
The lock-routing suite covered 4 singleton-conn helpers but not these two
withConnLock-wrapped delete helpers (lbug-adapter.ts), which also run during the
incremental --pdg writeback window — so a revert of either wrapper would have gone
uncaught. Add the two routing assertions to complete the coverage the file's header
claims (every singleton-conn helper reachable during the WAL-driver window).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* test(lbug): assert temp-conn deleteNodesForFile skips withConnLock (negative gate, #2264)
The positive case (singleton deleteNodesForFile locks each per-table count) was
covered, but not the negative branch of the targetConn === conn gate: a per-file/temp
connection (dbPath provided) must NOT take the singleton lock, or temp-conn callers
would needlessly contend with it. Add the negative-gate assertion so a regression
that unconditionally locks is caught.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* test(cli): assert the force-exit forwards process.exitCode, not a hardcoded 1 (#2264)
The existing cases asserted process.exit(1), but since the error catch always sets
exitCode=1 they couldn't distinguish forwarding (process.exit(process.exitCode ?? 1))
from a hardcoded 1. Add a case on the alreadyUpToDate path — which returns without
setting exitCode or calling process.exit — with a pre-set exitCode=2 and isLbugReady
forced true, asserting the wrapper force-exits with 2. Proves the exitCode-forwarding
branch.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* refactor(lbug): replace skipNativeClose flag with a dedicated closeLbugBeforeExit() (#2264)
The "skip the native close only when a process.exit is guaranteed to follow"
invariant was enforced by convention across ~4 call sites via a boolean option on
closeLbug — the exact foot-gun the review flagged. Encode the contract in the name
instead:
- New closeLbugBeforeExit() (CHECKPOINT via flushWAL, then return without the native
close); closeLbug() drops the option and is the plain real-close again.
- run-analyze success + error paths: options.skipNativeCloseOnExit ?
closeLbugBeforeExit() : closeLbug(). CLI SIGINT + worker SIGTERM call
closeLbugBeforeExit() directly. skipNativeCloseOnExit stays on AnalyzeOptions as the
caller's "I will exit" signal.
- lbug-checkpoint.test: assert closeLbugBeforeExit exists + has no native close, and
match `closeLbug =` precisely so it doesn't prefix-match the new function.
- Retarget the conn-serialization integration case to closeLbugBeforeExit(); add the
new export to the 12 analyze-*.test.ts lbug-adapter mocks.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* refactor(lbug): extract isSharedSingletonConn predicate for the lock gate (#2264)
The targetConn === conn object-identity gate (decides whether an op takes
withConnLock) was duplicated inline in queryAndDrain and deleteNodesForFile with
its own explanatory comments. Extract a single isSharedSingletonConn(c) predicate
with the rationale in one place; both sites route through it. Behavior unchanged —
covered by the lock-routing tests' positive (singleton locks) and negative
(temp-conn skips) cases.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* refactor(lbug): share the bounded checkpoint-then-exit cleanup (SIGINT/SIGTERM) (#2264)
The CLI SIGINT handler (analyze.ts) and the worker SIGTERM handler (analyze-worker.ts)
had near-identical Promise.race([closeLbugBeforeExit, timeout]).finally(exit) blocks
with separately-hardcoded 2s timeouts. Extract boundedCheckpointBeforeExit into a
shared shutdown-helpers module — parameterized by exit code, an optional flush-error
reporter (worker reports over IPC), and an optional beforeExit hook (CLI flushes the
logger). checkpoint + exit are injectable test seams, so it's unit-tested without the
real LadybugDB close or process.exit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* refactor(storage): extract registryPathEquals for the registry case-fold compare (#2264)
The Windows case-insensitive / POSIX case-sensitive registry-path comparison was
duplicated across 6 sites (registerRepo dedup, the fresh-merge findIndex,
removeRepo/removeBranchIndex local 'matches' helpers, isRepoRegistered, and the
path-match lookup). Extract a single registryPathEquals(a, b) predicate so every
registry lookup/dedup/finalize check answers identically; route all 6 through it.
No behavior change — repo-manager + finalize-invariant suites pass (incl. the
Windows case-fold case).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* feat(lbug): runtime-guard streamQuery against the WAL-checkpoint driver (#2264)
streamQuery is deliberately not wrapped in withConnLock (its per-row callback can
re-enter the adapter), so its unlocked per-row reads could race a CHECKPOINT on the
shared connection — the corruption window the lock serializes everything else
against. That invariant was comment-only, safe today only because the serve/read
path forks analyze workers. Make it enforced:
- lbug-adapter: a walDriverActive flag + markWalDriverActive(bool); streamQuery
throws an actionable error when the driver is active.
- wal-checkpoint-driver: arm the flag on start, disarm in stop() AFTER the in-flight
CHECKPOINT drains (clearing earlier would briefly allow a race).
A future in-process analyze overlapping a stream now fails loud instead of
corrupting native state. (reentrancy test's lbug-adapter mock gains the new export.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* docs(lbug): explain why closeLbugBeforeExit skips finalizeLbugSidecarsAfterClose (#2264)
Document the deliberate trade-off: the skip-close path intentionally does NOT run
the sidecar-finalize step that safeClose runs after a real close. It's designed for
released WAL handles; running it with the connection still open risks a Windows
file-lock on the in-use WAL. The CHECKPOINT already made the index durable and the
next run's preflightLbugSidecars reconciles residual WAL — the deferral is the
accepted cost of skipping the native close to dodge the destructor double-free.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
* fix(lbug): move WAL-driver-active flag to its own module (fix mock ripple, #2264)
The streamQuery guard (4449bb4a) put markWalDriverActive in lbug-adapter, and the
wal-checkpoint-driver imported it from there. That broke every test mocking
lbug-adapter while loading the real driver — CI's ubuntu lane caught
run-analyze-fts-repair.test.ts ('No markWalDriverActive export on the mock').
Move the one-bit shared flag to a dedicated wal-driver-state module: the driver
toggles markWalDriverActive there, streamQuery reads isWalDriverActive there, and
lbug-adapter no longer carries it — so mocking lbug-adapter no longer has to stub
the toggle. run-analyze-fts-repair now passes untouched; the reentrancy test's mock
addition is reverted (no longer needed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBJomjoTdBV2eveDVq4JMm
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bump gitnexus, plugin.json, and marketplace.json to 1.6.8 and add the
CHANGELOG section. Headline: opt-in PDG-backed impact analysis plus the
full PDG/taint substrate (CFG → reaching-defs → intra/inter-procedural
taint → control dependence) across the language matrix, multi-branch
indexing, private GitHub PAT + Azure DevOps support, and MCP trace/HTTP.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: fail tree-sitter summary parse drift
* docs(ci): cross-reference readiness regexes to their test mirror
The two report.match() literals in the upsert-issue github-script step are
duplicated as _ISSUE_READY_RE / _ISSUE_BLOCKER_RE in
test_check_tree_sitter_upgrade_readiness.py, and only the Python copy is
asserted against the rendered report. Since a stale regex now throws via
requireMatch (instead of the old silent '?' fallback), add a reciprocal
keep-in-sync note at the workflow site so a future prose edit can't desync
the JS literal from the asserted mirror undetected.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): route read-phase network failures to fetch_failed, not a crash
npm_view_json and fetch_text caught only (URLError, HTTPError[, JSONDecodeError]).
urllib wraps connect-phase OSErrors into URLError, but a failure during
resp.read() AFTER urlopen returns (ConnectionResetError, ssl.SSLError,
socket.timeout, http.client.IncompleteRead) is not a URLError subclass — it
escaped the helper, crashed main(), and left stdout empty. main() is unguarded
(the only top-level except wraps just stdout.reconfigure), and the report print
is its last statement, so an empty report then makes the workflow's requireMatch
throw on a non-drift scheduled run.
Broaden both except tuples with OSError + http.client.IncompleteRead so a
transient mid-body network blip yields None, routing the grammar to the existing
fetch_failed blocker bucket (a complete report) — preserving the fail-loud intent
for real drift while removing the crash-to-empty-stdout path. JSONDecodeError
stays explicit (it is a ValueError, not an OSError). Adds read-phase regression
tests that fail on the old narrow tuple.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(ci): document the regex-contract assertion counts
test_issue_update_summary_regex_matches_current_report asserts hardcoded
capture groups ("9","10") and "2" with no explanation. Document the
derivation from _render_report()'s mock corpus — 9 of 10 npm grammars Ready
(tree-sitter-cpp is the intentional pin), 2 blockers (pinned tree-sitter-cpp +
held vendored tree-sitter-c) — so a future grammar or pin change is an obvious
two-step update (mock + counts) rather than a mystery failure. Assertions
unchanged; comment only.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: name _Resp method receivers 'self' to clear py/not-named-self
The _Resp stub inside _patch_urlopen named its method receivers
`self_inner`, which CodeQL flags as py/not-named-self (PEP 8) — three
alerts on this PR's merge ref (lines 230/233/236). _patch_urlopen is a
@staticmethod, so there is no outer `self` to collide with; rename the
receivers to the conventional `self`. Pure rename, no behavior change.
All 25 tests in test_check_tree_sitter_upgrade_readiness still pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(impact-pdg): run mutation oracle's analyze child from built dist, not tsx-over-src
The nightly Impact PDG Mutation Report workflow failed at the first fixture with
ERR_MODULE_NOT_FOUND for src/cli/lazy-action.js. The harness shelled the real CLI
out as `node --import tsx src/cli/index.ts analyze …`; on the CI runner's Node
22.22.3, native TypeScript type-stripping is enabled by default and handles the
.ts entry instead of tsx, and native stripping does NOT remap the `./lazy-action.js`
import specifier to lazy-action.ts the way tsx does — so CLI startup crashes
before analyze even runs.
The workflow already builds dist/ (build: 'true'). Prefer the shipped
dist/cli/index.js (plain compiled JS — no tsx, no strip-types, and the parse
workers it spawns also resolve from dist/) for the analyze child, falling back to
tsx's own CLI over src only for build-free local runs. Production-faithful and
version-agnostic across the engines range (node >=22.0).
Verified on a real Node 22.22.3: the dist child starts cleanly with no
lazy-action resolution error; the full `--mutation --only=inter-dispatcher-thin`
run scores realized recall 1.0 and gate-mutation-recall passes. Workers are
independently confirmed green on 22.22.3 in CI (run 27874383902).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(impact-pdg): declare the mutation oracle's @babel/* deps
`bench/impact-pdg/mutation-oracle.mjs` imports @babel/parser, @babel/traverse,
@babel/generator and @babel/types to instrument + value-diff the fixture AST,
but none were declared in package.json. @babel/parser and @babel/types happen to
be hoisted into gitnexus/node_modules transitively, but @babel/traverse and
@babel/generator are only present at the monorepo root — so a fresh `npm ci` in
gitnexus/ (CI) can't resolve them and the oracle dies at module load with
`Cannot find package '@babel/traverse'` right after analyze succeeds.
Declare all four as devDependencies (they're already lazily imported only on the
--mutation path, so they stay out of the unit-test module graph). Verified the
oracle resolves them from gitnexus/node_modules and scores recall 1.0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(impact-pdg): gate only recall-gated mutation checks (honor recallGated)
The recall gate filtered checks by `typeof c.recall === 'number'`, which includes
the UPSTREAM fixtures. The mutation oracle is a FORWARD value-diff: it mutates the
criterion line and observes which downstream lines' values change, so its
behavioral AIS can never intersect a reverse (upstream) PDG slice — recall is 0
by construction. measure.mjs already marks these `recallGated: false` (alongside
id-discrimination corroboration cases) and excludes them from its own internal
gate; the standalone gate just didn't honor that flag, so `intra-control-loop`
(direction: upstream, recall 0) tripped the floor even though the oracle ran the
full suite cleanly (mean recall 0.923).
Filter on `c.recallGated === true` so the floor applies only to the downstream
cases the forward oracle can fairly validate. Verified locally: an
upstream+downstream report now scores 1 of 2 and the gate passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(impact-pdg): fail the mutation gate when it has no recall signal + fix README drift
Tri-review hardening of this PR's own changes:
- gate-mutation-recall.mjs: the floor check passed vacuously when `scored` was
empty (`min === null` short-circuits `min !== null && min < floor`). Narrowing
the filter to `recallGated === true` made an empty `scored` set reachable in
more inputs (a degenerate corpus, or a harvest that silently emptied every
behavioral AIS). Now fail loudly when checks exist but none are recall-gated,
so a hollow gate is red rather than a green "scored cases: 0 of N". A genuinely
empty report (0 checks) still passes — it's not a degenerate-corpus signal.
- README.md: the harness substrate section still documented the old
`node --import tsx src/cli/index.ts …` child invocation this PR replaced;
update it to the dist-preferred form to match `cliChildArgs`.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(routes): persist HTTP method on Route nodes
Part 1 of 2 for issue #2138 (skip redundant HTTP provider source-scan).
The ingestion routes phase already knows each route's HTTP verb —
`ExtractedRoute.httpMethod` (Spring/Laravel framework routes) and
`ExtractedDecoratorRoute.httpMethod` (decorator routes) — but dropped it
when creating the Route graph node. As a result `HttpRouteExtractor`'s
graph-assisted path could not recover the verb for `framework-route`
sources (whose edge `reason` is undecodable by `methodFromRouteReason`)
and had to fall back to re-scanning the handler source.
Changes:
- routes phase: carry `httpMethod` into `RouteEntry` and persist it as
`Route.method` (filesystem-derived Next.js/Expo/PHP routes have no
structural verb, so they stay method-less).
- HttpRouteExtractor: HANDLES_ROUTE query now returns `route.method`;
`extractProvidersGraph` prefers it and falls back to the edge reason
for older indexes / method-less routes (fail-open, fully backward
compatible).
- tests: graph-method precedence, multi-verb handler disambiguation via
the persisted verb, case normalization, and old-index fallback.
This change is intentionally NOT a performance optimization on its own:
the graph path still parses handler files to recover the handler *name*.
Eliminating that parse (and thus the redundant source-scan #2138 targets)
requires linking HANDLES_ROUTE to the handler symbol, which lands in
Part 2. This PR is the data-completeness groundwork for that.
Refs #2138
* test: account for new Route.method in blade route-registry assertion
The routes phase now persists httpMethod onto RouteEntry/Route nodes, so
the strict toEqual on the framework-route registry entry must include the
new method field.
* fix(routes): persist Route.method end-to-end + real-lbug round-trip test
Addresses review on #2234 (magyargergo + tri-review): the prior commit
read `route.method` in HANDLES_ROUTE_QUERY but never added the column to
the schema/persistence path, so against a real LadybugDB the query failed
to bind (`Cannot find property method for r.`) and the `catch { return [] }`
silently swallowed it — regressing the graph-assisted HTTP provider path.
- schema: add `method STRING` to ROUTE_SCHEMA.
- csv-generator: write `method` in the Route CSV row (header + row, column
order aligned with the COPY statement).
- lbug-adapter: add `method` to getCopyQuery('Route').
- routes phase: normalizeRouteMethod() canonicalizes the verb to upper-case
and skips non-verbs — Laravel resource/apiResource carry httpMethod
values like `resource`/`apiResource`, which must not land a junk method.
- http-route-extractor: log at debug when the HANDLES_ROUTE / FETCHES graph
query throws, so a total graph-provider outage is observable instead of
silently swallowed. Export HANDLES_ROUTE_QUERY for the round-trip test.
- tests: add a real-lbug round-trip (graph -> CSV -> COPY -> HANDLES_ROUTE_QUERY)
asserting the verb persists and reads back; update the blade registry
assertion for the normalized (upper-case) method.
Refs #2138
* fix(csv): coerce Route.method to string for escapeCSVField typecheck
node.properties.method is typed unknown (not a declared property), so
`x || ''` stayed unknown and failed tsc against escapeCSVField's
string|number param. Coerce explicitly with String(... ?? '').
* test(bench): regenerate emit-persistence fingerprint for Route.method column
Adding the method column to route.csv changes the byte-identity
fingerprint of the emit-persistence benchmark (the synthetic graph's
route.csv header now includes 'method'). scaling_ratio unchanged (~0.9,
linear); this is the documented regenerate-on-legitimate-emit-change
path. Streaming baseline (BasicBlock/PDG) is unaffected.
---------
Co-authored-by: henry <zhangwei2017@unipus.cn>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* fix(server): resolve clone/upload/mapping roots from GITNEXUS_HOME
CLONE_ROOT (git-clone.ts), UPLOAD_ROOT (upload-paths.ts), and the
server-mapping file (embeddings/server-mapping.ts) computed their root
from os.homedir() directly, ignoring GITNEXUS_HOME. The Docker image
sets GITNEXUS_HOME=/data/gitnexus (the persistent volume that also holds
the registry and indexes), so cloned/uploaded repos and their .gitnexus
indexes instead landed in the container's ephemeral ~/.gitnexus and were
lost on container recreation, while registry.json (which honors
GITNEXUS_HOME) kept pointing at the now-dead path. That also defeated
incremental re-analysis: a recreated container re-clones from scratch and
full-rebuilds instead of git pull + incremental update.
Source all three from the existing getGlobalDir() helper — the same
GITNEXUS_HOME-aware primitive the registry and groups already use.
Behavior is unchanged when GITNEXUS_HOME is unset (CLI / local installs):
it falls back to ~/.gitnexus exactly as before. No signatures change and
no UPLOAD_ROOT consumers are touched.
Adds test/unit/gitnexus-home-roots.test.ts covering both the
GITNEXUS_HOME-set and unset paths for all three roots.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(server): make clone-root assertions GITNEXUS_HOME-aware; cover server-mapping fallback
Addresses PR review (#2229):
- git-clone.test.ts hardcoded path.join(os.homedir(), '.gitnexus', 'repos')
for the clone root, so the "direct child of the clone root" and containment
assertions failed when GITNEXUS_HOME was set in the ambient env (e.g. a CI
runner). CLONE_ROOT now derives from getGlobalDir(); mirror that derivation
via EXPECTED_CLONE_ROOT (GITNEXUS_HOME || ~/.gitnexus), computed at module
load — the same point CLONE_ROOT is frozen — so the two always agree.
- The GITNEXUS_HOME-unset fallback test covered clone + upload roots but not
server-mapping (one of the three changed modules). Add a fallback case for
readServerMapping. It redirects HOME/USERPROFILE to a tmp dir (os.homedir()
honors them) so it exercises the real ~/.gitnexus fallback without writing
into the developer's actual ~/.gitnexus/server-mapping.json.
Both files pass with and without GITNEXUS_HOME set.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(server): use mkdtemp for GITNEXUS_HOME path-root test temp dirs
CodeQL js/insecure-temporary-file flagged the two writes that used a fixed
os.tmpdir() name (gitnexus-home-roots-test, gitnexus-fallback-home): a
co-located process could pre-create or symlink the predictable path before
the test write lands. Switch both to fs.mkdtemp(), which atomically creates a
uniquely-named directory — the canonical sanitizer this repo already uses
(see core/group/storage.ts). Behavior is unchanged; tests still pass with and
without GITNEXUS_HOME set.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(embeddings): resolve server-mapping path for parity with clone/upload roots
MAPPING_FILE now wraps path.resolve() like CLONE_ROOT (git-clone.ts) and
UPLOAD_ROOT (upload-paths.ts), so a relative GITNEXUS_HOME yields an absolute
path. No-op for the supported absolute-GITNEXUS_HOME config (Docker) and for
readServerMapping's only caller (run-analyze.ts).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(server): derive EXPECTED_CLONE_ROOT from getGlobalDir() to avoid drift
The test mirror now imports getGlobalDir() and computes the expected clone root
the same way production CLONE_ROOT does, instead of re-deriving the
GITNEXUS_HOME || ~/.gitnexus fallback by hand. Byte-identical today; future-proof
if getGlobalDir() grows a branch. (os import retained — still used by os.tmpdir().)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(server): rename shadowed savedHome in server-mapping fallback test
The inner savedHome (saves process.env.HOME) shadowed the describe-scope
savedHome (saves process.env.GITNEXUS_HOME). Renamed the inner one to
savedProcessHome at its declaration and both restore sites in the finally
block. Pure rename, no behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(server): scope customHome temp dir to the GITNEXUS_HOME-set cases
beforeEach created a customHome temp dir for all five tests, but the two
fallback cases never use it (one manages its own fakeHome, the other needs
none). Moved the mkdtemp/rm into a nested describe('with GITNEXUS_HOME set')
wrapping the three set-cases; the shared outer afterEach still restores
GITNEXUS_HOME and resets modules for all five.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
* test(lbug): lock PARALLEL=false as a tested correctness invariant (#2203)
The parallel CSV reader (Kuzu-derived, default PARALLEL=true) cannot parse
quoted fields with embedded newlines (kuzudb/kuzu#5778); our content/text
columns hold source code, so PARALLEL=false is mandatory for correctness.
Add a live-DB multiline-quoted round-trip that fails if it is ever flipped,
plus a static guard on the generated COPY queries. Export COPY_CSV_OPTS /
getCopyQuery for the static assertion; document the invariant at the source.
* feat(lbug): expose node/rel phase boundary via onNodePhaseComplete hook (#2203)
streamAllCSVsToDisk now fires an optional onNodePhaseComplete(nodeFiles)
callback right after node CSVs are flushed and before the relationship pass
writes any rel_*.csv — the boundary the COPY-overlap leg needs. The node-file
manifest construction is hoisted above the rel pass and reused in the return,
so output is byte-for-byte identical when no callback is supplied (verified by
the emit bench fingerprint and the splitRelCsvByLabelPair differential oracle).
The callback is not awaited, so the rel pass runs concurrently with the
caller's node COPY.
* perf(lbug): overlap node COPY with relationship emit (#2203)
The deferred parallelism leg of #2203. LadybugDB is single-writer and its
parallel CSV reader is unsafe for our multiline content (kuzudb/kuzu#5778), so
the only safe parallelism is pipeline-overlap: node COPY (uses conn, never the
rel files) runs concurrently with the relationship emit pass (writes rel_*.csv,
never conn). Node COPY now starts at streamAllCSVsToDisk's onNodePhaseComplete
boundary while the rel pass keeps writing; the relationship COPY still waits for
node COPY (FK precondition), so DB load order and content are unchanged.
- Extract copyNodeCSVs; start it in the hook (overlap) or after emit (serial).
- GITNEXUS_SERIAL_LBUG_LOAD=1 forces the legacy strictly-sequential path
(operator escape hatch + differential-test oracle).
- Settle the in-flight node-COPY promise on emit failure (no unhandled
rejection); rethrow node-COPY errors at the FK barrier.
- Preserve the PDG manifest merge + collision guards (node merge at the hook,
rel merge before rel COPY) and all retry/fallback/cleanup behavior.
- PROF_LBUG_LOAD gains mode=overlap|serial; copy-nodes becomes the residual
node-COPY time after emit (trends to 0 as overlap hides it).
* test(lbug): differential gate — overlap load === serial load (#2203)
Loads one fixture (multiple node tables, multiple edge pairs, multiline File
content + BasicBlock text) into two fresh DBs — once via the default node-COPY
‖ rel-emit overlap, once via GITNEXUS_SERIAL_LBUG_LOAD=1 — and asserts the two
databases are content-equivalent: identical per-table node counts, per-type
edge counts, byte-for-byte multiline content/text, and identical
insertedRels/skippedRels/warnings. This is the issue's byte-identical-content
acceptance gate for the parallelism leg.
* fix(review): apply autofix feedback
- csv-generator: onNodePhaseComplete doc-contract now matches reality (a sync
throw is allowed and is how loadGraphToLbug surfaces the manifest collision
guard) — drops the inaccurate 'must not throw synchronously' line.
- lbug-adapter: copyNodeCSVs totalSteps is the node-table count (drop the +1
rel-step holdover; the rel COPY has its own progress line).
- lbug-adapter: on emit+node-COPY double-failure, log the swallowed node-COPY
error before rethrowing the emit error (diagnosability).
- lbug-load-prof test: assert mode=overlap on the default path.
* fix(test): use mkdtemp for secure temp dirs (CodeQL js/insecure-temporary-file)
CodeQL flagged lbug-load-overlap.test.ts writing a file into a predictable
os.tmpdir() path. Create the base temp dir with fs.mkdtemp (atomic, random
suffix) in both new live-DB tests, and switch to the gitnexus-lbug- prefix that
TEST_FIXTURE_PREFIXES recognizes so the Windows stale-sidecar sweep covers
these fixtures.
* fix(lbug): check PDG manifest rel-pair collision before node COPY (#2203)
Found by Codex in tri-review. The manifest rel-pair collision guard ran after
the FK barrier (after node COPY committed), so on that should-never-happen
error branch the overlap path left orphan node rows AND the
GITNEXUS_SERIAL_LBUG_LOAD escape hatch diverged from the legacy 'validate
manifest before any COPY' behavior. Move the rel merge + collision check ahead
of beginNodeCopy/the barrier: the serial path now detects a collision before
committing any node rows (legacy parity restored — the escape hatch is a
faithful oracle again), and the overlap path detects it as early as csvResult
is available. The node-collision guard already ran before node COPY (in the
hook).
* test(lbug): cover rel-emit failure with node COPY in flight (#2203)
Resolves a P1 review gap on PR #2226: the overlap's catch(emitErr) branch
(settle the in-flight node-COPY promise, then rethrow the emit error) was
untested. Fault-injects via a vi.mock of streamAllCSVsToDisk that fires
onNodePhaseComplete (starting a real node COPY on a live DB) then throws,
asserting loadGraphToLbug rejects with the emit error and no unhandled
rejection leaks. Also covers the both-fail case (node COPY error is logged,
emit error still wins). Listener removed in finally; macrotask queue flushed
before the assertion so it can't pass vacuously.
* test(lbug): cover node-COPY hard-failure rethrow at the FK barrier (#2203)
Resolves the second P1 review gap on PR #2226. Mocks emit to fire
onNodePhaseComplete with a nodeFiles entry pointing at a missing CSV (a
bind-time COPY error that IGNORE_ERRORS does not suppress) and otherwise
succeed, so copyNodeCSVs throws, the error is captured in nodeCopyError, and
loadGraphToLbug rethrows it at the FK barrier — asserted via rejects /COPY
failed for File/.
* test(lbug): cover PDG manifest rel-pair collision in overlap + serial (#2203)
Resolves the P2 gap behind the Codex tri-review finding: the manifest rel-pair
collision guard (moved ahead of node COPY in ad195582) had no test. A leaky
graph with a structural BasicBlock->BasicBlock edge (routed by id-prefix, no
BasicBlock nodes — isolating the rel-pair clash from the node-CSV one) plus a
PdgEmitSink manifest declaring the same pair makes loadGraphToLbug reject with
the rel-pair collision error, asserted on both the overlap (default) and serial
(GITNEXUS_SERIAL_LBUG_LOAD=1) paths.
* perf(lbug): yield the event loop periodically during relationship emit (#2203)
Resolves a P2 review finding on PR #2226: the relationship-emit loop ran long
synchronous stretches between write-stream drain awaits, which could starve the
overlapped node-COPY callbacks on fast I/O and erode the node-COPY-||-rel-emit
overlap. Yield via setImmediate every REL_YIELD_EVERY (5000) edges so the node
COPY and drains get scheduling time. Scheduling-only — emit bench fingerprint
unchanged (byte-identical), csv-pipeline determinism + overlap differential
green.
* refactor(lbug): extract shared copyCsvWithRetry helper (#2203)
Resolves a P2 maintainability finding on PR #2226: the COPY + IGNORE_ERRORS
retry block was duplicated in copyNodeCSVs and the inline relationship-COPY
loop. Extract copyCsvWithRetry(conn, query, onError); the callback receives the
RAW retry error so each site keeps its own message shape + slice length (node
throws, slices 200; relationship warns + records the failed pair, slices 80).
Behavior-preserving — guarded by the live-DB round-trips plus the new
node-COPY-failure and overlap error-path tests.
* docs(lbug): document loadGraphToLbug non-transactionality (#2203)
Resolves the advisory review finding on PR #2226: loadGraphToLbug runs
independent COPYs with no surrounding transaction, so a mid-load failure leaves
a partial DB and recovery is a --force re-analyze. Make that contract explicit
on the function so callers don't assume atomicity.
* perf(lbug): add PROF_LBUG_LOAD persistence-path timing breakdown (#2203 U1)
loadGraphToLbug is un-timed today; the analyze 'emit' number is the
scope-resolution emit bucket, not the CSV->COPY persistence path. Add a
zero-cost-when-off per-stage breakdown (csv-emit/copy-nodes/rel-split/
copy-rels/fallback/total + node/rel counts) gated by PROF_LBUG_LOAD=1,
mirroring the PROF_SCOPE_RESOLUTION pattern. Document the flag in README.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(lbug): route relationships to per-pair CSVs in the emit pass (#2203 U2)
Relationships were written once to a monolithic relations.csv, then re-read
line-by-line (regex per edge) and re-split into per-FROM->TO-label-pair files
before COPY — writing and reading the entire ~1M-edge set twice. Route each
edge to its pair file directly during the single emit pass via a shared
RelPairRouter, eliminating the monolithic write + re-read + per-edge regex.
The router applies the SAME getNodeLabel + validTables filter as the legacy
splitRelCsvByLabelPair, which is retained as a differential oracle. A new
differential test asserts the direct-emit per-pair files are byte-for-byte
identical to the oracle's, with identical skip/total accounting. The prof
line (U1) drops its rel-split stage (routing now folds into csv-emit).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(lbug): skip per-row microtask tick in BufferedCSVWriter (#2203 U3)
addRow awaited an already-resolved promise on every buffered row, scheduling
a microtask per node even when nothing flushed (millions at scale). It now
returns a promise ONLY when it flushes; the node-emit loop awaits once per
iteration after the switch. Flush/drain semantics are unchanged, so
backpressure on the rows that actually write is preserved and the emitted
CSV bytes are byte-identical (covered by the determinism + differential tests).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* bench(lbug): emit throughput + byte-identity gate for the persistence path (#2203 U4)
Build-free bench (bench/emit-persistence/measure.mjs) times streamAllCSVsToDisk
on a synthetic graph at two scales and gates: (1) an order-independent sha256
fingerprint over every emitted CSV line — the byte-identity guard for the U2/U3
emit optimisations — and (2) a scaling-ratio budget catching an O(n^2) emit
re-regression. Wired into ci-tests.yml alongside the cfg/scope-capture benches.
The LadybugDB COPY half needs a real DB, so its timing stays in PROF_LBUG_LOAD
+ the integration round-trip tests (documented in the bench README, with the
deferred COPY-parallelism follow-up).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(review): apply autofix feedback (#2203)
- P1: router backpressure drain-await rejected with a generic AbortError,
masking the real EMFILE/disk-full error. Expose RelPairRouter.lastError and
rethrow it in the emit catch — mirrors the oracle's throw streamError ?? err.
- P1: cover RelPairRouter error + backpressure + teardown paths with a new
unit test (test/unit/rel-pair-routing.test.ts) using an injected mock stream.
- P2: wrap streamAllCSVsToDisk body in try/finally so the setMaxListeners bump
is always restored (the U2 rel-routing throw path could leak it).
- P2: dedup WriteStreamFactory — re-export the canonical type from
rel-pair-routing instead of a second identical declaration.
- P2: annotate splitRelCsvByLabelPair @internal as the retained differential
oracle so a future dead-code sweep doesn't delete the byte-identity guard.
- P3: differential test now covers the proc_ prefix + clears
GITNEXUS_SORT_GRAPH_OUTPUT to prevent env-leak desync.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(lbug): scope byte-identity to quote-free ids + lock the quote-in-id divergence (#2215 review)
The 'byte-identical' claim was unconditional, but the router derives labels from
the raw id while the retained splitRelCsvByLabelPair oracle re-derives them via a
regex over the escaped row — so for an id containing a double-quote they diverge
(the router is the more-correct path). Soften the wording in rel-pair-routing.ts,
the bench README, and the differential-test comment to document the exception,
and add a differential test asserting the intended divergence (router routes the
quote-in-id edge; oracle drops it) so a future change can't silently revert to
the buggy regex semantics.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* bench(lbug): per-file fingerprint so the gate catches pair-file mis-routing (#2215 review)
fingerprintEmit flattened every line of every per-pair file into one array,
sorted globally, and hashed — losing file boundaries, so a row routed to the
WRONG pair file produced an identical fingerprint. Hash a per-file digest
(filename + sha256(file bytes)) and combine the sorted entry list, so mis-routing
(and within-file row reordering) now changes the fingerprint. Baseline
regenerated; the new scheme yields a different hash on byte-identical emit,
confirming it is sensitive to file structure the old flatten ignored.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* bench(lbug): add absolute large-scale wall-time backstop to the emit gate (#2215 review)
The scaling-ratio gate only compares large/small, so a uniform Nx slowdown at
both scales passes with ratio ~1.0. Add an opt-in max_ms_large ceiling (1000ms
vs observed ~200ms — generous, host-noise-tolerant) that --check enforces
alongside the ratio, catching a gross absolute regression the ratio misses.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(lbug): cover the sorted-output path in the byte-identity differential (#2215 review)
The differential test only exercised the default insertion-order emit path. Add
a case under GITNEXUS_SORT_GRAPH_OUTPUT=1 that feeds the oracle the same
id-sorted order orderedRelationships() uses and asserts per-pair byte-identity,
so within-pair row reordering on the sorted path can't slip past the gate.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(lbug): cover the invalid-TO-label skip branch (#2215 review)
Only an invalid-FROM label was exercised; the validTables skip is an OR over
both endpoints, so the invalid-TO branch was untested (an inverted && would
have slipped through). Add a valid-FROM/invalid-TO edge to the differential
test and the router unit test, asserting it's skipped identically by router and
oracle.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(lbug): exercise the BufferedCSVWriter FLUSH_EVERY boundary in vitest (#2215 review)
The U3 addRow change (returns a flush promise only on flush; undefined when
buffered) and the loop's `if (pending) await pending` were only crossed by the
bench, never vitest (all fixtures are <500 nodes). Add a 600-node graph through
streamAllCSVsToDisk asserting all rows land exactly once across the 500-row
flush boundary — no drops, dups, or corruption.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(lbug): drop redundant step cast in buildRelRow (#2215 review)
GraphRelationship.step is already typed number?, so (rel as { step?: number }).step
was a no-op structural cast that obscured the shared-type coupling. Use rel.step
directly. Byte-identical — bench fingerprint unchanged, differential test green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(lbug): make the unknown-label node drop explicit (#2215 review)
With the U3 `let pending` switch idiom, a node whose label matches neither
codeWriterMap nor multiLangWriters left `pending` undefined and was silently
dropped — a footgun for a future node type. Add an explicit else with a comment
documenting that unknown labels are intentionally not persisted and that a new
type must be wired into a writer map. No behavior change (byte-identity + tests
unchanged).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(lbug): drop the unused WriteStreamFactory re-export (#2215 review)
The type was re-exported from lbug-adapter 'to preserve this module's surface,'
but no external code imports it by name from here (the only test reference is a
comment). Keep the import from rel-pair-routing.ts (its canonical home, still
used by splitRelCsvByLabelPair's signature) and drop the dead re-export.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): retain dense reaching-defs as differential oracle + fuzz harness (#2201 U1)
* refactor(cfg): extract shared harvest/adjacency/sweep + swappable in-set computer (#2201 U2)
* perf(cfg): sparse change-driven reaching-defs solver + canonical truncation (#2201 U3,U4)
* perf(cfg): switch production reaching-defs to the sparse solver (#2201 U5)
* perf(cfg): true SSA-sparse reaching-defs solver with auto-dispatch (#2201 U3)
Replace the per-variable worklist (correct but no faster — it still walks
pass-through blocks per binding) with Cytron SSA: CHK dominators + dominance
frontiers + phi-placement + stack renaming over a synthetic entry, answering
block-entry reaching queries by walking the SSA def-use graph (SCC-condensed,
cycle-safe). Pass-through blocks carry the dominating def via the rename stack
and phi-nodes statically capture loop merges, so dense-bindings drops from
O(n^2) to O(n) (5-23x faster, asymptotic) and deep nests are depth-independent.
The sweep now queries a lazy reachingAt accessor with a sparse intra-block
overlay (no full per-block lattice copy). Production auto-dispatches: SSA for
looping functions >=16 blocks (where it pays off, incl. the deep nests the
dense ceiling used to truncate -> ceiling stops firing), dense elsewhere (small
/ loop-free functions, 1.0x — no regression). Throw-edge and unreachable-block
functions fall back to dense (byte-identical). Held byte-identical to the dense
oracle across a 300k-CFG (~1.2M-comparison) differential fuzz.
* test(cfg): R5 contrast — dense ceiling fires, SSA solver converges (#2201 U6)
* bench(cfg): deep-nest scenario + tighten dense-bindings rd budget 10->2 (#2201 U7)
dense-bindings rd_scaling drops 5.2->0.86 (SSA linear); budget tightened to 2.0.
New deep-nest scenario (N nested loops, one carried var) measures rd under the
production blocks×64 ceiling and asserts the SSA solver still COMPUTES full
facts (facts_large_min) where the dense worklist would truncate — the
ceiling-stops-firing acceptance. CFG fingerprints unchanged.
* docs(cfg): document SSA-sparse solver + resolve the WTO no-go note (#2201 U8)
* fix(review): apply autofix feedback (#2201)
- Close the production SSA-dispatcher fuzz-coverage gap: the generator's
maxBlocks=14 was below SSA_MIN_BLOCKS=16, so the auto-dispatcher's SSA branch
was never differentially fuzzed. Raise to 36, add a hadLargeLoop coverage
assertion + a back-edge-into-entry canonical CFG. Validated byte-identical on
100k random CFGs incl. >=16-block looping shapes via both entry points.
- Correct stale function JSDocs + @internal annotations (dispatch/fallback roles).
- Add an independent rd_all_computed bench gate (catches partial truncation).
- maxBlockVisits comment, SSA_MIN_BLOCKS calibration note, nx->next rename.
* fix(cfg): gate out-of-range binding indices to the dense fallback (#2201 review)
Tri-review (adversarial lane, reproduced) found the SSA path less tolerant than
the dense oracle it replaced: an out-of-range binding index in defs/uses/mayDefs
(a corrupted/stale durable store) crashed the nBindings-sized arrays
(defBlocks[v]/stacks[u]), where dense tolerated it as a Map key. The throw
escaped the unguarded taint/harvest call sites and lost a whole file's taint
layer. Add a malformed-input gate that falls back to the dense solver (which
handles any index), preserving byte-identity AND the graceful per-function
degradation. Add an OOB canonical CFG to the differential fuzz + a production-
entry no-throw unit test (the generator only ever emitted in-range indices, so
this divergent input was structurally invisible).
* perf(cfg): bound the SSA value-graph, fall back to dense when oversized (#2201 review R1)
maxFacts bounds fact materialization in sweepFacts, but nothing bounded the
SSA-sparse solver's φ/value-graph construction. A high-binding-density deep
loop routed to SSA (≥16 blocks + a reachable loop) builds an O(blocks×bindings)
value graph the dense path would have truncated at its maxBlockVisits ceiling
(~1.5 GB measured on a 3000-block × 300-binding function).
Cap the value graph: after φ-placement (where nodeKeys.length == the φ count,
the input-superlinear term) plus a 2×Σgen bound on the renaming nodes, fall
back to computeInSetsDense before paying for renaming + Tarjan SCC. The fallback
is byte-identical (dense is the equivalence oracle) and bounded (dense honors
maxBlockVisits). Mirrors the existing throw/unreachable/OOB-binding gates.
The ceiling is DEFAULT_MAX_SSA_VALUE_GRAPH_NODES (1e6 — far above any real or
benchmarked function; dense-bindings/deep-nest build <1e4), overridable per call
via ReachingDefsLimits.maxSsaValueGraphNodes. The new unit test makes the
otherwise-invisible routing flip observable by pairing the cap with a tight
maxBlockVisits (dense truncates, SSA computes). Equivalence fuzz unchanged
(byte-identical, 20k CFGs green); tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(cfg): alias single-source SCC reaching-sets in reachByScc (#2201 review R2)
The SCC-condensation pass built a fresh Set for every SCC and copied each
cross-SCC operand's reaching-set element-by-element — O(defs²) at wide-fan-in φ
merges (a φ over many predecessors, each carrying a large reaching-set).
Add an alias fast path: an SCC with no own leaf keys whose cross-SCC operands
all resolve to ONE source SCC has exactly that source's reaching-set, so share
it by reference instead of copying. This is the common shape (pass-through φ /
single-operand value node). The full union is still built when an SCC has own
keys or genuinely merges ≥2 distinct sources.
Safe to share: reachByScc sets are read-only after construction (operand SCCs
are numbered before s in Tarjan's reverse-topological order and are only
iterated), and contents are identical — set iteration order is irrelevant
because sweepFacts sorts each use's keys before emission (KTD6). Byte-identical
to the dense oracle (30k-CFG fuzz green); tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(cfg): fold the SSA reachability gate into the RPO pass (#2201 review R8)
computeInSetsSparse ran a standalone reachability BFS to gate unreachable-block
functions to the dense oracle, then immediately computed a reverse-post-order
over the synthetic-entry graph — two traversals of the same successor structure.
reversePostOrder now returns the reachability bitmap its DFS already builds, and
the sparse path reuses it for the unreachable-block gate (S→entry is S's only
edge, so reachX[b] for b<n is exactly "reachable from entry" — identical to the
removed BFS). One traversal instead of two on every SSA-dispatched function.
The dispatcher's hasReachableLoop pass is left in place: it decides SSA-vs-dense
BEFORE the solver is entered, and computeInSetsSparse must stay self-contained
(the equivalence fuzz drives it directly, bypassing the dispatcher), so the two
cannot share a traversal without coupling the InSetsComputer contract.
Routing and facts unchanged — byte-identical to the dense oracle (30k-CFG fuzz,
including unreachable-block shapes, green); tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(cfg): trim per-statement/per-use/per-block allocations (#2201 review R9)
Three transient allocations in the hot paths, all behavior-preserving:
- sweepFacts: replace the per-statement `new Set([...defs, ...mayDefs])` with a
direct `includes()` scan over the (1–3 element) def/mayDef arrays, guarded by a
cheap hasSelfDefs flag that short-circuits pure-use statements.
- sweepFacts: reuse a single scratch array for each use's reaching def-keys
instead of spreading a fresh array per use. The KTD6 pre-sort still runs in
place (load-bearing for truncated byte-identity).
- computeInSetsSparse: build dPredsX by skipping consecutive-equal `from` values
(preds[b] is pre-sorted by buildAdjacency, so duplicates are adjacent) instead
of a per-block Set + spread + sort; the synthetic entry S = n exceeds every
block index so it appends in order.
The sweep is shared with the dense oracle, so these stay byte-identical on both
paths — 50k-CFG fuzz (incl. maxFacts truncation, the order-sensitive case)
green; tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(cfg): correct the sweepFacts truncation byte-identity mechanism (#2201 review R6)
The outer sweepFacts JSDoc attributed a truncated result's cross-solver
byte-identity to the two solvers producing "identical inSets — insertion order
included". That is wrong: the dense (RPO fixpoint) and SSA (renaming/SCC)
solvers deliberately build a loop-carried use's reaching set in DIFFERENT
insertion orders — same set, different order. The actual mechanism is the KTD6
per-use sort that canonicalizes each use's keys by defKey BEFORE the maxFacts
cutoff (already documented correctly on the inner comment). Rewrite the outer
doc to say so. Documentation only.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cfg): extract pure graph sub-stages to reaching-defs-graph.ts (#2201 review R4)
reaching-defs.ts had grown to ~1190 lines with the #2201 SSA rewrite. Move the
self-contained, pure (plain-array) algorithms into a sibling module:
- reversePostOrder
- buildDominators (Cooper-Harvey-Kennedy)
- buildDominanceFrontiers (Cytron)
- tarjanScc + condenseReachingSets (SCC condensation, alias fast path)
- hasReachableLoop (dispatcher loop check)
- unionSets / latticeEquals (def-set / lattice primitives)
The new module has a STRICT one-way dependency (it imports nothing from
reaching-defs.ts — every helper is parameterized over plain arrays/Sets), so
there is no import cycle and each stage is independently testable. reaching-defs.ts
now holds the orchestrator, the two solver bodies, harvest, adjacency, the
statement sweep, and the dispatcher: 1190 → 988 lines.
Pure mechanical extraction — behavior is preserved by the differential
equivalence fuzz (40k CFGs byte-identical) + the reaching-defs unit/snapshot
suites; tsc clean. The helpers are @internal (kept out of the shipped .d.ts by
the stripInternal change).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(pdg): stamp the reaching-defs solver identity for incremental re-analysis (#2201 review R3)
The SSA-sparse rewrite computes full REACHING_DEF facts for deep-loop functions
the old dense worklist truncated to empty at the blocks×64 ceiling. But an
existing `--pdg` index carries those stale-truncated rows, and nothing forced a
re-analysis: RepoMeta.pdg had no solver-identity key, so an upgraded run over an
unchanged file kept the incremental fast path and never recomputed.
Add a constant `reachingDefSolver: 'ssa-sparse-v1'` to the resolved pdg stamp
(and to the RepoMeta['pdg'] type). It rides the existing key-union
pdgModeMismatch comparator: a pre-#2201 stamp lacks the key, so
'ssa-sparse-v1' !== undefined trips one full writeback that recomputes the
fuller coverage — no `--force` needed — exactly like the M2 REACHING_DEF cap and
M5 CDG cap upgrade paths. A matching post-#2201 stamp compares equal, so there
is no spurious re-analysis churn on steady-state re-runs.
Tests: new pre-#2201→SSA upgrade block in pdg-mode-flip.test.ts (stamp present,
absent-key mismatch, identical-stamp no-churn) + the persisted-stamp shape
assertions and resolvePdgConfig DEFAULTS updated for the new key. tsc clean;
pdg-mode-flip + run-analyze suites green (55/55).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* build(ts): stripInternal so @internal test-only exports stay out of the shipped .d.ts (#2201 review R5)
computeReachingDefsDense/computeReachingDefsSparse are exported only for the
equivalence fuzz and tagged @internal, but `declaration: true` emitted them into
the public dist/**/*.d.ts. stripInternal removes any @internal-tagged export from
the declaration output.
This is repo-wide, which is the intended behavior: the same applies to every
other test-only @internal export (hf-env's withDownloadTimeout etc., worker-pool's
buildDispatchMessage/crashSignature, parse-impl's handleWorkerStartupFailure, the
logger/safe-parse test resets, and the new reaching-defs-graph SSA helpers) — all
of which are documented as not-public.
Verified:
- declaration emit succeeds with no TS4094/TS9006 ("cannot be named") errors;
- the @internal functions are gone from the emitted .d.ts (reaching-defs-graph.d.ts
is now `export {};`), while public symbols (computeReachingDefs) remain;
- gitnexus-web — the only cross-package consumer — typechecks clean and imports
only from gitnexus-shared, never from gitnexus internals;
- runtime .js and the vitest/tsx tests are source-based, so unaffected.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(bench): add wide-merge scenario + tighten deep-nest facts floor (#2201 review R7)
wide-merge: N bindings, each assigned in a 3-way branch (a wide multi-operand φ
per binding) inside a loop, then all used after the merge. Unlike dense-bindings
(one chained redef per `if`), every binding fans into its own wide φ, so the
scenario exercises φ-placement + renaming + the reachByScc condensation across
many independent wide merges. N bindings × constant arms ⇒ O(N) facts, so the
gate is rd_scaling LINEARITY (measured ~1.07; budget 2.0 catches a regression to
the per-binding-rescan O(N²) class the reachByScc alias path guards against). It
runs the production SSA path (10007 blocks + a loop) and computes all facts under
the blocks×64 budget (facts_large_min 24000 of a measured 26008 + the
rd_all_computed gate).
deep-nest: tighten facts_large_min 100 → 150 (measured 164) so a partial-
truncation regression that still cleared 100 — but lost facts — now fails, with
~9% headroom for noise.
bench --check PASS (9 scenarios) under --expose-gc; all existing CFG fingerprints
unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style(cfg): drop trailing blank line in reaching-defs.ts (prettier)
Whitespace-only — a stray trailing newline left by the U4 extraction. `prettier
--check` (the root format CI gate) now passes on every changed file. No behavior
change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): model Java value-position switch as control flow (#2207)
A value-position `switch` expression with ≥2 arms is now modeled as a
CFG dispatch in the two highest-value carriers, instead of collapsing
the owning statement to a single inline block:
- `var x = switch (k) { … }` — the arms become real blocks reached by
`switch-case` edges and rejoin at a binding continuation that carries
the declared name's def (uses stay on the arm blocks).
- `return switch (k) { … }` — each arm returns the function result,
threading every active finalizer.
This makes the arms control-dependent on the dispatch (the point of
#2207 — they previously produced zero CDG), mirroring the Kotlin / Rust
value-position binding pattern. `breaksBlock` routes a value-switch
declaration out of `visitSeq` coalescing; `visitReturn` and `visitStmt`
gain the carrier handling; `java-harvest` gains `bindingDefFacts`.
An assignment RHS (`x = switch …`), a call argument, and a multi-
declarator decl remain inline (documented gap). Java has no value-
position `if` (the ternary is excluded, like Kotlin's elvis).
Verified: 51 java-visitor tests (4 new), full CFG unit+integration
suites (661) green, CDG snapshot byte-identical, bench --check PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): model C# value-position switch expression as control flow (#2207)
A C# `switch_expression` (`k switch { p => v, … }`) with ≥2 arms is now
modeled as a CFG `switch-case` dispatch — a discriminant block, each arm
value a block reached by a dispatch edge, all arms rejoining at one exit —
in the three value-position carriers, instead of collapsing the owning
construct to a single inline block:
- `var x = k switch { … }` — arms rejoin at a binding continuation that
carries the declared name's def (discriminant + arm uses on the arms).
- `return k switch { … }` — each arm returns the function result,
threading every active finalizer.
- `=> k switch { … }` expression-bodied member — each arm returns.
The arms are now control-dependent on the discriminant (the point of
#2207 — they previously produced zero CDG). Arm patterns / `when` guards
are harvested as conditional uses on the dispatch; an unguarded `_`/`var`
arm is the exhaustive catch-all (a non-exhaustive switch keeps EXIT
reachable via a no-match edge). `switch_expression` is distinct from
`switch_statement`, so this adds a dedicated `visitSwitchExpr`.
An assignment RHS (`x = k switch …`), a call argument, and a multi-
declarator decl remain inline (documented gap). `csharp-harvest` gains
`bindingDefFacts`.
Verified: 47 csharp-visitor tests (4 new), full CFG unit+integration
suites (664) green, CDG snapshot byte-identical, bench --check PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): model PHP value-position match expression as control flow (#2207)
A PHP `match($v) { c => v, default => v }` with ≥2 arms is now modeled as
a CFG `switch-case` dispatch — a discriminant block, each arm value a
block reached by a dispatch edge, all arms rejoining at one exit (no
fallthrough) — in the two value-position carriers, instead of collapsing
the owning statement to one inline block:
- `$x = match($v) { … }` — the dominant PHP idiom (no typed local decl):
arms rejoin at a binding continuation carrying the assignment target's
def (condition + arm uses on the arms).
- `return match($v) { … }` — each arm returns the function result,
threading every active finally.
The arms are now control-dependent on the discriminant (the point of
#2207). Arm `match_condition_list`s are harvested as conditional uses on
the dispatch; a `default` arm is the catch-all (a defaultless `match`
throws UnhandledMatchError, kept EXIT-reachable via a no-match edge).
`php-harvest` gains `assignmentDefFacts`.
A `match` in a call argument / nested subexpression stays inline; the
ternary `?:` is excluded by design (a micro-branch, like elvis).
Verified: 37 php-visitor tests (3 new), CDG snapshot byte-identical,
bench --check PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): model Dart value-position switch expression as control flow (#2207)
A Dart 3 value-position `switch (v) { p => e, _ => e }` with ≥2 arms is
now modeled as a CFG `switch-case` dispatch — a discriminant block, each
arm value a block reached by a dispatch edge, all arms rejoining at one
exit (no fallthrough) — in the two value-position carriers, instead of
collapsing the owning statement to one inline block:
- `var x = switch (v) { … }` — single-binding decl; arms rejoin at a
binding continuation carrying the declared name's def.
- `return switch (v) { … }` — each arm returns the function result,
threading every active finalizer.
The arms are now control-dependent on the discriminant (the point of
#2207). A Dart call value parses as `identifier` + `selector` (multiple
children, not one node), so the arm-value facts come from a dedicated
`switchExprArmValueFacts`; arm patterns harvest conditionally onto the
dispatch; a `_` arm is the catch-all (a non-exhaustive switch keeps EXIT
reachable via a no-match edge). `dart-harvest` gains `bindingDefFacts` +
the arm value/pattern fact helpers.
A `switch_expression` in a call argument / multi-binding decl stays inline
(its conditional arm sub-evaluation remains the #2206 harvest may-def
path, re-pointed in the regression test). `?:`/`??`/`?.` excluded by
design.
Verified: 39 dart-visitor tests (4 new), CDG snapshot byte-identical,
bench --check PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): model Swift value-position if/switch as control flow (#2207)
A Swift 5.9 value-position `if`/`switch` is now modeled as control flow in
the two value-position carriers, instead of collapsing the owning
statement to one inline block:
- `let x = if … else … / switch v { … }` — arms rejoin at a binding
continuation carrying the declared name's def (condition + arm uses on
the branch blocks).
- `return if … / switch …` — each arm returns the function result,
threading every active finalizer.
The arms are now control-dependent on the branch (the point of #2207).
tree-sitter-swift reuses `if_statement` / `switch_statement` for the value
form (no separate `if_expression`/`switch_expression`), so the existing
`visitIf`/`visitSwitch` are reused — this mirrors the Kotlin carrier
exactly. `swift-harvest` gains `bindingDefFacts`.
A value-position `if` requires an `else`; a value `switch` needs ≥2
entries. A value branch in a call argument / interpolation stays inline;
`?:`/`??` are excluded by design.
Verified: 31 swift-visitor tests (4 new), CDG snapshot byte-identical,
bench --check PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): model Kotlin assignment-RHS and value-position try as control flow (#2205)
Completes the value-position branch carriers #2205 left deferred after the
initial val/var-binding + return + expr-body work:
- `x = when (k) { … }` / `x = if (c) a else b` / `x = try { … }` — a plain
`=` assignment whose RHS is a modelable branch now models the arms as
control flow and binds the LHS target at the rejoin (a compound `+=`
and a plain-call RHS stay inline).
- `val x = try { … } catch { … }` — a value-position `try` is now a
modelable value branch (reusing visitTry), so the binding/assignment
carriers route it through control flow too.
The arms are now control-dependent on the branch (the point of #2205).
`isModelableValueBranch` gains `try_expression`; `visitBranchExpr` routes
it to `visitTry`; `isControlFlow`/`visitStmt` gain the `assignment`
carrier; `kotlin-harvest` gains `assignmentDefFacts`.
A branch nested in a call argument (`f(when …)`) stays inline (the direct
value is the call); `?:`/`?.` micro-branches excluded by design. The
`return try { … }` carrier is intentionally left out (finalizer-threading
in return position is risky and was not requested).
Verified: 43 kotlin-visitor tests (5 new), full CFG unit+integration
suites (676) green, CDG snapshot byte-identical, bench --check PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): Java colon-form value switch — yield ends the arm, no fallthrough (#2211)
Tri-review (adversarial + correctness lanes) found that a value-position
colon-form switch expression — `int x = switch(k){ case 1: yield a();
case 2: yield b(); }` (valid Java 14+) — reused the statement `visitSwitch`
fallthrough logic, wiring a spurious `fallthrough` edge between the
yield-terminated colon groups. A switch EXPRESSION never falls through
between arms; the false edge dropped an arm's control-dependence edge and
added a false reaching-defs propagation edge (verified by a real-parser
probe). Arrow-form value switches were already correct.
Root cause: `visitYield` modeled `yield e;` as a block that CONTINUES to
the next statement. Semantically `yield` produces the switch-expression's
value and EXITS the switch. Fix: `visitYield` now terminates the arm,
jumping to the enclosing switch's exit and threading any finalizer it
crosses — exactly like a `break` out of the switch, but carrying the
yielded value's facts. Adds `ControlFlowContext.resolveYield()` (nearest
SWITCH frame, never an intervening loop). `yield` is Java-only here (C#
`yield return` is iterator semantics, untouched).
Tests: a colon-form value switch asserting NO `fallthrough` edge and that
BOTH arms are control-dependent on the dispatch (specific controller→
dependent pairs), plus a `return switch(…)` inside `try/finally` asserting
`finally-return` threading per arm.
Verified: full CFG unit+integration suites green, CDG snapshot byte-
identical, bench --check PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): Dart value switch — guarded `_` is not a catch-all; guard is a dispatch test (#2211)
Tri-review (adversarial + correctness lanes) found two issues in the Dart
`visitSwitchExpr` value-position modeling:
1. Catch-all detection was `pattern.text === '_'`, ignoring guards. A
guarded `_ when c => …` is NOT exhaustive (Dart throws at runtime if
no arm + guard matches), so falsely treating it as a catch-all
suppressed the conservative no-match edge — asserting an exhaustive
switch that isn't. The sibling C# visitor already gated catch-all on
`!guard`.
2. A `when` guard parses as a bare sibling between the pattern and the
value (no wrapper node), so it fell into the arm-VALUE children and
was harvested as an unconditional arm-value use instead of a
conditional dispatch test.
Fix: new `armParts()` splits a `switch_expression_case` at the `=>` token
into pattern / guard(s) / value(s). The pattern AND guard are harvested
conditionally onto the dispatch block (they evaluate before the body, only
when earlier arms missed); the arm-value facts come from the post-`=>`
children only; the catch-all is gated on an unguarded `_`. Removes the now-
unused dart-harvest `switchExprArm{Value,Pattern}Facts` (the visitor
harvests per-child via the existing `facts`/`factsConditional`).
Tests: a guarded value switch asserting the no-match edge is present (3
switch-case successors from the dispatch) and EXIT stays reachable, plus a
test that the guard's use is recorded on the dispatch block, not an arm.
Verified: full CFG suites green, CDG snapshot byte-identical, bench --check PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): Kotlin return try {…} models the value-position try (#2205, #2211)
Tri-review (maintainability + testing lanes) caught a doc-vs-code mismatch:
the visitor header documents a `return <branch>` carrier that includes
`try`, but `visitReturn` only matched `when_expression`/`if_expression`, so
`return try { … } catch { … }` fell through to the single inline-block path
— its arms were not modeled.
Fix: `visitReturn` now also matches `try_expression`, making the `return`
carrier uniform with the binding / assignment / expression-body carriers
(all route a value-position `try` through `isModelableValueBranch` →
`visitTry`, threading the active finalizers per arm). No new machinery —
just completes the carrier set the docstring already claimed.
Tests: `return try {…} catch {…}` (throw + return edges, CDG-bearing,
EXIT reachable) and `x = try {…} catch {…}` assignment-RHS (the assignment
carrier's try path, previously untested).
Verified: full CFG suites green, CDG snapshot byte-identical, bench --check PASS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): cover the no-match edge of non-exhaustive C#/PHP value switches (#2211)
The testing lane noted the `hasCatchAll`/`hasDefault === false` branch —
where a value-position `switch`/`match` with no catch-all/default arm adds
a conservative no-match edge so EXIT stays reachable — was untested (every
existing fixture used a `_`/`default` arm). Adds a C# `x switch { 1 => …,
2 => … }` and a PHP `match($x){ 1 => …, 2 => … }` (no default) test, each
asserting the dispatch fans to (arms + 1) `switch-case` successors and
`isExitReachableFromAllBlocks` holds.
Verified: full CFG suites green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(cfg): clarify Kotlin value-position try gate covers the finally-only path (#2211)
Tri-review (maintainability lane) noted the inline comment at the
`try_expression` branch of `isModelableValueBranch` said only "a catch's
value", but the gate fires on `catch_block || finally_block`. Reword to
acknowledge that a value-position `try` with a `catch` OR a `finally` is a
modelable branch. Comment-only; no behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(cfg): document the deliberate C# declaratorInit duplication (#2211)
Tri-review (maintainability lane) flagged the byte-identical `declaratorInit`
helper in `csharp.ts` (visitor) and `csharp-harvest.ts` (harvester). The two
are standalone classes with no shared base (repo convention) and the only
module both import is the generic `utils/ast-helpers` (types only) — not a
home for a C#-grammar-specific helper. Resolve the lowest-risk way: a
cross-reference comment at each definition noting the deliberate duplication
and the keep-in-sync requirement. No new shared module; no behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cfg): unwrap the paren in Dart visitSwitch for dispatch consistency (#2211)
Tri-review (maintainability lane) noted `visitSwitch` (statement form) used
the raw parenthesized `condition` (dispatch text `switch (x)`) while the new
`visitSwitchExpr` unwraps it (`switch x`). Probed the vendored tree-sitter-dart:
the `switch_statement` condition IS a `parenthesized_expression`, so apply
`unwrapParen` in `visitSwitch` too. The harvest walks into the paren either
way, so the discriminant's def/use facts are unchanged — only the dispatch
block's text string normalizes. Verified byte-identical (cdg-snapshot +
bench --check unchanged; the Dart unit tests assert topology, not block text).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): Swift single-entry value switch stays inline (#2211)
Tri-review (testing lane) noted the existing Swift "stays inline" test used
a plain call (`let x = g()`), which never exercises the single-entry switch
gate. Add a real one-entry value switch (`let x = switch v { default: g() }`)
asserting it coalesces (no switch-case edge), pinning the `>= 2` switch_entry
threshold in `isModelableValueBranch`.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): cover the Kotlin expression-body try carrier (#2205, #2211)
Tri-review (testing lane) noted the `fun f() = try { … } catch { … }`
expression-body carrier (visitExprBody -> isModelableValueBranch accepting
try_expression) existed but was untested. Add a regression asserting the
expr-body try is modeled (throw + return edges, CDG-bearing, EXIT reachable).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): pin the value-branch carriers never throw on a truncated AST (R4) (#2211)
Tri-review (adversarial lane) noted the new value-position branch carriers
must return undefined / never throw on a malformed AST, or a single bad
function would drop the whole file's CFG group (the R4 invariant). Add a
per-language regression feeding a TRUNCATED value-branch carrier (an
unterminated `var x = switch/match/if/when (…)`) through the existing
`collectFunctions` + `buildFunctionCfg(...).not.toThrow()` graceful-undefined
harness, for all six languages whose value-branch path is new
(Java/C#/PHP/Dart/Swift/Kotlin).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): validate cfg/visitors literals + drop 3 dead TS node types
Extend the grammar-literal CI gate (test/helpers/literal-collectors.ts)
to scan cfg/visitors/*.ts, mapping each visitor file to its grammar via
the existing basename rule (c-cpp -> C/C++, csharp -> C#, java -> Java,
go -> Go, typescript -> TS). Closes the gap where the gate never
validated CFG visitor node-type literals -- the prerequisite for adding
C-family visitors safely (#2195 U1).
The newly-scanned TS visitor surfaced 3 dead literals absent from every
grammar it serves (typescript/javascript/tsx all = 0): for_of_statement
(for-of parses as for_in_statement), async_function_declaration and
async_arrow_function (async functions are function_declaration /
arrow_function + an async child). Removed them; behavior-preserving --
the cases never matched, bench --check fingerprints unchanged, TS
visitor unit tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): language-agnostic CFG unit-test harness (#2195 U1)
Extract the grammar-agnostic engine from ts-cfg-harness into
makeCfgHarness(grammar, visitor, filePath) at test/helpers/cfg-harness.ts.
Function discovery delegates to visitor.isFunction, so the harness carries
no language-specific node-type knowledge -- each C-family visitor's unit
tests can drive the real worker-side builder against real source.
ts-cfg-harness becomes a thin TS binding re-exporting the same
parse/collectFunctions/cfgOf/cfgsOf (behavior-preserving: all 5 existing
consumers -- taint propagate/model-match/summary-harvest/taint-emit + cfg
harvest -- pass unchanged, 223 tests green). New harness.test.ts proves
TS-faithfulness and isFunction-delegation via a stub visitor.
The bench parameterization (measure.mjs) is sequenced into U7, where the
first C-family scaling scenario makes the {grammar, visitorFactory} seam
validatable against a real non-TS language.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): C and C++ CFG visitor + def/use harvest (#2195 U2)
Add createCCfgVisitor/createCppCfgVisitor over a shared CCfgWalk core.
Grammar introspection confirmed tree-sitter-c and tree-sitter-cpp share
every control-flow node type/field, so CppCfgWalk extends CCfgWalk with
only the C++-only nodes (try/catch/throw/for_range_loop/lambda) via a
visitExtra hook -- no language conditionals (AGENTS no-language-naming).
Wire both into c-cpp.ts providers.
Harvest (c-cpp-harvest.ts): two-phase binding table + per-statement
defs/uses/mayDefs (no sites[] yet -- U6). Edge kinds match the TS
contract; functionStartColumn populated; non-terminating loops (for(;;),
while(1)) emit the structural exit-escape edge so EXIT stays
reverse-reachable and CDG is not silently skipped -- verified against the
production post-dominator + control-dependence solvers (for(;;) -> 3 CDG
edges). buildFunctionCfg returns undefined rather than throwing.
23 real-parser regression tests; grammar-literal gate green (literals
validated against both grammars). Documented gaps: C++ RAII destructors,
setjmp/longjmp, computed goto (route to EXIT + warn).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): C# CFG visitor + def/use harvest (#2195 U3)
Add createCsharpCfgVisitor + csharp-harvest over the shared CfgBuilder /
ControlFlowContext, modeling the C# statement taxonomy: if/else,
for/foreach/while/do, switch_section (+ switch_expression arms),
try/catch/catch_filter/finally, using + lock (deterministic finalizers --
dispose/release runs on normal AND exception exit, finally-* completion
edges on crossing jumps), goto/labeled, yield (surface only), return/
throw/break/continue. Wire into csharpProvider.
Every literal validated against tree-sitter-c-sharp via the introspection
probe (record_declaration, no else_clause, switch_section, positional
access where no field exists). Edge kinds match the contract;
functionStartColumn populated; while(true) keeps EXIT reverse-reachable
(production CDG probe: 3 edges). buildFunctionCfg returns undefined
rather than throwing.
34 real-parser regression tests; grammar-literal gate green; no
regression (cfg unit dir 256/256, tsc clean). Documented gaps: yield
iterator state machine, goto case/default, async suspension points.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): Java CFG visitor + def/use harvest (#2195 U4)
Add createJavaCfgVisitor + java-harvest over the shared CfgBuilder /
ControlFlowContext: if/else, classic for, enhanced-for, while, do-while,
classic-vs-arrow switch (switch_block_statement_group fallthrough vs
switch_rule no-fallthrough), try/catch/finally + try-with-resources
(auto-close synthesized as a finalizer, closes on normal AND exception
exit) + synchronized (monitor-release finalizer), labeled break/continue
to the labeled frame, yield, return/throw/break/continue. Wire into
javaProvider.
Every literal validated against tree-sitter-java via the probe
(switch_expression covers both switch forms, generic_type, line_comment,
for init field). Edge kinds match the contract; functionStartColumn
populated; while(true)/for(;;) keep EXIT reverse-reachable (production
CDG probe: 3 edges; hazard fixture: 34 CDG edges). buildFunctionCfg
returns undefined rather than throwing.
43 real-parser regression tests; grammar-literal gate green; no
regression (cfg unit suite 304, tsc clean). Documented gaps: switch-as-
expression-value inline, yield state machine, async/field-write defs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): Go CFG visitor + def/use harvest (#2195 U5)
Add createGoCfgVisitor + go-harvest, the highest-divergence target:
for_statement (all four shapes -- for_clause C-style, while-style,
range_clause, bare for{}), expression/type switch (no implicit
fallthrough) + explicit fallthrough_statement, select_statement, defer
(LIFO finalizer legs at function exit), go (call is straight-line; the
closure body is its own CFG via isFunction), labeled break/continue/goto,
multiple-return assigns (a, b := f() defines each LHS). Wire into
goProvider.
CRITICAL (review A2): every non-terminating shape -- for{}, for cond{},
select{} with no default -- emits a structural exit-escape edge so EXIT
stays reverse-reachable and the production CDG is not silently skipped.
Verified: for{} -> CDG=3, select{} -> CDG=1, for-range -> CDG=2, all
exitReachable=true.
Every literal validated against tree-sitter-go via the probe. 32
real-parser regression tests; grammar-literal gate green; no regression
(186 across all 5 visitors + gate, full cfg unit 331, tsc clean).
Documented gaps: panic/recover unwind, goroutine happens-before.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): call-site sites[] taint substrate for C-family (#2195 U6)
Extend the C/C++/C#/Java/Go harvests with the call-site sites[] taint
substrate (SiteRecord/SiteArgOccurrence), mirroring the TS shape so the
shared taint matcher consumes all languages uniformly. Extract the
grammar-agnostic site machinery into cfg/visitors/call-site-harvest.ts
(CallSiteFactAccumulator -- names no language); each harvest adds only its
per-grammar visitCall/walkChain over its call node (C/C++ call_expression,
C# invocation_expression, Java method_invocation, Go call_expression).
INERT BY DESIGN: no C-family taint model exists (registerBuiltinTaintModels
is TS/JS only), so getSourceSinkConfig returns undefined for these
languages and the harvested sites produce ZERO TAINTED edges -- the
positive source->sink->TAINTED path is deferred with the model authoring.
sites emitted only when non-empty; facts-only attachment, block/edge
topology unchanged (pre-existing topology + def/use tests byte-identical).
23 new substrate tests; 574 green across the cfg/taint/emit suites; gate
green; tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): worker-mode PDG integration + bench parameterization (#2195 U7)
Prove the five C-family visitors build PDG through the REAL worker
pipeline. pipeline-pdg.test.ts: per-language (C/C++/C#/Java/Go) temp repo
run with pdg:true asserts BasicBlock+CFG+REACHING_DEF+CDG all > 0 (CDG>0
proves EXIT stays reverse-reachable end-to-end through the worker, incl.
each fixture's non-terminating loop/select); a paired run with pdg off
asserts == 0, the two flag-off graphs byte-identical (R3), no PDG types
leak, pinned by a golden snapshot. Counts e.g. Go 151 BB / 56 CDG.
Parameterize bench/cfg/measure.mjs by a per-language LANGS registry
resolved generically via getLanguageGrammar + getProvider(X).cfgVisitor
(no static import table). Default TS byte-identical -- all 6 TS
fingerprints unchanged under --check; taint-dense stays TS-only
(TS_JS_TAINT_MODEL never runs against model-less C-family CFGs). Add a
go:branchy scenario+baseline (namespaced) -- its fingerprint shape
(32 blocks/46 edges) matches TS branchy, cross-validating the Go visitor.
15 pipeline tests + bench --check PASS (7 scenarios); 354 unit cfg green;
dist rebuilt clean. Absorbs the bench parameterization deferred from U1.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): Python CFG visitor + def/use harvest (#2195 U8)
Add createPythonCfgVisitor + python-harvest -- the most structurally
divergent target (indentation blocks, elif, for/while-else, with, try/
except/except-group/else/finally, match/case, comprehensions, walrus),
confirming the shared CfgBuilder/ControlFlowContext core carries no
brace-family assumptions. for/while else-clause sits on the normal-
completion edge (not break); with modeled as try/finally dispose; match
has no fallthrough. Wire into pythonProvider.
Every literal validated against tree-sitter-python via the probe.
while True: keeps EXIT reverse-reachable (production CDG probe: 3 edges;
fixture: 42 CDG edges). 37 real-parser tests; gate green; no regression
(cfg unit 391, tsc clean). Gaps: async/generator suspension, comprehension
scope over-approximation. No sites[] (taint substrate, separate).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): PHP CFG visitor + def/use harvest (#2195 U9)
Add createPhpCfgVisitor + php-harvest: if/elseif/else (+ alt colon
syntax), for/foreach/while/do-while, switch (fallthrough) + match (no
fallthrough), try/catch/finally, break N/continue N (N-th enclosing
loop), goto, return/throw. Wire into phpProvider.
Every literal validated against tree-sitter-php (php_only) via the probe
(for_statement initialize/condition/update; throw_expression not
throw_statement; break/continue integer child). while(true) keeps EXIT
reverse-reachable (production CDG probe: 3 edges; break 2 escapes the
outer loop). 35 real-parser tests.
Also repoint worker-roundtrip's "non-CFG language" gate test from Python
(which now has a cfgVisitor) to COBOL (the permanent non-goal of the
rollout) -- a stale assertion the Python commit invalidated. Full
in-process sweep green (452 across 18 files). Gaps: match inline value,
goto plain-block.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): Ruby CFG visitor + def/use harvest (#2195 U10)
Add createRubyCfgVisitor + ruby-harvest: if/unless/elsif/else +
statement-modifier forms (x if c, x while c), while/until/for (until
inverts the sense), case/when + case/in (pattern, no fallthrough),
begin/rescue/else/ensure (ensure=finally, rescue=catch) + retry
(loop-back into begin), return/break/next/redo, blocks/lambdas as their
own closure CFGs. Wire into rubyProvider.
Every literal validated against tree-sitter-ruby via the probe (case vs
case_match, modifier nodes, typed rescue/ensure children). loop do /
while true keep EXIT reverse-reachable (production CDG probe: 3 edges).
34 real-parser tests; comprehensive sweep green (486). Gaps: yield,
expression-position if/case/begin inline, ivar/gvar non-local defs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): Rust CFG visitor + def/use harvest (#2195 U11)
Add createRustCfgVisitor + rust-harvest for the expression-oriented Rust:
if/else + if-let, loop (infinite -- structural escape edge), while/
while-let/for, match (no fallthrough) + guards, labeled break/continue
('outer), break-with-value, ? operator (try_expression) as an
early-return throw edge to EXIT, let-else (diverging else). visitLet
handles control-flow in value position (let x = loop/if/match). Wire into
rustProvider.
Every literal validated against tree-sitter-rust via the probe (label is
a named child not a field; line_comment; _ pattern). loop {} keeps EXIT
reverse-reachable (production CDG probe: 3 edges). 33 real-parser tests;
comprehensive sweep green (519). Gaps: panic, async/.await, macro bodies.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): Swift CFG visitor + def/use harvest (#2195 U12)
Add createSwiftCfgVisitor + swift-harvest (vendored tree-sitter-swift via
requireVendoredGrammar): if/else + optional binding (if let), guard...else
(diverging early exit), for-in/while/repeat-while (bottom-test), switch
(no implicit fallthrough; explicit fallthrough keyword; where guards),
do/catch + try/try?/try!, defer (LIFO finalizer at scope exit), labeled
break/continue, control_transfer_statement (one node for break/continue/
return/throw). Wire into swiftProvider.
Every literal validated against the vendored grammar via the probe (no
block node; if-let folds into condition+bound_identifier; defer parses as
a call_expression with trailing closure). while true keeps EXIT
reverse-reachable (production CDG probe: 3 edges). 24 real-parser tests;
comprehensive sweep green (543). Gaps: computed properties, defer
block-scope approx, fatalError traps.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): Kotlin CFG visitor + def/use harvest (#2195 U13)
Add createKotlinCfgVisitor + kotlin-harvest (vendored tree-sitter-kotlin):
if/else, when (subject + subjectless, no fallthrough), for/while/do-while,
try/catch/finally, jump_expression (return/return@/break/break@/continue/
continue@/throw), labeled loops, control_structure_body unwrapping,
expression-body functions. The grammar is field-less for control flow, so
the visitor navigates by child type+position. Wire into kotlinProvider.
Every literal validated against the vendored grammar via the probe
(line_comment/multiline_comment, not comment). while (true) keeps EXIT
reverse-reachable (production CDG probe: 3 edges; worker-mode fixture:
BB=82, CDG=41). 28 real-parser tests; comprehensive sweep green (571).
Gaps: value-position if/when/try inline, inline-fun non-local return,
getters/setters.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): Dart CFG visitor + def/use harvest (#2195 U14)
Add createDartCfgVisitor + dart-harvest (vendored tree-sitter-dart):
if/else, C-for/for-in/while/do-while, switch (empty-case fallthrough +
explicit continue-label) + switch_expression, try/on/catch/finally +
rethrow + assert (throw edges), return/break/continue/throw, labeled
loops, arrow bodies, closures. Dart splits a function into sibling
signature + function_body nodes, so the body (or function_expression) is
the CFG-bearing node. Wire into dartProvider.
Every literal validated against the vendored grammar via the probe (only
constant_pattern exists; removed speculative relational/logical pattern
names). while (true) keeps EXIT reverse-reachable (production CDG probe:
3 edges). 34 real-parser tests; comprehensive sweep green (605). Gaps:
labeled-loop grammar quirk (read via ERROR sibling), async straight-line,
value-position if/switch inline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): Vue (reuse TS visitor) + worker-mode proof for all langs (#2195 U15)
Vue SFC <script> blocks are extracted and parsed with the TS grammar
(parse-worker languageMap[Vue] = TypeScript.typescript), so wire
vueProvider.cfgVisitor = createTypeScriptCfgVisitor() -- pure reuse, no
Vue-specific visitor. vue-visitor.test.ts replicates the worker path
(extractVueScript -> TS parse -> CFG) and confirms branch edges + EXIT
reverse-reachable + CDG>0.
Extend pipeline-pdg.test.ts with a worker-mode block covering all eight
remaining languages (Python/PHP/Ruby/Rust/Swift/Kotlin/Dart/Vue): per-
language temp repo, real worker pool, BasicBlock+CFG+REACHING_DEF+CDG all
> 0 with --pdg (CDG>0 proves EXIT reverse-reachable end-to-end through the
worker despite each fixture's non-terminating loop), == 0 without. Counts
e.g. Ruby 122 BB/45 CDG, Vue 49 BB/11 CDG. 30 pipeline tests green.
COBOL: documented as the deliberate PDG non-goal (no grammar, exotic
PERFORM/GO-TO control flow) in cobol.ts + the worker-roundtrip gate.
This completes PDG language coverage: every supported language except
COBOL now builds CFG/REACHING_DEF/CDG under --pdg.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): surface skippedUnsoundFunctions in per-language stats (#2195 U2)
emitFileCdg computes skippedUnsoundFunctions (functions whose CDG is
withheld because EXIT isn't reverse-reachable from all blocks) but run.ts
dropped it on the floor — only cdgEdges/cdgDropped were aggregated. Add
the aggregation + a stats-line segment so CDG coverage gaps are an
explicit signal, not silent. Establishes the baseline skip count that
makes the U1 synthetic-escape pass's effect (the drop to genuine
anomalies only) measurable.
Additive; no emit-logic change. The emit-side field is covered by
cfg-emit.test.ts (asserts skippedUnsoundFunctions===1 + the warn on a
disconnected-block CFG); the run.ts aggregation is a thin pass-through.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): synthetic-escape pass restores CDG for exit-unreachable cycles (#2195 U1)
Unconditional goto-cycles (C/C++/C#/Go) wire a backward seq edge with no
structural exit-escape edge, so EXIT becomes non-reverse-reachable and
emitFileCdg silently skipped ALL control-dependence for the function.
New cfg/synthetic-escape.ts: a pure deterministic SCC routine (iterative
Tarjan, sorted adjacency) + augmentForPostDom(cfg). No-op when EXIT is
already reverse-reachable (terminating fns + visitor-escaped loops are
byte-identical — returns the same object). Otherwise it batch-bridges
every exit-less SCC by adding an ANALYSIS-ONLY escape edge from the SCC's
controlling block (highest out-degree branch; lowest-index tie-break) to
EXIT, on a shallow-cloned FunctionCfg — never mutating persisted
cfg.edges. emitFileCdg threads that augmented view through BOTH
isExitReachableFromAllBlocks AND computeControlDependence (the Ferrante
walk re-reads cfg.edges, so a tree-only augmentation would be wrong).
Precision (anti-masking): only a trapped region containing a control
point (>=2-successor block) is bridged — a branch-less trapped region
carries no recoverable control-dependence and is indistinguishable from a
genuine construction anomaly, so it stays on the skip path (the existing
disconnected-block skip test still skips, skippedUnsoundFunctions===1). A
residual non-cycle dangling block is never bridged.
repro `void handler(int a){ start: if(a>0){work();} goto start; }`:
before exitReachable=false/CDG=0 → after one synthetic 2->1 edge,
exitReachable=true, exact CDG = {2->2:T,2->2:F,2->3:T,2->4:T,2->4:F}
(pinned exactly, not CDG>0 — catches a wrong representative). AC2 property
test extended to the augmented graph; per-language goto-cycle regressions
(C/C++/C#/Go). 199 cfg tests green; bench --check fingerprints unchanged
(analysis-only, zero persisted drift).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): isolate the non-terminating-loop hazard in worker CDG asserts (#2195 U3)
The pipeline-pdg worker-mode blocks asserted a whole-fixture cdg>0
aggregate (satisfied by any branching fn) while the comment claimed it
proved the non-terminating-loop EXIT-reachability end-to-end. Add a per-
language `hazard` marker + isolate the assertion: locate the hazard
function's BasicBlocks by its anchor and assert >=1 CDG edge is sourced
within it (a marker mutation now fails the test — non-vacuous). C# keeps
the aggregate (its fixture has no infinite loop). Comments corrected.
Switch the 7 visitor unit tests (java/csharp/dart/kotlin/php/swift/c-cpp)
from the local exitReachableFromAll CFG-shape helper to the production
isExitReachableFromAllBlocks + computeControlDependence on the hazard
function, matching go/python/ruby/rust/vue. 241 unit + 30 pipeline green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): gate vendored-grammar worker assertions on isLanguageAvailable (#2195 U4)
The Swift/Kotlin/Dart worker-mode pipeline-pdg cases require a vendored
grammar prebuild that may be absent on a CI platform — they'd go red
there. Mark those three REMAINING_LANGS entries `vendored` and gate both
the --pdg-on and --pdg-off `it`s on isLanguageAvailable(SupportedLanguages
[lang]) → it.skip when the grammar can't load. Installed-grammar
languages stay unconditional. Grammars present here, so all 30 run green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cfg): remove dead useCount() from swift + rust harvests (#2195 U5)
useCount() was declared on the local FactAccumulator in swift-harvest.ts
and rust-harvest.ts but never called (a copy-paste artifact; ruby's copy
IS used in an emit guard, so it stays). Pure deletion — the swift/rust
visitor suites stay green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cfg): standardize harvester API table()->bindingTable() (#2195 U9)
The binding-table accessor was named table() in the C/C++/C#/Go harvests
but bindingTable() in the other 7. Rename the 4 (definitions + their
visitor call sites) to the majority name bindingTable(). Pure rename; the
4 visitor suites stay green and tsc confirms no call site was missed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cfg): consolidate scope-tree substrate into ScopeTreeHarvester (#2195 U6)
The Go/Java/C#/C-C++ def/use harvesters each carried a byte-identical copy
of the lexical scope-tree machinery (Scope record, two-phase resolution
cache, openScope/nearestScopeOf/resolve/def/use/conditional/bindingTable,
~270 lines total). Extract it into an abstract ScopeTreeHarvester base; the
four harvesters now extend it and supply only their genuine per-language
variation (the prescan switch, plus Go's _-blank-identifier overrides of
declare/def/use). Net -422 lines. Mechanical and byte-equivalent: cfg unit
suite 613 passed, bench --check fingerprints unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cfg): consolidate no-site def/use accumulator into DefUseAccumulator (#2195 U7)
The Kotlin/Python/Ruby/Rust/Dart/Swift harvesters each carried a
byte-identical copy of the no-site def/use accumulator (~270 lines total;
only Ruby's adds the live useCount() emit-guard helper). Extract it as an
exported DefUseAccumulator beside CallSiteFactAccumulator in
call-site-harvest.ts (the PR's own model for the with-site superset); the six
harvesters import it under their existing local FactAccumulator name. Pure
byte-equivalent move, no logic change: cfg unit suite 613 passed, tsc clean,
bench --check fingerprints unchanged (TS/Go paths untouched).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): consolidate copied visitor-test helpers into cfg-harness (#2195 U8)
The 13 *-visitor.test.ts files each copied a byte-identical set of CFG-shape
helpers (edgeKinds/block/reaches/reachable/bindingIdx/allSites/hasAnySites,
~380 lines total). Export them once from test/helpers/cfg-harness.ts and import
per file (only the subset each references). Also drop each file's local
exitReachableFromAll — a re-implementation of the production
isExitReachableFromAllBlocks (semantically identical: false iff some
entry-reachable non-EXIT block can't reach EXIT) — and point its live call
sites at the already-imported production function. Pure test-only mechanical
move, behavior-preserving: tsc clean, test/unit/cfg/ 613 passed unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): pin *-harvest.ts literals to their own grammar in the gate (#2195 U10)
The grammar-literal validation gate scans cfg/visitors/, but a <lang>-harvest.ts
basename was not in BASENAME_LANGS, so fileLanguages() fell it through to the
weak ALL_LANGS valid-if-any bucket — a node-type literal dead in its own grammar
but valid in some other grammar would pass undetected. Strip the -harvest suffix
and reuse the visitor basename map so go-harvest -> Go, c-cpp-harvest -> C+C++,
typescript-harvest -> TS, etc. The two language-agnostic harvesters
(call-site-harvest, scope-tree-harvest) name no grammar and stay valid-if-any.
Also corrects the now-inaccurate mode2Files comment. Adds a fileLanguages unit
test; the existing gate stays green (no harvest file has a dead literal), and a
scratch probe confirmed a bogus go-harvest literal is now caught.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): defensive per-statement cap on harvested taint sites (#2195 U11)
A statement's harvested sites[] had no explicit bound — a pathological or
machine-generated statement (hundreds of nested calls) could grow it without
limit. Add DEFAULT_PDG_MAX_SITES_PER_STATEMENT (512, mirroring the PDG edge/fact
cap style): openCallSite/addMemberRead check-before-push and stop at the cap,
keeping the first 512 sites fully intact and setting an observable
sitesTruncated flag. A cap-dropped openCallSite returns a -1 sentinel that
pushFrame/setSite*/the occurrence fan-out all tolerate (no dangling parent/via,
no clobber of kept sites). Generous enough that no real statement is affected:
bench --check fingerprints unchanged, cfg unit suite 617 passed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci(codeql): exclude nested test fixtures from the CodeQL gate (#2195)
The CodeQL results gate failed on test/integration/cfg/fixtures/python-hazards.py
('total' may be used before init, unused vars) — but that file is an intentional
CFG/PDG hazard fixture, exactly the synthetic broken-code the existing
'**/test/fixtures/**' exclusion is meant to skip. That glob does not match the
deeper test/integration/cfg/fixtures/ path, so the hazard fixtures leaked into
the scan. Add '**/test/**/fixtures/**' to cover fixtures nested anywhere under a
test tree. Analyze (python) and Analyze (javascript-typescript) both already pass
— production code is clean; this only silences fixture noise.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style(cfg): apply prettier + drop unused imports across the PDG files (#2195)
The merge of main into this branch pulled in the stricter quality gates
(prettier --check . and eslint .), which surfaced pre-existing formatting in the
PDG/CFG rollout (line-width wrapping across the visitor + harvest files, bench,
tests) plus 9 no-unused-imports errors. Mechanical autofix only — npm run
format + lint:fix equivalent, scoped to gitnexus/: removes unused FunctionCfg/
SiteRecord type imports left by the U8 helper consolidation and stale
FinalizerFrame imports in python.ts/ruby.ts. No behavior change: tsc clean, cfg
unit suite 617 passed, eslint 0 errors. (gitnexus-web class-order noise is a
local tailwind-plugin artifact CI does not flag — left untouched.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): harvest C++ structured-binding defs (auto [a,b]=e) (#2195)
The C++ def/use harvester only recorded a def when an init_declarator's
declarator was a plain identifier, so a structured binding (auto [a,b] = mk(),
incl. the auto& reference form whose binding sits under a reference_declarator)
declared only the first name in phase 1 and emitted ZERO defs in phase 2 — a,b
were walked as spurious uses and later use(a)/use(b) resolved to a synthetic
module binding, silently corrupting REACHING_DEF/taint for an idiomatic C++17
shape. Unwrap the structured_binding_declarator in both phases and def every
identifier leaf; result-of-initializer flows to the whole list. Inert for C
(no structured bindings). Characterization tests added (plain + reference form).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): model C++ co_return as a return terminator to EXIT (#2195)
co_return_statement was neither in CPP_CONTROL_FLOW_TYPES nor dispatched, so a
coroutine's co_return coalesced into a straight-line block and emitted a
spurious seq fallthrough to the following statement instead of an edge to EXIT
— statements after co_return looked reachable and the terminator edge was
missing, corrupting CFG/CDG for coroutines. Add the node type to the C++
control-flow set and dispatch it through visitReturn (block -> EXIT 'return',
no fallthrough). C path untouched; co_await/co_yield remain plain expressions.
Characterization test added; c-cpp suite + grammar-literal gate green, bench
--check unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): harvest C# out-var and deconstruction-declaration defs (#2195)
Two idiomatic C# write shapes recorded ZERO defs, silently breaking
REACHING_DEF/taint:
- out-var (G(out var n) / G(out int n)) parses as a declaration_expression;
it was neither declared (phase 1) nor def'd (phase 2), so n resolved to a
synthetic module binding and the callee-written value had no reaching def.
- deconstruction declaration (var (a, b) = T()) has a variable_declarator whose
name slot is a tuple_pattern (null name field), so declareVariableDeclaration
+ the variable_declaration walk skipped it entirely (only the assignment form
(a,b)=T() was handled). Both a and b were dropped.
Declare + def the declaration_expression's identifier (must-def: out params are
definitely-assigned), and route a null-name variable_declarator through the
tuple_pattern via the existing declareForeachTarget/defTupleTargets helpers.
Characterization tests added; csharp suite 42 passed, grammar gate + bench
--check green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): put embedded-script CFGs in file coordinates via lineOffset (#2195)
A Vue SFC <script> block parses at row 0 but lives at lineOffset in the .vue
file. Every other worker-emitted graph node adds lineOffset to reach file
coordinates, but collectFunctionCfgs built FunctionCfgs from the extracted
script's raw rows and never offset them. Two consequences for .vue files:
- inter-procedural taint silently resolved NOTHING — the summary-harvest join
keys graph Function/Method nodes by their (offset) startLine but looked up the
CFG's (unoffset) functionStartLine, missing by exactly lineOffset, so no
FunctionSummary was ever produced;
- persisted BasicBlock startLine/endLine (and the id's functionStartLine
segment) pointed at the wrong .vue line, breaking source mapping.
Thread lineOffset into collectFunctionCfgs and shift every CFG source-line field
(functionStartLine/End, block start/end, statement + non-synthetic binding
lines) into file coordinates at the one production chokepoint. A 0 offset
returns the CFG unchanged, so .ts/.js/etc. stay byte-identical (bench --check
fingerprints unchanged; worker-roundtrip + pipeline-pdg green). Unit tests for
the shift + the 0-offset no-op added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): surface CDG soundness skips at warn, not just debug (#2195)
skippedUnsoundFunctions (a function whose EXIT is not reverse-reachable from
all blocks, so control dependence is withheld) was only reported inside the
per-language logger.debug stats line — while the taint/RD coverage-gap and
cap-drop counts surface unconditionally at warn. A language that systematically
trapped EXIT (an unmodeled non-terminating / multi-terminal shape the
synthetic-escape pass can't bridge) would silently lose all CDG. Add a parallel
unconditional warn (R8) alongside the R4 taint-gap warn. Observability only —
no graph change; emit-layer skip counting stays covered by cfg-emit's
skippedUnsoundFunctions test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): harvest Kotlin x++/--x as a def (#2195)
The Kotlin harvester had no postfix_expression/prefix_expression case, so an
increment/decrement fell to the default descent and recorded its operand as a
use only — never a def. Every sibling harvester (Java/C#/C++/Dart/TS/PHP) models
inc/dec, so a Kotlin counting loop (while/for using i++) silently dropped the
loop-carried reaching-def of the counter. Add the case: def AND use the operand
when it is a plain simple_identifier and the operator is ++/-- (other pre/postfix
forms — -x, !x, x!!, x? — stay pure reads, byte-identical to the old descent).
Characterization tests for postfix + prefix added; kotlin suite 30 passed,
grammar gate + bench --check green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): harvest Go select channel-receive binding as a def (#2195)
walkValue had no receive_statement case, so a select receive (case v := <-ch:)
fell to the default descent: v was recorded as a USE of an uninitialized var
and the channel-sourced definition was invisible to REACHING_DEF/taint —
channels are a primary taint source in Go. Add the case mirroring
short_var_declaration: def each left identifier, use the <-ch right, attach
resultDefs for the := short form. prescan already declared the binding; this
completes the phase-2 fact. go:branchy bench fingerprint unchanged; go suite
40 passed, grammar gate green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): harvest all names of a Dart multi-variable declaration (#2195)
`var a = 1, b = 2;` is one initialized_variable_definition whose first binding
is the name/value field pair and whose subsequent bindings are trailing
initialized_identifier children. Both prescan (declareInitializedVar) and the
walkValue case read only the name/value fields, so every name after the first
was never declared or def'd — `b` resolved to a synthetic module binding and
its REACHING_DEF/taint flow was lost. Iterate the trailing initialized_identifier
nodes in both phases. (Dart-3 record/list pattern declarations `var (a,b)=pair`
remain a separate follow-up.) dart suite 35 passed, tsc + grammar gate green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): bind Swift switch-case value patterns (case let n) (#2195)
A switch value-binding (case let n where …, case .some(let v)) was never
declared in prescan, so n/v resolved to a synthetic module binding and a body
use(n) did not link to any def — a very common Swift idiom silently lost its
data dependence. Declare the switch_pattern's bindings (prescan, reusing
declarePattern) and emit them as MAY-defs on the dispatch block (a case may not
match) via a new switchPatternFacts, propagated into the case body. swift suite
25 passed, tsc + grammar gate green. (The rare ?? / ternary-arm may-def — Swift
assignment-as-expression — remains a separate follow-up.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): wire Swift multi-catch throw edges to every handler (#2195)
visitDo routed the protected body's throw edge only to handlerEntries[0], so a
do { try r() } catch A {} catch {} left the 2nd..Nth catch handlers UNREACHABLE
from ENTRY — orphaned blocks whose error bindings + def/use facts were stranded
in a dead component (a soundness gap for idiomatic Swift typed multi-catch).
Swift tries the catch clauses in order and the thrown type is unknown at CFG
time, so every protected block may reach ANY clause: edge each protected block
to every handlerEntry. Found by the per-language CFG/CDG verification swarm
(reproduced: 2-catch=1, 3-catch=3 unreachable blocks). swift suite 26 passed,
tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): synthesize a protected block for an empty Kotlin try {} (#2195)
An empty `try {}` body produced zero protected blocks, so visitTry's throw-edge
loop wired nothing to the catch and the try's entry fell through to the finally
— leaving the catch handler block + its error binding orphaned (unreachable from
ENTRY), a malformed CFG with stranded def/use facts. Mirror the existing
empty-`catch` synthesis: when the try body is empty and there is a catch or
finally, synthesize one protected block so the catch handler(s) are wired and
the try entry is the body, not the finally. Found by the per-language CFG
verification swarm. Non-empty try is byte-identical; kotlin suite 31 passed,
tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(cfg): bound the reaching-defs fixpoint with a per-block visit ceiling (#2195)
The per-language verification swarm reproduced, AT PRODUCTION DEFAULTS, a
reaching-defs blow-up: a machine-generated ~2000-line all-loops function (under
DEFAULT_PDG_MAX_FUNCTION_LINES) reaches ~10k basic blocks because loops emit ~5
blocks/line, and the dataflow fixpoint is O(blocks^2.3) on deep loop nests —
measured 62s (C/C++) and 2.05s + 810MB (Go) for ONE function. maxFacts does not
help: the fact count stays LINEAR, so it never fires.
Iterative reaching-defs on a reducible CFG converges in O(loop-nesting-depth)
passes, so a worklist re-visits each block a small multiple of times for real
code. Add a maxBlockVisits ceiling (emit passes blocks.length × 64 — far beyond
any hand-written nesting depth, ~15) that bails when the fixpoint has not
converged. An unconverged fixpoint's in/out sets are not sound, so it returns
NO facts (status 'truncated', like the existing 'overflow' guard) — a per-
function coverage gap, never wrong facts. Real code is byte-identical: full cfg
suites 725 passed, bench --check fingerprints unchanged.
NOTE: computeControlDependence's O(N²) up-walk on deep post-dom chains is the
sibling concern but stays ~13ms in production (bounded by the line cap + the
CDG materialization cap); a CDG work-budget is a documented follow-up.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): raise the parse-worker stack limit for deep CFG recursion (#2195)
The CFG visitors build per-function control-flow graphs by recursive descent
over the tree-sitter AST, so deeply-nested source overflows the worker thread's
call stack (~1.5k nesting levels) — caught per-function (R4 try/catch) but the
function silently gets no PDG. A worker thread's stack is governed by
resourceLimits.stackSizeMb (Node default 4 MB); the main process's
--stack-size=4096 flag does NOT propagate to worker threads (confirmed by prior-
art research on Node worker_threads). Raise it to 16 MB, pushing the overflow
threshold to several-thousand nesting levels — far beyond any hand-written code,
so only machine-generated/obfuscated nesting can still hit it (and that stays a
caught per-function skip, never a crash). Complements a future proactive depth
guard. pipeline-pdg worker tests 30 passed, tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(cfg): compute control dependence as a reverse-CFG dominance frontier (#2195)
The Ferrante §3.1.1 up-walk re-climbed the ipdom chain once per CFG edge,
which is Θ(N²) on a deep post-dom chain (a single branch fanning into a
shared spine took ~7.1s at 16k blocks). Replace it with the reverse-CFG
post-dominance-frontier formulation (Cytron, Ferrante, Rosen, Wegman &
Zadeck 1991): control dependence IS the dominance frontier of the reverse
CFG, computed bottom-up over the post-dom tree (PDF_local from a node's CFG
in-edges + PDF_up from its post-dom-tree children) in O(N + E + output).
LLVM (ReverseIDFCalculator), Joern (CdgPass) and WALA use the same form.
Output is the IDENTICAL deduped/sorted (controller, dependent, label) set:
verified byte-identical across all cfg unit+integration suites, the
cdg-snapshot oracle, and bench --check fingerprints (unchanged). The PDF
unions a label SET per (controller, dependent) pair, preserving the
multi-label rows the old per-row dedup kept on opposite-sense (goto-cycle)
arms. buildArmSenses, labelFor, the final sort and the maxEdges truncation
cap are kept verbatim; the post-order walk is iterative so a chain-deep
post-dom forest cannot overflow the stack.
Adds three regressions: multi-label-per-pair preservation, the literal
self-edge / NO_IPDOM seed guard (a !== x), and a fan-into-chain perf
tripwire (linear vs the former quadratic up-walk).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(cfg): record the reaching-defs WTO no-go decision (#2195)
Weak-topological-order / loop-aware iteration (Bourdoncle 1993) was
evaluated as the fix for the O(blocks²) deep-loop-nest blow-up and
rejected: a faithful WTO solver was 104/104 byte-identical to the RPO
worklist but 0% faster — the cost is inherent dense-set propagation +
lattice merges, not visitation order, and the loop-body-skip shortcut is
unsound on irreducible (goto) CFGs. Document this at the RPO-order site
and the emit.ts revisit-ceiling constant so the shipped blocks×64 bound
reads as the sound backstop it is, with SSA-sparse reaching-defs named as
the deferred real fix. Comment-only; no behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cfg): proactive visitor nesting-depth guard + observable CFG skips (#2195)
The CFG visitors are recursive-descent with no shared base, so a
pathologically nested function (machine-generated / adversarial) could
overflow the worker's native stack — a nondeterministic RangeError that
escaped to the language-group catch and silently dropped EVERY remaining
file's CFG.
Guard it proactively: CfgBuilder tracks live recursive-descent nesting
depth via enterNesting/exitNesting, called at each visitor's visitBody and
visitSeq choke points (visitBody covers nested control constructs incl.
else-if ladders; visitSeq covers deeply-nested bare blocks). Exceeding
MAX_CFG_NESTING_DEPTH (500, far below the ~1.2k+ native limit and far above
real code's ≤~50) throws a typed, DETERMINISTIC CfgNestingDepthError instead
of waiting for the engine's nondeterministic overflow.
collectFunctionCfgs now isolates the build PER FUNCTION: the depth bail or
any other throw is caught, counted, and skipped — one bad function no longer
loses the whole file's CFGs. CollectedCfgs.skipped widens from a bare number
to reason-counted buckets (tooManyLines / tooDeeplyNested / buildError). The
worker stops discarding that count (parse-worker.ts), aggregates it
per-language onto ParseWorkerResult.cfgSkipped (survives the parse cache via
slim's `...result`), and mergeChunkResults merges + warns per-language so a
CFG coverage gap is observable, not silent.
Behavior-preserving on normal code: the guard never fires below 500 nesting,
so the cfg unit+integration suites (731), the CDG/RD/CFG snapshots and the
bench --check fingerprints are all byte-identical. The worker stackSizeMb
4→16MB bump shipped earlier (de4c43a4).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cfg): unify the nesting guard behind CfgBuilder.withNesting (#2195)
Tri-review flagged the visitBody/visitSeq guard asymmetry (visitBody used
try/finally; visitSeq placed a bare exitNesting before its tail return). It
is not a bug today — the CfgBuilder is per-function and discarded on a bail,
so a leaked counter is never read — but a future mid-loop return in visitSeq
would silently corrupt the depth count. Replace both hand-paired sites in all
12 visitors with a single `CfgBuilder.withNesting(fn)` helper that enters on
the way in and exits in a finally, so the pair can never drift. Also document
that block-bodied constructs pass through BOTH choke points, so the effective
lexical ceiling is ~MAX_CFG_NESTING_DEPTH/2 (~250).
Behavior-preserving: 733 cfg tests + bench --check fingerprints byte-identical.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): include cfgSkipped in all parse-worker result initializers (#2195)
Tri-review found cfgSkipped omitted from the two reset/fallback
ParseWorkerResult initializers (only the main one carried it). The field is
optional with `?? {}` reads so there is no runtime bug, but the zero-state
initializers should be complete and consistent. Also correct the field's
doc-comment: the per-language merge + warn lives in `dispatchChunkParse`
(alongside skippedLanguages), not `mergeChunkResults` — and, like that
sibling telemetry, the warn fires for freshly-parsed chunks, not on a warm
cache hit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(cfg): scope the CDG byte-identical claim to the untruncated output (#2195)
Tri-review noted the rewrite's `maxEdges` truncation now trims the SORTED
edge set, whereas the old up-walk broke mid-walk in CFG-edge-iteration order
— so a truncated PREFIX can differ at the cap boundary (only when a single
function exceeds maxEdges; the FULL untruncated set is byte-identical, now
also confirmed by ~1M-case differential fuzz). Clarify the module doc and the
maxEdges param doc: the cap bounds OUTPUT count (peak working set ≈ output in
the DF formulation, not the old pre-dedup spike), and the byte-identical
guarantee is scoped to the untruncated output. Comments only.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cfg): strengthen depth-guard, perf-tripwire and buildError coverage (#2195)
Tri-review test-quality findings:
- cfg-builder: capture the thrown CfgNestingDepthError unconditionally (a
catch-only assertion silently passes if a future change stops throwing);
add a withNesting test that the counter balances on the THROW path too.
- control-dependence perf tripwire: assert controller IDENTITY (every edge
controlled by block 0, distinct dependents in range), not just length, so a
fast-but-wrong reimplementation can't pass on the M-1 count alone.
- worker-roundtrip: add the missing buildError test — a generic (non-depth)
buildFunctionCfg throw is caught per function, counted under buildError, and
does NOT drop the file's sibling CFGs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): model Kotlin value-position when/if as control dependence (#2205)
Idiomatic Kotlin uses `if`/`when` as EXPRESSIONS (`val x = when (k) { … }`,
`return if (c) a else b`, `fun f() = when (k) { … }`), but the visitor only
modeled them as control flow in statement position (the `isStatementPosition`
gate) — value-position branches collapsed into one straight-line block, so
their arms emitted no control dependence. Cross-language PDG validation
(#2195, PR #2197) measured the result: two Kotlin repos at 3% / 6% CDG per
BasicBlock vs 18–76% for every other language (incl. the other
expression-conditional languages, Rust 30% / Swift 24%).
Model value-position `when` (≥2 arms) and `if`/`else` as control flow in the
three dominant carriers — `property_declaration` (rejoin the arms at a
binding continuation carrying the bound name's def), `return`, and the
`fun f() = …` expression body (each arm returns) — mirroring the Rust
visitor's value-position `let` handling. `visitWhen`/`visitIf` are reused
unchanged; `isControlFlow` now routes a value-branch `val`/`var` decl to the
branch handler instead of coalescing it.
Measured: Exposed CDG 3636→5644 (+55%, 6%→9% of BasicBlocks), turbine
67→82 (+22%). Argument-position branches, assignment RHS, and value-position
`try` are left inline — a remaining gap tracked on #2205.
Behavior-preserving for non-Kotlin (bench --check byte-identical; 739 cfg
tests incl. 6 new value-position regressions).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): model Ruby value-position if/case on assignment RHS as control dependence (#2195)
`if`/`case` are expressions in Ruby; as an `assignment` RHS
(`x = if c then a else b end`, `x = case k … end`) the visitor previously
left the arms INLINE in one coalesced block (a gap documented at ruby.ts
§"if/case/begin are EXPRESSIONS"), emitting no control dependence. Model the
RHS branch as control flow and bind the LHS at the rejoin: `assignmentBranch`
detects the carrier, `visitSeq` routes it out of the coalescing path, and
`visitBindBranch` reuses `visitIf`/`visitCase` + a facts-only continuation
carrying the LHS def (new `harvest.assignmentDefFacts`). Mirrors the Kotlin
(#2205) and Rust value-position handling.
Honest impact: SMALL in practice — rack CDG 1269→1303, sinatra 1212→1232
(~+2–3%). `x = if/case` is far rarer in idiomatic Ruby than the Kotlin
analog (Ruby favors ternary / `||=` / guard modifiers), and Ruby's low
CDG/BB is mostly structural (micro-branches `&&`/`||`/`?:`/`&.` excluded by
design, plus many straight-line `.each`/`.map` block CFGs). This closes the
documented gap correctly; it is not a large ratio mover. Explicit
`return if … end` is NOT a carrier — tree-sitter-ruby drops that value; the
idiomatic implicit-last-expression conditional was already modeled.
Behavior-preserving for non-Ruby (bench --check byte-identical; 744 cfg
tests incl. 5 new value-position regressions).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): keep an empty-arm Kotlin when wired to the join (#2195)
An all-empty-arm `when` with an `else` — `when(k){0->{};else->{}}`, idiomatic
`else -> {}` "do nothing" — left the dispatch block with ZERO successors:
empty arms got no `switch-case` edge, and the `else` suppressed the no-match
edge. The dispatch and its join then became orphaned, so
isExitReachableFromAllBlocks returned false and emitFileCdg silently dropped
the ENTIRE function's control dependence (counted cdgSkippedUnsound). The
#2205 value-position fix newly routes `val x = when(…)` / `return when(…)` /
`fun f() = when(…)` through visitWhen, exposing it in those carriers too.
Wire every arm (empty or not) to the join, so the dispatch always has a
successor. Found by the per-language CFG verification swarm. Behavior-
preserving elsewhere (bench --check byte-identical; cfg suite green).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): wire C/C++ throw edges to every catch handler, not just the first (#2195)
`visitTry` edged every protected-region block only to `handlerEntries[0]`, so
in a multi-`catch` (`catch(int e){…} catch(double d){…} catch(...){…}`) the
2nd..Nth handlers were orphaned — unreachable from ENTRY, their catch-param
binding and body control/data flow silently lost. The runtime catch that
matches a thrown type is not statically known, so over-approximate: edge each
protected block to EVERY handler entry (mirrors the Swift multi-catch
handling). Found by the per-language CFG verification swarm; the two existing
exception tests both used a single catch, so it was never exercised.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): keep a bare continue in a Dart switch case — it targets the loop (#2195)
`caseStatements` stripped EVERY `continue_statement` from a case body, but only
a LABELED `continue LABEL;` is a switch fallthrough-spill (handled via
caseContinueLabel). A bare `continue;` targets the ENCLOSING LOOP (valid Dart);
dropping it removed the jump and fabricated a false case → next-statement
fall-through edge (e.g. `case 1: tainted(); continue; default: sink();` made
tainted() flow directly into sink()). Only strip the labeled form; a bare
`continue;` stays in the body and routes to the loop. Found by the per-language
CFG verification swarm.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): harvest Rust match-arm pattern bindings as defs (taint propagation) (#2206)
`visitMatch` visited arm bodies but never harvested the arm PATTERN's bindings,
so `match x { Some(n) => sink(n) }` left `n` with a use and no def/may-def —
taint from the matched subject could not propagate into the arm. Add
`matchArmPatternFacts` (the binders as MAY-defs, since only the matching arm
binds) and attach it to the dispatch block, co-located with the subject's use.
Found by the per-language CFG verification swarm; the match tests asserted
`hasUse` but never `hasDef`.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): harvest Swift guard case / if case enum-pattern bindings as locals (#2206)
`guard case .some(let v) = e` / `if case let .x(n) = e` nest the binder inside a
`pattern` condition child, not a direct `bound_identifier`. Both the declaration
(declareOptionalBindings) and the def-facts (conditionFacts, which ran walkValue
= a USE on the pattern) missed it, so the binding resolved to a synthetic
`@module` global with a use and no def — breaking taint propagation from the
subject. Declare the `pattern` child and def its leaves (may-def when
conditional). Found by the per-language CFG verification swarm.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): treat Dart switch-expression arm writes as may-defs, not hard kills (#2206)
`DartHarvester.walkValue` had no `switch_expression` case, so an assignment in an
arm value — `var y = switch(x){ 1 => z = 10, _ => z = 20 }` — became an
unconditional def that KILLED the prior `z`, even though only one arm runs (the
module docstring claimed it was a may-def, but the code didn't implement it).
Walk the subject always and each `switch_expression_case` under `conditional(…)`,
so arm writes are may-defs — mirroring `conditional_expression`. Found by the
per-language CFG verification swarm.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): model the C# `using var` declaration-form dispose finalizer (#2206)
`using var f = Open();` (C# 8) parses as a `local_declaration_statement` with a
leading `using` keyword, not a `using_statement`, so visitUsing never ran — no
dispose block, and a `return`/`break`/`continue` in its scope got no
`finally-*` completion edge. Unlike the delimited block form, its dispose runs
at ENCLOSING-SCOPE exit, so visitSeq now treats the REST of the sequence as the
protected body: the acquisition (`var f = e`) is a normal block outside the
dispose region (a throw there means the resource was never acquired), and
`buildUsingDeclScope` wraps the remainder in a synthetic dispose finalizer
(normal + exception exit, early exits thread through) — mirroring
buildProtectedSynthetic. Closes the last #2206 item. Found by the per-language
CFG verification swarm.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(mcp): add trace tool for shortest call path between symbols (#1821)
Implement the \ race\ MCP tool and \gitnexus trace\ CLI command that finds
the shortest directed call path between two symbols using BFS over CALLS +
HAS_METHOD edges.
- MCP tool definition in tools.ts with READ_ONLY annotations
- Directed BFS in local-backend.ts with parent-map path reconstruction
- Symbol resolution via resolveSymbolCandidates (name/UID/file-hint)
- Gap reporting with furthest reachable node and depth tracking
- CLI wiring: gitnexus trace <from> <to> [--from-uid] [--to-uid] [--depth]
- i18n keys in en.ts and zh-CN.ts + help-i18n.ts registration
- ARCHITECTURE.md tools table entry
- 16 unit tests (11 BFS core + 5 CLI wiring)
* test(mcp): account for trace tool in tools.test.ts count
The trace tool makes GITNEXUS_TOOLS length 15; update the hardcoded
count, add 'trace' to the expected-names list, and refresh the stale
"13 tools" comment and it() title.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): sanitize trace maxDepth to reject 0/NaN/negative
`Math.min(params.maxDepth ?? 10, 30)` had no lower bound and `??` does
not recover 0 or NaN, so `--depth 0|-5|abc` made the BFS loop run zero
iterations and return a false `no_path`. Clamp at the real boundary with
a `Number.isInteger && > 0` guard (the MCP inputSchema minimum is
advisory only), and reject a non-numeric `--depth` in the CLI up front
rather than forwarding NaN.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): check trace target before applying test-file filter
The `isTestFilePath` filter ran before the target-equality check, but
resolveSymbolCandidates does not exclude test-file symbols. A target (or
a required hop) that lives in a test file was therefore skipped under the
default includeTests=false and produced a false no_path with a
misleading dynamic-dispatch suggestion. Match the explicitly-requested
target first; non-target test-file nodes are still filtered.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* perf(mcp): bound trace BFS with per-level LIMIT and visited cap
The per-level query had no LIMIT and the visited set was uncapped, so a
high-fanout hub could materialize an unbounded frontier. Cap per-level
rows (interpolated LIMIT — Kuzu does not bind LIMIT) and the total
visited set; either cap sets a `truncated` flag so a resulting no_path
reports that the search was cut short rather than implying the graph was
exhausted.
Note: the sibling impact BFS shares the same unbounded pattern; applying
the cap there is deferred (out of scope for this PR).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(mcp): clarify trace traverses call + class-member edges
trace was advertised as a "shortest call path" but also traverses
HAS_METHOD (class→member) containment edges so a class-rooted trace can
descend into its methods. Keep that capability (consistent with impact/
context) and make the docs honest: rename EDGE_TYPES→TRAVERSAL_EDGE_TYPES,
state the call + class-member traversal in the MCP/CLI/i18n/ARCHITECTURE
descriptions, and note each hop's edge type is reported in edges[]. No
behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): set status:'error' on trace failure responses
Every trace return path sets a `status` discriminator except the
caught-error path, so a consumer switching on `result.status` saw
undefined on failure. Add status:'error' to both the backend trace()
catch and the CLI traceCommand catch.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): return a friendly error for non-string trace from/to
A non-string from/to reaching resolveSymbolCandidates surfaced a
low-level "x.includes is not a function" via name.includes. Guard the
four name/uid params at the top of _traceImpl and return a structured
status:'error' with a clear message instead.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(mcp): single row-decode + drop dead field in trace BFS
Decode each BFS row once into named locals instead of repeating
`(row.x ?? row[N])` across the two parent.set calls and the
furthest-tracking. Drop the `type` field from the parent map value (it
was written but never read), and rename the internal `deepestInfo` to
`lastReached` for accuracy (the output field `furthest` is unchanged).
Pure refactor — no behavior change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cli): dedicated trace includeTests i18n key + guard coverage
`trace|--include-tests` reused the impact help key, so rewording the
impact option would silently change trace's help text. Add a dedicated
help.option.trace.includeTests key in en + zh-CN and repoint it. Add CLI
coverage for the (already symmetric) --from-uid/--to-uid flag-value guard
and for --include-tests forwarding.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(mcp): faithful BFS mock + expand trace coverage
Fix makeResolveMock: concatenate neighbors across ALL frontier ids (it
returned only the first node's, so a multi-node frontier was unmodelled)
and key the UID branch on params.uid (the old query-text match never
fired). Add coverage: shortest path through the second frontier node
(proves the mock fix), confidence floor fallback, HAS_METHOD traversal
with a mixed edge-type chain, no_path furthest:null, and from_file
disambiguation.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style(trace): apply root prettier formatting to trace files
The root `quality / format` gate (prettier --check, printWidth 100) runs
on the full repo and flagged the trace sources/tests (the local config
masks it). Reformat to root style — no behavior change; trace + tools
suites and tsc stay green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(skills): document the trace tool for AI agents
Add `trace` to the GitNexus skill docs so agents reach for it instead of
hand-chaining context/impact hops. The guide gains a Tools Reference row
and a "shortest path between two symbols" subsection (params, result
shape, status/furthest/truncated semantics); the debugging skill gains a
"how does A reach B?" pattern row and a trace tool example. Mirrored to
the .claude and claude-plugin copies (byte-identical) and the cursor copy
(compact style).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(mcp): drop unused trace test fixtures (CodeQL js/unused-local-variable)
CodeQL flagged two unused locals in the trace BFS tests: the top-level
SYMBOL_C and a SYMBOL_D inside the maxDepth test (both defined, never
referenced). Remove them. No behavior change — 58 trace/tools tests stay
green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Gergő Magyar <gergomagyar@icloud.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(lbug): pin repos to exempt them from automatic pool eviction [#2189]
Add a pinnedRepos set and pinRepo/unpinRepo to the LadybugDB pool adapter.
evictLRU and the idle-timeout sweep skip pinned repos; closeOne clears the
pin on teardown so explicit close always wins and pins never leak across
operations. Behavior is byte-identical when nothing is pinned.
Bounded multi-repo callers (group sync) can now keep more than MAX_POOL_SIZE
repos resident through deferred cross-repo resolution.
* fix(group): pin repos during sync so >MAX_POOL_SIZE groups resolve [#2189]
syncGroup now pins each repo immediately after initLbug and releases the pin
(unpin then close) in the finally. This keeps every group member resident
through the deferred manifest/workspace resolution that runs after the init
loop, so cross-links anchor to real graph symbols instead of falling back to
synthetic UIDs when a group has more than MAX_POOL_SIZE repos.
Release is unpin-before-close plus closeOne's own pin-clear, so pins never
leak across syncs in the long-lived MCP server even on error.
* style(test): apply prettier formatting to #2189 test files
* fix(review): apply autofix feedback
Clarify the pinRepo docstring: the pin does not survive teardown (closeOne
clears it) and the repoId must match the key passed to initLbug. Addresses a
code-review finding that the prior 'or later holds' wording contradicted
closeOne's unconditional pin-clear.
* refactor(lbug): reference-count pool pins so overlapping holders are safe [#2189]
Change pinnedRepos from Set<string> to Map<string,number>. pinRepo
increments the lease count; unpinRepo decrements and deletes the key at 0
(flooring at zero, unknown-id no-op). evictLRU, the idle sweep, and closeOne
are transparent to the swap (has()/delete() keep their semantics: skip while
count>=1, force-clear on teardown).
A boolean Set could not represent two simultaneous holders, so the first
release wrongly cleared a pin another holder still needed — the concurrent
overlapping group_sync teardown race from the PR #2191 review (Finding 1).
Reference counts let two windows of one sync, or two concurrent syncs sharing
a repo, coexist safely: the repo stays exempt until the last lease releases.
* refactor(lbug): pinRepo returns a leak-proof release disposer [#2189]
pinRepo now returns a release() disposer (mirroring addPoolCloseListener)
that releases its own lease exactly once — a double-call is a guarded no-op,
so it can never over-decrement a sibling holder's reference count. Callers
can use the leak-proof pattern `const release = pinRepo(id); try { … }
finally { release(); }`. unpinRepo stays exported for explicit pairing.
Addresses the PR #2191 review's P3: the exported pin primitive had no
built-in pairing, so a caller that forgot to unpin would disable eviction
for a repo permanently.
* refactor(group): windowed manifest resolution bounds sync pool residency [#2189]
Replace whole-sync pinning with windowed deferred resolution. The init loop
extracts contracts without pinning (repos evict naturally); manifest links are
pre-sorted and partitioned into windows whose referenced in-group repos number
<= getMaxResidentRepos(), and each window re-inits + leases only its own repos,
resolves, then RELEASES the leases (not closeLbug — released repos stay
evictable for the LRU, which avoids stomping a concurrent MCP reader).
Peak per-sync pool residency is now bounded by getMaxResidentRepos() distinct
repos regardless of group size, removing the unbounded-mmap crash risk the PR
#2191 review flagged (Finding 3) — without a new magic-number threshold (it
reuses MAX_POOL_SIZE via an intent-named accessor). #2189 stays fixed: each
window resolves against live, freshly-leased pools, so cross-links anchor to
real graph symbols.
partitionManifestWindows is a pure, unit-tested function (every link in
exactly one window — the contract-dedup invariant). New
sync-windowed-resolution.test.ts asserts the partition bound and, through the
real pool, that concurrently-open Databases never exceed the resident cap for a
group larger than it. Rewrote the sync.test.ts pinning block (init loop no
longer pins; per-window lease/release; release-not-close).
* feat(pdg): add CDG + POST_DOMINATE edge types (M5 #2085)
* feat(pdg): post-dominator tree on reverse CFG (M5 #2085)
* feat(pdg): Ferrante control-dependence over the post-dom tree (M5 #2085)
* feat(pdg): emitFileCdg + optional POST_DOMINATE debug edges (M5 #2085)
* feat(pdg): wire CDG emission in-phase + pdgModeMismatch CDG-cap stamp (M5 #2085)
* test(pdg): CDG snapshot + end-to-end pipeline answerability (M5 #2085)
* fix(review): apply autofix feedback (M5 #2085)
* fix(pdg): label CDG edges by controller arm sense, not edge kind (#2188 F1/F2/F4)
Tri-review (with Codex as the independent engine) found the CDG 'T'/'F' label
was wrong for the commonest control flow: the M1 TS visitor wires a condition's
fall-through FALSE arm as `seq`/`loop-back`, but `branchSense` mapped both to
'T', so guard clauses, if-no-else, and loop `break` got 'T' instead of 'F' (F1,
P1). The structural CDG edges were correct; only the label — the AC3 "under what
condition does X run?" answer — was wrong.
- F1: replace edge-kind `branchSense` with controller-arm-sense `labelFor`. An
ambiguous fall-through edge (seq/loop-back) takes the COMPLEMENT of its source
block's explicit cond-true/cond-false sibling arm. This correctly handles
do/while (loop-back = TRUE arm) and inner-if-in-loop (loop-back = FALSE arm) —
the ambiguity a kind→label table cannot resolve. Adds real-parser regression
tests (the hand-built tests used a fictional cond-false edge and missed it).
- F2: correct the false "sound over-approximation that never drops a real
dependence" claim in post-dominators.ts — exit-unreachable regions both drop
and invent control dependences (latent for the current TS visitor, which keeps
EXIT reverse-reachable). Reframe the exit-less-loop test to characterize, not
bless, the degenerate behavior.
- F4: make the AC2 property-test reference compute post-dominance INDEPENDENTLY
(node-removal reachability, no shared code with post-dominators.ts), so a
post-dom direction bug can no longer pass both the impl and the reference.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(ci): root-prettier format + run-analyze pdg stamp gains maxCdgEdgesPerFunction (#2085)
Two deterministic CI failures from the M5 CDG work:
- quality/format: basicblock-roundtrip.test.ts failed CI's root `prettier --check .`
(the pre-commit hook uses the gitnexus-local prettier config, which differs);
reformatted with the root config.
- tests/ubuntu/coverage: run-analyze.test.ts pinned the resolved RepoMeta.pdg
shape (DEFAULTS) and the all-zero cap override without the new
maxCdgEdgesPerFunction key (default 5000); added it so resolvePdgConfig
toEqual and pdgModeMismatch(DEFAULTS) pass. (The stale-test sweep missed this
file in PR #2188 — same trap M2 hit.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(mcp): add pdg_query tool definition (controls/flows modes) [M6 #2086]
* feat(mcp): pdg_query backend — controls (CDG) + flows (REACHING_DEF) + e2e test [M6 #2086]
* feat(mcp): document PDG edges + pdg_query (schema, cypher, skill, --pdg-gated ai-context) [M6 #2086]
* fix(mcp): correct pdg_query symbol-anchor lower bound + harden inputs [PR #2188 review]
Tri-review (Codex + adversarial + correctness lanes) of the M6 pdg_query
surface found the symbol-anchor window over-includes a neighbor function's
block. The upper bound was widened to the 1-based BasicBlock basis (symEnd+1)
but the lower bound was left 0-based, so a block on the line directly above the
target function leaked into the result. Shift both bounds +1 ([symStart+1,
symEnd+1]) so the window is the function's true block span.
Also from the same review:
- pdg_query no longer throws on a no-arguments MCP call: the dispatch passes
raw `params`, so default it to {} → a clean mode-validation error instead of
a TypeError. (`explain` shares this latent pattern — pre-existing follow-up.)
- tools.ts: the controls-mode description no longer hard-codes the 'F' branch
sense for guards — `if (!ok) return;` rides the predicate's 'T' arm; the
guard:true flag is label-agnostic (regex on the dependent block text).
Tests: a hand-seeded adjacency regression (verified failing without the
lower-bound +1) + a no-arguments validation test. Skill doc updated to document
the two-sided [symStart+1, symEnd+1] window.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): drop always-true anchor conditional in pdg_query [CodeQL #2188]
CodeQL alert 756 flagged `...(anchor ? { anchor } : {})` in _pdgQueryImpl as a
useless conditional: `anchor` is unconditionally assigned in both the file-path
and symbol branches before the return (the not-found/ambiguous/no-layer paths
return earlier), so it is always truthy. Drop `| undefined` from the declaration
(TypeScript definite-assignment holds across both branches) and emit `anchor`
directly.
No runtime change — the `anchor` field was already present on every result.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cli): add hasPdg to the noStats bridge expectation [#2188]
The M6 work threaded `hasPdg: options.pdg === true` into the AIContextOptions
passed to generateAIContextFiles on the --skills regeneration path, but this
test's strict .toEqual expectation predated it (4 keys vs 3 → CI failure). Add
`hasPdg: false` (the value on this non---pdg path). The assertion stays strict;
the #1477 noStats bridging it guards is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(cli): collapse generateGitNexusContent params to an options bag [#2188]
The function had grown to 9 positional params; reaching `hasPdg` meant passing
six `undefined`s (the M6 review's maintainability flag). Collapse params 3-9
(generatedSkills, groupNames, noStats, skipSkills, runnerPath, defaultBranch,
hasPdg) into a `GitNexusContentOptions` object with the defaults moved to
destructuring. The body is unchanged (same local names); the single production
caller and the test calls become self-documenting named fields.
Pure refactor — generated AGENTS.md/CLAUDE.md content is byte-identical.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): skip CDG for exit-unreachable CFGs (unsound post-dominance) [#2188]
M5 review P2: computePostDominators roots only at cfg.exitIndex and nothing
enforced that EXIT is reachable from every block. For an entry-reachable region
that cannot reach EXIT (a non-terminating loop, or a multi-terminal CFG a future
visitor might emit) the EXIT-rooted reverse walk degenerates — it both drops
real control dependences and invents spurious ones.
Add a pure precondition predicate `isExitReachableFromAllBlocks` (co-located with
the algorithm it guards) and gate it in emitFileCdg: a CFG that violates it is
skipped for CDG (counted as skippedUnsoundFunctions + one onWarn), while its CFG
and REACHING_DEF projections — which do not depend on post-dominance — are kept.
A CDG-specific gate, not a widening of isEmitSafeCfg, so the blast radius is
exactly the unsound CDG. The current TS visitor always satisfies the
precondition (every loop gets a structural header→loopExit edge), so CDG output
for real fixtures is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cfg): bound computeControlDependence materialization (heap parity) [#2188]
M5 review P2: unlike computeReachingDefs (maxFacts) and the emit-side edge cap,
computeControlDependence materialized the full deduped seen/out before
emitFileCdg's per-function cap could trim it — O(edges × post-dom depth) heap
for a deeply nested function.
Add a `maxEdges` ceiling (default 0 = unbounded) returning {edges, truncated},
mirroring computeReachingDefs's {facts, truncated}. The ceiling is checked
before pushing a new unique edge, so `truncated` means a genuine overflow (not
merely "reached cap"). emitFileCdg passes a FIXED materialization ceiling (8× the
default edge cap) — deliberately NOT derived from the runtime edge cap, because
CDG's materialization IS the deduped-edge quantity the cap reports on (deriving
it would pre-truncate that set and lose the exact dropped count). A ceiling hit
is surfaced via onWarn + the truncated flag — never silent.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(mcp): share resolveBlockAnchor; fix explain's anchor off-by-one [#2188]
M6 review P2 (duplication) + the flagged pre-existing _explainImpl correctness
follow-up. _pdgQueryImpl and _explainImpl each carried a near-identical
symbol↔block anchor resolver that had DRIFTED: pdg_query used the corrected
[symStart+1, symEnd+1] window (BasicBlock startLine is 1-based, the symbol span
0-based) while _explainImpl still used [symStart, symEnd] — dropping a taint
source on the function's final line AND leaking a neighbor's block on the line
directly above.
Extract one `resolveBlockAnchor` helper, used by both, that applies the correct
window and a single (bare) clause convention (callers compose their own WHERE).
This removes ~50 duplicated lines and fixes explain's anchor in one place.
A hand-seeded characterization test (taint-explain Block 4) pins both bounds —
verified to FAIL on the pre-fix window (it returned the line-10 neighbor instead
of the line-15 final-line source). Existing taint-explain + pdg-query suites are
unchanged (their fixtures have interior sources/sinks).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(mcp): pdg_query reports "status unknown" when the layer can't be confirmed [#2188]
M6 review P3 (Codex): when meta is UNREADABLE and the bounded global existence
probe returns zero rows of the edge type, _pdgQueryImpl asserted "no PDG layer"
— but a genuinely edge-free layer (all-linear functions) is indistinguishable
from a missing one via that probe. Soften only that fallback path to an
inconclusive "PDG layer status unknown — was this repo indexed with --pdg?"
note. The meta-stamped path (stamp present, cap absent ⇒ layer truly missing)
keeps the definitive "no PDG layer" wording.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(mcp): cover pdg_query ambiguous / pagination / Windows-path gaps [#2188]
M6 review test-gap follow-ups, all hand-seeded with controlled data:
- ambiguous symbol name → status:'ambiguous' + ranked candidates shape
(uid/name/filePath/score), never a silent guess;
- total/truncated page boundary in both directions (limit below the match count
sets truncated with the full total; limit above it omits truncated);
- a Windows-style filePath containing ':' resolves and fnLineOf decodes the
function-line segment correctly (split-from-right past the drive letter).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(skills): ship gitnexus-pdg-query skill mirrors + add pdg_query to the guide [#2086]
M6 bundled pdg_query into this PR, but the skill shipped only in the canonical
gitnexus/skills/ root. Mirror it (byte-identical) to the two hand-maintained
roots the sibling taint skill uses — .claude/skills/gitnexus/ and the plugin —
so Claude Code + plugin users get it too.
Also extend the gitnexus-guide tool reference (all 3 copies, now byte-identical):
add a `pdg_query` row + a "Control & data dependence" section mirroring the
taint/`explain` section, and reconcile the pre-existing drift where only the
.claude copy carried the `check` tool row (a real registered tool) — all three
now list it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(architecture): refresh CFG/PDG section for the full M1–M6 stack [#2086]
The PR body had deferred the "ARCHITECTURE docs refresh" to #2086; now that M6
ships here, do it:
- MCP tools table gains `explain` and `pdg_query` (were absent).
- "Optional CFG/PDG emission" was M1-only; rewrite to cover the whole opt-in
stack — M1 CFG, M2 REACHING_DEF, M3/M4 taint, M5 CDG (Ferrante over CHK
post-dominators, with the exit-unreachable skip), M6 read surface (pdg_query +
explain, anchored + LIMIT-bounded, shared resolveBlockAnchor) — and note the
no-Function→BasicBlock-edge join.
- LadybugDB schema notes the `--pdg` additions: the `BasicBlock` node table and
the CFG/REACHING_DEF/CDG/TAINTED/SANITIZES/TAINT_PATH relation types, kept out
of the default VALID_RELATION_TYPES / web schema.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(ci): add shared vendored-grammars manifest; monitor reads it
.github/vendored-grammars.json is the single source of truth for the vendored
tree-sitter grammars (c/swift/kotlin/dart/proto): name, upstream coords, and
policy holds. update-vendored-grammars.mjs now builds its GRAMMARS map from the
manifest (behavior-preserving — same exported shape). Adds manifest-agreement
tests so the loader can't silently skew from the file.
* fix(ci): classify vendored grammars from manifest, drop bare "?" (#858)
The readiness report decided "is this vendored?" via is_vendored_pin (a file:
package.json spec) — but the 5 vendored grammars aren't in package.json, so they
were misrouted through the npm path and rendered bare "?" for ABI (read from an
empty node_modules), plus a spurious "? (fetch failed)" for github-only proto.
Now vendored grammars are classified by membership in the shared manifest and
their ABI is read from gitnexus/vendor/<name>/src/parser.c (always in a
checkout). github-only vendored grammars skip the npm peer-dep fetch; the
tree-sitter-c hold is surfaced from the manifest (held, not plain "Ready"); and
every remaining unintrospectable value renders a labeled token, never a bare
"?". --assert-current now covers the vendored grammars too instead of skipping
them. Adds a stdlib unittest suite incl. a manifest⇄vendor-dir consistency
guard.
* docs(ci): document the shared vendored-grammars manifest
Both tree-sitter workflow headers now point at .github/vendored-grammars.json as
the shared source of truth; the readiness workflow gains a PR-path trigger on the
manifest + test, and runs the readiness unit tests on validation events.
CONTRIBUTING.md documents the manifest contract under CI automation contracts.
* fix(review): apply autofix feedback
- Guard manifest reads in both scripts with a clear error (was an opaque
module-import traceback that crashed the script and test collection).
- Never render a bare "?": relabel the npm-path ABI/version/peer sentinels and
the vendored upstream-ABI miss to labeled tokens; the report is now ?-free
regardless of node_modules/network, and the test is hermetic.
- Add a VENDORED_NAMES ⊆ GRAMMARS guard + manifest-missing error test.
- Drop now-dead is_vendored_pin/is_vendored/_(vendored)_.
- Compose held + out-of-range vendored blocker reasons instead of overwriting.
- Reword the shared-manifest docs to not over-claim shared upstream coords.
* fix(ci): apply root prettier formatting to mjs + ts test
The quality/format gate runs root `prettier --check .` (printWidth 100, the
gitnexus-local config differs and falsely passed locally).
* test(ci): make both tree-sitter scripts testable offline
The scripts hit live npm/GitHub, which makes the report run flaky and the
monitor's detect/apply logic untestable. Add hermetic seams:
- readiness: --offline flag (+ GITNEXUS_TS_READINESS_OFFLINE env) no-ops the npm
registry + upstream fetches; the report renders deterministically (vendored
ABIs from the repo, npm columns marked 'offline', no bare '?'). 3 tests assert
an offline run touches ZERO network (urlopen patched to raise).
- monitor: detect() and apply() accept injected deps (vendoredVersion/
resolveUpstream/fetchSource/readAbi) so the newer/ABI/hold gating runs offline
with fixtures; apply gains --dry-run (validates but writes nothing). 6 tests
cover newer/same-version/held-c/ABI-15/applicable + a no-mutation dry-run.
* fix(review): keep --assert-current hermetic + harden the no-bare-? invariant
Tri-review findings (PR #2187):
- P2 REGRESSION: --assert-current (documented 'hermetic and offline', run in CI
without --offline) routed the 5 vendored grammars through vendored_drift_summary,
which fetches upstream parser.c + commit sha — 10 discarded network calls per run.
Fix: read the vendored ABI locally via a new vendored_abi_from_repo() helper (also
used by vendored_drift_summary). Now verifiably network-free.
- Unify the upstream-ABI miss sentinel: prose said 'n/a (generated at build)' while
the matrix said 'n/a' — and 'generated at build' is a wrong cause (swift HAS a
committed parser.c). Both now render neutral 'n/a'.
- Fix the stale assert_current docstring claiming swift is prebuilt-only/no parser.c.
- Guard the last latent bare-? path (vendor package.json missing 'version').
Tests: AssertCurrent (network-free guard + out-of-range via the new injection
point), malformed-JSON manifest, detect() error-path, explicit npm/github
undefined assertions. 17 Python + 15 vitest, all hermetic.
* fix(review): use a single unittest import style (CodeQL 753)
CodeQL py/import-and-import-from flagged `import unittest` + `from unittest
import mock`. Collapse to `from unittest import TestCase, main, mock`.
* fix(review): explicit raise in _matrix_row (CodeQL 754)
CodeQL py/mixed-returns flagged the implicit fall-through after self.fail()
(which it doesn't model as NoReturn). End with an explicit raise AssertionError.
* test(review): replace non-null assertions with a must() guard
@typescript-eslint/no-non-null-assertion flagged 4 `!` operators. Add a
narrowing must<T>(value, message) helper (throws on undefined) and a named
baseResolveUpstream, removing every non-null assertion.
* fix(review): unguessable heredoc delimiter for the report output
The report embeds the manifest `hold` field (fork-PR-editable); a fixed
DRIFT_EOF delimiter in a hold value could close the $GITHUB_OUTPUT heredoc
early and inject output keys. Use DRIFT_EOF_$(openssl rand -hex 16) — a value
the report cannot contain. (Randomized delimiter over base64: keeps REPORT raw
markdown, no consumer-side decode.)
* fix(review): scope issues:write to scheduled runs (two-job split)
GitHub Actions has no step-level permissions, so the only way to keep PR runs
(incl. forks) from receiving `issues: write` is to split the job. A `report`
job (contents:read, all events) renders the report + the PR `::warning::` and
exposes report/exit_code as job outputs; a schedule-only `upsert-issue` job
(needs: report, issues:write, no checkout) consumes them for the issue upsert +
close. The 'Check upgrade readiness' check name is preserved.
* fix(review): launder npm-version '?' in disposition prose
The disposition bucket prose interpolated r['npm_version'] raw, so a successful
200 npm /latest response lacking a 'version' key would render a bare '?' (the
matrix cell already laundered it). Add npm_version_label ('unknown' for '?') and
use it in all five bucket renderers. Test a version-less npm response.
* refactor(review): load_vendored_manifest returns only the consumed 'hold'
The readiness script reads only the grammar names + 'hold'; the 'key' and
'upstream' fields were phantom data (upstream-drift coords live in the script's
own GRAMMARS map). Narrow the return to {hold}.
* fix(review): unify detect()/apply() 'newer' check for github grammars
detect() compared the bare sha7 while apply() compared up.version (the full
<base>-g<sha7> provenance string apply() also writes). After the bot re-vendored
a github grammar once, detect() reported a perpetual false 'update available'
while apply() correctly saw 'already current' — a noisy job summary + wasted
--apply subprocess (the PR-exists guard absorbed it before any duplicate PR).
Extract a shared isNewer(up, have) helper used by both. Tests cover equal-
provenance (false), first-vendoring plain-version (true, not suppressed), and
sha-advanced (true). Coupled with U12 (the detect⇄apply agreement assertion
lives there once apply()'s not-newer path returns instead of process.exit).
* test(review): cover main()'s out-of-range + prebuilt-only vendored ABI branches
main()'s vendored-ABI classification reads through vendored_abi_from_repo (the
local-read seam --assert-current uses), so patching it drives the
'Vendored (ABI out of range)' blocker branch and the prebuilt-only (vendored_abi
None → 'prebuilt' cell, not '?') branch — neither reachable today since all 5
vendor dirs ship parser.c at ABI 14.
* test(review): monitor-side manifest⇄vendor-dir consistency guard
Mirror the Python consistency guard on the monitor side — the monitor consumes
the same manifest and is the side that WRITES files from manifest `name`, so
manifest/vendor-dir drift must fail CI here too.
* fix(review): validate grammar names at manifest load (path-traversal guard)
The manifest `name` is joined into gitnexus/vendor/<name> paths in both scripts
(and apply() WRITES there), so reject any name not matching tree-sitter-[a-z0-9-]+
at the single load chokepoint — defense-in-depth even though the live trust
boundary already prevents exploitation. loadManifestGrammars gains an injectable
`raw` arg + export for testing; tests reject a '../etc' name in both scripts.
* refactor(review): apply() throws ApplyExit; CLI maps to exit codes
apply()'s 4 process.exit calls killed the vitest worker, blocking in-process
tests of its error branches. Replace them with a thrown ApplyExit{code}; the
not-newer (already-current) path returns `have` instead of exit(0). The isMain
CLI block try/catches and maps ApplyExit.code → process.exit, so the monitor's
subprocess contract (exit 0/2/3) is byte-identical (verified via subprocess
smoke). Tests cover unknown-key=2, held=3, ABI-reject=3, and not-newer (returns
current, no throw, no write).
* refactor(review): extract vendored render helper; trim docstrings (<1000 lines)
Extract the 'Vendored parsers' prose render into _render_vendored_section() so
main() coordinates named phases rather than inlining a ~450-line monolith, and
condense the most verbose docstrings/comments. The script drops from 1092 to 999
lines (under the 1000 bar the maintainability review flagged). Behavior-preserving:
the deterministic --offline render is byte-identical before/after (verified
in-place), --assert-current still passes, and the full unit suite is green.
* fix(review): row-diff regex captures only the Status cell
The change-detection regex captured the whole row tail as group 2, so any
non-status cell drift (e.g. an upstream-ABI bump) emitted a false-positive
'change' line. Capture only the Status cell ([^|]+? before the final |$).
The workflow parseRows regex and the Python _ROW_DIFF_RE stay byte-identical;
the stability test now asserts group 2 is the status string (e.g. c →
'Vendored — held') and contains no pipe.
* fix(ci): hoist intro string out of the list literal (CodeQL 755)
The U13 extraction moved the 'Vendored parsers' intro paragraph (implicitly
concatenated string literals) INTO a list literal, tripping CodeQL
py/implicit-string-concatenation-in-list (reads as a possibly-missing comma
between elements). Hoist it into a parenthesized `intro` variable. Render is
byte-identical.
* feat(cli): add --embeddings-baseurl/-model/-auth-token/-dims flags to analyze
Add four CLI flags to `gitnexus analyze` that configure a custom
OpenAI-compatible HTTP embedding endpoint by setting the
GITNEXUS_EMBEDDING_URL / _MODEL / _API_KEY / _DIMS env vars the HTTP
embedding client already reads. Flags override env vars; env vars keep
working as before. URLs are validated (http/https) and dims must be a
positive integer. Prints "Using custom embedding endpoint: <url>" when
a URL+model pair is configured, and warns when the flags are passed
without --embeddings. The new env keys are added to the analyze
snapshot/restore set so programmatic callers don't leak state. The
non-secret flags are also accepted from .gitnexusrc; the auth token is
intentionally CLI/env-only.
* fix(analyze): set GITNEXUS_EMBEDDING_DIMS from CLI flags before module import
schema.ts reads EMBEDDING_DIMS at module-load time via the static-import
chain (analyze.ts -> run-analyze.ts -> schema.ts). The previous approach
of setting the env var inside analyzeCommandImpl ran AFTER schema.ts had
already loaded with the default 384, causing "Expected: 384, Actual: 4096"
errors when using --embeddings-dims 4096.
Fix: use Commander's preAction hook to set GITNEXUS_EMBEDDING_* env vars
before the lazy import of analyze.ts triggers the schema.ts module load.
* fix(analyze): use hook callback arg instead of this in preAction
Commander v14 passes the command as first argument, not as this binding.
* refactor(cli): rename --embeddings-* analyze flags to singular --embedding-*
Aligns the custom embedding endpoint flags with the existing singular
tuning flags (--embedding-threads/--embedding-device): --embedding-base-url,
--embedding-model, --embedding-auth-token, --embedding-dims. Renames the
derived AnalyzeOptions fields and the .gitnexusrc KEY_SPECS keys to match.
Behavior-preserving; the GITNEXUS_EMBEDDING_* env vars are unchanged.
Refs #2140 review.
* fix(cli): validate and normalize --embedding-dims before module-load reads it
The preAction hook wrote GITNEXUS_EMBEDDING_DIMS unvalidated, so an invalid
value (abc/0/-5/0x10) threw from schema.ts during the lazy import — surfacing
as a raw unhandled rejection on the synchronous program.parse path instead of
a friendly error. And '1e3' slipped through: schema.ts parseInt froze the
vector column at FLOAT[1] while the impl's Number-based check accepted 1000,
so http-client requested 1000-dim vectors against a 1-dim column.
Extract a dependency-free normalizeEmbeddingDims helper (strict /^\d+$/ +
positive, trim-then-validate, canonicalized) shared by both the hook (CLI
path, before module-load) and analyzeCommandImpl (direct-call path). All three
readers — schema.ts, http-client, and this helper — now agree on one value,
and invalid input gets a clean message instead of a crash or a silent mismatch.
Refs #2140 review.
* fix(cli): mask credentials in the custom embedding endpoint confirmation
A base URL with userinfo (http://user:pass@host/v1) or a query token
(?api_key=…) passed the new-URL + http/https validation and was printed
verbatim in the 'Using custom embedding endpoint:' line, leaking the secret
to terminal scrollback and CI logs. Route it through the existing safeUrl()
(now exported from http-client) which strips userinfo + query, keeping
protocol/host/path. Single source of truth — no second sanitizer.
Refs #2140 review.
* fix(cli): drop the ineffective embeddingDims .gitnexusrc key
embeddingDims as a .gitnexusrc key silently did nothing: .gitnexusrc loads in
analyzeCommandImpl, AFTER the lazy import already ran schema.ts's module-load
read of GITNEXUS_EMBEDDING_DIMS, so a config value never sized the vector
column. Remove it (config now fails closed on the key, like the auth token);
URL/MODEL stay as config keys because they're read lazily at runtime. Dims
remains available via --embedding-dims or GITNEXUS_EMBEDDING_DIMS.
Refs #2140 review.
* refactor(cli): narrow the analyze preAction hook to GITNEXUS_EMBEDDING_DIMS
Only DIMS is read at module-load (schema.ts), so only it must be set before the
lazy import. URL/MODEL/API_KEY are read lazily at runtime, so analyzeCommandImpl
is their sole setter — and because the impl's env snapshot is taken AFTER this
hook ran, leaving those three in the hook leaked them past restore. Drop them
from the hook (the impl already sets+restores them), and capture/restore the
pre-hook DIMS baseline via a postAction hook so a CLI --embedding-dims override
no longer leaks into a later in-process program.parseAsync.
Refs #2140 review.
* fix(cli): gate the custom-endpoint confirmation on the embedding flags
The confirmation collapsed into one if/else chain that emits at most one
message reflecting the run's intent. Gating on embeddingsEnabled stops the
'Using custom embedding endpoint' line from printing on every analyze run when
GITNEXUS_EMBEDDING_URL+MODEL merely happen to be set in the environment, and
ordering the '--embeddings absent' note first removes the contradiction where
it printed alongside 'Using custom embedding endpoint'.
Refs #2140 review.
* test(cli): cover the custom embedding endpoint flags
Adds direct-call (analyzeCommandImpl path) coverage the original PR lacked:
URL validation (empty/invalid/non-http), model/token emptiness, dims
validation incl. the 1e3 regression, credential masking in the confirmation
line, confirmation gating (absent --embeddings; ambient env must not trigger
it), CLI-over-env precedence, and the GITNEXUS_EMBEDDING_* snapshot/restore
round-trip. Complements embedding-dims.test.ts and http-client-safe-url.test.ts.
Refs #2140 review.
* test(cli): e2e-cover the --embedding-dims crash path on the real CLI
The dims-validation fix lives in the commander preAction hook, which only
fires on the program.parse path; the direct analyzeCommand() unit tests bypass
it. Add a subprocess e2e (run via tsx, no build) asserting that invalid
--embedding-dims (abc/0/-5/1e3/3.5) produces the friendly flag-named error and
exit 1 — NOT the raw schema.ts module-load throw that the original bug
surfaced. Cases exit inside the hook (no repo/import/pipeline), so they're
deterministic and fast. Updates the unit-suite comment to point at it.
Refs #2140 review.
---------
Co-authored-by: Gergo Magyar <gergomagyar@icloud.com>
2026-06-13 14:01:12 +01:00
1166 changed files with 1086261 additions and 672119 deletions
"description":"Code intelligence powered by a knowledge graph. Provides execution flow tracing, blast radius analysis, and augmented search across your codebase."
@@ -24,6 +24,7 @@ Run from the project root. This parses all source files, builds the knowledge gr
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
| `--pdg` | Build the program-dependence layers used by `explain` and `pdg_query` (taint, CDG, and REACHING_DEF). |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Claude Code, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
When no path exists, `trace` reports the furthest reachable node — exactly where the chain breaks (dynamic dispatch, reflection, or an external boundary).
@@ -75,13 +83,34 @@ Notes: `offset` ≥ `total` returns an empty page (with `total` still reported).
### Taint findings (`explain`)
`explain` returns intra-procedural taint findings (`TAINTED` edges) recorded by `gitnexus analyze --pdg` — each with a sink category (command-injection, code-injection, path-traversal, sql-injection, xss), source/sink lines, and the ordered hop path with the variable carried on each hop.
`explain` returns taint findings recorded by `gitnexus analyze --pdg` — intra-procedural `TAINTED` edges plus cross-function `TAINT_PATH` hops where the interprocedural taint phase found a function-level source→sink chain. Each finding includes a sink category (command-injection, code-injection, path-traversal, sql-injection, xss), source/sink lines, and the ordered hop path with the variable carried on each hop.
- `explain {}` — enumerate all findings for the repo (bounded by `limit`, deterministic order)
- `explain { target: "src/vuln.ts" }` — findings in a file (suffix path match accepted)
- `explain { target: "runUserCommand" }` — findings in a function (resolved like `context`; ambiguous names return ranked candidates)
A repo indexed without `--pdg` returns a clear "no taint layer" note. Caveats: findings are intra-procedural only — cross-function, closure/callback, property/field, and implicit flows are not modeled, so the absence of a finding is **not** proof of safety. `SANITIZES` (sanitizer-kill) edges are queryable via `cypher`.
A repo indexed without `--pdg` returns a clear "no taint layer" note. Caveats: closure/callback, property/field, and implicit flows are not modeled, and interprocedural findings are function-level `TAINT_PATH` hops rather than statement-level path proof, so the absence of a finding is **not** proof of safety. `SANITIZES` (sanitizer-kill) edges are queryable via `cypher`.
### Control & data dependence (`pdg_query`)
`pdg_query` reads the control/data-dependence layers `gitnexus analyze --pdg` records (CDG + REACHING_DEF, basic-block granular) — the control/data analog of `explain`. It is **always anchored** (a `target` file path or symbol, resolved like `context`) and has two modes:
- `pdg_query { mode: "controls", target: "..." }` — CDG: "under what condition does X run?". Each edge is a controlling predicate block → dependent block with the branch sense (`'T'`/`'F'`) in `reason`; an edge into an early `return`/`throw` is flagged `guard: true` (guard-clause discovery — the sense depends on the predicate, so don't filter guards by a fixed label).
- `pdg_query { mode: "flows", target: "...", variable?: "..." }` — REACHING_DEF def→use edges within the function; pass `variable` to trace one binding.
A repo indexed without `--pdg` returns a "no PDG layer" note (or "status unknown" when the layer can't be confirmed). Intra-procedural only — cross-function flow is taint's domain (`explain`). The raw CDG/REACHING_DEF edges are also queryable via `cypher`. See the `gitnexus-pdg-query` skill for the full query surface.
### Shortest path between two symbols (`trace`)
`trace` answers "how does A reach B?" in one call — the shortest directed path over `CALLS` (plus `HAS_METHOD`, so a class-rooted trace descends into its methods) instead of chaining 3–8 `context`/`impact` hops by hand.
- `trace { from: "validateUser", to: "executeQuery" }` — shortest path between two symbols.
- Disambiguate common names with `from_uid`/`to_uid` (zero-ambiguity) or `from_file`/`to_file`; an ambiguous name returns ranked candidates.
- `maxDepth` (default 10, max 30) bounds the search; `includeTests` (default false) lets the traversal pass through test-file symbols.
Returns ordered `hops` (each `{ name, filePath, startLine }`) and an aligned `edges[]` of `{ relType, confidence }`, so call hops and containment (`HAS_METHOD`) hops stay distinguishable. When no path exists it reports the **furthest** reachable node (where the chain breaks) and sets `truncated: true` if a traversal cap was hit first. Every result carries a `status`: `ok` / `no_path` / `ambiguous` / `not_found` / `error`.
Cross-repo (experimental): pass `repo: "@groupName"` to trace across a group's member repos — the path may cross **one**`ContractLink` boundary (reported as a `CONTRACT_LINK` hop with the bridged contract in `crossings[]`). Omit `to` entirely to follow `from`'s outgoing HTTP call to whatever provider endpoint it lands on. Groups are configured via `group_list` / `group_sync`.
## Resources Reference
@@ -98,8 +127,10 @@ Lightweight reads (~100-500 tokens) for navigation:
## Graph Schema
**Nodes:** File, Function, Class, Interface, Method, Community, Process
description:"Use when the user wants the GitNexus engineering pipeline run end-to-end on a task: gitnexus-plan (plan depth chosen up front), a blocking gate to execute with gitnexus-work or stop, finishing with a gitnexus-review of the result. Examples: \"/gitnexus-lfg Add retry support to the ingestion pipeline\", \"run the gitnexus pipeline on this\", \"plan, build and review this feature\"."
---
# gitnexus-lfg — plan → gate → work → review
Thin orchestrator over three existing skills. It adds no engineering logic of
its own — it sequences `gitnexus-plan`, `gitnexus-work`, and
`gitnexus-review`, with the user deciding at the plan gate. Run every lane
by actually invoking the named skill (read its SKILL.md and follow it);
never inline a summary of what the skill would have done.
```
/gitnexus-lfg <task description>
/gitnexus-lfg docs/plans/<existing-plan>.md # skip lane 1, start at the gate
```
## Lane 1 — Plan
**Boundary triage first.** If the task is plainly below the planning
boundary — trivial or small-bounded work an agent finishes in well under ~35
turns (the measured regime where a planning pass costs more than it returns;
measured in the GitNexus repository's `eval/workflow_bench/`) — say so and
offer `gitnexus-work` direct mode as an alternative to the full pipeline
before spending the plan lane. Honor the user's choice.
The threshold is a promoted benchmark policy measured offline, not a
timeless heuristic — never self-edit it from a live task. Its re-evaluation
governance lives in this skill's README.
Otherwise invoke `gitnexus-plan` with the task (knob overrides pass through
verbatim; `gitnexus-plan` owns the up-front depth question — never ask it
again here). If the input is already a plan file path, skip to Lane 2. The
plan lands in `docs/plans/` — record its path; every later lane consumes it.
## Lane 2 — The plan gate (user choice, blocking)
Present the plan's chat summary (objective, proposed changes, sequence, top
risks, open questions, plan path), then ask the user — as a blocking
question (`AskUserQuestion` in Claude Code; a numbered list in chat on CLIs
without a blocking tool):
1.**Proceed to work** — continue to Lane 3.
2.**Stop here** — the plan file is the deliverable; end the pipeline.
Depth was the user's up-front choice in Lane 1, so deepening is not offered
by default — but honor an explicit request for it at the gate: run
`gitnexus-plan` Deepen mode on the plan file and return here with the
strengthened plan, as many times as the user asks. Do not proceed past the
gate without an explicit choice — the gate is the pipeline's only checkpoint
and exists precisely because execution is expensive to unwind.
**Headless / non-interactive runs:** no one can answer the gate, so end the
pipeline after Lane 1 — the plan file is the deliverable (gate option 2) —
and say so in the final report. Never auto-proceed to execution.
## Lane 3 — Work
Invoke `gitnexus-work` with the plan path. It re-anchors the plan at HEAD,
executes the Implementation Sequence as verified atomic commits, refreshes
the knowledge graph when done (its Phase 4), and reports deviations. If it routes back for re-planning (structural drift), run the
Deepen pass and return to the Lane 2 gate rather than pushing through.
## Lane 4 — Review
Invoke `gitnexus-review` on the completed work. Pass an open PR URL/number
when one exists; otherwise pass the current branch. The review skill owns
target resolution, exact-SHA checkout/index alignment, and merge-base
selection. Do not duplicate that logic here. If work left local changes,
pass `local` as a second, separately labeled review surface.
Surface the review verdict and findings to the user. Findings the user
wants fixed: those within `gitnexus-work`'s direct-mode bounds (1–2 files,
no architectural decisions) → hand to `gitnexus-work` direct mode; anything
larger → offer the plan gate instead (Deepen the plan with the findings, or
stop). Then re-run this lane's review once. On that re-run, do not start
another fix cycle even if findings remain — report them and point the user
at `/gitnexus-work` (or the plan gate) to continue deliberately.
## Final report
One message: plan path, deepen cycles run, commits produced, verification
status, review verdict with unresolved findings, and what (if anything) was
explicitly left undone. The pipeline does not push or open a PR on its own —
| **Codex CLI** | Ask: "run gitnexus-plan for <task>" (Codex reads `AGENTS.md`) — or install the user-level prompt below | `AGENTS.md` § Engineering planning & execution |
| **Any AGENTS.md-aware agent** | Ask it to "read `.claude/skills/gitnexus-plan/SKILL.md` and follow it for <task>" | `AGENTS.md` § Engineering planning & execution |
```
/gitnexus-plan Add retry support to the ingestion pipeline
/gitnexus-plan Fix the stale warm-cache invalidation bug in exportedTypeMap
/gitnexus-plan depth:deep impact_depth:3 Migrate the emit phase to streaming COPY
```
Output: `docs/plans/YYYY-MM-DD-gitnexus-plan-<slug>.md` — a 13-section plan whose
section 11 is a machine-readable **implementation context pack** that a
follow-up agent can consume without re-investigating the repository. Compact
and full packs both include versioned evidence provenance: a canonical global
dirty digest and a sorted, per-layer cited-path manifest. An npm-dependency-free,
versioned Node helper shared byte-for-byte with `gitnexus-work` is the only
supported serializer, so planner and executor hash identical bytes. The same
helper is the only supported existing-plan reader and plan writer. Its
descriptor-anchored `read-plan` receipt binds the canonical path, exact base64
bytes, and SHA-256 digest before Deepen or execution. The writer accepts a repo-relative
description:'Use when you need a deep, implementation-ready engineering plan for a code change — built from GitNexus graph intelligence, statement-level PDG analysis, and targeted source verification, compact enough that an implementation agent can start without re-investigating. Also strengthens existing plans via Deepen mode. Examples:"/gitnexus-plan Add retry support to the ingestion pipeline","/gitnexus-plan deepen docs/plans/<plan>.md","plan this change using the knowledge graph".'
| `depth` | by category | `narrow` = `impact_depth` 1, PDG only if one function is clearly central; `default` = this table; `deep` = `impact_depth` 3 + clusters/processes read |
| `form` | by category | `compact` (core sections + mini-pack, ≤80 lines excl. pack — see `references/plan-template.md`) or `full` (all 13 sections) |
| `impact_depth` | 2 | `maxDepth` for `impact` |
| `pdg_data_depth` | 2 | Data-dependence hops in the PDG slice |
| `pdg_control_depth` | 2 | Control-dependence hops in the PDG slice |
| `max_snippet_lines` | 30 | Longest source excerpt quoted in the plan |
| `freshness` | by category | `strict` (full-plan categories) = refresh a stale index (and a missing PDG layer) with `analyze --index-only [--pdg]` before relying on the graph; `accept` (compact categories) = plan on the current graph, source-weighted and labelled, refreshing only if a graph claim becomes load-bearing |
## Fallback mode (GitNexus or PDG unavailable)
1. Say so, first thing, in chat and in the plan.
2. Use targeted repo exploration (grep/glob/reads) to approximate callers,
dependencies, execution flow, state changes, and related tests.
3. Label every such finding **source-derived** in the plan — never present it
as graph-derived, and never fabricate statement-level edges.
4. Recommend `analyze --index-only` (add `--pdg` for the PDG layers) via
the resolved runner — `node .gitnexus/run.cjs`, installed `gitnexus`, or
`npx gitnexus` — when it would materially raise confidence.
## Skill feedback
If this run exposed friction in the instructions, include concise feedback in
the final response. Feedback is chat-only: do not append evaluation learnings,
edit benchmark data, or modify this skill during a live planning task.
description:'Review code changes with GitNexus from a GitHub PR URL or number, a branch/ref or commit range, or local staged, unstaged, and untracked changes. Use when the user asks for a code review, merge-risk assessment, regression hunt, missing-test analysis, or a verdict on whether a PR, branch, commit range, or local diff is safe.'
---
# GitNexus review
Review the requested change surface without editing source, committing, pushing,
posting, or resolving threads. A later explicit request may authorize those
actions. Use GitNexus for structural evidence and source inspection for proof;
description:CI review swarm lane. Assumes the change is broken and constructs concrete failure scenarios — races, hostile inputs, state corruption, abuse of new surfaces — verified against source and the GitNexus graph. Read-only; reports findings only.
description:CI review swarm lane. Hunts logic errors, edge cases, contract breaks, and state bugs in the changed symbols of a PR, grounded in the GitNexus graph. Read-only; reports findings only.
description:CI review swarm gate. Audits the orchestrator's draft review before publication — every finding anchored and concrete, severities calibrated, sections and verdict wording conformant, no generic filler. Returns PASS or a defect list; never rewrites the review.
description:'Use when executing an engineering plan produced by gitnexus-plan (or a small bounded task directly) — implements step by step with GitNexus impact checks before every symbol edit, tests from the plan''s scenarios, and detect_changes gating every commit. Examples:"/gitnexus-work docs/plans/2026-07-11-gitnexus-plan-ingestion-retry.md","/gitnexus-work"(latest plan), "execute the plan".'
---
# gitnexus-work — execute a gitnexus-plan
Execute an implementation plan produced by `gitnexus-plan`, shipping it as a
sequence of verified, atomic commits. The plan's section 11
(`implementation_context` pack) is the primary machine-readable input; the
prose sections are its rationale. This skill **does** edit code — it is the
executor counterpart to the planning-only `gitnexus-plan`.
```
/gitnexus-work <plan path> # execute this plan
/gitnexus-work # newest docs/plans/*gitnexus-plan*.md here
/gitnexus-work <small task text> # direct mode, see Input triage
```
## Input triage
- **Plan path** (or blank → the newest `docs/plans/*gitnexus-plan*.md` under
the current repo root): the normal mode; continue to Phase 1. Schema-2
description:"Use when querying or extending GitNexus's PDG control/data-dependence surface (the `pdg_query` MCP tool, CDG/REACHING_DEF edges), or reasoning about \"what controls X\" / \"where does Y flow\" / guard clauses. Examples: \"what guards this statement?\", \"trace this variable within the function\", \"why is the pdg_query result empty?\", \"add a CDG query\"."
---
# PDG query surface with GitNexus
Expert knowledge for the `pdg_query` MCP tool and the control/data-dependence
edges it reads — the opt-in `--pdg` program-dependence layers. Read this before
touching `gitnexus/src/mcp/local/local-backend.ts` (`_pdgQueryImpl`) or the
`pdg_query` tool def, or when explaining a `pdg_query` result.
## When to Use
- "Under what condition does this statement run?" (guarding predicates).
- "Where does this variable flow inside the function?" (def→use).
- Guard-clause discovery (early-return guards — subsumes the #559 heuristic).
- Extending or reviewing `pdg_query` / the CDG / REACHING_DEF read path.
- Debugging an empty or surprising `pdg_query` result.
## The layered substrate (build order)
`pdg_query` runs **on** the same graph taint runs on. Each layer is opt-in
behind `--pdg`; a default `analyze` run records none of them (byte-identical).
description:"Use when the user wants to review a pull request, understand what a PR changes, assess risk of merging, or check for missing test coverage. Examples: \"Review this PR\", \"What does PR #42 change?\", \"Is this PR safe to merge?\""
---
# PR Review with GitNexus
## When to Use
- "Review this PR"
- "What does PR #42 change?"
- "Is this safe to merge?"
- "What's the blast radius of this PR?"
- "Are there missing tests for this PR?"
- Reviewing someone else's code changes before merge
- NEVER rename symbols with find-and-replace — use `gitnexus_rename`.
- NEVER commit without running `gitnexus_detect_changes()`.
- NEVER ignore HIGH/CRITICAL risk warnings from impact analysis.
- NEVER run `npx gitnexus analyze` without `--embeddings` if `.gitnexus/meta.json` shows stored embeddings.
- NEVER run `npx gitnexus analyze` without `--embeddings` if the index metadata (`.gitnexus/gitnexus.json` / legacy `meta.json`) shows stored embeddings.
Full rules: **[AGENTS.md](../AGENTS.md)** (`gitnexus:start` block, Cursor Cloud section).
@@ -10,6 +10,8 @@ A cross-platform Dev Container that pre-installs Claude Code, OpenAI Codex CLI,
>
> The trade-off of the copy model: host and container config **diverge after first create.** A skill or plugin you add on the host later won't appear in the container until you wipe the config volume and rebuild (see [§ Rebuild / reset](#rebuild--reset)). Edits you make inside the container persist across rebuilds but never reach the host.
"_comment":"Single source of truth for the VENDORED SET + policy holds, read by BOTH .github/scripts/update-vendored-grammars.mjs (weekly auto-PR bot) and .github/scripts/check-tree-sitter-upgrade-readiness.py (daily readiness report -> issue #858). The monitor also resolves each grammar's upstream from the `upstream` field here; the readiness report reads vendored ABIs from gitnexus/vendor/<name>/src/parser.c and keeps its own upstream-drift coords. A consistency-guard test asserts this set equals the gitnexus/vendor/tree-sitter-* directories. See CONTRIBUTING.md.",
"grammars":{
"c":{
"name":"tree-sitter-c",
"upstream":{"npm":"tree-sitter-c"},
"hold":"ABI-pinned at 0.21.4 (#1242/#858) — needs a tree-sitter runtime upgrade before bumping"
},
"swift":{
"name":"tree-sitter-swift",
"upstream":{"npm":"tree-sitter-swift"}
},
"kotlin":{
"name":"tree-sitter-kotlin",
"upstream":{"npm":"tree-sitter-kotlin"},
"hold":"pinned to unreleased fwcd main commit c8ac3d26 for `fun interface` support (fwcd/tree-sitter-kotlin#169, closes #87) — npm latest (0.3.8) lacks the fix, so the monitor must NOT auto-revert (isNewer is strict-inequality: 0.3.8 != 0.4.0). Drop this hold and bump when upstream cuts a release that includes the fix"
echo "::notice::Release GitHub App secrets (RELEASE_APP_ID / RELEASE_APP_PRIVATE_KEY) are not configured — prebuilds will build and upload as artifacts, but the auto-PR is skipped. Provision the App, or run with open_pr=false to suppress this notice."
fi
# ── Fork PRs: emit the PR identity so the trusted `commit-fork-prebuilds`
# workflow_run job can push the rebuilt prebuilds back onto the fork's
# branch. That job has no PR context of its own (workflow_run.pull_requests
# is empty for forks), so it reads this. Same-repo PRs don't need it — the
# aggregate job below commits straight onto their branch. This artifact is
# untrusted producer output: every field is allowlist-validated again on
# the consumer side AND cross-checked against the workflow_run authority.
✅ **Rebuilt native prebuilds pushed to this PR branch.** A grammar source change re-cut the vendored \`tree-sitter\` prebuilds for all 6 platforms and they're now committed on your branch. ([builder run](${RUN_URL}))" ;;
nothing-to-commit)
body="${marker}
✅ Native prebuilds are already up to date on this branch — nothing to push." ;;
loop-prevented)
body="${marker}
🔁 Skipping prebuild push: the branch HEAD is already an automated prebuild commit." ;;
lease-failed)
body="${marker}
⏳ The PR head moved while the prebuilds were building, so they weren't pushed. Push another commit (or wait for the next build) and they'll be re-cut. ([builder run](${RUN_URL}))" ;;
push-failed)
body="${marker}
⚠️ Rebuilt native prebuilds are ready but **couldn't be pushed to your fork branch**. Tick **Allow edits by maintainers** in the PR sidebar so CI can commit them — or download them from the [builder run](${RUN_URL}) artifacts (\`ts-prebuild-*\`) and commit them under \`gitnexus/vendor/<grammar>/prebuilds/\` yourself." ;;
*)
body="${marker}
❓ Prebuild delivery finished in an unexpected state (\`${RESULT:-unknown}\`). See the [builder run](${RUN_URL})." ;;
esac
# Strip the YAML block indent so the rendered comment starts at column 0.
body="$(printf '%s\n' "$body" | sed 's/^ //')"
# Upsert a single sticky comment keyed by the marker; only ever edit our
# own bot comment (PATCH on someone else's 403s and would abort).
existing=$(gh api "repos/${GH_REPO}/issues/${PR}/comments" --paginate \
echo '::error::GITNEXUS_BENCH_AUTH_TOKEN is not configured. The evolution loop runs real benchmark sessions and needs an Anthropic API key (not the Claude Code OAuth token).'
Automated skill-evolution promotion. The deterministic gate passed; this PR is the human-review step — inspect the diff and the evidence before merging.
- **Call & inheritance resolution (RFC #909 Ring 3):** See ARCHITECTURE.md § Scope-Resolution Pipeline. All languages resolve calls and inheritance through the scope-resolution pipeline (`Registry.lookup`, `preEmitInheritanceEdges`, `emitHeritageEdges`, `buildMro` → `MethodDispatchIndex`). **Shared code in `gitnexus/src/core/ingestion/` must not name languages** — plug language behavior in via `LanguageProvider` / `ScopeResolver` hooks. A language plugs in by implementing `ScopeResolver` (`scope-resolution/contract/scope-resolver.ts`) and registering it in `SCOPE_RESOLVERS`. (The legacy call-resolution DAG + `@heritage` capture path were removed in RING4-1 #942.)
Four canonical, CLI-neutral skill specs under `.claude/skills/` (Claude Code invokes
them as slash commands; Codex or any other agent reading this file should read the
named SKILL.md and follow it directly — user-level Codex prompts are documented in the
plan/work/lfg skill READMEs):
- **`gitnexus-plan/SKILL.md`** — deep, implementation-ready plan for a code change:
GitNexus graph intelligence for navigation, statement-level PDG slices for behavioral
constraints, targeted source reads for verification. Output lands in `docs/plans/`
with a reusable implementation context pack (section 11). Planning-only — it never
edits code (index freshness refreshes via `analyze --index-only` are the one
permitted state change). Interactive runs ask up front how deep to go
(quick / standard / deep); Deepen mode strengthens an existing plan in place.
- **`gitnexus-work/SKILL.md`** — executes a gitnexus-plan as verified atomic commits:
drift-checks the plan's evidence pin against HEAD, `impact` before every symbol
edit, tests from the plan's scenarios, `detect_changes` before every commit.
- **`gitnexus-review/SKILL.md`** — read-only GitNexus review of a PR URL/number,
branch or commit range, or local staged/unstaged/untracked changes. It pins exact
SHAs, aligns the graph and checkout, runs a PDG-backed taint pass on trust-boundary
diffs, scales to per-domain expert lenses from the graph's clusters (dispatched as
parallel swarm lanes — `ci-personas/` — when the CI review agent runs it), and
reports evidence-backed findings.
- **`gitnexus-lfg/SKILL.md`** — pipeline orchestrator: plan (depth asked up front) →
blocking user gate (proceed or stop) → work → `gitnexus-review`.
The family ships with the npm package (`gitnexus/skills/`, installed to editor targets
by `gitnexus setup`) and the Claude Code plugin; review also has a standalone Cursor
mirror. `gitnexus/test/unit/shipped-skills-sync.test.ts` guards the copies. Token savings of the workflow are measurable with
`eval/workflow_bench/` (real headless CLI runs, free-model routing supported — see its README).
## Changelog
| Date | Version | Change |
|------|---------|--------|
| 2026-07-20 | 1.14.0 | `gitnexus-review` gains a coordinated swarm: six `ci-personas/` lanes the CI review agent dispatches as subagents (via the `Agent` tool), with a bounded critic gate and sidechain-excluded evidence. |
| 2026-07-16 | 1.13.0 | `gitnexus-plan` asks plan depth up front (quick/standard/deep) in interactive runs; `gitnexus-lfg` gate slimmed to proceed/stop (Deepen stays as the route-back mechanism). |
| 2026-07-16 | 1.12.0 | Renamed `gitnexus-pr-review` to `gitnexus-review`; added PR URL/number, branch/range, and local-change targets plus install migration (setup warns on a legacy `gitnexus-pr-review` dir and leaves it in place; uninstall removes it). |
| 2026-07-11 | 1.11.0 | Skill family shipped via npm skills/ + plugin (sync-guarded); added eval/workflow_bench token-savings benchmark. |
| 2026-05-22 | 1.8.0 | Kotlin added to `MIGRATED_LANGUAGES` (registry-primary call resolution by default). Closes #1756 (companion-vs-instance dispatch) and #1757 (lambda scopes); refs #1746. RFC §6.4 corpus criterion waived (corpus-mode wiring is #927-scope); fixture criterion met. |
| 2026-04-23 | 1.7.0 | TypeScript added to `MIGRATED_LANGUAGES` (registry-primary call resolution by default). |
| 2026-04-20 | 1.6.0 | Added scope-resolution pipeline pointer (RFC #909 Ring 3); Python migrated to registry-primary. |
@@ -74,17 +111,18 @@ commits, or posts.
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (20319 symbols, 54304 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> Index stale? Run `node .gitnexus/run.cjs analyze` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? `npx gitnexus analyze` (npm 11 crash → `npm i -g gitnexus`; #1939).
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST run `detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows. For regression review, compare against the default branch: `detect_changes({scope: "compare", base_ref: "main"})`.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `query({search_query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `context({name: "symbolName"})`.
1.**Ingestion** — `analyze.ts` → `runFullAnalysis` (`run-analyze.ts`) → `runPipelineFromRepo` (`pipeline.ts`). DAG of 14 phases builds a `KnowledgeGraph` in memory, then loads into LadybugDB under `.gitnexus/`. Repo registered in `~/.gitnexus/registry.json` for MCP discovery.
1.**Ingestion** — `analyze.ts` → `runFullAnalysis` (`run-analyze.ts`) → `runPipelineFromRepo` (`pipeline.ts`). DAG of 15 phases builds a `KnowledgeGraph` in memory, then loads into LadybugDB under `.gitnexus/`. Repo registered in `~/.gitnexus/registry.json` for MCP discovery.
| `group_list` | List repo groups or details for one group |
| `group_sync` | Rebuild group Contract Registry (`contracts.json`) and bridge graph |
`query`, `context`, and `impact` are group-aware: pass `repo: "@<groupName>"` (or `"@<groupName>/<memberPath>"` to scope to one member) plus optional `service: "<monorepo/path>"`. Group-mode `query` merges per-repo results via Reciprocal Rank Fusion; group-mode `impact` runs the local walk in the chosen member and fans out across boundaries via the Contract Bridge (`gitnexus/src/core/group/cross-impact.ts`). The previously-planned `group_query`, `group_context`,`group_impact`, `group_contracts`,`group_status` MCP tools are intentionally not introduced — group-level state is exposed via resources instead:
`query`, `context`, and `impact` are group-aware: pass `repo: "@<groupName>"` (or `"@<groupName>/<memberPath>"` to scope to one member) plus optional `service: "<monorepo/path>"`. Group-mode `query` merges per-repo results via Reciprocal Rank Fusion; group-mode `impact` runs the local walk in the chosen member and fans out across boundaries via the Contract Bridge (`gitnexus/src/core/group/cross-impact.ts`). `trace` is also group-aware via `repo: "@<groupName>"` — but, unlike the others, it resolves`from`/`to` across **all** members (a`@<groupName>/<memberPath>` suffix is advisory for trace, not a scope); pass `from_uid`/`to_uid` to disambiguate a symbol name that occurs in more than one member.
Group-mode `trace` (`gitnexus/src/core/group/cross-trace.ts`) stitches a path that crosses repositories: it resolves `from`/`to` across all members, and when they live in different repos it joins the home-repo segment to the target-repo segment over a single `ContractLink` boundary (an HTTP consumer→provider link, joined on `Contract.symbolUid`), reported as a `CONTRACT_LINK` hop in `crossings[]`. The crossing is clamped to one boundary (`MAX_SUPPORTED_CROSS_DEPTH`, shared with cross-impact); deeper `crossDepth` is reported via `notes[]`. With `pdg: true` (experimental, opt-in), each boundary-adjacent segment is enriched with its intra-procedural REACHING_DEF data-flow when that repo was indexed with `--pdg` (reusing the same anchored `flows` query as `pdg_query`); data flow never crosses the repo boundary, and a missing PDG layer degrades to call-level hops with a note. Two stores meet only at the `symbolUid` grain — the per-repo PDG/call graph and the group bridge — so this is the documented join; full cross-program (SDG-like) data flow across the boundary remains deferred (see `docs/plans/2026-06-18-002-feat-unified-pdg-impact-evaluation-plan.md`). The previously-planned `group_query`, `group_context`, `group_impact`, `group_contracts`, `group_status` MCP tools are intentionally not introduced — group-level state is exposed via resources instead:
- **Binding accumulator lifecycle** — created in `parse`, disposed by `crossFile` (in `finally`). No other phase should take ownership.
- **Skippable phases** — `skipGraphPhases` omits MRO/communities/processes (faster tests); `pruneLocalSymbols` still runs (it is graph cleanup, not analysis). `skipWorkers` is no longer a sequential escape hatch — it (like `--workers 0` / `GITNEXUS_WORKER_POOL_SIZE=0`) is rejected with an actionable error, since the worker pool is the sole parse path (§ Chunked parse-and-resolve).
- **Skippable phases** — `skipGraphPhases` omits MRO/di/communities/processes (faster tests); `pruneLocalSymbols` still runs (it is graph cleanup, not analysis). `skipWorkers` is no longer a sequential escape hatch — it (like `--workers 0` / `GITNEXUS_WORKER_POOL_SIZE=0`) is rejected with an actionable error, since the worker pool is the sole parse path (§ Chunked parse-and-resolve).
- **Local-symbol pruning** — `pruneLocalSymbols` removes inert block-local value symbols after scope resolution has consumed them. Opt out per-call with `PipelineOptions.keepLocalValueSymbols` or globally with the `GITNEXUS_KEEP_LOCAL_VALUE_SYMBOLS` env var.
### How to add a new phase
@@ -195,7 +201,9 @@ Language-agnostic scope-resolution resolver. This is the resolution path for eve
ReferenceIndex
│ emitReceiverBoundCalls ── FIRST
│ emitFreeCallFallback ── THEN
│ emitReferencesViaLookup ── LAST (uses handledSites)
│ emitReferencesViaLookup ── uses handledSites + deferred-site skip set
@@ -204,9 +212,30 @@ Language-agnostic scope-resolution resolver. This is the resolution path for eve
Orchestrator: `runScopeResolution(input, provider)` in `scope-resolution/pipeline/run.ts`.
Pipeline phase: `scopeResolutionPhase` in `scope-resolution/pipeline/phase.ts` — iterates the registered `SCOPE_RESOLVERS` over the worker-serialized `ParsedFile`s. (Per-language `emitScopeCaptures` hooks may reuse a cached Tree via the orchestrator's `treeCache`, but in worker-pool runs that cache is empty — Trees can't cross MessageChannels — so they consume the pre-extracted `ParsedFile` instead; § Performance notes.)
### Optional CFG/PDG emission (`--pdg`, #2081 M1)
### Callable-value flow
On a `--pdg` run, the parse worker builds a per-function control-flow graph from the tree-sitter AST (`LanguageProvider.cfgVisitor`; TypeScript/JavaScript in M1) and serializes it onto `ParsedFile.cfgSideChannel` as plain data. Scope-resolution then emits `BasicBlock` nodes + `CFG` edges from that side-channel **inside Phase 4 of `runScopeResolution`, while the disk-backed ParsedFile store is still live** — the only window where the worker-built CFGs are loaded (the store is cleared right after the phase returns). A standalone post-`mro` phase would read an empty store, so the CFG emit deliberately lives in-phase, mirroring the `applyCaptureSideChannel` pattern. The opt-in is off by default (graph byte-identical), folded into the parse-cache key (a pdg-off warm cache is never reused on a `--pdg` run), and bounded by a per-function edge cap that logs any dropped edges. Edge *kind* (`seq`/`cond-true`/`loop-back`/…) rides in the `CFG` relationship's `reason` (CFG is a single `CodeRelation` type, not one type per kind). See `core/ingestion/cfg/`.
First-class callable values use a language-neutral inclusion analysis in `passes/callable-value-flow.ts`. Providers recognize their own syntax and emit JSON-safe `CallableFlowSite` facts (`seed`, `copy`, `alias`, `address`, `load`, `store`, `formal`, `argument`, and `invoke`) into `ParsedFile`; shared ingestion never branches on a language name. These always-on facts cross workers and the durable parse store, whose schema is bumped whenever their semantic shape changes.
The emit stage defers only invocation sites proven to reference a flow cell. Ordinary receiver/free/reference passes still resolve direct callees first and record exact callee IDs by file/line/column. Property dispatch then runs before callable flow because a property-dispatched wrapper call can seed actual-to-formal propagation. The callable solver consumes those direct targets, propagates callable sets through lexical cells and formals, and emits `CALLS` at the real indirect invocation site with reason `callable-value-flow` (confidence 0.8 for a singleton, 0.7 for a bounded multi-target set).
The solver is flow-insensitive but bounded: dependency-indexed work items rerun only when a cell they read changes; target/address sets cap at 32; a hostile fact graph has a finite work budget; overflow or budget exhaustion emits no partial `CALLS` and produces a structured warning. Lexical shadowing is function/block aware, invocation/constructor results are not reinterpreted as callable designators, and overload selection uses provider-supplied signature metadata. C/C++ additionally associate visible prototypes with unique definitions so actual-to-formal flow crosses translation units; the provider-owned `hasFileLocalCallableLinkage` hook prevents `static` declarations or definitions from leaking across files. C++ member-function pointers preserve parameter/cv shape, keep non-virtual targets exact, and expand virtual targets through `MethodDispatchIndex`/MRO.
Property-key dispatch remains a separate conservative fallback. Its per-key fan-out cap is 32; capped keys synthesize no partial calls and are reported at warning level with language, skipped-key count, dropped key names (bounded), and cap; the count also travels in `RunScopeResolutionStats.propertyDispatchSkippedKeys`.
Standalone (regex-based) providers such as COBOL participate via `ScopeResolver.scopeResolutionEdgeMode: 'callable-flow-only'`: `runScopeResolution` runs for them, but every ordinary emission path — heritage, interface implementations, receiver-bound, free-call fallback, reference/import edges, post-resolution hooks — is gated off, so their legacy phase (e.g. `cobolPhase`) remains the sole owner of structural edges and the callable solver's `CALLS` are purely additive. A callable-flow-only provider whose files emitted no callable facts exits early, before finalize, keeping the opt-in proportional to source scanning.
On a `--pdg` run the parse worker builds a per-function control-flow graph from the tree-sitter AST (`LanguageProvider.cfgVisitor`; TypeScript/JavaScript today) and serializes it onto `ParsedFile.cfgSideChannel` as plain data. Scope-resolution then emits the program-dependence layers from that side-channel **inside Phase 4 of `runScopeResolution`, while the disk-backed ParsedFile store is still live** — the only window where the worker-built CFGs are loaded (the store is cleared right after the phase returns). A standalone post-`mro` phase would read an empty store, so the emit deliberately lives in-phase, mirroring the `applyCaptureSideChannel` pattern. The opt-in is off by default (graph byte-identical), folded into the parse-cache key (a pdg-off warm cache is never reused on a `--pdg` run), and each layer is bounded by a per-function edge cap that logs any dropped edges. All layers are `BasicBlock → BasicBlock` edges in the single `CodeRelation` table, keyed by `type`; there is **no**`Function → BasicBlock` edge — the symbol↔block join is reconstructed from the BasicBlock id prefix + line span. The layers build on each other:
- **M1 — CFG** (#2081): `BasicBlock` nodes + `CFG` edges. Edge *kind* (`seq`/`cond-true`/`loop-back`/…) rides the `reason` column (CFG is one `CodeRelation` type, not one per kind).
- **M2 — REACHING_DEF** (#2082): GEN/KILL def→use data dependence from a pure fixpoint solver; the variable name rides `reason`.
- **M3/M4 — TAINTED / SANITIZES / TAINT_PATH** (#2083–#2084): intra- and inter-procedural taint (source→sink) — the `explain` tool's data.
- **M5 — CDG** (#2085): Ferrante control dependence over a Cooper–Harvey–Kennedy post-dominator tree (the EXIT-rooted reverse CFG); branch sense (`'T'`/`'F'`) rides `reason`. A CFG whose EXIT is unreachable from some block is skipped for CDG (post-dominance would be unsound) while its CFG/REACHING_DEF layers are kept.
- **M6 — read surface** (#2086): the `pdg_query` MCP tool answers "what gates X?" (CDG, `mode: controls`) and "where does Y flow?" (REACHING_DEF, `mode: flows`); `explain` is the taint consumer. Both are always anchored + `LIMIT`-bounded (LadybugDB has no rel-property index) and share one `resolveBlockAnchor` helper. These PDG edge types are deliberately kept out of the default `VALID_RELATION_TYPES` / web schema.
- **Cross-repo trace enrichment**: group-mode `trace` (`pdg: true`) reuses the same anchored REACHING_DEF `flows` query to annotate a boundary-adjacent segment with how a value reaches the cross-repo call — strictly intra-procedural (data flow never crosses the repo boundary). See the group-aware tools note above.
See `core/ingestion/cfg/` (emit + the pure CFG / post-dominator / control-dependence / reaching-defs / taint passes) and `mcp/local/local-backend.ts` (`_pdgQueryImpl`, `_explainImpl`, the shared `resolveBlockAnchor`).
### `ScopeResolver` contract
@@ -227,6 +256,7 @@ Single interface a language implements to plug into the pipeline. Contract fully
| `collapseMemberCallsByCallerTarget?` | One CALLS edge per (caller, target) instead of per-site — default off |
| `hoistTypeBindingsToModule?` | Walk up to Module scope when looking up a method's return-type typeBinding — default off; enable only when bindings are stored at module level |
| `hasFileLocalCallableLinkage?` | Precise internal-linkage predicate used only when joining callable declarations/prototypes to cross-file definitions; C/C++ use it for `static` free functions |
### Per-language registration
@@ -258,6 +288,7 @@ CI auto-discovers the set via `tsx`. No workflow edit required.
- **Cross-phase Tree cache**: the orchestrator's `treeCache` (`RunScopeResolutionInput.treeCache`) lets a scope-resolution per-language hook (`emitScopeCaptures`) reuse a tree instead of re-parsing. Workers leave it empty — Trees can't cross MessageChannels — so in normal (worker-pool) runs scope-resolution does NOT rely on it: workers serialize each file's `ParsedFile` (+ capture side-channel) and stream them in, so scope-resolution consumes the pre-extracted artifact rather than re-parsing on the main thread (§ Chunked parse-and-resolve). `PROF_SCOPE_RESOLUTION=1` emits hit/miss counters and a worker-engaged warning.
- **Typed relationship iteration**: heritage + MRO walk only the EXTENDS / IMPLEMENTS / HAS_METHOD edges via `iterRelationshipsByType`, not the full relationship map.
- **Workspace-resolution-index**: O(1) `findOwnedMember` / `findExportedDef` / `classScopeByDefId` built once per run.
- **Callable-value worklist**: dependency-indexed inclusion propagation is linear in a reverse-ordered copy-chain fixture; target/address sets cap at 32 and the whole worklist has a finite budget with no partial edge emission on exhaustion.
- **SCC-ordered cross-file return-type propagation** (PR #1050): `propagateImportedReturnTypes` walks `indexes.sccs` in reverse-topological order (leaves first), so multi-hop alias chains like `models.User → service.user → app.user` collapse to the terminal class in a single linear pass. Within each importer, the source module's `typeBindings` is chain-followed BEFORE mirroring (so we mirror terminal types, not intermediate refs), and the importer's own `typeBindings` is chain-followed AFTER mirroring (so local `const x = importedFn()` resolves before downstream importers run). Cyclic SCCs reach a partial fixpoint within a single pass without iterating to convergence — see the `ts-circular` cross-file-binding fixture which only asserts pipeline-no-throw. PROF output (`PROF_SCOPE_RESOLUTION=1`) splits `finalize` from `propagate` so quadratic regressions in the chain-follow surface independently.
---
@@ -289,6 +320,7 @@ Each language implements `LanguageProvider` (`language-provider.ts`). Key fields
| `exportChecker` | Public/exported symbol detection |
| `typeConfig` | Type annotation extraction rules |
| `mroStrategy` | `first-wins` / `c3` / `none` |
| `descriptionExtractor` | Optional hook returning a symbol's doc-comment text as its `description`; feeds the embedding metadata header so doc-only terms are semantically searchable (issue #2270). Most languages register `createLeadingDocDescriptionExtractor` (shared, language-neutral; per-language comment/wrapper config passed at the call site) |
16 providers in `languages/index.ts` via `satisfies Record<SupportedLanguages, LanguageProvider>` — missing a language is a compile error.
**Optional `--pdg` additions** (off by default, opt-in via `gitnexus analyze --pdg`; see _Optional CFG/PDG emission_ above): a `BasicBlock` node table, plus the PDG relation types `CFG`, `REACHING_DEF`, `CDG`, `TAINTED`, `SANITIZES`, and `TAINT_PATH` on the same `CodeRelation` table. These are deliberately kept out of the default `VALID_RELATION_TYPES` / web graph schema — query them via `cypher`, `explain`, or `pdg_query`.
## Embeddings and search
**Embeddings** (`src/core/embeddings/`): Snowflake arctic-embed-xs (384D). Embeddable: File, Function, Class, Method, Interface. Incremental via SHA1 content hash. Separate `Embedding` table.
- **Call & inheritance resolution:** See ARCHITECTURE.md § Scope-Resolution Pipeline. Shared pipeline code in `gitnexus/src/core/ingestion/` must not name languages — use `LanguageProvider` / `ScopeResolver` hooks instead (see AGENTS.md). (The legacy call-resolution DAG was removed in #942.)
- **GitNexus:** `.claude/skills/gitnexus/`; MCP and indexed-repo rules live only in [AGENTS.md](AGENTS.md) (`gitnexus:start` … `gitnexus:end`). See **GitNexus rules** below.
- **GitNexus:** standard skills in`.claude/skills/gitnexus-*/`; MCP and indexed-repo rules live only in [AGENTS.md](AGENTS.md) (`gitnexus:start` … `gitnexus:end`). See **GitNexus rules** below.
- **Engineering plans, execution & review:** `/gitnexus-plan <task>` (implementation-ready plans via GitNexus + statement-level PDG + source verification; Deepen mode for existing plans), `/gitnexus-work [plan]` (executes a plan as impact-checked, detect_changes-gated atomic commits), `/gitnexus-review [PR|branch|range|local]` (read-only graph-backed review), `/gitnexus-lfg <task>` (plan with depth asked up front → proceed/stop gate → work → review pipeline). Specs in `.claude/skills/gitnexus-{plan,work,review,lfg}/SKILL.md` (see AGENTS.md § Engineering planning & execution).
## Changelog
| Date | Version | Change |
|------|---------|--------|
| 2026-07-20 | 1.8.0 | The CI review agent runs `gitnexus-review` as a coordinated swarm — six `ci-personas/` lanes dispatched via the `Agent` tool with a bounded critic gate. |
| 2026-07-16 | 1.7.0 | `/gitnexus-plan` asks depth up front in interactive runs; `/gitnexus-lfg` gate slimmed to proceed/stop. |
| 2026-07-16 | 1.6.0 | Renamed `/gitnexus-pr-review` to `/gitnexus-review` and added PR, branch/range, and local-change targets. |
| 2026-07-11 | 1.5.0 | Added `/gitnexus-work` and `/gitnexus-lfg` to the engineering plans & execution pointer. |
| 2026-04-13 | 1.3.0 | Updated GitNexus index stats after DAG refactor. |
| 2026-03-24 | 1.2.0 | Removed duplicated gitnexus:start block and scope table; replaced with pointers to AGENTS.md. |
| 2026-03-23 | 1.1.0 | Updated agent instructions to match AGENTS.md. |
@@ -56,17 +62,18 @@ See the `<!-- gitnexus:start --> … <!-- gitnexus:end -->` block in **[AGENTS.m
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **GitNexus** (26675 symbols, 35395 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **GitNexus** (20319 symbols, 54304 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> Index stale? Run `node .gitnexus/run.cjs analyze` from the project root — it auto-selects an available runner. No `.gitnexus/run.cjs` yet? `npx gitnexus analyze` (npm 11 crash → `npm i -g gitnexus`; #1939).
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST run `detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows. For regression review, compare against the default branch: `detect_changes({scope: "compare", base_ref: "main"})`.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `query({search_query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `context({name: "symbolName"})`.
@@ -13,12 +13,17 @@ This project uses the [PolyForm Noncommercial License 1.0.0](https://polyformpro
## Development setup
**Prerequisites:** Node.js — `gitnexus/` requires `>=22.0.0` and `gitnexus-web/` requires `^20.19.0 || >=22.12.0` (enforced via the `engines` field in each package). Use `nvm install` to match the local version.
**Prerequisites:** Node.js — `gitnexus/` requires `^22.18.0 || >=24.11.0` and `gitnexus-web/` requires `^20.19.0 || >=22.12.0` (enforced via the `engines` field in each package). Use `nvm install` to match the local version.
5. Run tests as described in [TESTING.md](TESTING.md).
The CLI build imports `gitnexus-shared`, so a fresh clone must install and build
the shared package before running `npm install` in `gitnexus/`. This is the same
order used by the repository's `setup-gitnexus` CI action.
### Containerized development (optional)
@@ -144,6 +149,15 @@ Re-invoking `/autofix` after a successful apply is a safe no-op — the workflow
**Sensitive paths.** The apply workflow refuses any patch that touches `.github/` (workflow files, CODEOWNERS, dependabot config). A malicious PR could ship a custom prettier or ESLint config that reformats workflow YAML; if accepted, those edits would be pushed under `contents: write` without human review. Apply formatter changes to files under `.github/` manually in a normal commit so they get the same review every other workflow change gets.
### Vendored tree-sitter grammars
`.github/vendored-grammars.json` is the **single source of truth** for the vendored tree-sitter grammar **set** and each grammar's policy `hold` (the ones shipped from `gitnexus/vendor/<name>` rather than installed from npm). It lists each grammar's name, upstream coords (`npm` or `github`), and any `hold`. The monitor resolves upstreams from it; the readiness report keeps its own upstream-drift coords and reads vendored ABIs from `gitnexus/vendor/`. Two workflows read it:
- `tree-sitter-upgrade-readiness.yml` (`.github/scripts/check-tree-sitter-upgrade-readiness.py`) — daily; renders the tree-sitter-0.25 readiness report (issue #858), reading each vendored grammar's ABI from `gitnexus/vendor/<name>/src/parser.c`.
Sharing the manifest keeps the two aligned: a consistency-guard test asserts the manifest set equals the `gitnexus/vendor/tree-sitter-*` directories. **When you vendor a new grammar (or remove one), update `.github/vendored-grammars.json` in the same change** — otherwise that guard fails CI and the readiness report regresses to `?` placeholders.
## AI-assisted contributions
If you use coding agents, follow project context files (e.g. `AGENTS.md`, `CLAUDE.md`) and avoid drive-by refactors unrelated to the issue. Prefer incremental, test-backed changes.
@@ -160,7 +174,9 @@ routes between two modes based on the triggering event:
not enforce branch reachability. No Docker build (RC-only). Before cutting a
@@ -134,7 +134,7 @@ Run the commands relevant to the touched area. If something cannot be run in the
### 4.5 If CI workflows or release pipelines changed
- [ ] The workflow passes a dry-run or triggered run before merge; concurrency (`cancel-in-progress`) and the `setup-gitnexus` action remain wired correctly.
- [ ] The workflow passes a dry-run or triggered run; concurrency (`cancel-in-progress`) and the `setup-gitnexus` action remain wired correctly. Workflows that only execute once registered on the default branch (an `issue_comment` trigger, or a newly added `workflow_dispatch`) cannot be dry-run pre-merge — merge them **registered but disabled**, then validate same-repo and fork execution post-merge before enabling.
- [ ]`CHANGELOG.md` is **not** edited here — it is owned by the release process.
This guide shows how to connect GitNexus to the Kilo Code VS Code extension using Kilo’s MCP support, based on a setup that has been tested successfully.
## Prerequisites
GitNexus should already be installed globally and working on the target repository, and the repository should be indexed successfully with `gitnexus analyze` before testing inside Kilo.
## Tested Versions
| Component | Version |
| --- | --- |
| VS Code | 1.125.1 (user setup) |
| Node.js | 24.15.0 |
| Kilo Code | 7.3.50 |
| OS | Windows 11 25H2 / Windows_NT x64 10.0.26200 |
| GitNexus | 1.6.7 |
## Where Kilo Stores MCP Config
Kilo Code stores MCP server configuration in its main config file. For the VS Code extension, config can be stored at either the global or project level.
| Scope | Config path |
| --- | --- |
| Global | `~/.config/kilo/kilo.jsonc` |
| Project | `kilo.jsonc` or `.kilo/kilo.jsonc` in the project root |
Kilo supports local MCP servers through STDIO, and GitNexus should be added as a local server under the `mcp` key in `kilo.jsonc`. Use this configuration:
```jsonc
{
"mcp":{
"gitnexus":{
"type":"local",
"command":["npx","-y","gitnexus@latest","mcp"],
"enabled":true,
"timeout":10000
}
}
}
```
## Check It Through the Kilo UI
1. restart kilo code extension or vs code
2. open kilo code settings
3. select mcp server section
#### From there, Kilo allows adding, editing, enabling, disabling, and deleting MCP servers, and it writes changes directly to the appropriate config file.

## Test the Connection
After configuration, Kilo automatically detects the tools exposed by the MCP server and can use them from chat once the server is available.
@@ -19,7 +19,7 @@ Maintainer may widen scope per task.
2.**Never rename with find-and-replace** in GitNexus-indexed projects — use `rename` MCP tool with `dry_run: true` first, review `graph` vs `text_search` edits. No separate `gitnexus rename` CLI exists.
3.**Run impact analysis before editing shared symbols** — `impact` (upstream) for functions/classes/methods others call. Do not ignore HIGH/CRITICAL without maintainer sign-off.
4.**Run `detect_changes` before commit** — confirm diffs map to expected symbols/processes when the graph is available.
5.**Preserve embeddings** — plain `npx gitnexus analyze` now preserves any embeddings recorded in `.gitnexus/meta.json`(the previous behavior wiped them). Use `--embeddings` to also generate vectors for new/changed nodes; use `--drop-embeddings` only when an explicit wipe is intended (e.g., model swap).
5.**Preserve embeddings** — plain `npx gitnexus analyze` now preserves any embeddings recorded in the index metadata (`.gitnexus/gitnexus.json`, mirrored to the legacy `meta.json`) — the previous behavior wiped them. Use `--embeddings` to also generate vectors for new/changed nodes; use `--drop-embeddings` only when an explicit wipe is intended (e.g., model swap).
---
@@ -30,20 +30,20 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
### Stale graph after edits
- **Trigger:** MCP warns index is behind `HEAD`, or search doesn't match latest commit.
- **Do:** `npx gitnexus analyze` (plus `--embeddings` if used). Runs incrementally by default — the pipeline parses every file every run (cross-file resolution requires it), but tree-sitter dispatch is skipped for unchanged file chunks via the content-addressed cache, and only changed-file rows (plus their importers, transitively) are rewritten in LadybugDB.
- **Do:** `npx gitnexus analyze` (plus `--embeddings` if used). Runs incrementally by default — the pipeline parses every file every run (cross-file resolution requires it), but tree-sitter dispatch is skipped for unchanged file chunks via the content-addressed cache, and only changed-file rows (plus their importers, transitively) are rewritten in LadybugDB. When the effective write set exceeds ~50% of the repo's files (minimum 50 files), the run transparently switches to the full wipe + bulk-COPY write plan and logs "switching to a full DB write" — expected behavior, not a bug, and file-level bookkeeping stays incremental.
- **Why:** Tools query LadybugDB from last analyze; git changes are invisible until re-indexed.
### Index seems corrupt or "incremental" is misbehaving
- **Trigger:** `analyze` produces unexpected results, or `meta.json.incrementalInProgress` is set, or the index is in a half-state after a crash.
- **Do:** `npx gitnexus analyze --force` to rebuild from scratch. The dirty-flag check forces this automatically when a previous incremental run didn't complete cleanly, but `--force` is the manual escape hatch. Safe to delete the `.gitnexus/parse-cache/` directory (and any legacy `.gitnexus/parse-cache.json`) at any time — content-addressed, will be regenerated.
- **Trigger:** `analyze` produces unexpected results, or `incrementalInProgress` is set in the index metadata (`.gitnexus/gitnexus.json` / legacy `meta.json`), or the index is in a half-state after a crash.
- **Do:** `npx gitnexus analyze --force` to rebuild from scratch. The dirty-flag check forces this automatically when a previous incremental run didn't complete cleanly, but `--force` is the manual escape hatch. A dirty-flag recovery rebuild parks the interrupted run's sidecars beside the DB as `lbug.wal.dirty-recovery` / `lbug.shadow.dirty-recovery` for post-mortem debugging — harmless, and removable with `npx gitnexus clean --lbug-sidecars`. Safe to delete the `.gitnexus/parse-cache/` directory (and any legacy `.gitnexus/parse-cache.json`) at any time — content-addressed, will be regenerated.
- **Why:** Incremental writeback is selective DB row replacement; if the on-disk state is inconsistent for any reason, a full rebuild is the cheapest path back to a known-good index.
### Embeddings vanished after analyze
- **Trigger:** Semantic search quality drops; `stats.embeddings` in `meta.json` is 0 after refresh.
- **Trigger:** Semantic search quality drops; `stats.embeddings` in the index metadata (`gitnexus.json` / legacy `meta.json`) is 0 after refresh.
- **Do:** Re-run `npx gitnexus analyze --embeddings` to regenerate. Check the analyze log for a `Warning: could not load cached embeddings` line — if present, the cache restore failed (corrupt DB / schema mismatch) and the rebuild had nothing to preserve. If you intentionally passed `--drop-embeddings`, this is expected.
- **Why:** Plain `analyze` preserves prior vectors by re-inserting them after the rebuild; the only ways to end up at zero are an explicit `--drop-embeddings`, a cache-load failure (now logged), or a model/dimension change that invalidates the cache.
- **Why:** Plain `analyze` preserves prior vectors by re-inserting them after the rebuild; the only ways to end up at zero are an explicit `--drop-embeddings`, a cache-load failure (now logged), or a model/dimension change that invalidates the cache. A dirty-recovery run that cannot move the crashed WAL aside now either discards it (logged: forensics lost, embeddings still preserved) or fails fast with a lock error naming the holder — it never silently zeroes embeddings.
### MCP lists no repos
@@ -61,7 +61,7 @@ Format: **Trigger → Instruction → Reason**. Append new Signs when the same m
- **Trigger:** Errors opening `.gitnexus/lbug` while MCP and analyze both run.
- **Do:** Stop overlapping processes (one writer at a time). Retry analyze or restart MCP.
- **Why:** Embedded DB expects single-process ownership.
- **Why:** Embedded DB expects single-process ownership.`@ladybugdb/core` 0.18.0 also reports this contention as `"Only one write transaction at a time is allowed in the system."` — our busy/lock retry matcher (`isDbBusyError` in `src/core/lbug/lbug-config.ts`) recognizes this exact string too, so it's auto-retried the same as any other lock error. If you see that exact message, it's the same "one writer at a time" issue above, not a new failure mode.
**Important:** If you already had embeddings, **always** pass `--embeddings` on later analyzes, or they can be dropped. See `stats.embeddings` in `.gitnexus/meta.json`(0 means none).
**Important:** If you already had embeddings, **always** pass `--embeddings` on later analyzes, or they can be dropped. See `stats.embeddings` in `.gitnexus/gitnexus.json` (or its legacy `meta.json`mirror; 0 means none).
**Large repos:** Analyze may skip or limit embedding work when node counts are very high; watch CLI output.
@@ -152,7 +152,9 @@ Analyze re-execs Node with a **large old-space heap** when needed (`analyze.ts`)
## LadybugDB / lock errors
Only one process should open a repo’s `.gitnexus/lbug` store at a time. If MCP and a second `analyze` run conflict, stop one process, then retry `analyze` or restart MCP.
Only one process should open a repo's `.gitnexus/lbug` store at a time. If MCP and a second `analyze` run conflict, stop one process, then retry `analyze` or restart MCP.
If the error text is `"Only one write transaction at a time is allowed in the system."` instead of a lock/busy message, it's the same underlying conflict — our retry matcher (`isDbBusyError` in `src/core/lbug/lbug-config.ts`) recognizes this exact string and auto-retries it. The fix if it still surfaces after retries is the same: stop the overlapping process.
@@ -16,18 +16,19 @@ Evaluate whether GitNexus code intelligence improves AI agent performance on rea
> **Recommended**: Use `native_augment` mode. It mirrors the Claude Code model — the agent gets both explicit GitNexus tools (fast bash commands) AND automatic enrichment of grep results with callers, callees, and execution flows. The agent decides when to use explicit tools vs rely on enriched search output.
**Models supported:**
**Models supported** (see `configs/models/` for the current list):
- Claude 3.5 Haiku, Claude Sonnet 4, Claude Opus 4
- MiniMax M1 2.5
- Claude Haiku 4.5, Claude Sonnet 4, Claude Opus 4
- MiniMax M1 2.5, MiniMax M2.5
- GLM 4.7, GLM 5
- DeepSeek
- Any model supported by litellm (add a YAML config)
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.