Skip to content
Development
Skill

/gitnexus-taint-analysis

Use when working on, reviewing, or extending GitNexus's CFG/taint/PDG subsystem (the `--pdg` layers), or when reasoning about source→sink data-flow findings. Examples: \"How does taint analysis work here?\", \"Why didn't explain find this flow?\", \"Add a new sink/source\",

From plugin
gitnexus
47k13 skills7 agents1 MCP
Install
$ npx -y skills add abhigyanpatwari/GitNexus --skill gitnexus-taint-analysis --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/gitnexus-taint-analysis

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when working on, reviewing, or extending GitNexus's CFG/taint/PDG subsystem (the `--pdg` layers), or when reasoning about source→sink data-flow findings. Examples: \"How does taint analysis work here?\", \"Why didn't explain find this flow?\", \"Add a new sink/source\",

SKILL.md

gitnexus-taint-analysis.SKILL.md
name: gitnexus-taint-analysis
description: "Use when working on, reviewing, or extending GitNexus's CFG/taint/PDG subsystem (the `--pdg` layers), or when reasoning about source→sink data-flow findings. Examples: \"How does taint analysis work here?\", \"Why didn't explain find this flow?\", \"Add a new sink/source\", \"Review the interprocedural taint code\"."

CFG & Taint Analysis with GitNexus

Expert knowledge for the opt-in `--pdg` program-analysis subsystem: control-flow graphs, reaching definitions, and intra- + inter-procedural taint. Read this before touching `gitnexus/src/core/ingestion/cfg/**` or `gitnexus/src/core/ingestion/taint/**`, or when explaining a finding.

When to Use

  • "How does the taint engine work / why is this flow (not) reported?"
  • Adding a source, sink, or sanitizer to the model.
  • Extending or reviewing the CFG / reaching-defs / taint / summary code.
  • Understanding the `explain` MCP tool's findings (intra- vs inter-procedural).
  • Debugging a false positive or false negative in `--pdg` output.

The layered substrate (build order)

Taint runs **on** the graph, not beside it. Each layer is opt-in behind `--pdg` and a default `analyze` run is **byte-identical** (the golden parity gate is the hard floor for every change here).

L1  CFG            per-function basic blocks + control-flow edges   (M1 #2081)
L2  REACHING_DEF   GEN/KILL def→use data dependence (pure solver)   (M2 #2082)
L3  Taint (intra)  source→sink over RD facts, minus sanitizers      (M3 #2083)
L4  Taint (inter)  per-function summaries composed over CALLS       (M4 #2084)
  • **Worker-built, main-thread-solved.** The parse worker builds each function's

CFG + harvests def/use + call-site facts onto `ParsedFile.cfgSideChannel` (plain, structured-clone-safe data — never AST nodes). The main thread runs the pure solvers. NEVER re-parse on the main thread (re-introduces the #1983 OOM).

  • **In-phase emit (KTD1).** L1–L4-harvest all run INSIDE the scope-resolution

pdg window (`scope-resolution/pipeline/run.ts`, gated `input.pdg === true`), because the disk-backed ParsedFile store is cleared when that phase ends — a standalone post-`mro` phase would read empty data. The cross-function fixpoint (L4) is the exception: it runs in its OWN registered phase (`taintSummaries`) AFTER scope-resolution, because it needs the COMPLETE call graph, and consumes small plain summary data threaded out via `ScopeResolutionOutput`.

  • **Pure-solver contract.** `computeReachingDefs`, `computeTaintFlows`,

`harvestFunctionSummary`, and `solveInterprocTaint` are pure and deterministic (no graph, no I/O, no logger; sorted outputs). Snapshot tests and content-derived edge ids depend on it.

Intra-procedural taint (L3)

Forward reachability over RD facts from matched **sources** to matched **sinks**, killed by **sanitizers**. Key design points worth internalizing:

  • **Occurrence-tagged sites.** A flat per-arg binding set cannot tell

`exec(escape(x))` (safe) from `exec(x)` (finding); the harvest records nested call structure (`SiteRecord.parent`/via-tags) so sanitizer interposition is precise.

  • **Kind-set sanitizer model.** A taint carries a set of *neutralized*

`SinkKind`s; a sink fires unless its kind is in the set. So `escape(req.body)` suppresses `res.send` (xss) but STILL fires `db.query` (sql) — a kind-blind kill would be a suppressed live injection (the forbidden FN direction). `path.basename(t)` neutralizes path-traversal only, not command-injection.

  • **Statement-level finding identity.** NOT block-pair (block conflation drops

distinct findings; `exec(req.body, req.query)` is two findings).

  • Persisted as `TAINTED` edges (BasicBlock→BasicBlock); the path rides the

`reason` column via the shared versioned codec (`taint/path-codec.ts`).

Interprocedural taint (L4) — the functional/summary method

The production approach (Sharir-Pnueli 1981; the same shape as Meta's Pysa and Mariana Trench, and FB Infer) — NOT full IFDS tabulation. Each function is reduced to a compact **summary**, and summaries are composed over the already- resolved `CALLS` graph.

**Summary shape** (`taint/summary-model.ts`, whole-parameter granularity):

| Edge | Meaning | Analogue | |------|---------|----------| | `param→return` | a param flows to the return value | TITO — **reserved** (the floor already covers its recall; precision pass deferred) | | `param→callee-arg` | a param flows into arg *j* of a call (carries the path's neutralized sink kinds) | TITO into callee | | `param→sink` | a param reaches a modelled sink | partial/triggered sink | | `source→return` | the function generates+returns a source | generative — **composed** via the caller's `callResults` | | `source→callee-arg` | a generated source flows into a call | fixpoint SEED | | `callResults` | a user-function call's result flows to a sink/return/callee-arg in the caller | composes with callee `source→return` |

**The fixpoint** (`taint/interproc-solver.ts`): the unit is `(function, parameter, source)`. Seed from `source→callee-arg`, propagate via `param→callee-arg`, fire a finding when a tainted param meets `param→sink`.

  • **Cycle-safe by monotonicity.** The tainted-set is monotone over a finite

lattice (`fn × param × source`), so the worklist converges — a recursive call just re-proposes an already-visited entry. SCC condensation would only refine processing order; correctness/termination don't require it.

  • **Source-discriminated state (load-bearing).** Key the state by the SOURCE

too. Keying only by `(fn, param)` collapses multi-source flows: a sink param tainted by source A is marked visited and a later flow from source B is dropped before firing — the recurring multi-source bug class. (Bit M3; bit M4 U9.)

  • **Name-based call join.** Match a summary's call-arg edge to a `CALLS` edge by

CALLEE NAME, not call-site line — line-base parity (CFG 1-based vs reference site) is fragile; the callee ide

Read more
Ships withgitnexus

The context engine for Enterprise Codebases Indexes any codebase into a knowledge graph — every dependency, call chain, cluster, and execution flow — then exposes it through smart MCP tools so AI agents never miss code.

Get the whole plugin

Other skills on gitnexus.