brain-ingest-gate
Pre-write quality gate for content entering the brain. No raw copies: a bare cp/mv into the brain repo is a bug. Before any new page lands, resolve named…
Deduplicate and synthesize raw concept stubs into a tiered intellectual map (T1 Canon to T4 Riff), tracing idea evolution across sources over time. Transforms thousands of raw concept pages into a curated intellectual fingerprint. Includes a reversible curation cull pass (Phase
$ npx -y skills add garrytan/gbrain --skill concept-synthesis --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/concept-synthesisContext preview
The summary Claude sees to decide when to auto-load this skill.
Deduplicate and synthesize raw concept stubs into a tiered intellectual map (T1 Canon to T4 Riff), tracing idea evolution across sources over time. Transforms thousands of raw concept pages into a curated intellectual fingerprint. Includes a reversible curation cull pass (Phase
name: concept-synthesis version: 0.2.0 description: Deduplicate and synthesize raw concept stubs into a tiered intellectual map (T1 Canon to T4 Riff), tracing idea evolution across sources over time. Transforms thousands of raw concept pages into a curated intellectual fingerprint. Includes a reversible curation cull pass (Phase 5) with hard keep/delete/merge verdicts, substance gates, grounding labels, cluster budgets, and merge-with-backlinks salience promotion. triggers: - "concept synthesis" - "synthesize my concepts" - "find patterns across my notes" - "build my intellectual map" - "trace idea evolution" - "canon vs riff" - "cull my concepts" - "which concepts to keep" - "concept quality rubric" mutating: true writes_pages: true writes_to: - concepts/
> **Convention:** see [conventions/quality.md](../conventions/quality.md) for > back-link enforcement and quote-fidelity requirements. > > **Convention:** see [_brain-filing-rules.md](../_brain-filing-rules.md) — > output files under `concepts/` per the primary-subject rule.
Many ingestion pipelines (signal-detector, idea-ingest, voice-note-ingest) create a concept page for every idea mentioned. Over months this produces:
This skill transforms that raw material into a curated intellectual map.
Phase 1: Dedup + merge (deterministic)
N stubs → ~N/4 canonical concepts
├── Jaccard dedup (word-overlap on titles + first-paragraph)
├── Substring dedup ("founder mode" vs "founder mode vs manager mode")
├── Semantic dedup (LLM: "are these the same idea?")
└── Merge timelines + aliases from duplicates into the canonical page
Phase 2: Score + tier (deterministic + heuristic)
Each canonical concept → scored and tiered
├── Frequency: distinct sources referencing this concept
├── Timespan: first mention → last mention in days
├── Breadth: distinct months it appears in
├── Engagement: avg engagement on concept-bearing sources (if available)
└── Tier: T1 Canon | T2 Developing | T3 Speculative | T4 Riff
Phase 3: Synthesize (LLM, T1+T2 only)
T1 + T2 concepts → rich synthesis
├── Evolution narrative: how the idea sharpened over time
├── Best articulation: highest-engagement or most precise quote
├── Related concepts: cross-links to other concepts
├── Context: what was happening when this idea emerged / evolved
└── Counter-positions: what this idea argues against
Phase 4: Cluster + map (LLM)
All tiered concepts → intellectual clusters
├── Group related concepts into domains (auto-named via LLM)
├── Generate cluster summary pages
├── Build a master concepts/README.md with the full map
└── Identify idea genealogies (concept A → evolved into concept B)
Phase 5: Curation cull (rubric + reversible merge)
Each concept → hard verdict: ELITE | KEEP | MERGE/REWRITE | DELETE
├── 6-axis rubric (substance 2x, packaging 1x) + minimum substance gate
├── Grounding labels (VERIFIED / OPINION / NEEDS_SOURCE / UNSAFE)
├── Cluster budgets + reputational-risk gate
├── Merge-with-backlinks into cluster canonicals (fully reversible)
└── merge_count / independent_sources → emergent tier promotionThe skill is markdown agent instructions. The agent uses gbrain's existing operations + LLM passes:
# 1. List all concept pages gbrain query "type:concept" --limit 10000 --json # 2. Phase 1 dedup — agent applies Jaccard + substring locally, # then LLM passes to identify semantic duplicates. # 3. Phase 2 tier — agent scores each canonical concept based on # frequency / timespan / breadth and writes tier into frontmatter. # 4. Phase 3 synthesis — for each T1/T2, agent reads the timeline # + associated source pages and writes a synthesis section # onto the concept page via put_page. # 5. Phase 4 clustering — agent reads the tiered concept list # and writes concepts/README.md with the full intellectual map.
--- title: "concept name" type: concept tier: 1 tier_label: "Canon" mention_count: 18 distinct_months: 8 first_mention: "YYYY-MM-DD" last_mention: "YYYY-MM-DD" composite_score: 78.4 aliases: ["alternate phrasing 1", "alternate phrasing 2"] related: ["sibling-concept-1", "sibling-concept-2"] --- # concept name **Tier 1 — Canon** | 18 mentions across 8 months ## Synthesis [2-4 paragraph narrative tracing how the idea evolved, what it means in the user's worldview, why it matters. Third-person analytical voice.] ## Best Articulation > "Verbatim quote from a source — the most precise or highest-engagement > expression of this idea." — [Date](source-url) ## Evolution | Period | Expression | Signal | |--------|-----------|--------| | YYYY-MM | "First articulation" | First use — aspiration frame | | YYYY-MM | "Sharpening" | Anti-pattern emerges | | YYYY-MM | "Peak form" | Cleanest expression | ## Related Concepts - [sibling concept](sibling-concept.md) — relationship description - [sibling concept](sibling-concept.md) — relationship description ## Timeline [Full timeline with deduped entries, quotes, source links]
--- title: "concept name" type: concept tier: 4 tier_label: "Riff" mention_count: 1 --- # concept name **Tier 4 — Riff** | 1 mention > "Quote from the source" — [Date](URL)
# Intellectual Universe ## Canon (T1) — N concepts The permanent intellectual fingerprint. Ideas that recur across years.
Give the agent you already use a memory you control. GBrain stores explicit facts with their sources, supports corrections and withdrawal, and makes the same memory available across your agents.
Repo: garrytan/gbrain
Pre-write quality gate for content entering the brain. No raw copies: a bare cp/mv into the brain repo is a bug. Before any new page lands, resolve named…
When you report a brain page to the user — created, edited, committed, or relayed from a subagent — a working link is part of the deliverable, in the SAME…
Brain knowledge base operations. The core read/write cycle: brain-first lookup, read-enrich-write loop, source attribution, ambient enrichment, back-linking.…
Token-hygiene audit of the always-loaded context stack — CLAUDE.md, AGENTS.md, auto-memory MEMORY.md, and the bootstrap-rendered identity files (SOUL.md,…
When the user corrects a factual error, root-cause it immediately. Don't just note the correction — trace the error to its source, fix the source, and prevent…
Confirmation gate before any bulk delete, cleanup, or destructive operation that could result in data loss — shell-level (rm -rf, git rm, bulk sed) or…