Skip to content
Data
Skill

/sigma_ingest

Extract durable ktx wiki knowledge from staged Sigma data model specs and workbook summaries. Load for WorkUnits with unitKey sigma-data-models or sigma-workbooks.

From plugin
ktx
1.6k17 skills
Install
$ npx -y skills add Kaelio/ktx --skill sigma_ingest --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/sigma_ingest

Context preview

The summary Claude sees to decide when to auto-load this skill.

Extract durable ktx wiki knowledge from staged Sigma data model specs and workbook summaries. Load for WorkUnits with unitKey sigma-data-models or sigma-workbooks.

SKILL.md

sigma_ingest.SKILL.md
name: sigma_ingest
description: Extract durable ktx wiki knowledge from staged Sigma data model specs and workbook summaries. Load for WorkUnits with unitKey sigma-data-models or sigma-workbooks.
callers: [memory_agent]

Sigma Ingest

Sigma ingest turns staged data model specs and workbook summaries into durable ktx wiki knowledge. The deterministic `project()` step has already written semantic-layer YAML for all warehouse-table data model elements before this skill runs — do not re-write those SL sources.

Work unit structure

Sigma produces at minimum two work units per ingest run:

  • `sigma-data-models` or `sigma-data-models-N`
  • `rawFiles`: `data-models/<id>.json` files (one per data model in this batch)
  • `peerFileIndex`: `workbooks/<id>.json` files + `sigma-manifest.json` + `sigma-projection-config.json`
  • When the workspace has more than 50 data models, split into batches: `sigma-data-models-0`, `sigma-data-models-1`, … with `displayLabel` like `"Sigma: data models (1/8)"`. When ≤50 data models, the unitKey is simply `sigma-data-models` with no suffix.
  • `sigma-workbooks` or `sigma-workbooks-N`
  • `rawFiles`: `workbooks/<id>.json` files (one per workbook in this batch)
  • `peerFileIndex`: `data-models/<id>.json` files + `sigma-manifest.json` + `sigma-projection-config.json`
  • When the workspace has more than 2000 workbooks, split into batches: `sigma-workbooks-0`, `sigma-workbooks-1`, … with `displayLabel` like `"Sigma: workbooks (1/4)"`. When ≤2000 workbooks, the unitKey is simply `sigma-workbooks` with no suffix.

`sigma-manifest.json` and `sigma-projection-config.json` are never in `rawFiles`. They live at the staged dir root and always appear in `peerFileIndex`.

Staged file shapes

**`data-models/<id>.json`** — one per data model (in `rawFiles` for data-model units):

{
  "sigmaId": "abc-123",
  "name": "Revenue Model",
  "path": "Finance/Revenue Model",
  "latestVersion": 3,
  "updatedAt": "2026-01-15T00:00:00Z",
  "isArchived": false,
  "spec": {
    "name": "Revenue Model",
    "pages": [{
      "id": "p1",
      "name": "Main",
      "elements": [{
        "id": "elem1",
        "kind": "table",
        "name": "Opportunities",
        "hidden": false,
        "source": {
          "kind": "warehouse-table",
          "connectionId": "<sigma-internal-uuid>",
          "path": ["DATABASE", "SCHEMA", "OPPORTUNITIES"]
        },
        "columns": [
          { "id": "c1", "name": "Deal Amount", "formula": "[OPPORTUNITIES/Amount]", "description": "Net contract value in USD" },
          { "id": "c2", "name": "Total ARR", "formula": "Sum([OPPORTUNITIES/ARR])", "description": "Annualised recurring revenue" }
        ]
      }]
    }]
  }
}

`source.kind` discriminates:

  • `warehouse-table` — element maps directly to a warehouse table. Has `connectionId` and `path` (array of path segments forming the fully-qualified table name). `project()` writes an SL source when `connectionMappings` covers this `connectionId`.
  • `table` — element is a derived view layered on top of another element; identified by `source.elementId`. No warehouse path. Wiki-only.

**`workbooks/<id>.json`** — one per workbook, in `rawFiles` for workbook units (summary only; no spec endpoint exists):

{
  "sigmaId": "wb-abc",
  "name": "ARR Tracker",
  "path": "Finance/Dashboards",
  "latestVersion": 2,
  "updatedAt": "2026-01-16T00:00:00Z",
  "isArchived": false,
  "workbookUrlId": "57a96EMo3G...",
  "description": "Tracks ARR by segment and cohort for the finance team"
}

**Peer files (available via `peerFileIndex`, not `rawFiles`):**

**`sigma-manifest.json`** — fetch summary; use for provenance only.

**`sigma-projection-config.json`** — written by `fetch()`, contains two fields the skill must read:

  • `connectionMappings`: `{sigmaInternalUuid: ktxWarehouseConnectionId}`. Use the mapped warehouse connection ID for `entity_details` when verifying warehouse identifiers found in data model specs.
  • `workbookFilter`: the filter settings that were active when workbooks were last fetched:
  • `includeArchived` (default `false`) — when `false`, archived workbooks are not in `workbooks/`; `isArchived: true` files will only appear when this was `true`.
  • `includeExplorations` (default `false`) — when `false`, exploration-type workbooks (unsaved analyses) are excluded; treat present workbooks as intentional, curated reports.
  • `updatedSince` (optional ISO 8601 string) — when set, only workbooks updated on or after this date are staged; the set is a recent-changes slice, not the full workspace. Do not infer that absent workbooks were deleted.

`sigma-manifest.json` also reflects any active `dataModelFilter`. When `dataModelFilter.updatedSince` was set during fetch, `dataModelCount` reflects only matching models, not the full workspace. Do not infer that absent data models were deleted.

Read `sigma-projection-config.json` first and keep `workbookFilter` in scope while processing the WorkUnit.

Required workflow

1. Read every `rawFiles` entry for the WorkUnit. 2. Read `sigma-projection-config.json` from the staged dir to get `connectionMappings`. 3. For each data model file: extract business semantics from element names, column descriptions, and the domain context of the model. Skip hidden elements and hidden columns. 4. For each workbook file: extract business domain knowledge from the name and description. When `workbookFilter.updatedSince` is set, treat the staged set as a recent-changes slice — absent workbooks were not deleted, they were simply outside the filter window. 5. Use `discover_data` before writing to find existing wiki pages on the same topic. 6. Write wiki candidates with `context_candidate_write`. Do not call `wiki_write` directly from a Sigma WorkUnit; Stage 4 reconciliation promotes candidates. 7. Do not write or edit SL sources. The `project()` step owns all SL output for Sigma.

Identifier Verification Protocol

Before writing a wiki page or

Read more
Ships withktx

ktx is an executable context layer for data and analytics agents 🐙 Allow Claude Code, Codex, or other AI agents to query analytical databases accurately and with full context of your company

Get the whole plugin
Stats
1,588
Stars
103
Forks
Active
Maintenance
TypeScript
Language
Apache-2.0
License
4d ago
Last commit
4mo ago
Created

Repo: Kaelio/ktx

Other skills on ktx.

analytics
Skill

analytics

Use when answering a question that needs data from a ktx-connected database - investigating, analyzing, "how many", "show me", "what's the breakdown of",…

@kaelio@kaelioView Skill
dbt_ingest
Skill

dbt_ingest

Map dbt `schema.yml` / `properties.yml` models and sources into ktx semantic-layer overlays and column notes. Covers `sources:` vs `models:`, column…

@kaelio@kaelioView Skill
ingest_triage
Skill

ingest_triage

Classify and resolve conflicts detected during bundle ingest (structural duplicates, definitional contradictions, near-duplicate clusters, re-ingest changes,…

@kaelio@kaelioView Skill