Skip to content
Agent Orchestration
Skill

/omh-data-pipelines

[omh] Data pipeline work -- an ETL or streaming job, a backfill or replay, duplicate events, a schema change downstream, a lineage question, a data-quality regression: make every rerun idempotent, bound every replay, and gate each load on observed checks. Use when the user says:

BOOST
From plugin
oh-my-hermes
3.3k145 skills
Install
$ npx -y skills add rlaope/oh-my-hermes --skill omh-data-pipelines --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/omh-data-pipelines

Context preview

The summary Claude sees to decide when to auto-load this skill.

[omh] Data pipeline work -- an ETL or streaming job, a backfill or replay, duplicate events, a schema change downstream, a lineage question, a data-quality regression: make every rerun idempotent, bound every replay, and gate each load on observed checks. Use when the user says:

SKILL.md

omh-data-pipelines.SKILL.md
name: "omh-data-pipelines"
description: "[omh] Data pipeline work -- an ETL or streaming job, a backfill or replay, duplicate events, a schema change downstream, a lineage question, a data-quality regression: make every rerun idempotent, bound every replay, and gate each load on observed checks. Use when the user says: data-pipelines, data pipeline, data pipelines, etl, elt, etl pipeline, etl job, etl backfill."
metadata:
  hermes:
    tags: [workflow, oh-my-hermes, planning]
    category: planning
    phase: data-pipelines
    role: planner
    quality_tier: idempotent-replay-gated

Data Pipelines

This is a Hermes-native `data-pipelines` workflow skill.

Why This Exists

`data-pipelines` exists because pipeline work had no owner: `backend` owns a service's schema migration, `data-analysis` analyzes data it is handed, and `relational-db` owns a database's locks and indexes, while a backfill that duplicated events or a schema change with unknown readers reached memory and event lanes with no idempotency contract or replay bound at all.

First Steps

  • Ask what makes a row unique at the sink before planning any rerun.
  • Bound the window and the targets before ordering any replay or backfill step.

Do Not Use When

  • The ask is a service's own database migration, API, or queue design; use `backend`.
  • The ask is analyzing, charting, or summarizing a dataset that was handed over; use `data-analysis`.
  • The ask is a slow query, an index, or DDL locking a live table; use `relational-db`.
  • The ask is remembering or syncing what the assistant knows about the user; use `memory-sync`.

Examples

Good example:

  • Prompt: our airflow etl backfill is producing duplicate events
  • Expected behavior: Find the sink's unique key, name the append that duplicated rows, write the idempotency contract (event-id dedupe or partition overwrite), then bound the backfill window and gate it on key uniqueness and row count against the prior window.
  • Why: Rerunning an appending backfill doubles the duplicates it was meant to fix.

Bad example:

  • Prompt: just delete the duplicates and rerun the whole history
  • Expected behavior: Refuse the unbounded rerun: fix the write to be idempotent first, then backfill a bounded window behind a quality gate.
  • Why: Deleting duplicates without fixing the write guarantees the next rerun duplicates again.

Completion Checklist

  • The sink's unique key and the idempotency contract are stated.
  • Every replay or backfill is bounded by window and target.
  • Every downstream reader of a schema change is named with its impact.
  • Every load names its data-quality gate and the value that stops it.
  • OMH ran nothing, and every count cites observed output or is marked unverified.

Recovery Notes

  • If no unique key exists at the sink, the first step is defining one; say so before any rerun.
  • If lineage is unavailable, list readers found by search and mark the map incomplete.

Workflow Lane

  • Current lane: **Coding handoff** (`idea-to-deploy`, `llm-app-dev`, `cto-loop`, `deploy-and-monitor`, `code-review`, `build-failure-triage`, `verification-gate`, `security-safety-review`, `+28 more`) - coding owners, handoffs, review, CI, and merge evidence.
  • If intent belongs to another lane, hand back to `oh-my-hermes` or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: `omh-routing/references/skill-common-rail.md`.

Use When

Use when a batch or streaming data pipeline needs planning or repair: an ETL, ELT, Airflow, dbt, Spark or Kafka job; a backfill or a replay of past events; duplicate or missing rows; a schema change whose downstream readers are unknown; a lineage question; or a data-quality regression. The output is the lineage, the schema change's downstream impact, an idempotency contract, a bounded replay or backfill plan, and the data-quality gate each load must pass; OMH runs no job and reads no warehouse.

Strong routing signals: `data-pipelines`, `data pipeline`, `data pipelines`, `etl`, `elt`, `etl pipeline`, `etl job`, `etl backfill`, `airflow dag`, `airflow etl`, `airflow backfill`, `dagster`, `dbt model`, `dbt run`, `spark job`, `kafka topic`, `kafka events`, `kafka consumer`, `backfill`, `data backfill`, `replay events`, `replay the events`, `event replay`, `idempotent`, `idempotency`, `exactly once`, `exactly-once`, `duplicate events`, `lineage`, `data lineage`, `data quality`, `data quality check`, `schema evolution`, `late arriving data`, `dead letter queue`, `batch job`

Catalog Metadata

Category: `planning` Phase: `data-pipelines` Hermes role: `planner` Quality tier: `idempotent-replay-gated` Reasoning demand: `standard`

Quality bar:

  • Find what makes a row unique at the sink before proposing any rerun.
  • Load `references/pipeline-method.md` for the idempotency patterns, the schema compatibility table, the replay and backfill procedure, and the quality checks instead of recalling them.
  • Map lineage from the orchestrator's graph first and mark anything found only by search.
  • Treat duplicates as an idempotency defect, not a cleanup task: fix the write, then repair the rows.
  • Keep prepared, run, and verified as separate states for every load and check.

Handoff policy:

Keep the lineage map, schema impact, idempotency contract, replay or backfill plan, and quality gates in Hermes. Row counts, job runs, query results and check outcomes are recorded only from executor, operator, or wrapper observed output; OMH never runs a pipeline, triggers a backfill, or queries a warehouse.

Required inputs:

  • the pipeline: its orchestrator, its sources, its sinks, and its schedule or trigger
  • the unit of the problem: the table, topic, or model, and the time window affected
  • what makes a row unique at the sink: the natural key, the event id, or the partition
  • the downstream readers already known: models, dashboards, exports, services
  • observed counts, job logs, or check results for any claim about what
Read more
Ships withoh-my-hermes

English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.

Get the whole plugin
Stats
3,264
Stars
244
Forks
Active
Maintenance
Python
Language
MIT
License
10h ago
Last commit
4mo ago
Created
4d ago
Added

Repo: rlaope/oh-my-hermes

Other skills on oh-my-hermes.