/write-pipeline
MUST use when creating or modifying a data pipeline, a set of scripts marked `pipeline` and wired together by `on` / `materialize` annotations.
$ npx -y skills add windmill-labs/windmill --skill write-pipeline --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/write-pipeline
Context preview
The summary Claude sees to decide when to auto-load this skill.
MUST use when creating or modifying a data pipeline, a set of scripts marked `pipeline` and wired together by `on` / `materialize` annotations.
SKILL.md
write-pipeline.SKILL.mdname: write-pipeline
description: MUST use when creating or modifying a data pipeline, a set of scripts marked `pipeline` and wired together by `on` / `materialize` annotations.
Data pipeline authoring
A **data pipeline** is NOT a flow. A flow is one runnable that orchestrates steps internally. A data pipeline is a set of **independent scripts**, each deployed on its own, that form a DAG by reading and writing shared **storage assets** (DuckLake tables, data tables, S3 objects, volumes, resources) and by declaring execution **triggers**. The pipeline is visualized and edited at `/pipeline/<folder>`; every node is a normal workspace script that happens to carry pipeline annotations. When the user asks for a "data pipeline" (or to "ingest / transform / materialize" data across steps), build pipeline-annotated scripts — do NOT build a flow.
Default to DuckDB + DuckLake
A pipeline node that produces a table should almost always be a **`duckdb`** node that materializes its output into a **DuckLake** table with `-- materialize ducklake://<name>/<table>` (in a DuckDB node the annotation uses SQL `--` comment syntax; write the body as a bare `SELECT` and let the runtime do the write). DuckLake is the default lakehouse store for pipelines and is the shape the pipeline editor is built around, so prefer it unless the work specifically calls for something else:
- `postgresql` / data tables — only for row-level, OLTP-style mutations against an existing Postgres data table (frequent single-row upserts/updates, transactional reads that an app queries live).
- `bun` / `python3` — only for non-tabular work that doesn't map to SQL: calling an external API, wrangling files, arbitrary glue. When such a node still produces tabular data for downstream steps, land it in DuckLake (write it with the wmill SDK / ducklake helpers) rather than inventing a parallel store.
Do not spread a pipeline across postgres, S3, and DuckLake when one DuckLake lake would do; a consistent DuckLake lakehouse is the goal.
Storage prerequisites
A DuckLake pipeline only runs once the workspace has **object storage** (S3 / Azure Blob / GCS) **and a DuckLake catalog** configured — DuckLake tables and `s3://` assets can't be materialized or read without it. Check which DuckLake catalogs the workspace has before you build. `wmill ducklake list` lists them. Drafting the annotated scripts does not require storage, but the pipeline can't ingest, materialize, or read its assets until it exists. So if there is none (or the user hits "storage not configured" errors), say so and give the right next step **by role**:
- a workspace **admin** sets it up in Workspace settings → Object Storage (add an S3/Azure/GCS storage), then adds a DuckLake catalog on top of it;
- anyone **without admin rights** should ask a workspace admin to configure object storage + a DuckLake catalog.
Never hand back a DuckLake pipeline that cannot run without flagging the missing storage and pointing to who sets it up.
What makes a script a pipeline node
A script joins the pipeline when its source begins with the `pipeline` annotation as a top-of-file comment, **written in the script's own comment syntax** — `//` for TS/JS (bun), `--` for SQL (DuckDB/Postgres), `#` for Python/Bash. So it's `-- pipeline` in a DuckDB node, `# pipeline` in a Python node, `// pipeline` in a bun node. Every annotation below uses that same prefix (the `//` shown is the TS form). All other wiring is expressed as annotation comments near the top of the file:
- `// on <ref>` — declares an execution-DAG **input** (what triggers/feeds this node). `<ref>` is either:
- an **asset URI** (the node runs when that asset is produced upstream): `ducklake://main/orders`, `datatable://main/users`, `$res:f/folder/my_resource`, `volume://name/path`, or an S3 object (see the S3 storage-form rule below).
- a **native trigger kind**: `schedule`, `webhook`, `email`, `kafka`, `mqtt`, `amqp`, `nats`, `postgres`, `sqs`, `gcp`, or `data_upload` (a user-uploaded S3 file). For these the actual trigger row (cron, topic, …) is created separately; the annotation only declares the binding. **`data_upload` is special**: there is no trigger row — the node instead declares an **`S3Object` input parameter** fed by the auto-generated upload picker; it never hard-codes a key. Any language can be the `data_upload` node:
- Python (has `import wmill`): `def main(file: wmill.S3Object):` then `wmill.load_s3_file(file)`; TS (has `import * as wmill from "windmill-client"`): `export async function main(file: wmill.S3Object)`. Qualify the type as `wmill.S3Object` (or add `from wmill import S3Object` / `import { S3Object } from "windmill-client"`) — a bare `S3Object` is undefined.
- **DuckDB** takes the s3object arg via a `-- $<name> (s3object)` declaration and reads it directly, so a single DuckDB node can ingest **and** materialize:
-- pipeline
-- on data_upload
-- materialize ducklake://main/raw_uploads
-- $file (s3object)
SELECT * FROM read_csv($file)- **Outputs** are inferred from what the body writes — a `CREATE TABLE`, a `wmill.writeS3File(...)`, a DuckLake/datatable write. To declare a managed output explicitly, use `// materialize <asset-uri>`.
- Optional badges: `// partitioned <daily|hourly|weekly|monthly|dynamic>`, `// freshness <duration>` (e.g. `1h`), `// tag <worker-tag>`, `// retry <count> [delay]`, `// data_test <kind> ...` (managed DuckLake targets only — deploy rejects it beside a `dbt://` one), `// measure <name> = <agg> [where <pred>]` and `// dimension <name> = <expr>` (see "Declared metrics" below).
S3 object wiring (storage form matters)
An `s3://` URI's first slashes select the **storage**, not part of the key — get this wrong and the producer/consumer edge silently won't connect:
- `s3:///<key>` (**triple** slash, empty first segment) = the **default** workspace storage. A downstream node reading or triggering on tha
Read more
name: write-pipeline description: MUST use when creating or modifying a data pipeline, a set of scripts marked `pipeline` and wired together by `on` / `materialize` annotations.
Data pipeline authoring
A **data pipeline** is NOT a flow. A flow is one runnable that orchestrates steps internally. A data pipeline is a set of **independent scripts**, each deployed on its own, that form a DAG by reading and writing shared **storage assets** (DuckLake tables, data tables, S3 objects, volumes, resources) and by declaring execution **triggers**. The pipeline is visualized and edited at `/pipeline/<folder>`; every node is a normal workspace script that happens to carry pipeline annotations. When the user asks for a "data pipeline" (or to "ingest / transform / materialize" data across steps), build pipeline-annotated scripts — do NOT build a flow.
Default to DuckDB + DuckLake
A pipeline node that produces a table should almost always be a **`duckdb`** node that materializes its output into a **DuckLake** table with `-- materialize ducklake://<name>/<table>` (in a DuckDB node the annotation uses SQL `--` comment syntax; write the body as a bare `SELECT` and let the runtime do the write). DuckLake is the default lakehouse store for pipelines and is the shape the pipeline editor is built around, so prefer it unless the work specifically calls for something else:
- `postgresql` / data tables — only for row-level, OLTP-style mutations against an existing Postgres data table (frequent single-row upserts/updates, transactional reads that an app queries live).
- `bun` / `python3` — only for non-tabular work that doesn't map to SQL: calling an external API, wrangling files, arbitrary glue. When such a node still produces tabular data for downstream steps, land it in DuckLake (write it with the wmill SDK / ducklake helpers) rather than inventing a parallel store.
Do not spread a pipeline across postgres, S3, and DuckLake when one DuckLake lake would do; a consistent DuckLake lakehouse is the goal.
Storage prerequisites
A DuckLake pipeline only runs once the workspace has **object storage** (S3 / Azure Blob / GCS) **and a DuckLake catalog** configured — DuckLake tables and `s3://` assets can't be materialized or read without it. Check which DuckLake catalogs the workspace has before you build. `wmill ducklake list` lists them. Drafting the annotated scripts does not require storage, but the pipeline can't ingest, materialize, or read its assets until it exists. So if there is none (or the user hits "storage not configured" errors), say so and give the right next step **by role**:
- a workspace **admin** sets it up in Workspace settings → Object Storage (add an S3/Azure/GCS storage), then adds a DuckLake catalog on top of it;
- anyone **without admin rights** should ask a workspace admin to configure object storage + a DuckLake catalog.
Never hand back a DuckLake pipeline that cannot run without flagging the missing storage and pointing to who sets it up.
What makes a script a pipeline node
A script joins the pipeline when its source begins with the `pipeline` annotation as a top-of-file comment, **written in the script's own comment syntax** — `//` for TS/JS (bun), `--` for SQL (DuckDB/Postgres), `#` for Python/Bash. So it's `-- pipeline` in a DuckDB node, `# pipeline` in a Python node, `// pipeline` in a bun node. Every annotation below uses that same prefix (the `//` shown is the TS form). All other wiring is expressed as annotation comments near the top of the file:
- `// on <ref>` — declares an execution-DAG **input** (what triggers/feeds this node). `<ref>` is either:
- an **asset URI** (the node runs when that asset is produced upstream): `ducklake://main/orders`, `datatable://main/users`, `$res:f/folder/my_resource`, `volume://name/path`, or an S3 object (see the S3 storage-form rule below).
- a **native trigger kind**: `schedule`, `webhook`, `email`, `kafka`, `mqtt`, `amqp`, `nats`, `postgres`, `sqs`, `gcp`, or `data_upload` (a user-uploaded S3 file). For these the actual trigger row (cron, topic, …) is created separately; the annotation only declares the binding. **`data_upload` is special**: there is no trigger row — the node instead declares an **`S3Object` input parameter** fed by the auto-generated upload picker; it never hard-codes a key. Any language can be the `data_upload` node:
- Python (has `import wmill`): `def main(file: wmill.S3Object):` then `wmill.load_s3_file(file)`; TS (has `import * as wmill from "windmill-client"`): `export async function main(file: wmill.S3Object)`. Qualify the type as `wmill.S3Object` (or add `from wmill import S3Object` / `import { S3Object } from "windmill-client"`) — a bare `S3Object` is undefined.
- **DuckDB** takes the s3object arg via a `-- $<name> (s3object)` declaration and reads it directly, so a single DuckDB node can ingest **and** materialize:
-- pipeline
-- on data_upload
-- materialize ducklake://main/raw_uploads
-- $file (s3object)
SELECT * FROM read_csv($file)- **Outputs** are inferred from what the body writes — a `CREATE TABLE`, a `wmill.writeS3File(...)`, a DuckLake/datatable write. To declare a managed output explicitly, use `// materialize <asset-uri>`.
- Optional badges: `// partitioned <daily|hourly|weekly|monthly|dynamic>`, `// freshness <duration>` (e.g. `1h`), `// tag <worker-tag>`, `// retry <count> [delay]`, `// data_test <kind> ...` (managed DuckLake targets only — deploy rejects it beside a `dbt://` one), `// measure <name> = <agg> [where <pred>]` and `// dimension <name> = <expr>` (see "Declared metrics" below).
S3 object wiring (storage form matters)
An `s3://` URI's first slashes select the **storage**, not part of the key — get this wrong and the producer/consumer edge silently won't connect:
- `s3:///<key>` (**triple** slash, empty first segment) = the **default** workspace storage. A downstream node reading or triggering on tha
Windmill is fully open-sourced (AGPLv3) and Windmill Labs offers dedicated instances and commercial support and licenses.
Repo: windmill-labs/windmill

