databricks-agent-brick…
Create Agent Bricks: Knowledge Assistants (KA) for document Q&A and Supervisor Agents for…
Develop Lakeflow Spark Declarative Pipelines (formerly Delta Live Tables) on Databricks. Use when building batch or streaming data pipelines with Python or SQL. Invoke BEFORE starting implementation.
$ npx -y skills add databricks/databricks-agent-skills --skill databricks-pipelines --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/databricks-pipelinesContext preview
The summary Claude sees to decide when to auto-load this skill.
Develop Lakeflow Spark Declarative Pipelines (formerly Delta Live Tables) on Databricks. Use when building batch or streaming data pipelines with Python or SQL. Invoke BEFORE starting implementation.
name: databricks-pipelines description: Develop Lakeflow Spark Declarative Pipelines (formerly Delta Live Tables) on Databricks. Use when building batch or streaming data pipelines with Python or SQL. Invoke BEFORE starting implementation. compatibility: Requires databricks CLI (>= v1.0.0) metadata: version: "0.3.0" parent: databricks-core
**FIRST**: Use the parent `databricks-core` skill for CLI basics, authentication, profile selection, and data discovery commands.
Use this tree to determine which dataset type and features to use. Multiple features can apply to the same dataset — e.g., a Streaming Table can use Auto Loader for ingestion, Append Flows for fan-in, and Expectations for data quality. Choose the dataset type first, then layer on applicable features.
User request → What kind of output? ├── Intermediate/reusable logic (not persisted) → Temporary View │ ├── Preprocessing/filtering before Auto CDC → Temporary View feeding CDC flow │ ├── Shared intermediate streaming logic reused by multiple downstream tables │ ├── Pipeline-private helper logic (not published to catalog) │ └── Published to UC for external queries → Persistent View (SQL only) ├── Persisted dataset │ ├── Source is streaming/incremental/continuously growing → Streaming Table │ │ ├── File ingestion (cloud storage, Volumes) → Auto Loader │ │ ├── Message bus (Kafka, Kinesis, Pub/Sub, Pulsar, Event Hubs) → streaming source read │ │ ├── Existing streaming/Delta table → streaming read from table │ │ ├── CDC / upserts / track changes / keep latest per key / SCD Type 1 or 2 → Auto CDC │ │ ├── Multiple sources into one table → Append Flows (NOT union) │ │ ├── Historical backfill + live stream → one-time Append Flow + regular flow │ │ └── Windowed aggregation with watermark → stateful streaming │ └── Source is batch/historical/full scan → Materialized View │ ├── Aggregation/join across full dataset (GROUP BY, SUM, COUNT, etc.) │ ├── Gold layer aggregation from streaming table → MV with batch read (spark.read / no STREAM) │ ├── JDBC/Federation/external batch sources │ └── Small static file load (reference data, no streaming read) ├── Output to external system (Python only) → Sink │ ├── Existing external table not managed by this pipeline → Sink with format="delta" │ │ (prefer fully-qualified dataset names if the pipeline should own the table — see Publishing Modes) │ ├── Kafka / Event Hubs → Sink with format="kafka" + @dp.append_flow(target="sink_name") │ ├── Custom destination not natively supported → Sink with custom format │ ├── Custom merge/upsert logic per batch → ForEachBatch Sink (Public Preview) │ └── Multiple destinations per batch → ForEachBatch Sink (Public Preview) └── Data quality constraints → Expectations (on any dataset type)
Error → cause/fix mappings agents hit constantly. For DAB-bundle vs CLI-iteration deploy issues, see the workflow-specific reference files.
| Error / symptom | Cause / fix | |-----------------|-------------| | Rejection of `CREATE OR REPLACE STREAMING TABLE` / `MATERIALIZED VIEW` | `CREATE OR REPLACE` is standard SQL, NOT SDP. Use `CREATE OR REFRESH STREAMING TABLE` / `CREATE OR REFRESH MATERIALIZED VIEW`. | | CLI errors on `databricks fs ls /Volumes/...` | The `dbfs:` prefix is required even for UC Volume paths: `databricks fs ls dbfs:/Volumes/<catalog>/<schema>/<volume>/<path>`. | | `DELTA_CLUSTERING_COLUMNS_DATATYPE_NOT_SUPPORTED` at first write | A `CLUSTER BY` column is BOOLEAN / ARRA
Build on Databricks with AI coding agents such as Claude Code, Cursor, Codex, and GitHub Copilot. This repository provides the skills and agent plugins for Databricks AI Tools.
Repo: databricks/databricks-agent-skills
Create Agent Bricks: Knowledge Assistants (KA) for document Q&A and Supervisor Agents for…
Use Databricks built-in AI Functions (ai_classify, ai_extract, ai_summarize, ai_mask,…
Create Databricks AI/BI dashboards. Must use when creating, updating, or deploying Lakeview…
Design the UX of custom-code Databricks Apps (AppKit/React) data screens — KPI/overview…
Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio,…
Build apps on Databricks Apps platform. Use when asked to create data apps, analytics tools,…