databricks-agent-brick…
Create Agent Bricks: Knowledge Assistants (KA) for document Q&A and Supervisor Agents for…
Generate realistic synthetic data using Spark + Faker (strongly recommended). Supports serverless execution, multiple output formats (Parquet/JSON/CSV/Delta), and scales from thousands to millions of rows. For small datasets (<10K rows), can optionally generate locally and
$ npx -y skills add databricks/databricks-agent-skills --skill databricks-synthetic-data-gen --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/databricks-synthetic-data-genContext preview
The summary Claude sees to decide when to auto-load this skill.
Generate realistic synthetic data using Spark + Faker (strongly recommended). Supports serverless execution, multiple output formats (Parquet/JSON/CSV/Delta), and scales from thousands to millions of rows. For small datasets (<10K rows), can optionally generate locally and
name: databricks-synthetic-data-gen description: "Generate realistic synthetic data using Spark + Faker (strongly recommended). Supports serverless execution, multiple output formats (Parquet/JSON/CSV/Delta), and scales from thousands to millions of rows. For small datasets (<10K rows), can optionally generate locally and upload to volumes. Use when user mentions 'synthetic data', 'test data', 'generate data', 'demo dataset', 'Faker', or 'sample data'." compatibility: Requires databricks CLI (>= v1.0.0) metadata: version: "0.1.0" parent: databricks-core
> Catalog and schema are **always user-supplied** — never default to any value. If the user hasn't provided them, ask. For any UC write, **always create the schema if it doesn't exist** before writing data.
Generate realistic, story-driven synthetic data for Databricks using **Spark + Faker + Pandas UDFs** (strongly recommended).
Synthetic data should demonstrate how Databricks helps solve real business problems.
**The pattern:** Something goes wrong → business impact ($) → analyze root cause → identify affected customers → fix and prevent.
**Key principles:**
**Why no flat distributions:** Uniform data has no story — no spikes, no anomalies, no cohort, no 20/80, no skew, nothing to investigate. It can't show Databricks' value for root cause analysis.
| When | Guide | |------|-------| | User mentions **ML model training** or complex time patterns | [references/1-data-patterns.md](references/1-data-patterns.md) — ML-ready data, time multipliers, row coherence | | Errors during generation | [references/2-troubleshooting.md](references/2-troubleshooting.md) — Fixing common issues |
1. **Data tells a story** — Something goes wrong, impacts $, can be analyzed and fixed. Show Databricks value. 2. **All data serves the story** — Every table and column must be coherent and usable in dashboards or ML models. No orphan data, no random noise — if it doesn't help explain or plot a futur dashboard or predict, don't generate it. 3. **Industry terms, simple schema** — Use domain-specific vocabulary but keep it easy to understand (few tables, clear relationships) 4. **Never uniform distributions** — Skewed categories, log-normal amounts, 80/20 patterns. Flat = no story = useless 5. **Enough data for trends** — ~100K+ rows for main tables so patterns survive aggregation 6. **Ask for catalog/schema** — Never default, always confirm before generating 7. **Present plan for approval** — Show tables, distributions, assumptions before writing code 8. **Master tables first** — Generate parent tables, write to Delta, then create children with valid FKs 9. **Use Spark + Faker + Pandas UDFs** — Scalable, parallel. Polars only if user explicitly wants local + <30K rows 10. **Use Databricks Connect Serverless by default to generate data** — Update databricks-connect on python 3.12 if required (avoid using execute_code unless instructed to not use Databricks Connect) 11. **No `.cache()` or `.persist()`** — Not supported on serverless. Write to Delta, read back for joins 12. **No Python loops or `.collect()`** — Use Spark parallelism. No driver-side iteration, avoid Pandas↔Spark conversions
**Before generating any code, you MUST present a plan for user approval.**
**You MUST explicitly ask the user which catalog to use.** Do not assume or proceed without confirmation.
Example prompt to user: > "Which Unity Catalog should I use for this data?"
When presenting your plan, always show the selected catalog prominently:
📍 Output Location: catalog_name.schema_name Volume: /Volumes/catalog_name/schema_name/raw_data/
This makes it easy for the user to spot and correct if needed.
Ask the user about:
**If user doesn't specify a story:** Propose one. Don't generate bland data — suggest an incident, anomaly, or trend that shows Databricks value (e.g., "I'll include a system outage that causes ticket spike and churn — this lets you demo root cause analysis").
Show a clear specification with **the business story and your assumptions surfaced**:
📍 Output Location: {user_catalog}.support_demo
Volume: /Volumes/{user_catalog}/support_demo/raw_data/
📖 Story: A payment system outage causes support ticket spike. Resolution times
degrade, enterprise customers churn, revenue drops $2.3M. With Databricks we
identify the root cause, affected customers, and prevent future impact.| Table | Description | Rows | Key Assumptions | |-------|-------------|------|-----------------| | customers | Customer profiles with tier, MRR | 10,000 | Enterprise 10% but 60% of revenue | | tickets | Support tickets with priority, resolution_time | 80,000 | Spike during outage, SLA breaches | | incidents | System events (outages, deployments) | 50 | Payment outage mid-month | | churn_events | Customer cancellations with reason | 500 |
Build on Databricks with AI coding agents such as Claude Code, Cursor, Codex, and GitHub Copilot. This repository provides the skills and agent plugins for Databricks AI Tools.
Repo: databricks/databricks-agent-skills
Create Agent Bricks: Knowledge Assistants (KA) for document Q&A and Supervisor Agents for…
Use Databricks built-in AI Functions (ai_classify, ai_extract, ai_summarize, ai_mask,…
Create Databricks AI/BI dashboards. Must use when creating, updating, or deploying Lakeview…
Design the UX of custom-code Databricks Apps (AppKit/React) data screens — KPI/overview…
Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio,…
Build apps on Databricks Apps platform. Use when asked to create data apps, analytics tools,…