Skip to content
Data
Skill

/databricks-serverless

Serverless compute for Databricks jobs, Lakeflow pipelines and Declarative Automation Bundles (DABs), with STANDARD or PERFORMANCE_OPTIMIZED and classic only where serverless cannot run the workload. Use when creating, deploying, scheduling or editing a job or pipeline, or

BOOST
From plugin
databricks-agent-skills
34534 skills4 commands3 hooks
Install
$ npx -y skills add databricks/databricks-agent-skills --skill databricks-serverless --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/databricks-serverless

Context preview

The summary Claude sees to decide when to auto-load this skill.

Serverless compute for Databricks jobs, Lakeflow pipelines and Declarative Automation Bundles (DABs), with STANDARD or PERFORMANCE_OPTIMIZED and classic only where serverless cannot run the workload. Use when creating, deploying, scheduling or editing a job or pipeline, or

SKILL.md

databricks-serverless.SKILL.md
name: databricks-serverless
description: "Serverless compute for Databricks jobs, Lakeflow pipelines and Declarative Automation Bundles (DABs), with STANDARD or PERFORMANCE_OPTIMIZED and classic only where serverless cannot run the workload. Use when creating, deploying, scheduling or editing a job or pipeline, or deciding its compute. Invoke BEFORE writing a job spec. For migrating existing classic workloads, use databricks-serverless-migration."
compatibility: Requires databricks CLI (>= v0.292.0)
metadata:
  version: "0.3.2"
parent: databricks-core

Deploying Databricks jobs and pipelines on serverless

**FIRST**: Use the parent `databricks-core` skill for CLI basics, authentication and profile selection.

This skill decides compute. Where another skill or example (for example `databricks-jobs`) shows cluster configuration, follow this skill instead.

Serverless is the default compute for every job, pipeline and bundle you create. If the user explicitly asks for a cluster, do what they ask and mention the serverless option in one line. When editing an existing job, keep its compute unless the user asks to change it; tasks you add follow this skill.

1. Serverless or classic

Use classic compute only for:

  • R code, or Scala in notebook cells (`%scala`): serverless notebooks run Python and SQL only;
  • custom Docker images;
  • OS packages (`apt-get`) or native libraries with no pip equivalent;
  • custom Spark data source JARs;
  • a stated hard requirement for an instance type or Databricks Runtime version;
  • an explicit user request;
  • existing code over the budget in section 5;
  • a workspace without serverless: if creating or running the job fails because serverless is not

available there, put the whole job on classic and say so.

Then give only the blocked task a cluster (one small `job_clusters` entry with the latest LTS runtime and `autoscale: {min_workers: 1, max_workers: 4}`), keep every other task serverless, and state the blocker in one line. For a blocked pipeline, give the pipeline a `clusters` block instead of `serverless: true`.

Not reasons for classic: Scala or Java in a JAR task (JAR tasks run on serverless), GPUs (serverless GPU), large data, long runtimes, Kafka, cost (serverless STANDARD is on par with or cheaper than on-demand classic in most cases), or habit.

2. Put every task on serverless

  • **New jobs:** leave out `new_cluster`, `job_clusters`, `job_cluster_key`, `existing_cluster_id`,

`instance_pool_id`, `node_type_id`, `num_workers`, `autoscale`, `spark_version`. A task without a cluster runs on serverless. In an existing job, keep the cluster settings it has.

  • **Notebook tasks** need nothing else. **Python script, wheel and JAR tasks** take their libraries

from a job-level environment. Use the latest environment version the workspace offers, as a quoted string:

  environments:
    - environment_key: default
      spec:
        environment_version: "4"                # example; use the latest
        dependencies: ["requests==2.32.3"]   # pin versions; JARs go in java_dependencies
  tasks:
    - task_key: main
      spark_python_task: {python_file: /Workspace/.../main.py}
      environment_key: default
  • **SQL tasks:** a serverless SQL warehouse. **Pipelines:** `serverless: true`, no `clusters` block.
  • **Streaming:** `.trigger(availableNow=True)`, the trigger serverless supports (no `processingTime`

or `continuous` triggers). For always-on processing, run that stream in a continuous job (`continuous: {pause_status: UNPAUSED}`), which starts the next run as soon as one finishes, or use a continuous pipeline.

  • **Tag every job you create** for resource tracking, as other Databricks skills do: job-level

`tags: {"aidevkit_project": "ai-dev-kit"}`. Keep tags the user already has.

3. Choose the performance mode

Set the job-level `performance_target`:

  • **`STANDARD`** when nothing is time-critical and cost matters more than speed: scheduled batch and

ETL, nightly or weekly runs, or any job where the user states no latency need. Runs start within minutes instead of seconds and cost less. This is the default choice.

  • **`PERFORMANCE_OPTIMIZED`** only when the user names an SLA or a deadline that is tight relative

to the runtime, a person waits on the result, or the job runs every 30 minutes or more often.

`performance_target` applies to the whole job, notebook tasks included (only interactive notebooks always run performance-optimized). Say which mode you chose and why in one line, e.g. "STANDARD: nightly batch, no deadline." Serverless GPU tasks ignore `performance_target`.

Pipelines have no `performance_target` of their own. To schedule a pipeline, trigger it from a job with a `pipeline_task` and set the mode on that job; the job's mode applies to the pipeline update. A continuous pipeline runs in STANDARD only when a continuous job runs it.

4. Code you write runs on serverless from the start

Serverless reads Unity Catalog tables, `/Volumes/...` paths, cloud paths behind a Unity Catalog external location, and the sample data under `dbfs:/databricks-datasets/`. It cannot read DBFS mounts (`dbfs:/mnt/...`) or other DBFS root paths; only those count as a DBFS blocker.

DataFrame and SQL APIs only (no RDDs, `sc.*`, `spark.sparkContext`); Unity Catalog three-part table names; `/Volumes/...` paths instead of DBFS mounts (`dbfs:/mnt/...`); `availableNow` streaming triggers; no `.cache()` or `.persist()`; no `spark.conf.set` for executor, driver or memory settings; pinned dependencies in the environment, never init scripts.

5. Existing code: a quick check, not a migration

The user is waiting for the deployment. Spend one short pass, then deploy.

1. **Scan only the files the job runs** (a text search such as `grep` or `rg`; do not read the whole repo) for: `sc.`, `sparkContext`, `.rdd`, `parallelize`, `mapPartitions`, `reduceByKey`, `groupByKey`, `dbfs:/`, `/dbfs/`, `dbutils.fs.mount`, `.cache(`

Read more
Ships withdatabricks-agent-skills

Build on Databricks with AI coding agents such as Claude Code, Cursor, Codex, GitHub Copilot, and many others. This repository provides the skills and agent plugins for Databricks AI Tools.

Get the whole plugin

Other skills on databricks-agent-skills.