Skip to content
Data
Skill

/databricks-ai-runtime

Databricks AI Runtime, the `databricks air` CLI commands for submitting and managing GPU training workloads on Databricks serverless compute. Use for: writing and submitting `databricks air` workload YAML, passing hyperparameters and secrets, checking run status,

BOOST
From plugin
databricks-agent-skills
34534 skills4 commands3 hooks
Install
$ npx -y skills add databricks/databricks-agent-skills --skill databricks-ai-runtime --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/databricks-ai-runtime

Context preview

The summary Claude sees to decide when to auto-load this skill.

Databricks AI Runtime, the `databricks air` CLI commands for submitting and managing GPU training workloads on Databricks serverless compute. Use for: writing and submitting `databricks air` workload YAML, passing hyperparameters and secrets, checking run status,

SKILL.md

databricks-ai-runtime.SKILL.md
name: databricks-ai-runtime
description: "Databricks AI Runtime, the `databricks air` CLI commands for submitting and managing GPU training workloads on Databricks serverless compute. Use for: writing and submitting `databricks air` workload YAML, passing hyperparameters and secrets, checking run status, listing/cancelling runs, streaming a run's logs and watching its progress, custom Docker image setup, and environment configuration."
compatibility: Requires the Databricks CLI with the air command. Flattened code_source requires CLI 1.20.0 or newer; older CLIs can use the nested snapshot format documented below. See the [installation guide](https://docs.databricks.com/aws/en/machine-learning/ai-runtime/cli/installation).
metadata:
  version: "0.2.1"

Databricks AI Runtime (`databricks air`)

Databricks AI Runtime is a set of `databricks air` CLI commands for submitting GPU training workloads to Databricks serverless compute. It manages environment setup, distributed training configuration, and workload lifecycle, without requiring you to manage clusters or infrastructure.

This skill covers the three everyday flows: **starting a run**, **checking on a run**, and **monitoring a run** as it trains.

Reference discipline

The CLI's built-in help is the source of truth for field names, GPU types, on-cluster environment variables, and constraints. Prefer it at runtime over anything memorized. Look up a config field by passing its path to `-h` on the `run` command (the path is a separate argument):

databricks air run -h config                              # list the YAML fields
databricks air run -h config.compute                      # one field
databricks air run -h config.compute.accelerator_type     # a nested field

`databricks air <subcommand> --help` documents each command's flags. This skill covers workflow and the few things the help does not spell out.

Session setup

Check the CLI version, confirm the command is available, and list your Databricks config profiles:

databricks version
databricks air --help
databricks auth profiles --skip-validate    # available Databricks config profiles

The examples use flattened `code_source` fields, available in [Databricks CLI 1.20.0 and newer](https://github.com/databricks/cli/releases/tag/v1.20.0). On an older CLI, use the equivalent nested snapshot format shown under [Workload YAML](#workload-yaml). CLI 1.20.0 and newer accept both formats. Do not mix nested and flattened fields in the same configuration.

Pass `-p <profile>` on every command. Add `-o json` for scripted/programmatic calls so you get a structured envelope instead of human text:

  • **Success**: `{"v": 1, "ts": "...", "data": {...}}`
  • **Error**: `{"v": 1, "ts": "...", "error": {"code": "...", "kind": "...", "message": "...", "retryable": bool}}`

Config validation (input) errors are the exception: they print as plain text to stderr with a non-zero exit even under `-o json`, so check the exit code and stderr too, not just the parsed JSON. When you do get a structured error, its `kind` tells you what to do:

| kind | Meaning | Action | |---|---|---| | `PERMANENT` | Not authenticated / permission denied / non-retryable | If auth: `databricks auth login --host <workspace-url>` | | `NOT_FOUND` | Bad run ID / resource | Check the input | | `CONFLICT` | Idempotency key collision | Use a new `--idempotency-key` | | `TRANSIENT` | Retryable | Retry |

Subcommands

| Command | What it does | |---|---| | `databricks air run` | Submit a workload from a YAML file. Add `--watch` to stream logs until it finishes. | | `databricks air get <JOB_RUN_ID>` | Show status, config, and timing for one run. | | `databricks air list` | List your active runs (add `--all-status` for finished, `--all-users`, `--filter`, `--limit`). | | `databricks air cancel <JOB_RUN_ID> [<JOB_RUN_ID>...]` / `... --all` | Cancel one or more runs, or all of your active runs. | | `databricks air logs <JOB_RUN_ID>` | Stream a running run's logs, or fetch a finished run's logs. |

For flags and defaults, run `databricks air <subcommand> --help`.

Workload YAML

experiment_name: my-training-job     # alphanumeric, - and _ only
compute:
  num_accelerators: 1
  accelerator_type: GPU_1xA10        # run `databricks air run -h config.compute` for valid values
environment:
  version: "5"                       # client image version; prefix "databricks_ai_v" to also load the ML venv
  dependencies:
    - mlflow
code_source:
  root_path: "."                     # relative to workload.yaml; extracted to $CODE_SOURCE_PATH on each node
command: |-
  cd "$CODE_SOURCE_PATH"
  python train.py

Submit with `databricks air run --file workload.yaml -p <profile>`. Field details live in `databricks air run -h config.<field>` (the authoritative source). Today's accelerator types are `GPU_1xA10` and `GPU_8xH100`; confirm with `databricks air run -h config.compute`.

For an older CLI, replace only the `code_source` block above with:

code_source:
  type: snapshot
  snapshot:
    root_path: "."

The values and behavior are the same: add `type: snapshot` under `code_source` and put all snapshot fields (`root_path`, and any `remote_volume`, `git`, or `include_paths`) inside `code_source.snapshot`. Keep the rest of the workload unchanged. Check `databricks air run -h config.code_source` for the fields supported by the installed CLI.

Configuring the workload

These optional top-level fields cover the common needs. Run `databricks air run -h config.<field>` for the full reference on any of them.

Environment variables and secrets

Set plain values under `env_variables`; pull sensitive values from Databricks Secrets under `secrets` (each entry maps an env var name to a `scope/key`, and the value is materialized on the worker, never logged):

env_variables:
  BATCH_SIZE: "32"
secrets:
  HF_TOKEN: my_scope/hf_token          # -> $HF_TOKEN on every node
  WANDB_API_KEY: my_s
Read more
Ships withdatabricks-agent-skills

Build on Databricks with AI coding agents such as Claude Code, Cursor, Codex, GitHub Copilot, and many others. This repository provides the skills and agent plugins for Databricks AI Tools.

Get the whole plugin

Other skills on databricks-agent-skills.