databricks-agent-brick…
Create Agent Bricks: Knowledge Assistants (KA) for document Q&A and Supervisor Agents for…
Databricks Model Serving endpoint lifecycle and ops. Use when asked to: CRUD serving endpoints (CLI or MLflow Deployments client); configure traffic routing for A/B / canary deploys and zero-downtime version swaps; retrieve OpenAPI schemas; inspect logs, metrics, or permissions;
$ npx -y skills add databricks/databricks-agent-skills --skill databricks-model-serving --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/databricks-model-servingContext preview
The summary Claude sees to decide when to auto-load this skill.
Databricks Model Serving endpoint lifecycle and ops. Use when asked to: CRUD serving endpoints (CLI or MLflow Deployments client); configure traffic routing for A/B / canary deploys and zero-downtime version swaps; retrieve OpenAPI schemas; inspect logs, metrics, or permissions;
name: databricks-model-serving description: "Databricks Model Serving endpoint lifecycle and ops. Use when asked to: CRUD serving endpoints (CLI or MLflow Deployments client); configure traffic routing for A/B / canary deploys and zero-downtime version swaps; retrieve OpenAPI schemas; inspect logs, metrics, or permissions; manage AI Gateway rate limits; discover Foundation Model API endpoints at runtime; integrate endpoints into Databricks Apps; or stream from off-platform clients (Vercel AI SDK v6, standalone Node.js). NOT for: training, MLflow autologging, UC registration, custom PyFunc/ResponsesAgent authoring (databricks-ml-training); Knowledge Assistants/Supervisor Agents (databricks-agent-bricks); MLflow evaluation (databricks-mlflow-evaluation)." compatibility: Requires databricks CLI (>= v0.294.0) metadata: version: "0.4.0" parent: databricks-core
**FIRST**: Use the parent `databricks-core` skill for CLI basics, authentication, and profile selection.
Model Serving provides managed endpoints for serving LLMs, custom ML models, and external models as scalable REST APIs. Endpoints are identified by **name** (unique per workspace).
| Type | When to Use | Key Detail | |------|-------------|------------| | Pay-per-token | Foundation Model APIs (Llama, GPT-5, Claude, Gemini, etc.) | Uses `system.ai.*` catalog models, pre-provisioned in every workspace. Discover at runtime — see [Foundation Model API endpoints](#foundation-model-api-endpoints) below. | | Provisioned throughput | Dedicated GPU capacity | Guaranteed throughput, higher cost | | Custom model | Your own MLflow models or containers | Deploy any model with an MLflow signature |
Serving Endpoint (top-level, identified by NAME) ├── Config │ ├── Served Entities (model references + scaling config) │ └── Traffic Config (routing percentages across entities) ├── AI Gateway (rate limits, usage tracking) └── State (READY / NOT_READY, config_update status)
**Do NOT guess command syntax.** Discover available commands and their usage dynamically:
# List all serving-endpoints subcommands databricks serving-endpoints -h # Get detailed usage for any subcommand (flags, args, JSON fields) databricks serving-endpoints <subcommand> -h
Run `databricks serving-endpoints -h` before constructing any command. Run `databricks serving-endpoints <subcommand> -h` to discover exact flags, positional arguments, and JSON spec fields for that subcommand.
> **Do NOT list endpoints before creating.**
databricks serving-endpoints create <ENDPOINT_NAME> \
--json '{
"served_entities": [{
"entity_name": "<MODEL_CATALOG_PATH>",
"entity_version": "<VERSION>",
"min_provisioned_throughput": 0,
"max_provisioned_throughput": 0,
"workload_size": "Small",
"scale_to_zero_enabled": true
}],
"traffic_config": {
"routes": [{
"served_entity_name": "<ENTITY_NAME>",
"traffic_percentage": 100
}]
}
}' --profile <PROFILE>databricks serving-endpoints get <ENDPOINT_NAME> --profile <PROFILE> # Check: state.ready == "READY"
`mlflow.deployments.get_deploy_client("databricks").create_endpoint(name=..., config={...})` takes the same JSON shape as the CLI. Two gotchas:
To roll an endpoint to a new model version: repoint the alias **and** call `update_endpoint` with the new `served_entities` + matching `traffic_config`. Missing either half is the common bug — alias-only doesn't update the endpoint; `update_endpoint`-only leaves the alias pointing at the old version.
from mlflow.tracking import MlflowClient
from mlflow.deployments import get_deploy_client
registry = MlflowClient(registry_uri="databricks-uc")
deploy = get_deploy_client("databricks")
registry.set_registered_model_alias(FULL_NAME, "prod", new_version)
deploy.update_endpoint(endpoint=ENDPOINT_NAME, config={
"served_entities": [{"entity_name": FULL_NAME, "entity_version": new_version,
"workload_size": "Small", "scale_to_zero_enabled": True}],
"traffic_config": {"routes": [
{"served_model_name": f"{NAME}-{neBuild on Databricks with AI coding agents such as Claude Code, Cursor, Codex, and GitHub Copilot. This repository provides the skills and agent plugins for Databricks AI Tools.
Repo: databricks/databricks-agent-skills
Create Agent Bricks: Knowledge Assistants (KA) for document Q&A and Supervisor Agents for…
Use Databricks built-in AI Functions (ai_classify, ai_extract, ai_summarize, ai_mask,…
Create Databricks AI/BI dashboards. Must use when creating, updating, or deploying Lakeview…
Design the UX of custom-code Databricks Apps (AppKit/React) data screens — KPI/overview…
Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio,…
Build apps on Databricks Apps platform. Use when asked to create data apps, analytics tools,…