Skip to content
Data
Skill

/spark-cli

Manages Fabric Spark work, including notebook cell code with %%configure, %%sql, PySpark and notebookutils, named notebook runs, Livy sessions, triage of failed or OOM notebook and pipeline Spark runs, and the Materialized Lake View (MLV) lifecycle. Load it before writing MLV

BOOST
From plugin
skills-for-fabric
1.2k25 skills5 agents3 MCP
Install
$ npx -y skills add microsoft/skills-for-fabric --skill spark-cli --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/spark-cli

Context preview

The summary Claude sees to decide when to auto-load this skill.

Manages Fabric Spark work, including notebook cell code with %%configure, %%sql, PySpark and notebookutils, named notebook runs, Livy sessions, triage of failed or OOM notebook and pipeline Spark runs, and the Materialized Lake View (MLV) lifecycle. Load it before writing MLV

SKILL.md

spark-cli.SKILL.md
name: spark-cli
description: "Manages Fabric Spark work, including notebook cell code with %%configure, %%sql, PySpark and notebookutils, named notebook runs, Livy sessions, triage of failed or OOM notebook and pipeline Spark runs, and the Materialized Lake View (MLV) lifecycle. Load it before writing MLV SQL, since CREATE MATERIALIZED LAKE VIEW and its CONSTRAINT clause are Fabric-only. KQL materialized views belong to eventhouse-cli."

> **Telemetry — MANDATORY.** Every `api.fabric.microsoft.com` call must carry > `x-ms-fabric-skill: spark-cli` (`az rest`: `--headers "x-ms-fabric-skill=spark-cli"`), > including every LRO poll, `fabric_lro` and retry. Snippets omit it — add it anyway.

> **CRITICAL NOTES** > 1. To find the workspace details (including its ID) from workspace name: list all workspaces and, then, use JMESPath filtering > 2. To find the item details (including its ID) from workspace ID, item type, and item name: list all items of that type in that workspace and, then, use JMESPath filtering > 3. **Skill disambiguation**: use `spark-cli` for notebook cell code (including `%%sql` cells), named notebook runs, Livy sessions, Spark failure triage, and everything about a Materialized Lake View -- writing the definition, reviewing a query for incremental-refresh readiness, and scheduling, refreshing, monitoring or diagnosing an existing one. A KQL materialized view in an Eventhouse is `eventhouse-cli`; plain read-only T-SQL against a Warehouse or Lakehouse SQL endpoint is `sqldw-cli`. > 4. **Hard routing boundary**: never execute an Eventhouse/KQL materialized-view request from this skill. Route it to `eventhouse-cli`; if that skill is unavailable, state that the request cannot be completed in the current skill context and stop without calling Fabric APIs or creating artifacts.

Fabric Spark and Materialized Lake Views -- CLI Skill

This one skill owns Fabric Spark: notebook cell authoring, notebook runs, Livy-session analysis, Spark failure diagnostics, and the whole Materialized Lake View lifecycle.

It is a **mode dispatcher** and contains NO procedures. Pick the mode that matches the request from the table below, then **read the matching `references/<mode>.md` file end to end with your file-reading tool BEFORE issuing a single command**. That file holds the endpoints, payload shapes, templates and gotchas; acting without it produces wrong payloads and wrong results.

Mode selection

| Mode | Use when the request ... | Example triggers | Read this first | |---|---|---|---| | `authoring` | writes notebook cell code (PySpark, Scala, SparkR, `%%sql`, `%%configure`), runs a notebook by name and reports a NORMAL run status, or authors a Materialized Lake View definition or reviews an MLV query for incremental-refresh readiness | write notebook code, notebook cell code, %%sql cell, run notebook, execute notebook, notebookutils, create materialized lake view, is this MLV query incremental-refresh ready | [references/authoring.md](references/authoring.md) | | `consumption` | runs interactive ad-hoc PySpark in a Lakehouse Livy session -- never a notebook | create a Livy session, run calculation in Livy, PySpark, DataFrame analysis, join tables across lakehouses, Delta time-travel | [references/consumption.md](references/consumption.md) | | `operations` | diagnoses a FAILED, unhealthy, throttled or slow Spark notebook / pipeline / Livy run | failed notebook, Spark Livy health, Spark OOM, why is my notebook slow, job diagnostics, 430 throttling | [references/operations.md](references/operations.md) | | `mlv` | discovers or operates EXISTING Materialized Lake Views: Spark SQL discovery, refresh schedules, on-demand refresh, run history, cancellation, and refresh-failure classification | discover MLVs, list materialized lake views, schedule MLV, MLV run history, cancel refresh, trigger MLV refresh, diagnose MLV refresh failure | [references/mlv.md](references/mlv.md) |

Mode boundary rule

Mode is decided by the artefact and the outcome, not by the language. A notebook cell is always `authoring` even when the cell is `%%sql`. A Livy session is always `consumption`. A Spark run that FAILED or is unhealthy is `operations`; a run that succeeded is reported by `authoring`. For a Materialized Lake View the VERB decides: writing or reviewing the definition is `authoring`, while discovering, scheduling, refreshing, monitoring or diagnosing existing MLVs is `mlv`. If discovery must be executed, switch to `consumption` for Livy or `authoring` for a notebook only after reading the discovery command from `references/mlv.md`.

If a request genuinely spans modes, handle them one at a time and read each reference before you start that part. If the mode is ambiguous after reading this table, ask one short clarifying question instead of guessing.

Terminal write -- the step you must not skip

Reading the reference and planning the change is NOT completing the task. Each mutating mode ends with one state-changing call. If you did not issue it, nothing was persisted -- say so explicitly rather than reporting success.

| Mode | Terminal write | |---|---| | `authoring` | for an EXISTING notebook, `POST .../notebooks/{id}/updateDefinition` to save the cell; a NEW notebook needs `POST /v1/workspaces/{ws}/items` first; for a run request, trigger the job via the Jobs API. Printing cell code into the chat is not saving or running it. | | `consumption` | none -- this mode is read-only | | `operations` | none -- this mode is read-only | | `mlv` | `POST /v1/workspaces/{ws}/lakehouses/{lakehouse}/jobs/refreshMaterializedLakeViews/instances` for an on-demand refresh, or the schedule create/update/delete call for a scheduling request. Reporting what the schedule would be is not creating it. |

Before you report the task done, confirm the terminal call returned success and, where the reference documents a readback, read the artefact back to prove the change landed.

Shared essentials (all modes)

Re

Read more
Ships withskills-for-fabric

Microsoft Fabric Skills are reusable AI assistant instructions for working with Microsoft Fabric. They help GitHub Copilot CLI and compatible AI coding tools understand Fabric workloads, APIs, query patterns, and operational best practices.

Get the whole plugin

Other skills on skills-for-fabric.