Skip to content
Data
Skill

/hdinsight-migration

Ports HDInsight Spark and Hive workloads to Fabric, converting HiveContext and SparkContext to SparkSession, WASB and ABFS paths to OneLake shortcuts, Hive DDL to Delta tables in a Lakehouse, and Oozie coordinators and actions to pipeline activities and triggers. Use when the

BOOST
From plugin
skills-for-fabric
1.2k25 skills5 agents3 MCP
Install
$ npx -y skills add microsoft/skills-for-fabric --skill hdinsight-migration --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/hdinsight-migration

Context preview

The summary Claude sees to decide when to auto-load this skill.

Ports HDInsight Spark and Hive workloads to Fabric, converting HiveContext and SparkContext to SparkSession, WASB and ABFS paths to OneLake shortcuts, Hive DDL to Delta tables in a Lakehouse, and Oozie coordinators and actions to pipeline activities and triggers. Use when the

SKILL.md

hdinsight-migration.SKILL.md
name: hdinsight-migration
description: "Ports HDInsight Spark and Hive workloads to Fabric, converting HiveContext and SparkContext to SparkSession, WASB and ABFS paths to OneLake shortcuts, Hive DDL to Delta tables in a Lakehouse, and Oozie coordinators and actions to pipeline activities and triggers. Use when the source workload is HDInsight."

> **Telemetry — MANDATORY.** Every `api.fabric.microsoft.com` call must carry > `x-ms-fabric-skill: hdinsight-migration` (`az rest`: `--headers "x-ms-fabric-skill=hdinsight-migration"`), > including every LRO poll, `fabric_lro` and retry. Snippets omit it — add it anyway.

> **CRITICAL NOTES** > 1. To find workspace details (including its ID) from a workspace name: list all workspaces, then use JMESPath filtering > 2. To find item details (including its ID) from workspace ID, item type, and item name: list all items of that type in that workspace, then use JMESPath filtering > 3. HDInsight has no `mssparkutils` or `dbutils` equivalent — `notebookutils` is net-new capability being introduced > 4. `HiveContext` and `SQLContext` are legacy Spark 1.x/2.x APIs — Fabric uses Spark 3.x `SparkSession` exclusively > 5. `wasb://` paths are deprecated and require a Storage Account key or SAS — replace with OneLake shortcuts > 6. Fabric cannot read or shortcut `hdfs://` directly. Its Pipeline HDFS connector supports Anonymous authentication only; export or bridge Kerberos sources to supported storage first.

HDInsight → Microsoft Fabric Migration

Prerequisite Knowledge

Read these companion documents before executing migration tasks:

  • [COMMON-CORE.md](../../common/COMMON-CORE.md) — Fabric REST API patterns, authentication, token audiences, item discovery
  • [COMMON-CLI.md](../../common/COMMON-CLI.md) — `az rest`, `az login`, token acquisition, Fabric REST via CLI
  • [SPARK-AUTHORING-CORE.md](../../common/SPARK-AUTHORING-CORE.md) — Notebook deployment, lakehouse creation, Spark job execution

For notebook and Lakehouse creation, see [spark-cli](../spark-cli/SKILL.md). For Fabric Warehouse DDL/DML authoring, see [sqldw-cli](../sqldw-cli/SKILL.md).

---

Table of Contents

| Topic | Reference | |---|---| | Migration Workload Map | [§ Migration Workload Map](#migration-workload-map) | | SparkSession & Context API Changes | [§ SparkSession API Changes](#sparksession--context-api-changes) | | WASB / ABFS → OneLake Path Migration | [path-migration.md](resources/path-migration.md) | | Hive DDL → Delta Lake / Lakehouse Schemas | [hive-to-delta.md](resources/hive-to-delta.md) | | Oozie → Fabric Pipelines | [§ Oozie → Fabric Pipelines](#oozie--fabric-pipelines) | | Introducing `notebookutils` | [§ Introducing notebookutils](#introducing-notebookutils) | | Before/After Code Patterns | [code-patterns.md](resources/code-patterns.md) | | Spark Configuration Differences | [§ Spark Configuration Differences](#spark-configuration-differences) | | Must / Prefer / Avoid | [§ Must / Prefer / Avoid](#must--prefer--avoid) | | Authentication & Token Acquisition | [COMMON-CORE.md § Authentication](../../common/COMMON-CORE.md#authentication--token-acquisition) | | Lakehouse Management | [SPARK-AUTHORING-CORE.md § Lakehouse Management](../../common/SPARK-AUTHORING-CORE.md#lakehouse-management) |

---

Migration Workload Map

| HDInsight Component | Fabric Target | Notes | |---|---|---| | **Spark cluster** (notebooks, scripts) | Fabric Spark (Lakehouse / Notebooks / SJD) | No persistent cluster — Starter Pool or Custom Pool provides on-demand Spark | | **Hive / HiveServer2** | **Lakehouse SQL Endpoint** + Lakehouse schemas | Delta Lake replaces Hive metastore; schemas provide namespace equivalent | | **HBase** | **Fabric Warehouse** or **Azure Cosmos DB** (separate from Fabric) | HBase has no direct Fabric equivalent — assess workload access patterns | | **Oozie workflows** | **Fabric Data Pipelines** | Map Oozie actions to Fabric activities; see [§ Oozie → Fabric Pipelines](#oozie--fabric-pipelines) | | **YARN Resource Manager** | **Fabric Spark monitoring** (Spark UI, Monitoring Hub) | No YARN — Fabric manages compute automatically | | **Ambari** | **Fabric Monitoring Hub** + **Admin Portal** | Cluster health, capacity, and job monitoring | | **WASB / ABFS storage** | **OneLake Shortcuts** → `abfss://workspace@onelake.dfs.fabric.microsoft.com/` | See [path-migration.md](resources/path-migration.md) | | **Ranger policies** | **Fabric workspace roles** + **OneLake data access roles** | Map Ranger row/column filters to Lakehouse row-level security | | **Livy REST server** | **Fabric Livy API** | Compatible endpoint — see SPARK-AUTHORING-CORE.md |

---

SparkSession & Context API Changes

HDInsight Spark clusters often use legacy Spark 1.x / 2.x API styles. Replace all of these with the unified `SparkSession`:

| Legacy HDInsight Pattern | Fabric Spark 3.x Replacement | |---|---| | `from pyspark import SparkContext; sc = SparkContext()` | Not needed — `sc = spark.sparkContext` (pre-instantiated) | | `from pyspark.sql import HiveContext; hc = HiveContext(sc)` | Not needed — `spark` session has Hive-compatible SQL support via Delta schemas | | `from pyspark.sql import SQLContext; sqlc = SQLContext(sc)` | Not needed — use `spark.sql(...)` directly | | `SparkSession.builder.enableHiveSupport().getOrCreate()` | Not needed in Fabric — `spark` is pre-built and available | | `sc.textFile("wasb://container@account.blob.core.windows.net/path")` | `spark.read.text("abfss://workspace@onelake.dfs.fabric.microsoft.com/lh.Lakehouse/Files/path")` | | `sqlContext.sql("CREATE TABLE ... STORED AS ORC")` | See [hive-to-delta.md](resources/hive-to-delta.md) for Delta DDL equivalent |

> In Fabric notebooks, `spark` (SparkSession) and `sc` (SparkContext) are **pre-instantiated** — do not call `SparkContext()` or `SparkSession.builder...getOrCreate()` at the top of migrated notebooks.

---

Oozie → Fabric Pipelines

Map Oozie workflow actions to Fabric Data Pipeline activities:

| Oo

Read more
Ships withskills-for-fabric

Microsoft Fabric Skills are reusable AI assistant instructions for working with Microsoft Fabric. They help GitHub Copilot CLI and compatible AI coding tools understand Fabric workloads, APIs, query patterns, and operational best practices.

Get the whole plugin

Other skills on skills-for-fabric.