Skip to content
Development
Skill

/gke-ai-troubleshooting-tpu-mxla-hang

Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics

GuideBOOST
From plugin
google-skills
21k156 skills1 MCP
Install
$ npx -y skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/gke-ai-troubleshooting-tpu-mxla-hang

Context preview

The summary Claude sees to decide when to auto-load this skill.

Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics

SKILL.md

gke-ai-troubleshooting-tpu-mxla-hang.SKILL.md
name: gke-ai-troubleshooting-tpu-mxla-hang
description: >-
  Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED`
  logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA)
  Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute
  Cloud Monitoring multi-slice latency metrics
  (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO
  launch divergence and host data-input stalls from TPU chip, SparseCore, ICI,
  or network fabric faults. Use when multi-slice TPU training jobs freeze
  without progressing steps, emit `HANG_DETECTED`, or stall in collective
  operations. Don't use for gradual step-time throughput drops without hangs
  (use gke-ai-troubleshooting-tpu-performance-degradation) or pod
  preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption).
metadata:
  version: "1.0.0"
  category: AiAndMachineLearning

Troubleshoot GKE TPU multi-slice hangs with the MXLA Hang Analyzer

Diagnose Cloud TPU multi-slice training hangs on Google Kubernetes Engine (GKE) by correlating Megascale `HANG_DETECTED` logs and **ML Diagnostics Workload Monitoring** `Megascale XLA (MXLA) Hang Analyzer` reports with **1-minute multi-slice latency metrics** (`kubernetes.io/container/multislice/*`) and GKE node topology labels.

---

Prerequisites

  • **Tools**: Install the

[Google Cloud SDK](https://cloud.google.com/sdk/docs/install) (`gcloud` with `alpha` component for `gcloud alpha mldiagnostics`) and `kubectl`.

  • **Cloud Billing & Project Configuration**: Verify an active billing account is

linked (`gcloud billing projects describe {project_id}`), authenticate (`gcloud auth login`), set the target project (`gcloud config set project {project_id}`), and ensure `container.googleapis.com`, `logging.googleapis.com`, `monitoring.googleapis.com`, and `hypercomputecluster.googleapis.com` are enabled.

  • **Supported workloads and versions**: Google Cloud ML Diagnostics only

supports JAX on TPUs (see [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt)). Workload Monitoring is enabled by default, supports the `jobset` and `job` GKE job types, and is compatible with GKE versions `1.36.0-gke.4681000` and later ([Configure GKE for ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/gke.md.txt), which also covers the cluster setup needed for on-demand profiling in Step 4 Path B). If a workload uses another framework (such as PyTorch) or another custom resource type, `gcloud alpha mldiagnostics` won't list ML runs or monitored events for it. The Megascale XLA hang analyzer and Megascale XLA metrics require LibTPU `0.40.0` or later (see "Get started" in [Workload monitoring with ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/workload-monitoring.md.txt)).

  • **Required IAM Roles**:
  • Cluster Director Editor (`roles/hypercomputecluster.editor`), the role

listed in the "IAM permissions" section of [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt), for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions)

  • Monitoring Viewer (`roles/monitoring.viewer`) for the PromQL queries in Step

2

  • Logs Viewer (`roles/logging.viewer`) for the Cloud Logging query in Step 1
  • Kubernetes Engine Viewer (`roles/container.viewer`) for the `kubectl get

nodes` query in Step 3

  • For remediation (`[High Risk]` steps): Kubernetes Engine Cluster Admin

(`roles/container.clusterAdmin`)

  • **Reference Documentation**:
  • [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt)

(Sections: "IAM permissions")

  • [Configure GKE for ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/gke.md.txt)

(Sections: "Set up with gcloud CLI, Google Cloud console, or Terraform", "Manual installation", "Connection-operator")

  • [Workload monitoring with ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/workload-monitoring.md.txt)

(Sections: "Get started", "Megascale XLA (MXLA) Hang Analyzer", "Access Workload Monitoring information through the API", "System Metrics")

  • [Get started with the ML Diagnostics CLI](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/cli.md.txt)

(Sections: "List machine learning runs", "Monitored-events commands", "List profiler targets", "Capture on-demand profiler sessions")

  • [Get support](https://docs.cloud.google.com/kubernetes-engine/docs/getting-support.md.txt)

(Sections: "Before you contact support")

  • [Auto-repair nodes](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/node-auto-repair.md.txt)

(Sections: "Repair criteria", "Verify node auto-repair is enabled for a Standard node pool", "Enable auto-repair for an existing Standard node pool", "Node auto repair in TPU slice nodes")

  • [Deploy TPU workloads in GKE Standard](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/tpus.md.txt)

(Sections: "Configure auto repair for TPU slice nodes")

  • [Dump HLO Computations (OpenXLA)](https://openxla.org/xla/hlo_dumps)

> **Read-only rule**: Run read-only diagnostic commands only. Never drain, > delete, or re-create nodes, or run any other command that changes the cluster. > Give the user any fix to apply themselves.

When you recommend a fix, link the doc section that describes it.

---

MXLA Hang Analyzer routing overview

A Megascale hang occurs when a multi-slice worker has waited on a Megascale communication operation for a set timeout period. The TPU logs then show a Megascale `HANG_DETECTED` message. `HANG_DETECTED` is a catch-all signal that the workload isn't progressing, and the cause can be in software or in hardware. When a hang occurs, ML Diagnostics runs the `Megascale XLA (MXLA) Hang Analyzer`, wh

Read more
Ships withgoogle-skills

This repository contains Agent Skills for Google products and technologies, including Google Cloud.

Get the whole plugin

Other skills on google-skills.