finding-google-skills
Google platform decision and setup guidance, loaded on demand from Google's skill catalog.…
Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics
$ npx -y skills add google/skills --skill gke-ai-troubleshooting-tpu-mxla-hang --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/gke-ai-troubleshooting-tpu-mxla-hangContext preview
The summary Claude sees to decide when to auto-load this skill.
Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics
name: gke-ai-troubleshooting-tpu-mxla-hang description: >- Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption). metadata: version: "1.0.0" category: AiAndMachineLearning
Diagnose Cloud TPU multi-slice training hangs on Google Kubernetes Engine (GKE) by correlating Megascale `HANG_DETECTED` logs and **ML Diagnostics Workload Monitoring** `Megascale XLA (MXLA) Hang Analyzer` reports with **1-minute multi-slice latency metrics** (`kubernetes.io/container/multislice/*`) and GKE node topology labels.
---
[Google Cloud SDK](https://cloud.google.com/sdk/docs/install) (`gcloud` with `alpha` component for `gcloud alpha mldiagnostics`) and `kubectl`.
linked (`gcloud billing projects describe {project_id}`), authenticate (`gcloud auth login`), set the target project (`gcloud config set project {project_id}`), and ensure `container.googleapis.com`, `logging.googleapis.com`, `monitoring.googleapis.com`, and `hypercomputecluster.googleapis.com` are enabled.
supports JAX on TPUs (see [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt)). Workload Monitoring is enabled by default, supports the `jobset` and `job` GKE job types, and is compatible with GKE versions `1.36.0-gke.4681000` and later ([Configure GKE for ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/gke.md.txt), which also covers the cluster setup needed for on-demand profiling in Step 4 Path B). If a workload uses another framework (such as PyTorch) or another custom resource type, `gcloud alpha mldiagnostics` won't list ML runs or monitored events for it. The Megascale XLA hang analyzer and Megascale XLA metrics require LibTPU `0.40.0` or later (see "Get started" in [Workload monitoring with ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/workload-monitoring.md.txt)).
listed in the "IAM permissions" section of [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt), for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions)
2
nodes` query in Step 3
(`roles/container.clusterAdmin`)
(Sections: "IAM permissions")
(Sections: "Set up with gcloud CLI, Google Cloud console, or Terraform", "Manual installation", "Connection-operator")
(Sections: "Get started", "Megascale XLA (MXLA) Hang Analyzer", "Access Workload Monitoring information through the API", "System Metrics")
(Sections: "List machine learning runs", "Monitored-events commands", "List profiler targets", "Capture on-demand profiler sessions")
(Sections: "Before you contact support")
(Sections: "Repair criteria", "Verify node auto-repair is enabled for a Standard node pool", "Enable auto-repair for an existing Standard node pool", "Node auto repair in TPU slice nodes")
(Sections: "Configure auto repair for TPU slice nodes")
> **Read-only rule**: Run read-only diagnostic commands only. Never drain, > delete, or re-create nodes, or run any other command that changes the cluster. > Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
---
A Megascale hang occurs when a multi-slice worker has waited on a Megascale communication operation for a set timeout period. The TPU logs then show a Megascale `HANG_DETECTED` message. `HANG_DETECTED` is a catch-all signal that the workload isn't progressing, and the cause can be in software or in hardware. When a hang occurs, ML Diagnostics runs the `Megascale XLA (MXLA) Hang Analyzer`, wh
This repository contains Agent Skills for Google products and technologies, including Google Cloud.
Repo: google/skills
Google platform decision and setup guidance, loaded on demand from Google's skill catalog.…
Provides safety-critical validation, guardrails, and data reduction for gcloud CLI operations…
Provides expert guidance on Identity and Access Management (IAM) and authenticating and…
Guides a developer's first steps on Google Cloud, covering account creation, billing setup,…
Searches, retrieves, and synthesizes official Google developer documentation across Google…
Guides developers through managing (adding, removing, and clearing) audience members for…