Skip to content
Development
Skill

/gke-ai-troubleshooting-tpu-performance-degradation

Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system

GuideBOOST
From plugin
google-skills
21k156 skills1 MCP
Install
$ npx -y skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/gke-ai-troubleshooting-tpu-performance-degradation

Context preview

The summary Claude sees to decide when to auto-load this skill.

Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system

SKILL.md

gke-ai-troubleshooting-tpu-performance-degradation.SKILL.md
name: gke-ai-troubleshooting-tpu-performance-degradation
description: >-
  Diagnose GKE Cloud TPU training throughput drops and step-time regressions
  (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud
  alpha mldiagnostics monitored-events` /
  `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring
  system metrics (`kubernetes.io/node/accelerator/*`). Distinguishes hardware
  and network fabric throttling from workload resource bottlenecks (HBM
  capacity, host memory, or host CPU saturation). Use when TPU training
  throughput or duty cycle drops without crashing pods, when
  `PERFORMANCE_DEGRADATION` monitored events fire, or when triaging slow
  multi-slice training steps. Don't use for complete multi-slice XLA execution
  stalls with `HANG_DETECTED` logs (use gke-ai-troubleshooting-tpu-mxla-hang) or
  pod eviction/interruption restarts (use
  gke-ai-troubleshooting-jobset-interruption).
metadata:
  version: "1.0.0"
  category: AiAndMachineLearning

Troubleshoot GKE TPU performance degradation with ML Diagnostics Workload Monitoring

Diagnose and mitigate Cloud TPU training throughput drops and step-time regressions (`15%+` drop in TPU duty cycle) on Google Kubernetes Engine (GKE) by correlating **ML Diagnostics Workload Monitoring** `MonitoredEvent` analyzer reports with **1-minute Cloud Monitoring system metrics** and GKE node topology labels.

---

Prerequisites

  • **Tools**: Install the

[Google Cloud SDK](https://cloud.google.com/sdk/docs/install) (`gcloud` with `alpha` component for `gcloud alpha mldiagnostics`) and `kubectl`.

  • **Cloud Billing & Project Configuration**: Verify an active billing account is

linked (`gcloud billing projects describe {project_id}`), authenticate (`gcloud auth login`), set the target project (`gcloud config set project {project_id}`), and ensure `container.googleapis.com`, `monitoring.googleapis.com`, and `hypercomputecluster.googleapis.com` are enabled.

  • **Supported workloads and GKE versions**: Google Cloud ML Diagnostics only

supports JAX on TPUs (see [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt)). Workload Monitoring is enabled by default, supports the `jobset` and `job` GKE job types, and is compatible with GKE versions `1.36.0-gke.4681000` and later, as stated in [Configure GKE for ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/gke.md.txt). If a workload uses another framework (such as PyTorch) or another custom resource type, `gcloud alpha mldiagnostics` won't list ML runs or monitored events for it. On-demand profiling (Step 4 Path A) additionally requires the cluster setup described in that document.

  • **Required IAM Roles**:
  • Cluster Director Editor (`roles/hypercomputecluster.editor`), the role

listed in the "IAM permissions" section of [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt), for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions)

  • Monitoring Viewer (`roles/monitoring.viewer`) for the PromQL queries in Step

2

  • Kubernetes Engine Viewer (`roles/container.viewer`) for the `kubectl get

nodes` query in Step 3

  • For remediation (`[High Risk]` steps): Kubernetes Engine Cluster Admin

(`roles/container.clusterAdmin`)

  • **Reference Documentation**:
  • [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt)

(Sections: "IAM permissions")

  • [Configure GKE for ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/gke.md.txt)

(Sections: "Set up with gcloud CLI, Google Cloud console, or Terraform", "Manual installation", "Connection-operator")

  • [Workload monitoring with ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/workload-monitoring.md.txt)

(Sections: "Workload Monitoring and analyzers", "HBM Capacity Analyzer", "Host Memory Utilization Analyzer", "CPU Utilization Analyzer", "Access Workload Monitoring information through the API", "List all monitored events", "System Metrics")

  • [Get started with the ML Diagnostics CLI](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/cli.md.txt)

(Sections: "List machine learning runs", "Monitored-events commands", "List profiler targets", "Capture on-demand profiler sessions")

  • [Scale container resource requests and limits](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/vertical-pod-autoscaling.md.txt)

(Sections: "Identify workloads without resource requests or limits")

  • [Auto-repair nodes](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/node-auto-repair.md.txt)

(Sections: "Repair criteria", "Verify node auto-repair is enabled for a Standard node pool", "Enable auto-repair for an existing Standard node pool", "Node auto repair in TPU slice nodes")

  • [Deploy TPU workloads in GKE Standard](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/tpus.md.txt)

(Sections: "Configure auto repair for TPU slice nodes")

  • [Get support](https://docs.cloud.google.com/kubernetes-engine/docs/getting-support.md.txt)

(Sections: "Before you contact support")

> **Read-only rule**: Run read-only diagnostic commands only. Never drain, > delete, or re-create nodes, or run any other command that changes the cluster. > Give the user any fix to apply themselves.

When you recommend a fix, link the doc section that describes it.

---

Analyzer routing overview

By default, Workload Monitoring treats a 15% drop in TPU duty cycle as a performance degradation, raises a `MonitoredEvent` (`type: PERFORMANCE_DEGRADATION`), and runs its analyzers. Consult the [Workload Monitoring and analyzers](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/workload-monitoring.md.txt) section in the official Cloud TPU documentation for the com

Read more
Ships withgoogle-skills

This repository contains Agent Skills for Google products and technologies, including Google Cloud.

Get the whole plugin

Other skills on google-skills.