finding-google-skills
Google platform decision and setup guidance, loaded on demand from Google's skill catalog.…
Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system
$ npx -y skills add google/skills --skill gke-ai-troubleshooting-tpu-performance-degradation --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/gke-ai-troubleshooting-tpu-performance-degradationContext preview
The summary Claude sees to decide when to auto-load this skill.
Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system
name: gke-ai-troubleshooting-tpu-performance-degradation description: >- Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system metrics (`kubernetes.io/node/accelerator/*`). Distinguishes hardware and network fabric throttling from workload resource bottlenecks (HBM capacity, host memory, or host CPU saturation). Use when TPU training throughput or duty cycle drops without crashing pods, when `PERFORMANCE_DEGRADATION` monitored events fire, or when triaging slow multi-slice training steps. Don't use for complete multi-slice XLA execution stalls with `HANG_DETECTED` logs (use gke-ai-troubleshooting-tpu-mxla-hang) or pod eviction/interruption restarts (use gke-ai-troubleshooting-jobset-interruption). metadata: version: "1.0.0" category: AiAndMachineLearning
Diagnose and mitigate Cloud TPU training throughput drops and step-time regressions (`15%+` drop in TPU duty cycle) on Google Kubernetes Engine (GKE) by correlating **ML Diagnostics Workload Monitoring** `MonitoredEvent` analyzer reports with **1-minute Cloud Monitoring system metrics** and GKE node topology labels.
---
[Google Cloud SDK](https://cloud.google.com/sdk/docs/install) (`gcloud` with `alpha` component for `gcloud alpha mldiagnostics`) and `kubectl`.
linked (`gcloud billing projects describe {project_id}`), authenticate (`gcloud auth login`), set the target project (`gcloud config set project {project_id}`), and ensure `container.googleapis.com`, `monitoring.googleapis.com`, and `hypercomputecluster.googleapis.com` are enabled.
supports JAX on TPUs (see [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt)). Workload Monitoring is enabled by default, supports the `jobset` and `job` GKE job types, and is compatible with GKE versions `1.36.0-gke.4681000` and later, as stated in [Configure GKE for ML Diagnostics](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/gke.md.txt). If a workload uses another framework (such as PyTorch) or another custom resource type, `gcloud alpha mldiagnostics` won't list ML runs or monitored events for it. On-demand profiling (Step 4 Path A) additionally requires the cluster setup described in that document.
listed in the "IAM permissions" section of [ML Diagnostics platform](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/overview.md.txt), for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions)
2
nodes` query in Step 3
(`roles/container.clusterAdmin`)
(Sections: "IAM permissions")
(Sections: "Set up with gcloud CLI, Google Cloud console, or Terraform", "Manual installation", "Connection-operator")
(Sections: "Workload Monitoring and analyzers", "HBM Capacity Analyzer", "Host Memory Utilization Analyzer", "CPU Utilization Analyzer", "Access Workload Monitoring information through the API", "List all monitored events", "System Metrics")
(Sections: "List machine learning runs", "Monitored-events commands", "List profiler targets", "Capture on-demand profiler sessions")
(Sections: "Identify workloads without resource requests or limits")
(Sections: "Repair criteria", "Verify node auto-repair is enabled for a Standard node pool", "Enable auto-repair for an existing Standard node pool", "Node auto repair in TPU slice nodes")
(Sections: "Configure auto repair for TPU slice nodes")
(Sections: "Before you contact support")
> **Read-only rule**: Run read-only diagnostic commands only. Never drain, > delete, or re-create nodes, or run any other command that changes the cluster. > Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
---
By default, Workload Monitoring treats a 15% drop in TPU duty cycle as a performance degradation, raises a `MonitoredEvent` (`type: PERFORMANCE_DEGRADATION`), and runs its analyzers. Consult the [Workload Monitoring and analyzers](https://docs.cloud.google.com/tpu/docs/ml-diagnostics/workload-monitoring.md.txt) section in the official Cloud TPU documentation for the com
This repository contains Agent Skills for Google products and technologies, including Google Cloud.
Repo: google/skills
Google platform decision and setup guidance, loaded on demand from Google's skill catalog.…
Provides safety-critical validation, guardrails, and data reduction for gcloud CLI operations…
Provides expert guidance on Identity and Access Management (IAM) and authenticating and…
Guides a developer's first steps on Google Cloud, covering account creation, billing setup,…
Searches, retrieves, and synthesizes official Google developer documentation across Google…
Guides developers through managing (adding, removing, and clearing) audience members for…