/gke-tpu-metrics-monitoring
Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating
$ npx -y skills add google/skills --skill gke-tpu-metrics-monitoring --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/gke-tpu-metrics-monitoring
Context preview
The summary Claude sees to decide when to auto-load this skill.
Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating
SKILL.md
gke-tpu-metrics-monitoring.SKILL.mdname: gke-tpu-metrics-monitoring
description: >-
Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE
system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory,
node readiness, multi-host TPU node pool availability, host maintenance or
preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs.
Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.
metadata:
category: CloudObservabilityAndMonitoring
GKE TPU Metrics Monitoring Guide
This skill enables the agent to monitor GKE TPU workloads, nodes, and node pools using GKE system metrics. It helps diagnose if workload interruptions or performance issues are caused by underlying infrastructure.
Step 0: Mandatory Context
Independently gather required context (such as cluster details or node pool names) using available GKE and Cloud tools, or use the provided `{variable}` placeholders:
- `{project_id}`: The GCP Project ID.
- `{cluster_name}`: The GKE Cluster Name.
- `{location}`: The GKE Cluster Location (region or zone).
- `{node_name}`: (Optional) The name of the specific GKE node.
- `{node_pool_name}`: (Optional) The name of the GKE node pool.
---
Diagnostic Steps
Step 1: Verify TPU Runtime Metrics Configuration [Low Risk] [Auto]
Before analyzing runtime metrics, verify that the workload is configured to export them.
- **Action**: Verify that the Pod specification for the TPU workload includes:
- `containerPort: 8431`
- JAX version `0.4.14` or later (if using JAX).
- GKE version is `1.27.4-gke.900` or later.
- GKE System Metrics are enabled on the cluster.
Step 2: Monitor TPU Runtime Metrics [Low Risk] [Auto]
If configured correctly, the following metrics are available in Cloud Monitoring (monitored resources `k8s_node` and `k8s_container`):
- **Container Metrics**:
- `kubernetes.io/container/accelerator/duty_cycle`: Percentage of time over the past sampling period (60 seconds) during which the TensorCores were actively processing on a TPU chip.
- `kubernetes.io/container/accelerator/memory_used`: Amount of accelerator memory allocated in bytes.
- `kubernetes.io/container/accelerator/memory_total`: Total accelerator memory in bytes.
- **Node Metrics**:
- `kubernetes.io/node/accelerator/duty_cycle`
- `kubernetes.io/node/accelerator/memory_used`
- `kubernetes.io/node/accelerator/memory_total`
Step 3: Check Node Status Condition [Low Risk] [Auto]
Query the status condition of GKE nodes (GKE version `1.32.1-gke.1357001` or later).
- **PromQL Query (Check if a specific node is Ready)**:
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", node_name="{node_name}", condition="Ready", status="True"}- **PromQL Query (List nodes with non-Ready conditions that are True)**:
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition!="Ready", status="True"}- **PromQL Query (List nodes that are NOT Ready)**:
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition="Ready", status="False"}- **PromQL Query (Fleet-wide node status)**:
avg by (condition,status)(avg_over_time(kubernetes_io:node_status_condition{monitored_resource="k8s_node"}[5m]))Step 4: Check Node Pool Status [Low Risk] [Auto]
Query the status of multi-host TPU node pools.
- **PromQL Query (Verify if a specific node pool is Running)**:
kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}", node_pool_name="{node_pool_name}", status="Running"}- **PromQL Query (Monitor node pools grouped by status)**:
count by (status)(count_over_time(kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool"}[5m]))_Possible statuses_: `Provisioning`, `Running`, `Error`, `Reconciling`, `Stopping`.
Step 5: Check Node Pool Availability [Low Risk] [Auto]
Query if all nodes in a multi-host TPU node pool are available.
- **PromQL Query (Check availability over time)**:
avg by (node_pool_name)(avg_over_time(kubernetes_io:node_pool_multi_host_available{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[5m]))_Value_: `1` (True, all nodes available) or `0` (False, some nodes unavailable).
Step 6: Analyze Node Interruptions [Low Risk] [Auto]
Query the count of interruptions for GKE nodes.
- **PromQL Query (Breakdown of interruptions and causes)**:
sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node"}[5m]))_Interruption Types_: `TerminationEvent`, `MaintenanceEvent`, `PreemptionEvent`. _Interruption Reasons_: `HostError`, `Eviction`, `AutoRepair`.
- **PromQL Query (Filter for Host Maintenance events)**:
sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node", interruption_reason="HW/SW Maintenance"}[5m]))- **PromQL Query (Interruption count aggregated by node pool)**:
sum by (node_pool_name,interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_pool_interruption_count{monitored_resource="k8s_node_pool", interruption_reason="HW/SW Maintenance", node_pool_name="{node_pool_name}"}[5m]))Step 7: Calculate Recovery and Interruption Metrics [Low Risk] [Auto]
Calculate Mean Time to Recovery (MTTR) and Mean Time Between Interruptions (MTBI) over the last 7 days.
- **PromQL Query (MTTR - Mean Time to Recovery)**:
sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_sum{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[7d])) / sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_count{moniRead more
name: gke-tpu-metrics-monitoring description: >- Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging. metadata: category: CloudObservabilityAndMonitoring
GKE TPU Metrics Monitoring Guide
This skill enables the agent to monitor GKE TPU workloads, nodes, and node pools using GKE system metrics. It helps diagnose if workload interruptions or performance issues are caused by underlying infrastructure.
Step 0: Mandatory Context
Independently gather required context (such as cluster details or node pool names) using available GKE and Cloud tools, or use the provided `{variable}` placeholders:
- `{project_id}`: The GCP Project ID.
- `{cluster_name}`: The GKE Cluster Name.
- `{location}`: The GKE Cluster Location (region or zone).
- `{node_name}`: (Optional) The name of the specific GKE node.
- `{node_pool_name}`: (Optional) The name of the GKE node pool.
---
Diagnostic Steps
Step 1: Verify TPU Runtime Metrics Configuration [Low Risk] [Auto]
Before analyzing runtime metrics, verify that the workload is configured to export them.
- **Action**: Verify that the Pod specification for the TPU workload includes:
- `containerPort: 8431`
- JAX version `0.4.14` or later (if using JAX).
- GKE version is `1.27.4-gke.900` or later.
- GKE System Metrics are enabled on the cluster.
Step 2: Monitor TPU Runtime Metrics [Low Risk] [Auto]
If configured correctly, the following metrics are available in Cloud Monitoring (monitored resources `k8s_node` and `k8s_container`):
- **Container Metrics**:
- `kubernetes.io/container/accelerator/duty_cycle`: Percentage of time over the past sampling period (60 seconds) during which the TensorCores were actively processing on a TPU chip.
- `kubernetes.io/container/accelerator/memory_used`: Amount of accelerator memory allocated in bytes.
- `kubernetes.io/container/accelerator/memory_total`: Total accelerator memory in bytes.
- **Node Metrics**:
- `kubernetes.io/node/accelerator/duty_cycle`
- `kubernetes.io/node/accelerator/memory_used`
- `kubernetes.io/node/accelerator/memory_total`
Step 3: Check Node Status Condition [Low Risk] [Auto]
Query the status condition of GKE nodes (GKE version `1.32.1-gke.1357001` or later).
- **PromQL Query (Check if a specific node is Ready)**:
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", node_name="{node_name}", condition="Ready", status="True"}- **PromQL Query (List nodes with non-Ready conditions that are True)**:
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition!="Ready", status="True"}- **PromQL Query (List nodes that are NOT Ready)**:
kubernetes_io:node_status_condition{monitored_resource="k8s_node", cluster_name="{cluster_name}", condition="Ready", status="False"}- **PromQL Query (Fleet-wide node status)**:
avg by (condition,status)(avg_over_time(kubernetes_io:node_status_condition{monitored_resource="k8s_node"}[5m]))Step 4: Check Node Pool Status [Low Risk] [Auto]
Query the status of multi-host TPU node pools.
- **PromQL Query (Verify if a specific node pool is Running)**:
kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}", node_pool_name="{node_pool_name}", status="Running"}- **PromQL Query (Monitor node pools grouped by status)**:
count by (status)(count_over_time(kubernetes_io:node_pool_status{monitored_resource="k8s_node_pool"}[5m]))_Possible statuses_: `Provisioning`, `Running`, `Error`, `Reconciling`, `Stopping`.
Step 5: Check Node Pool Availability [Low Risk] [Auto]
Query if all nodes in a multi-host TPU node pool are available.
- **PromQL Query (Check availability over time)**:
avg by (node_pool_name)(avg_over_time(kubernetes_io:node_pool_multi_host_available{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[5m]))_Value_: `1` (True, all nodes available) or `0` (False, some nodes unavailable).
Step 6: Analyze Node Interruptions [Low Risk] [Auto]
Query the count of interruptions for GKE nodes.
- **PromQL Query (Breakdown of interruptions and causes)**:
sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node"}[5m]))_Interruption Types_: `TerminationEvent`, `MaintenanceEvent`, `PreemptionEvent`. _Interruption Reasons_: `HostError`, `Eviction`, `AutoRepair`.
- **PromQL Query (Filter for Host Maintenance events)**:
sum by (interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_interruption_count{monitored_resource="k8s_node", interruption_reason="HW/SW Maintenance"}[5m]))- **PromQL Query (Interruption count aggregated by node pool)**:
sum by (node_pool_name,interruption_type,interruption_reason)(sum_over_time(kubernetes_io:node_pool_interruption_count{monitored_resource="k8s_node_pool", interruption_reason="HW/SW Maintenance", node_pool_name="{node_pool_name}"}[5m]))Step 7: Calculate Recovery and Interruption Metrics [Low Risk] [Auto]
Calculate Mean Time to Recovery (MTTR) and Mean Time Between Interruptions (MTBI) over the last 7 days.
- **PromQL Query (MTTR - Mean Time to Recovery)**:
sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_sum{monitored_resource="k8s_node_pool", cluster_name="{cluster_name}"}[7d])) / sum(sum_over_time(kubernetes_io:node_pool_accelerator_times_to_recover_count{moniThis repository contains Agent Skills for Google products and technologies, including Google Cloud. This repository is under active development.
Repo: google/skills
Other skills on google-skills.
- /data-manager-api-audience-ingestion
Guides developers through managing (adding, removing, and clearing) audience members for Google products using the Data Manager API and its associated client libraries. Use this skill when the user wants to upload audience members, remove specific users, or clear/replace an
Open skill - /data-manager-api-event-ingestion
Guides developers through implementing event and conversion ingestion to Google products using the Data Manager API /v1/events/ingest endpoint and its associated client libraries. Use this skill when the user wants to upload offline conversions, enhanced conversions for leads,
Open skill - /data-manager-api-setup
Guides developers through client library installation and authentication setup steps for the Data Manager API. Use this skill when a user is getting started with the Data Manager API and needs to setup their local environment, install the client library, or setup access to the
Open skill - /google-ads-api-account-diagnostics
Diagnoses Google Ads account performance issues such as conversion loss (value or volume), low lead flow/volume, and lost impression share (opportunities) due to ad rank, bids, or budgets. Use when troubleshooting sudden performance drops, analyzing campaign impression share
Open skill - /google-ads-api-mcp-setup
Guides developers through downloading, configuring, and installing the official open-source Google Ads MCP Server. Use this skill when a user wants to connect their AI assistant (such as Gemini, Claude Code, or Cursor) to their Google Ads account to query campaigns or retrieve
Open skill - /google-ads-api-quickstart
Guides developers through Google Ads API quickstart: credential setup, choosing from 6 client libraries/REST, configuring environments, and running a "retrieve campaigns" script. Troubleshoots common setup errors: USER_PERMISSION_DENIED, login_customer_id issues, and
Open skill

