Skip to content
Development
Skill

/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring

Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Use when checking TPU slice lifecycle states, troubleshooting slice provisioning failures, validating single-slice or multi-slice (JobSet) workload manifests, or safely patching stuck finalizers and

From plugin
google-skills
20k146 skills1 MCP
Install
$ npx -y skills add google/skills --skill gke-ai-troubleshooting-tpu-dynamic-slices-monitoring --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring

Context preview

The summary Claude sees to decide when to auto-load this skill.

Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Use when checking TPU slice lifecycle states, troubleshooting slice provisioning failures, validating single-slice or multi-slice (JobSet) workload manifests, or safely patching stuck finalizers and

SKILL.md

gke-ai-troubleshooting-tpu-dynamic-slices-monitoring.SKILL.md
name: gke-ai-troubleshooting-tpu-dynamic-slices-monitoring
description: >-
  Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Use when checking TPU slice lifecycle states, troubleshooting slice provisioning failures, validating single-slice or multi-slice (JobSet) workload manifests, or safely patching stuck finalizers and disabling the slice controller. Don't use for generic GKE cluster node pool creation or standard non-TPU workload management (use gke-basics or gke-cluster-creation instead).
metadata:
  category: Containers

GKE TPU Dynamic Slices Monitoring & Management

Monitors the status of TPU Slice custom resources, troubleshoots provisioning failures, validates workload manifests on dynamic slices, and performs cleanups.

Prerequisites

  • Cloud Logging enabled for the project.
  • `kubectl` and `gcloud` CLIs configured to access the GKE cluster.

Diagnostic Workflow

Step 0: Context Acquisition & Time Window Definition

Gather project, cluster, and slice context using cluster tools or the following parameters:

  • **Project ID**: `{project_id}` (e.g., `my-gcp-project`)
  • **Cluster Name**: `{cluster_name}` (e.g., `tpu-cluster`)
  • **Region/Zone**: `{location}` (e.g., `us-central1-a`)
  • **Slice Name**: `{slice_name}` (e.g., `test-slice`)
  • **Issue Time**: `{timestamp}` (Optional; default to the last 30 minutes

window `[T - 30m]` to `[T + 30m]`)

--------------------------------------------------------------------------------

Step 1: Describe the Slice Custom Resource [Low Risk]

When asked to inspect, troubleshoot, or check a slice status, immediately execute `kubectl describe slice {slice_name}` using available cluster tools to perform the inspection. Parse the resulting `Status.Conditions` output against the condition table below to diagnose the exact state and provide concrete recommendations.

  • **Command**:
    kubectl describe slice {slice_name}

State & Reason Analysis

Analyze the `Status.Conditions` (especially `Type: Ready` and its `Reason` and `Status`):

| Lifecycle State / Reason | Meaning | Recommended Action | | :--- | :--- | :--- | | **`SliceNotCreated`** | GKE Slice Controller is initializing the slice and performing resource checks. | Wait a few minutes and re-check slice status. | | **`SliceCreationFailed`** | Prerequisites validation failed (e.g., selected nodes don't exist, nodes are already used by another slice, or the topology doesn't match the number of partitions). | Verify selected nodes exist, are unallocated, and topology matches partition count. | | **`ACTIVATING`** | GKE is actively forming and provisioning the TPU slice. | Monitor node provisioning. | | **`ACTIVE`** | The TPU slice is successfully formed and ready to host workloads. | Proceed to deploy or check workloads. | | **`ACTIVE_DEGRADED`** | The slice is usable, but one or more sub-blocks are degraded. | Monitor workload logs for interconnect or device errors. Check faulty node VMs. | | **`FAILED`** | GKE failed to form the TPU slice (e.g., selected nodes are not part of the same reservation block). | Ensure all selected nodes belong to the same reservation block. | | **`DEACTIVATING`** | The slice is dismantling (triggered by user deletion or a critical systemic failure). | Wait for dismantling to finish, or patch finalizers if stuck. | | **`INCOMPLETE`** | The terminal phase before the Slice CR is deleted from the cluster. | No action required; the resource will be removed shortly. |

Provisioning Failure Troubleshooting Checklist

When investigating slice creation or provisioning failures (`SliceCreationFailed` or `FAILED`), perform the following verification steps:

1. **Node Existence & Allocation Check**: Verify that the selected TPU nodes exist in the cluster and are not already allocated to another slice (`kubectl get nodes -l cloud.google.com/gke-tpu-slice`, `kubectl get slice -A`). 2. **Topology Alignment**: Confirm that the partition count matches the requested topology dimensions (e.g. topology `2x2` requires 4 nodes). 3. **Reservation Block Alignment Check**: Confirm that all selected TPU nodes belong to the same reservation and reservation block.

--------------------------------------------------------------------------------

Step 2: Verify Workload Specification [Low Risk]

Ensure workload manifests are configured correctly to target the dynamic slice.

1. Single-Slice Workload Requirements

Check that the Pod template contains the following annotations and selectors:

  • **Annotations**:
  • `cloud.google.com/gke-tpu-slice-topology: "{topology}"` (e.g.,

`"4x4x4"`)

  • **NodeSelector**:
  • `cloud.google.com/gke-tpu-topology: "{topology}"` (e.g., `"4x4x4"`)
  • `cloud.google.com/gke-tpu-accelerator: "{accelerator_type}"` (e.g.,

`"tpu7x"`)

  • `cloud.google.com/gke-tpu-slice: "{slice_name}"` (e.g., `"test-slice"`)

2. Multi-Slice (JobSet) Workload Requirements

If deploying a multi-slice JobSet, verify:

  • **JobSet Annotation**:
  • `alpha.jobset.sigs.k8s.io/exclusive-topology:

cloud.google.com/gke-tpu-slice`

  • **Pod Template Annotations**:
  • `cloud.google.com/gke-tpu-slice-topology: "{topology}"`
  • **Pod Template NodeSelector**:
  • `cloud.google.com/gke-tpu-topology: "{topology}"`
  • `cloud.google.com/gke-tpu-accelerator: "{accelerator_type}"`
  • *Note: Do NOT manually specify `cloud.google.com/gke-tpu-slice` in the

nodeSelector; JobSet handles slice assignment automatically.*

--------------------------------------------------------------------------------

Resolution & Management Workflow

Resolution 1: Force Delete a Stuck Slice [High Risk]

If a slice is stuck in `DEACTIVATING` or deletion hangs indefinitely due to stuck finalizers:

1. **Identify Cause**: Explain that finalizers on the slice resource (`metadata.finalizers`) are preventing Kubernetes from completing resou

Read more
Ships withgoogle-skills

This repository contains Agent Skills for Google products and technologies, including Google Cloud.

Get the whole plugin

Other skills on google-skills.