Skip to content
Development
Skill

/gke-ai-troubleshooting-node-unresponsive-timeout

Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair. Use when nodes stop heartbeating beyond the node auto-repair threshold and pods

GuideBOOST
From plugin
google-skills
21k156 skills1 MCP
Install
$ npx -y skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/gke-ai-troubleshooting-node-unresponsive-timeout

Context preview

The summary Claude sees to decide when to auto-load this skill.

Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair. Use when nodes stop heartbeating beyond the node auto-repair threshold and pods

SKILL.md

gke-ai-troubleshooting-node-unresponsive-timeout.SKILL.md
name: gke-ai-troubleshooting-node-unresponsive-timeout
description: >-
  Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady /
  NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel
  panics, hardware lockups, or disabled node auto-repair. Use when nodes stop
  heartbeating beyond the node auto-repair threshold and pods remain stuck in
  Terminating. Don't use for healthy nodes, pod-only application crashes, or
  routine GKE upgrades.
metadata:
  version: "1.0.0"
  category: Containers

Troubleshoot unresponsive GKE TPU and GPU nodes (`NodeStatusUnknown`)

When the Compute Engine host of a TPU or GPU node has a fatal hardware error, kernel panic, or non-maskable interrupt (NMI) lockup, the guest OS stops responding. The kubelet can no longer send heartbeats, so the node `Ready` condition becomes `Unknown` with `Reason: NodeStatusUnknown` (`Kubelet stopped posting node status.`). If node auto-repair is disabled on the node pool, GKE doesn't repair the node. The node can stay `NotReady`, and pods on it can stay in `Terminating`, which blocks multi-host `JobSet` workloads from recovering.

Prerequisites

  • **Tools**: Install the

[Google Cloud SDK](https://cloud.google.com/sdk/docs/install) (`gcloud`) and `kubectl`.

  • **Cloud Billing & Project Configuration**: Verify an active billing account is

linked (`gcloud billing projects describe {project_id}`), authenticate (`gcloud auth login`), set the target project (`gcloud config set project {project_id}`), and ensure `container.googleapis.com`, `compute.googleapis.com`, `logging.googleapis.com`, and `monitoring.googleapis.com` are enabled.

  • **Required IAM Roles**:
  • Kubernetes Engine Viewer (`roles/container.viewer`)
  • Compute Viewer (`roles/compute.viewer`)
  • Logs Viewer (`roles/logging.viewer`)
  • Monitoring Viewer (`roles/monitoring.viewer`)
  • For remediation (`[High Risk]` steps): Kubernetes Engine Cluster Admin

(`roles/container.clusterAdmin`)

  • **Documentation**:
  • [Troubleshoot nodes with the NotReady status in GKE](https://docs.cloud.google.com/kubernetes-engine/docs/troubleshooting/node-notready.md.txt)

(Sections: "Check the node's status and conditions", "Confirm node preemption", "Verify that the node has recovered")

  • [Deploy TPU workloads in GKE Standard](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/tpus.md.txt)

(Sections: "Monitor health metrics for TPU nodes and node pools", "Configure auto repair for TPU slice nodes")

  • [Troubleshoot OOM events](https://docs.cloud.google.com/kubernetes-engine/docs/troubleshooting/oom-events.md.txt)
  • [View GKE logs](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/view-logs.md.txt)

(Sections: "System logs")

  • [Viewing serial port output](https://docs.cloud.google.com/compute/docs/troubleshooting/viewing-serial-port-output.md.txt)
  • [Troubleshoot Linux VM boot issues due to kernel panic](https://docs.cloud.google.com/compute/docs/troubleshooting/kernel-panic.md.txt)
  • [Auto-repair nodes](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/node-auto-repair.md.txt)

(Sections: "Settings for Autopilot and Standard", "Verify node auto-repair is enabled for a Standard node pool", "Get information about recent automated repair events", "Enable auto-repair for an existing Standard node pool", "Repair criteria", "Node repair process", "Node auto repair in TPU slice nodes")

  • [Troubleshooting VM shutdowns and reboots](https://docs.cloud.google.com/compute/docs/troubleshooting/troubleshooting-reboots.md.txt)

(Sections: "Querying Cloud Audit Logs", "Reviewing Cloud Audit Logs")

> **Read-only rule**: Run read-only diagnostic commands only. Never drain, > delete, or re-create nodes, or run any other command that changes the cluster. > Give the user any fix to apply themselves.

When you recommend a fix, link the doc section that describes it.

---

Diagnostic workflow

Step 0: Collect context and set the investigation window `[Low Risk]`

Collect the target parameters. By default, query a 60-minute window `[T - 30m, T

  • 30m]` around `{issue_time}`:
  • `{project_id}`: Google Cloud project ID
  • `{cluster_name}`: GKE cluster name
  • `{location}`: Cluster region or zone
  • `{nodepool_name}`: Target TPU or GPU node pool name
  • `{node_name}`: Unresponsive GKE node name (and its Compute Engine `{zone}`)
  • `{issue_time}`: Incident timestamp in RFC3339 UTC
  • `{start_time}`: `{issue_time} - 30m`
  • `{end_time}`: `{issue_time} + 30m`

---

Step 1: Verify the `NodeStatusUnknown` heartbeat timeout `[Low Risk]`

1. **Check Kubernetes node conditions**: To inspect the node status and verify whether the `Ready` condition is `Unknown` with `Reason: NodeStatusUnknown` (`Kubelet stopped posting node status.`), follow the instructions in the section [Check the node's status and conditions](https://docs.cloud.google.com/kubernetes-engine/docs/troubleshooting/node-notready.md.txt). 2. **Query Cloud Logging (read-only LQL)**: Query `k8s_node` and `k8s_cluster` logs across `[{start_time}, {end_time}]` to confirm when the control plane lost heartbeat contact with `{node_name}`:

(resource.type="k8s_node" OR resource.type="k8s_cluster")
resource.labels.cluster_name="{cluster_name}"
("{node_name}" AND ("NodeNotReady" OR "NodeStatusUnknown" OR "Kubelet stopped posting node status"))
timestamp >= "{start_time}" AND timestamp <= "{end_time}"

3. **Query Cloud Monitoring (read-only PromQL)**: Correlate the duration of the `Unknown` state using the GKE system metric `kubernetes.io/node/status_condition` (`kubernetes_io:node_status_condition`, GKE `1.32.1-gke.1357001+`) documented in [Monitor health metrics for TPU nodes and node pools](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/tpus.md.txt), filtered by `condition="Ready"` and `status="Unknown"`:

kubernetes_io:node_status_condition{
  monitored_resou
Read more
Ships withgoogle-skills

This repository contains Agent Skills for Google products and technologies, including Google Cloud.

Get the whole plugin

Other skills on google-skills.