finding-google-skills
Google platform decision and setup guidance, loaded on demand from Google's skill catalog.…
Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair. Use when nodes stop heartbeating beyond the node auto-repair threshold and pods
$ npx -y skills add google/skills --skill gke-ai-troubleshooting-node-unresponsive-timeout --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/gke-ai-troubleshooting-node-unresponsive-timeoutContext preview
The summary Claude sees to decide when to auto-load this skill.
Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair. Use when nodes stop heartbeating beyond the node auto-repair threshold and pods
name: gke-ai-troubleshooting-node-unresponsive-timeout
description: >-
Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady /
NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel
panics, hardware lockups, or disabled node auto-repair. Use when nodes stop
heartbeating beyond the node auto-repair threshold and pods remain stuck in
Terminating. Don't use for healthy nodes, pod-only application crashes, or
routine GKE upgrades.
metadata:
version: "1.0.0"
category: ContainersWhen the Compute Engine host of a TPU or GPU node has a fatal hardware error, kernel panic, or non-maskable interrupt (NMI) lockup, the guest OS stops responding. The kubelet can no longer send heartbeats, so the node `Ready` condition becomes `Unknown` with `Reason: NodeStatusUnknown` (`Kubelet stopped posting node status.`). If node auto-repair is disabled on the node pool, GKE doesn't repair the node. The node can stay `NotReady`, and pods on it can stay in `Terminating`, which blocks multi-host `JobSet` workloads from recovering.
[Google Cloud SDK](https://cloud.google.com/sdk/docs/install) (`gcloud`) and `kubectl`.
linked (`gcloud billing projects describe {project_id}`), authenticate (`gcloud auth login`), set the target project (`gcloud config set project {project_id}`), and ensure `container.googleapis.com`, `compute.googleapis.com`, `logging.googleapis.com`, and `monitoring.googleapis.com` are enabled.
(`roles/container.clusterAdmin`)
(Sections: "Check the node's status and conditions", "Confirm node preemption", "Verify that the node has recovered")
(Sections: "Monitor health metrics for TPU nodes and node pools", "Configure auto repair for TPU slice nodes")
(Sections: "System logs")
(Sections: "Settings for Autopilot and Standard", "Verify node auto-repair is enabled for a Standard node pool", "Get information about recent automated repair events", "Enable auto-repair for an existing Standard node pool", "Repair criteria", "Node repair process", "Node auto repair in TPU slice nodes")
(Sections: "Querying Cloud Audit Logs", "Reviewing Cloud Audit Logs")
> **Read-only rule**: Run read-only diagnostic commands only. Never drain, > delete, or re-create nodes, or run any other command that changes the cluster. > Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
---
Collect the target parameters. By default, query a 60-minute window `[T - 30m, T
---
1. **Check Kubernetes node conditions**: To inspect the node status and verify whether the `Ready` condition is `Unknown` with `Reason: NodeStatusUnknown` (`Kubelet stopped posting node status.`), follow the instructions in the section [Check the node's status and conditions](https://docs.cloud.google.com/kubernetes-engine/docs/troubleshooting/node-notready.md.txt). 2. **Query Cloud Logging (read-only LQL)**: Query `k8s_node` and `k8s_cluster` logs across `[{start_time}, {end_time}]` to confirm when the control plane lost heartbeat contact with `{node_name}`:
(resource.type="k8s_node" OR resource.type="k8s_cluster")
resource.labels.cluster_name="{cluster_name}"
("{node_name}" AND ("NodeNotReady" OR "NodeStatusUnknown" OR "Kubelet stopped posting node status"))
timestamp >= "{start_time}" AND timestamp <= "{end_time}"3. **Query Cloud Monitoring (read-only PromQL)**: Correlate the duration of the `Unknown` state using the GKE system metric `kubernetes.io/node/status_condition` (`kubernetes_io:node_status_condition`, GKE `1.32.1-gke.1357001+`) documented in [Monitor health metrics for TPU nodes and node pools](https://docs.cloud.google.com/kubernetes-engine/docs/how-to/tpus.md.txt), filtered by `condition="Ready"` and `status="Unknown"`:
kubernetes_io:node_status_condition{
monitored_resouThis repository contains Agent Skills for Google products and technologies, including Google Cloud.
Repo: google/skills
Google platform decision and setup guidance, loaded on demand from Google's skill catalog.…
Provides safety-critical validation, guardrails, and data reduction for gcloud CLI operations…
Provides expert guidance on Identity and Access Management (IAM) and authenticating and…
Guides a developer's first steps on Google Cloud, covering account creation, billing setup,…
Searches, retrieves, and synthesizes official Google developer documentation across Google…
Guides developers through managing (adding, removing, and clearing) audience members for…