Skip to content
Development
Command

/debug-pod

Debug a failing or unhealthy Kubernetes pod by analyzing events, logs, and configuration.

From plugin
rohitg00-claude-code-toolkit
2.5k199 skills138 agents199 commands
Install
$ npx -y skills add rohitg00/awesome-claude-code-toolkit --agent claude-code

How it fires

How this command gets triggered: by you, by Claude, or both.

  • Fires itselfClaude auto-loads it when your prompt matches the work.
  • You can call itInvoke it directly when you want it.
  • Slash command/debug-pod

Context preview

What this command does when you run it.

Debug a failing or unhealthy Kubernetes pod by analyzing events, logs, and configuration.

Command definition

debug-pod.md

Debug a failing or unhealthy Kubernetes pod by analyzing events, logs, and configuration.

Steps

1. Get pod status: `kubectl get pod <name> -n <namespace> -o wide`. 2. Describe the pod for events and conditions: `kubectl describe pod <name> -n <namespace>`. 3. Analyze the pod state:

  • **Pending**: Check node resources, scheduling constraints, PVC binding.
  • **CrashLoopBackOff**: Check container logs for startup errors.
  • **ImagePullBackOff**: Verify image name, tag, and registry credentials.
  • **OOMKilled**: Check memory limits vs actual usage.
  • **Running but unhealthy**: Check probe configuration and endpoints.

4. Fetch container logs: `kubectl logs <pod> -n <ns> --previous` for crash logs. 5. Check resource usage: `kubectl top pod <name> -n <namespace>`. 6. Verify configuration:

  • ConfigMaps and Secrets are mounted correctly.
  • Environment variables are set.
  • Service account has required permissions.

7. Suggest fixes based on the diagnosis.

Format

Pod: <name> in <namespace>
Status: <status>
Restarts: <count>
Node: <node-name>

Diagnosis:
  Root cause: <description>
  Evidence: <log lines or events>

Fix:
  1. <action to take>
  2. <verification command>

Rules

  • Always check events first; they often reveal the root cause immediately.
  • Fetch logs from the previous container instance for crash analysis.
  • Check node-level issues if multiple pods on the same node are affected.
  • Verify DNS resolution if the pod cannot reach other services.
  • Check RBAC permissions if the pod gets authorization errors.
Read more
Ships withrohitg00-claude-code-toolkit

The most comprehensive toolkit for Claude Code -- 135 agents, 35 curated skills (+400,000 via SkillKit), 42 commands, 176+ plugins, 20 hooks, 15 rules, 7 templates, 15 MCP configs, 26 companion apps, 53 ecosystem entries, and more.

Get the whole plugin