amazon-location-servic…
Integrates Amazon Location Service APIs for AWS applications. Use this skill when users want to add maps (interactive MapLibre or static images); geocode…
Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters. Scope mirrors the HyperPod troubleshooting guide. Invoke when the user reports a Slurm node stuck in down/drain, "Node unexpectedly rebooted" after auto-repair, slurmd
$ npx -y skills add awslabs/agent-plugins --skill hyperpod-slurm-debugger --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/hyperpod-slurm-debuggerContext preview
The summary Claude sees to decide when to auto-load this skill.
Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters. Scope mirrors the HyperPod troubleshooting guide. Invoke when the user reports a Slurm node stuck in down/drain, "Node unexpectedly rebooted" after auto-repair, slurmd
name: hyperpod-slurm-debugger description: Diagnostic-only skill for Slurm scheduler and node-daemon issues on Amazon SageMaker HyperPod Slurm clusters. Scope mirrors the HyperPod troubleshooting guide. Invoke when the user reports a Slurm node stuck in down/drain, "Node unexpectedly rebooted" after auto-repair, slurmd not running, jobs stuck PENDING with REASON=Resources while sinfo shows idle nodes, jobs stuck COMPLETING after node replacement, GRES/GPU counts wrong, scontrol ping failing, slurmctld unresponsive, an Action:Reboot/Replace request that did not trigger HyperPod auto-recovery, or auto-resume not restarting a job. Also triggers on "drain before reboot", "diagnose a Slurm node", "investigate stuck jobs." metadata: version: "0.0.1"
Diagnostic-only. Identify and classify Slurm scheduler and node-daemon issues on HyperPod Slurm clusters. Do not run, recommend, or print any state-mutating command. For remediation, link to the official AWS or Slurm documentation.
Invoke when the user reports any of the symptoms in the [decision table](#decision-table).
drifts the live state from the IaC plan.
Canonical recovery URLs: [references/slurm-details.md → Authoritative recovery documentation](references/slurm-details.md).
installed locally.
returns empty stdout intermittently with `Cannot perform start session: EOF` and every check silently misreports. Install: `expect` package on Amazon Linux / RHEL / Debian / Ubuntu / macOS. Script exits at prerequisite check if missing.
Ask the user for:
1. HyperPod cluster name (not Slurm partition name). 2. AWS region. 3. Optional: a specific Slurm node name.
aws sagemaker describe-cluster --cluster-name <NAME/ARN> --region <REGION> \ --query 'Orchestrator' --output json
If `Orchestrator.Eks` is present, stop. Route per [When NOT to invoke](#when-not-to-invoke).
bash scripts/slurm-diagnose.sh --cluster <NAME> --region <REGION> # Scope to a node: bash scripts/slurm-diagnose.sh --cluster <NAME> --region <REGION> --node <SLURM_NODE>
Relay the script output to the user verbatim.
For each finding, look up the section in the [decision table](#decision-table) and link the user to the corresponding AWS / Slurm doc. Do not type out remediation commands.
| Symptom (`sinfo -o "%N %T %30E"` or script finding) | Section | | ----------------------------------------------------------- | ------------------------------------------------------ | | Node state = `down` or `down*`, reason other than below | [A: Node Down](#a-node-down) | | Node state = `down*`, Reason = `Node unexpectedly rebooted` | [B: Unexpected Reboot](#b-unexpected-reboot) | | Jobs `PENDING` with `REASON=Resources` while nodes are idle | [C: Controller State](#c-controller-state) | | Jobs stuck `COMPLETING` after node replacement | [C: Controller State](#c-controller-state) | | `scontrol ping` returns `DOWN` for the controller | [C: Controller State](#c-controller-state) | | GRES (GPU) counts incorrect or not released | [C: Controller State](#c-controller-state) | | `state=fail` issued but no recovery occurred | [D: Action Reason Mismatch](#d-action-reason-mismatch) | | Accounting errors or RPC errors mentioning `dbd` | [C: Controller State](#c-controller-state) (slurmdbd) | | `slurm.conf` edited; new partitions or nodes not visible | [C: Controller State](#c-controller-state) (config) | | Job exited on a hardware failure but did not restart | [E: Auto-resume](#e-auto-resume) |
| Behavior | Default | Override | | -------------------- | -------------------------------------------------------------------------------------------------- | -------------------------- | | Mode | read-only — always; no remediation flag exists | n/a | | Region | `$AWS_DEFAULT_REGION`, falling back to `us-east-1` | `--region <R>` | | Scope | all nodes in `down` / `drain` / `fail` / "unexpectedly rebooted" | `--node <SLURM_NODE_NAME>` | | Output | colorized terminal | `--no-color` | | SSM target format | `sagemaker-cluster:<clusterId>_<instanceGroupName>
Read this in other languages: 日本語 Generative AI can make mistakes. You should consider reviewing all output and costs generated by your chosen AI model and agentic coding assistant. See AWS Responsible AI Policy.
Integrates Amazon Location Service APIs for AWS applications. Use this skill when users want to add maps (interactive MapLibre or static images); geocode…
Build and deploy full-stack web and mobile apps with AWS Amplify Gen2
Build, manage, and operate APIs with Amazon API Gateway (REST, HTTP, and WebSocket). Triggers on phrases like: API Gateway, REST API, HTTP API, WebSocket API,…
Build resilient, long-running, multi-step applications with AWS Lambda durable functions with automatic state persistence, retry logic, and orchestration for…
Evaluate, configure, and migrate workloads to AWS Lambda Managed Instances (LMI). Triggers on: Lambda Managed Instances, LMI, capacity provider,…
Build, run, debug, and operate applications on AWS Lambda MicroVMs — Firecracker-isolated, snapshot-resumable serverless compute environments that run inside a…