finding-google-skills
Locates and loads the right Google product skill on demand from a remote catalog index, instead of preloading every skill. Use at the START of any request…
Runs batch and HPC workloads on GKE, utilizing job queues and parallel processing. Use when running GKE batch jobs, configuring GKE HPC, or setting up GKE job queues. Don't use for standard web application deployments (use gke-app-onboarding instead).
$ npx -y skills add google/skills --skill gke-batch-hpc --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/gke-batch-hpcContext preview
The summary Claude sees to decide when to auto-load this skill.
Runs batch and HPC workloads on GKE, utilizing job queues and parallel processing. Use when running GKE batch jobs, configuring GKE HPC, or setting up GKE job queues. Don't use for standard web application deployments (use gke-app-onboarding instead).
name: gke-batch-hpc description: >- Runs batch and HPC workloads on GKE, utilizing job queues and parallel processing. Use when running GKE batch jobs, configuring GKE HPC, or setting up GKE job queues. Don't use for standard web application deployments (use gke-app-onboarding instead). metadata: category: Containers
This reference covers running batch processing and high-performance computing (HPC) workloads on GKE.
> **MCP Tools:** `apply_k8s_manifest`, `get_k8s_resource`, > `describe_k8s_resource`, `get_k8s_logs`, `delete_k8s_resource`, > `list_k8s_events`
apiVersion: batch/v1
kind: Job
metadata:
name: batch-job
spec:
parallelism: 10
completions: 100
backoffLimit: 3
template:
spec:
containers:
- name: worker
image: <IMAGE>
resources:
requests:
cpu: "1"
memory: "2Gi"
restartPolicy: NeverThe golden path enables JobSet monitoring (`JOBSET` in monitoringConfig).
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: training-job
spec:
replicatedJobs:
- name: workers
replicas: 4
template:
spec:
parallelism: 1
completions: 1
template:
spec:
containers:
- name: worker
image: <IMAGE>
resources:
requests:
cpu: "4"
memory: "8Gi"Kueue manages job scheduling and resource allocation for batch workloads:
# Install Kueue kubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/latest/download/manifests.yaml
# Define a ClusterQueue
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
name: batch-queue
spec:
namespaceSelector: {}
resourceGroups:
- coveredResources: ["cpu", "memory"]
flavors:
- name: default
resources:
- name: "cpu"
nominalQuota: 100
- name: "memory"
nominalQuota: "200Gi"
---
# Allow a namespace to use the queue
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata:
name: batch-local
namespace: batch-jobs
spec:
clusterQueue: batch-queueFor tightly-coupled HPC workloads that need low-latency inter-node communication:
# Standard clusters: create node pool with compact placement gcloud container node-pools create hpc-pool \ --cluster <CLUSTER_NAME> --region <REGION> \ --machine-type c3-standard-44 \ --placement-type COMPACT \ --num-nodes 8 \ --enable-autoscaling --min-nodes 0 --max-nodes 16 \ --quiet
Use the MPI Operator for MPI-based HPC applications:
# Install MPI Operator kubectl apply -f https://raw.githubusercontent.com/kubeflow/mpi-operator/master/deploy/v2beta1/mpi-operator.yaml
apiVersion: kubeflow.org/v2beta1
kind: MPIJob
metadata:
name: hpc-simulation
spec:
slotsPerWorker: 4
mpiReplicaSpecs:
Launcher:
replicas: 1
template:
spec:
containers:
- name: launcher
image: <MPI_IMAGE>
command: ["mpirun", "-np", "32", "./simulation"]
resources:
requests:
cpu: "1"
memory: "2Gi"
limits:
cpu: "2"
memory: "4Gi"
Worker:
replicas: 8
template:
spec:
containers:
- name: worker
image: <MPI_IMAGE>
resources:
requests:
cpu: "4"
memory: "8Gi"
limits:
cpu: "8"
memory: "16Gi"Batch workloads are ideal Spot VM candidates (interruptible, can checkpoint). Use a ComputeClass with Spot-first priority and `activeMigration` to return to Spot when available. See the `gke-compute-classes` skill for the Spot-with-fallback pattern.
For batch clusters, allow node pools to scale to zero when no jobs are running:
scheduled
memory, and optionally GPU/TPU) for all batch/HPC manifests. This is critical for Kueue admission, autoscaling, and preventing resource starvation in the cluster.
VMs/TPUs, advise using **GKE maintenance exclusions** to block automatic cluster upgrades/reboots during the active training window to minimize unnecessary preemption.
distributed MPI applications via the `MPIJob` custom resource.
sharing; use **JobSet** for multi-component tightly coupled workloads.
application-level checkpointing (e.g., using Orbax or PyTorch checkpointing) to survive Spot VM preemption.
This repository contains Agent Skills for Google products and technologies, including Google Cloud.
Repo: google/skills
Locates and loads the right Google product skill on demand from a remote catalog index, instead of preloading every skill. Use at the START of any request…
Provides safety-critical validation, guardrails, and data reduction for gcloud CLI operations across Google Cloud Platform (GCP) services and infrastructure.…
Provides expert guidance on authenticating and authorizing to Google Cloud services and APIs, covering human users, service identities, Application Default…
Guides a developer's first steps on Google Cloud, covering account creation, billing setup, project management, and deploying a first resource. Use when a new…
Searches, retrieves, and synthesizes official Google developer documentation across Google Cloud, AI/Gemini, Android, Chrome, Web, Flutter, Go, Firebase, and…
Guides developers through managing (adding, removing, and clearing) audience members for Google products using the Data Manager API and its associated client…