finding-google-skills
Locates and loads the right Google product skill on demand from a remote catalog index, instead of preloading every skill. Use at the START of any request…
Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Use when configuring GKE workload reliability, setting up PDBs, or configuring GKE health probes (liveness, readiness, startup). Don't use for disaster recovery setup or full cluster
$ npx -y skills add google/skills --skill gke-reliability --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/gke-reliabilityContext preview
The summary Claude sees to decide when to auto-load this skill.
Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Use when configuring GKE workload reliability, setting up PDBs, or configuring GKE health probes (liveness, readiness, startup). Don't use for disaster recovery setup or full cluster
name: gke-reliability description: >- Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Use when configuring GKE workload reliability, setting up PDBs, or configuring GKE health probes (liveness, readiness, startup). Don't use for disaster recovery setup or full cluster backups (use gke-backup-dr instead). metadata: category: Containers
This reference covers high availability and reliability configuration for GKE clusters and workloads.
> **MCP Tools:** `get_cluster`, `get_k8s_resource`, `describe_k8s_resource`, > `apply_k8s_manifest`, `list_k8s_events`
| Setting | Golden Path Value | Notes | | ---------------- | --------------------- | -------------------------------- | | Cluster type | Regional (4 zones: | Control plane replicated across | : : us-central1-a/b/c/f) : zones : | Upgrade strategy | SURGE (`maxSurge: 1`) | Rolling upgrades with extra | : : : capacity : | Auto-repair | `true` | Unhealthy nodes replaced | : : : automatically : | Auto-upgrade | `true` | Nodes follow control plane | : : : version : | Release channel | REGULAR | Balanced freshness and stability | | Stateful HA | Enabled | Leader election for stateful | : : : workloads :
# MCP (preferred) get_cluster(name="projects/<PROJECT>/locations/<REGION>/clusters/<CLUSTER>", readMask="location,locations,nodePools.locations") # gcloud fallback gcloud container clusters describe <CLUSTER> --region <REGION> \ --format="json(location, locations)" \ --quiet
regional
PDBs ensure minimum pod availability during voluntary disruptions (node upgrades, autoscaler scale-down).
**Check existing PDBs:**
# MCP (preferred) get_k8s_resource(parent="...", resourceType="poddisruptionbudget") # kubectl fallback kubectl get pdb --all-namespaces
**Create PDB:**
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: my-app-pdb
namespace: default
spec:
minAvailable: 2 # Or use maxUnavailable: 1
selector:
matchLabels:
app: my-app> Every production Deployment with 2+ replicas should have a PDB.
Every production container should have liveness and readiness probes. Startup probes are recommended for slow-starting apps.
**Check existing probes:**
# MCP (preferred) describe_k8s_resource(parent="...", resourceType="deployment", name="<APP>", namespace="<NS>") # kubectl fallback kubectl get deployment <APP> -n <NS> -o yaml | grep -E "livenessProbe|readinessProbe|startupProbe"
**Recommended probe configuration:**
spec:
containers:
- name: app
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 15
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
readinessProbe:
httpGet:
path: /readyz
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
startupProbe: # For slow-starting apps
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 30 # 30 * 5s = 150s max startup timepremature restarts)
Ensure applications handle `SIGTERM` and drain in-flight requests:
spec:
terminationGracePeriodSeconds: 30 # Default; increase for long-running requests
containers:
- name: app
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 5"] # Allow LB to deregisterDistribute pods across zones and nodes to survive failures:
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: my-app
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: my-appacross zones
don't block scheduling
| Workload Type | Minimum Replicas | Reason | | -------------------- | -------------------- | ------------------------------ | | Stateless web/API | 2 | Survive single pod/node | : : : failure : | Critical services | 3 | Survive zone failure with zone | : : : spread : | Stateful (databases) | 3 (with replication) | Application-level quorum | | Batch/jobs | 1 | Ephemeral by nature |
1. **Regional clusters for production**: Always use regional clusters to survive zone failures. 2. **PDBs fo
This repository contains Agent Skills for Google products and technologies, including Google Cloud.
Repo: google/skills
Locates and loads the right Google product skill on demand from a remote catalog index, instead of preloading every skill. Use at the START of any request…
Provides safety-critical validation, guardrails, and data reduction for gcloud CLI operations across Google Cloud Platform (GCP) services and infrastructure.…
Provides expert guidance on authenticating and authorizing to Google Cloud services and APIs, covering human users, service identities, Application Default…
Guides a developer's first steps on Google Cloud, covering account creation, billing setup, project management, and deploying a first resource. Use when a new…
Searches, retrieves, and synthesizes official Google developer documentation across Google Cloud, AI/Gemini, Android, Chrome, Web, Flutter, Go, Firebase, and…
Guides developers through managing (adding, removing, and clearing) audience members for Google products using the Data Manager API and its associated client…