
GKE Reliability
FreeEnhance the reliability of your GKE workloads.
Free · Opens the source repo
What GKE Reliability does
The GKE Reliability skill provides a comprehensive reference for configuring high availability and reliability in Google Kubernetes Engine (GKE) clusters and workloads. It focuses on essential components such as Pod Disruption Budgets (PDBs), health probes, and topology spread constraints to ensure that applications remain resilient during maintenance and unexpected disruptions. This skill is particularly useful for developers and DevOps engineers who are responsible for deploying and maintaining applications in a Kubernetes environment.
By utilizing this skill, users can verify cluster high availability, create and manage PDBs, and set up health probes for their applications. The skill includes workflows that guide users through the process of checking existing configurations and implementing best practices for production deployments. For instance, the recommended configurations for liveness and readiness probes help ensure that applications are only receiving traffic when they are ready, thereby improving user experience and reliability.
Moreover, the skill emphasizes the importance of distributing workloads across multiple zones and nodes to withstand failures. By implementing topology spread constraints, users can effectively manage the placement of pods to enhance fault tolerance. The skill also provides guidelines on the minimum number of replicas needed for different types of workloads, ensuring that applications can handle failures without significant downtime.
Overall, the GKE Reliability skill is a valuable resource for anyone looking to improve the stability and resilience of their GKE deployments. It equips users with the knowledge and tools necessary to configure their workloads for optimal reliability, making it an essential addition to any Kubernetes toolkit.
When to use it
Use this skill when configuring GKE workloads to enhance their reliability through PDBs, health probes, and topology spread constraints.
When not to use it
This skill is not suitable for disaster recovery setups or full cluster backups; for those purposes, consider using gke-backup-dr instead.
What you can build with it
Configuring a New GKE Cluster
When setting up a new GKE cluster, use this skill to ensure that high availability settings and reliability features like PDBs and health probes are correctly configured.
Ensuring Application Resilience
Use this skill to implement best practices for application resilience, such as setting up health probes and topology spread constraints to minimize downtime.
Reviewing Existing GKE Workloads
When auditing existing GKE workloads, leverage this skill to verify that PDBs and health probes are in place and configured according to best practices.
How to install GKE Reliability
View source1. Install with the skills CLI
npx skills add google/skills/gke-reliability --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by googleGKE Reliability
This reference covers high availability and reliability configuration for GKE clusters and workloads.
MCP Tools:
get_cluster,get_k8s_resource,describe_k8s_resource,apply_k8s_manifest,list_k8s_events
Golden Path Reliability Defaults
| Setting | Golden Path Value | Notes |
|---|---|---|
| Cluster type | Regional (4 zones: | Control plane replicated across |
| : : us-central1-a/b/c/f) : zones : | ||
| Upgrade strategy | SURGE (maxSurge: 1) | Rolling upgrades with extra |
| : : : capacity : | ||
| Auto-repair | true | Unhealthy nodes replaced |
| : : : automatically : | ||
| Auto-upgrade | true | Nodes follow control plane |
| : : : version : | ||
| Release channel | REGULAR | Balanced freshness and stability |
| Stateful HA | Enabled | Leader election for stateful |
| : : : workloads : |
Workflows
1. Verify Cluster High Availability
# MCP (preferred)
get_cluster(name="projects/<PROJECT>/locations/<REGION>/clusters/<CLUSTER>",
readMask="location,locations,nodePools.locations")
# gcloud fallback
gcloud container clusters describe <CLUSTER> --region <REGION> \
--format="json(location, locations)" \
--quiet
- If
locationis a region (e.g.,us-central1), the control plane is regional - If
locationshas multiple entries, nodes span multiple zones
2. Pod Disruption Budgets (PDBs)
PDBs ensure minimum pod availability during voluntary disruptions (node upgrades, autoscaler scale-down).
Check existing PDBs:
# MCP (preferred)
get_k8s_resource(parent="...", resourceType="poddisruptionbudget")
# kubectl fallback
kubectl get pdb --all-namespaces
Create PDB:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: my-app-pdb
namespace: default
spec:
minAvailable: 2 # Or use maxUnavailable: 1
selector:
matchLabels:
app: my-app
Every production Deployment with 2+ replicas should have a PDB.
3. Health Probes
Every production container should have liveness and readiness probes. Startup probes are recommended for slow-starting apps.
Check existing probes:
# MCP (preferred)
describe_k8s_resource(parent="...", resourceType="deployment", name="<APP>", namespace="<NS>")
# kubectl fallback
kubectl get deployment <APP> -n <NS> -o yaml | grep -E "livenessProbe|readinessProbe|startupProbe"
Recommended probe configuration:
spec:
containers:
- name: app
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 15
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
readinessProbe:
httpGet:
path: /readyz
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
startupProbe: # For slow-starting apps
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 30 # 30 * 5s = 150s max startup time
- Readiness: Determines when a pod can accept traffic
- Liveness: Determines when to restart a container
- Startup: Disables liveness/readiness until the app is ready (prevents premature restarts)
4. Graceful Shutdown
Ensure applications handle SIGTERM and drain in-flight requests:
spec:
terminationGracePeriodSeconds: 30 # Default; increase for long-running requests
containers:
- name: app
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 5"] # Allow LB to deregister
5. Topology Spread Constraints
Distribute pods across zones and nodes to survive failures:
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: my-app
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: my-app
- Zone spread (
DoNotSchedule): Hard requirement -- pods must be balanced across zones - Node spread (
ScheduleAnyway): Best-effort -- prefer distribution but don't block scheduling
6. Replicas
| Workload Type | Minimum Replicas | Reason |
|---|---|---|
| Stateless web/API | 2 | Survive single pod/node |
| : : : failure : | ||
| Critical services | 3 | Survive zone failure with zone |
| : : : spread : | ||
| Stateful (databases) | 3 (with replication) | Application-level quorum |
| Batch/jobs | 1 | Ephemeral by nature |
Best Practices & Production Guidelines
- Regional clusters for production: Always use regional clusters to survive zone failures.
- PDBs for everything: Every production workload with 2+ replicas needs a PodDisruptionBudget (PDB) to protect against voluntary disruptions.
- Probes with Explicit Timeouts: Every production container must have both
liveness and readiness probes defined. Always explicitly define
initialDelaySeconds,periodSeconds, andtimeoutSecondsfor all probes. Never rely on the Kubernetes default timeout of 1 second if your application requires more, but always set a strict limit to prevent hanging connections. - Zone spreading: Use topology spread constraints to distribute pods across failure domains (zones and nodes).
- Graceful shutdown: Handle
SIGTERMand set appropriateterminationGracePeriodSecondswith apreStopsleep hook to allow load balancer deregistration. - Maintenance windows: Schedule upgrades during low-traffic periods (see
the
gke-upgradesskill).
Frequently asked questions about GKE Reliability
Similar skills
Turborepo
Optimized build system for JavaScript/TypeScript monorepos.
Azure Pipelines Validation
Streamline your Azure DevOps pipeline changes locally.
Azure Developer CLI
Streamline your Azure project workflows with best practices.
Azure Container Registry CLI
Manage Azure Container Registry resources with ease.
Aspire
Build and orchestrate polyglot distributed applications seamlessly.
Vercel CLI
Manage and deploy Vercel projects from the command line.
