
GKE Workload Troubleshooting
FreeDiagnose GKE workload failures with precision.
Free ยท Opens the source repo
What GKE Workload Troubleshooting does
The GKE Workload Troubleshooting skill is designed for developers and operations teams managing applications on Google Kubernetes Engine (GKE). It provides a systematic approach to diagnosing common workload failures such as CrashLoopBackOff, OOMKilled, and ImagePullBackOff. By leveraging Kubernetes logs and events, this skill enables users to identify the root cause of issues when pods fail to start or crash repeatedly, streamlining the troubleshooting process.
This skill operates in a non-interactive manner, ensuring that it can be executed autonomously without user intervention. It begins by extracting necessary context from the user's environment or prompt, including project ID, cluster name, and workload details. This context is crucial for the skill to function effectively, allowing it to fetch cluster credentials and analyze pod statuses accurately. The skill also includes fallback mechanisms for scenarios where the cluster is unreachable, providing users with diagnostic commands to run manually.
The diagnostic workflow is structured into clear steps, starting with the analysis of pod statuses and conditions. It identifies specific issues based on the phase and state of the pods, guiding users through the necessary commands to gather information about their workloads. By querying namespace events, it helps pinpoint infrastructure or configuration problems that may be causing failures. This structured approach not only aids in identifying issues but also suggests potential fixes based on the analysis of logs and events.
Overall, this skill is an essential tool for anyone working with GKE who needs to quickly diagnose and resolve workload issues. Its focus on read-only diagnostics ensures that users can safely analyze their environments without making unintended changes, making it ideal for both experienced Kubernetes users and those newer to the platform.
When to use it
Use this skill when you encounter pod failures or crashes in GKE, such as CrashLoopBackOff or OOMKilled errors.
When not to use it
This skill is not suitable for GKE cluster provisioning, node pool creation, or troubleshooting non-Kubernetes Google Cloud services.
What you can build with it
Diagnosing a CrashLoopBackOff Issue
When a pod is repeatedly crashing, use this skill to analyze the pod's status and logs to identify the root cause.
Resolving ImagePullBackOff Errors
If your application fails to start due to image pulling issues, this skill will help you diagnose and suggest fixes based on the event logs.
Investigating OOMKilled Events
Utilize this skill to determine if memory limits are causing your pods to be killed and receive recommendations for adjustments.
How to install GKE Workload Troubleshooting
View source1. Install with the skills CLI
npx skills add google/skills/gke-workload-troubleshooting --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by googleGKE Workload Troubleshooting Skill
Use this skill to systematically diagnose and resolve failures in application workloads deployed in GKE clusters. This skill operates non-interactively and enforces a read-only diagnostics boundary before proposing manifest or config corrections.
๐ Diagnostic Workflow
Step 0: Non-Interactive Context Discovery & Time Window Definition
-
Parameter Extraction: Extract required context (
project_id,cluster_name,cluster_location,workload_name,workload_namespace) non-interactively from the user prompt, activeSETTINGS.md, or active environment defaults:- Default
workload_namespacetodefaultif omitted. - Infer missing cluster parameters from active environment (
kubectl config current-contextorgcloud config get-value project). - Prioritize non-interactive context discovery from prompts and environment defaults to ensure autonomous execution flow.
- Default
-
Cluster Credentials & Fallback Mode:
- Attempt credential fetch:
gcloud container clusters get-credentials {cluster_name} --region/--zone {cluster_location} - Fallback / Dry-Run Mode: If the cluster is unreachable,
non-existent, or live command execution fails (such as in sandboxed
evaluations, dry-run mode, or offline analysis):
- Limit retry attempts to avoid resource exhaustion and context overflow in unreachable cluster scenarios.
- Immediately present the exact sequence of
kubectldiagnostic commands for the human operator to run. - Synthesize the root cause analysis and output the proposed GitOps manifest fix based on the reported symptoms.
- Attempt credential fetch:
-
Time Handling & Fallbacks:
- Determine Issue Timestamp ({issue_time}):
- Specific Time Provided: If the user provides a specific
timestamp, use it as
{issue_time}. - Relative Time Provided (e.g., "5 minutes ago"): Dynamically
calculate the corresponding UTC timestamp based on current system
time, and use it as
{issue_time}. - No Time Provided (Default): Use current system time as
{issue_time}.
- Specific Time Provided: If the user provides a specific
timestamp, use it as
- Window Calculation: Center a 1-hour query window around
{issue_time}(start_time={issue_time} - 30m,end_time={issue_time} + 30m).
- Determine Issue Timestamp ({issue_time}):
Step 1: Analyze Pod Status and Conditions
Inspect the workload's active pod states and controller status.
Diagnostic Commands:
# 1. Inspect the deployment's actual selector labels:
kubectl get deployment {workload_name} -n {workload_namespace} -o jsonpath='{.spec.selector.matchLabels}'
# 2. Query the pods using the returned labels, for example:
kubectl get pods -l {selector_labels} -n {workload_namespace}
kubectl get deploy/{workload_name} -n {workload_namespace} -o yaml
Diagnostic Decision Tree:
-
Phase: Pending:
- The Pod cannot schedule on any node. Proceed directly to Step 2 (Query Namespace Events).
-
State: CrashLoopBackOff / Error:
- Container is booting but exiting repeatedly. Check the terminated status using:
kubectl get pod {pod_name} -n {workload_namespace} -o jsonpath='{.status.containerStatuses[*].lastState.terminated}'- ExitCode: 137 (OOMKilled): Memory limit reached. Proceed to Step 3 (Inspect Logs) and inspect container startup command to differentiate between an application-level memory leak/loop vs an infrastructure capacity limit mismatch, then proceed to Step 5 to propose fixes.
- ExitCode: 1 or other non-zero codes: The application code crashed. Proceed directly to Step 3 (Inspect Logs).
-
State: ContainerCreating:
- The container is blocked during volume mount, networking setup, or image pulling. Proceed directly to Step 2 (Query Namespace Events).
Step 2: Query Namespace Events
Look for infrastructure, volume, image, or scheduling alerts in GKE.
Diagnostic Command:
kubectl get events -n {workload_namespace} --sort-by='.metadata.creationTimestamp'
# Or query Cloud Logging for historical GKE events within the time window:
gcloud logging read "resource.type=\"k8s_cluster\" AND logName=\"projects/{project_id}/logs/events\" AND jsonPayload.involvedObject.namespace=\"{workload_namespace}\"" --start-time="{start_time}" --end-time="{end_time}" --project="{project_id}"
Note: Retrieve the sorted events list and manually inspect the event timestamps
(CreationTimestamp/LastSeen) to identify failures occurring within the
{start_time} and {end_time} window.
Signature Identifiers:
FailedScheduling: Node resource exhaustion. Look for messages like0/3 nodes are available: 3 Insufficient memory.or missing node affinity tolerations (e.g. Spot VM taints).FailedMount:- Missing PersistentVolumeClaim (
PVC). - Missing Secret (
Secret "{secret_name}" not found). - Missing ConfigMap (
ConfigMap "{configmap_name}" not found).
- Missing PersistentVolumeClaim (
Failed/BackOff(Image Pull):- Wrong image tag, missing image registry authentication (e.g., ImagePullBackOff).
- Resolution Steps for Wrong Image Tag:
- Identify the failing container image name and the invalid tag.
- Check the Git repository history for the last known working image tag
for this workload. Run
git log -p -S "{image_name}" -- {manifest_file_path}(or usegit logon the folder containing manifests) to identify the previous working tag in Git. - If the invalid tag is a recent change in git history, compare it to the tag from the last successful commit.
- Propose reverting the image tag to the last working version, or correcting the tag version in the manifest patch.
Step 3: Inspect Application Logs
Extract exceptions and stack traces from the application runtime.
Diagnostic Commands:
# Check current active log stream (handles multi-container pods)
kubectl logs {pod_name} -n {workload_namespace} --all-containers --tail=100
# Check logs from previously terminated container instances (handles multi-container pods)
kubectl logs {pod_name} -n {workload_namespace} --all-containers -p --tail=100
Signature Identifiers:
- Out-of-Memory (OOM) Analysis: Inspect container logs and startup
commands (
spec.containers[*].command). Differentiate between an Application Code Leak/Loop (unbounded array appending, memory leak signatures) vs an Infrastructure Capacity Ceiling Mismatch (legitimate workload demand exceeding limits). - Stack Trace / Unhandled Exception: Look for language-specific stack
traces (e.g.,
panic:,NullPointerException,Traceback (most recent call)). This indicates an application bug. - Egress Network Timeout: Look for connection timeouts (e.g.,
Connection timed out,dial tcp: i/o timeout). Proceed to Step 4 (Verify Connectivity). - Permission Errors (ReadOnlyRootFilesystem): Look for write errors (e.g.,
Read-only file system,Permission deniedwhen writing to/tmpor/var/log). Propose adding anemptyDirvolume mount to that directory in the manifest.
Step 4: Verify Service Connectivity and Network Policies
Troubleshoot connection drops to other services.
Diagnostic Commands:
# Verify target endpoint is active
kubectl get endpoints {target_service_name} -n {target_namespace}
# Query network policies inside namespace
kubectl get networkpolicies -n {workload_namespace} -o yaml
Logic & Dry-Run Fallback:
-
Live Cluster Mode:
- If
kubectl get endpointsreturns an empty list, the target microservice itself is failing to schedule or boot (troubleshoot target service). - If endpoints exist but logs show timeouts, analyze
NetworkPolicyegress blocks to verify if egress traffic to the target service's IP/port is allowed.
- If
-
Sandboxed / Dry-Run Mode:
- If live
kubectlqueries fail or cluster connection is unavailable, do NOT retry live cluster access or enter repetitive connection attempts. - Immediately inspect the application source code (e.g.
worker.py,app.go, DB connection strings) or Deployment manifests to identify the target service hostname (e.g.account-db) and destination port (e.g.5432). - Present the exact
kubectl get endpointsandkubectl get networkpoliciescommands for the user, and synthesize the requiredNetworkPolicyegress patch allowing traffic to the target service and port.
- If live
Step 5: Propose GitOps Correction
Following the GitOps boundary, do not apply patches directly to the cluster.
- Synthesize the root cause analysis for the human operator (e.g. "payment-api is failing with exit code 137 because its memory limit is set to 256Mi while actual usage spiked to 270Mi").
- Generate the corrected YAML manifest patch (e.g. increase memory limits, add missing Secret mounts, or add tolerations for Spot nodes).
- Check if a branch or Pull Request (PR) already exists for this workload/failure. If so, update the existing branch/PR or notify the user instead of creating a duplicate. Otherwise, create a branch, commit the change, open a Pull Request (PR) on GitHub, and conclude the workflow (do not wait for human merge).
Frequently asked questions about GKE Workload Troubleshooting
Similar skills
Turborepo
Optimized build system for JavaScript/TypeScript monorepos.
Azure Pipelines Validation
Streamline your Azure DevOps pipeline changes locally.
Azure Developer CLI
Streamline your Azure project workflows with best practices.
Azure Container Registry CLI
Manage Azure Container Registry resources with ease.
Aspire
Build and orchestrate polyglot distributed applications seamlessly.
Vercel CLI
Manage and deploy Vercel projects from the command line.
