
AWS Resource Health Diagnosis
OfficialFreeDiagnose AWS resource issues with CloudWatch insights.
Free · Opens the source repo
What AWS Resource Health Diagnosis does
The AWS Resource Health Diagnosis skill provides a systematic approach to evaluating the health of AWS resources, leveraging CloudWatch logs and metrics to identify and diagnose issues. This skill is particularly useful for DevOps engineers, system administrators, and cloud architects who need to ensure their AWS infrastructure is running optimally. By following the outlined workflow, users can quickly assess the status of their resources and take informed action to resolve any detected issues.
The process begins with the retrieval of AWS diagnostic best practices, which inform the subsequent steps. Users must have the AWS CLI configured and authenticated, and they need to identify the specific resource they want to analyze. The skill supports a range of AWS services, including EC2, Lambda, RDS, ECS, ALB, DynamoDB, SQS, and API Gateway, allowing for comprehensive health checks tailored to each service type. This flexibility makes it suitable for diverse cloud environments.
Once the target resource is identified, the skill performs a health status assessment using service-specific commands to gather key health indicators. Following this, it analyzes logs and metrics to uncover recurring error patterns and performance trends. The skill classifies issues based on severity and conducts root cause analysis, categorizing problems into configuration issues, resource constraints, network issues, application issues, dependency issues, and security issues. This thorough analysis helps users pinpoint the source of problems effectively.
Finally, the skill generates a remediation plan that includes immediate actions, short-term fixes, and long-term improvements. Users receive a structured report summarizing the health assessment, identified issues, and recommended actions, making it easier to implement solutions and maintain resource health. This skill is essential for anyone managing AWS resources who needs to proactively address health and performance issues.
When to use it
Use this skill when you need to assess the health of AWS resources and diagnose issues based on CloudWatch data.
When not to use it
This skill is not suitable for environments without AWS resources or for users unfamiliar with AWS CLI commands.
What you can build with it
Diagnosing EC2 Instance Health
Use this skill to assess the health of your EC2 instances, checking for critical issues like instance status and performance metrics.
Troubleshooting Lambda Function Errors
Quickly analyze Lambda functions for error rates and performance issues, helping you identify bottlenecks and improve function reliability.
Evaluating RDS Database Performance
Assess RDS instances for CPU utilization and connection issues, generating actionable insights to optimize database performance.
How to install AWS Resource Health Diagnosis
View source1. Install with the skills CLI
npx skills add github/awesome-copilot/aws-resource-health-diagnose --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by githubAWS Resource Health & Issue Diagnosis
This workflow analyzes a specific AWS resource to assess its health status, diagnose potential issues using CloudWatch logs and metrics, and develop a comprehensive remediation plan for any problems discovered.
Prerequisites
- AWS CLI configured and authenticated
- Target AWS resource identified (name, type, and optionally region/account)
- CloudWatch logging and metrics enabled on the target resource
Workflow Steps
Step 1: Get AWS Diagnostic Best Practices
Fetch https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/ for monitoring and troubleshooting guidance to inform the diagnostic approach.
Step 2: Resource Discovery & Identification
Locate the target resource using the appropriate AWS CLI command for its type:
# EC2
aws ec2 describe-instances --filters "Name=tag:Name,Values=<name>"
# Lambda
aws lambda get-function --function-name <name>
# RDS
aws rds describe-db-instances --db-instance-identifier <name>
# ECS
aws ecs describe-services --cluster <cluster> --services <name>
# ALB
aws elbv2 describe-load-balancers --names <name>
# DynamoDB
aws dynamodb describe-table --table-name <name>
# SQS
aws sqs get-queue-attributes --queue-url <url> --attribute-names All
# API Gateway
aws apigatewayv2 get-apis
If multiple matches are found, prompt the user to specify region/account.
Step 3: Health Status Assessment
Run service-specific health checks:
# EC2
aws ec2 describe-instance-status --instance-ids <id>
# RDS
aws rds describe-db-instances --db-instance-identifier <name> \
--query 'DBInstances[0].DBInstanceStatus'
# Lambda - error rate over 24h
aws cloudwatch get-metric-statistics --namespace AWS/Lambda \
--metric-name Errors --dimensions Name=FunctionName,Value=<name> \
--start-time $(date -u -d '24 hours ago' +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
--period 3600 --statistics Sum
# ECS
aws ecs describe-services --cluster <cluster> --services <name> \
--query 'services[0].[status,runningCount,desiredCount,pendingCount]'
Key health indicators by service type:
- Lambda: Error rate, throttle rate, duration P99, concurrent executions
- RDS: CPU utilization, FreeStorageSpace, DatabaseConnections, ReadLatency/WriteLatency
- ECS: Running vs desired task count, task stop reason
- ALB: TargetResponseTime, HTTPCode_ELB_5XX_Count, UnHealthyHostCount
- SQS: ApproximateNumberOfMessagesNotVisible, ApproximateAgeOfOldestMessage
- DynamoDB: ConsumedReadCapacityUnits, ThrottledRequests, SuccessfulRequestLatency
Step 4: Log & Metrics Analysis
Find log groups and run CloudWatch Logs Insights queries:
# Find log groups
aws logs describe-log-groups --log-group-name-prefix /aws/<service>/<name>
# Start a query (last 24h errors)
aws logs start-query \
--log-group-name /aws/lambda/<name> \
--start-time $(date -u -d '24 hours ago' +%s) \
--end-time $(date -u +%s) \
--query-string 'filter @message like /ERROR/ | stats count(*) as errorCount by bin(1h)'
# Get results
aws logs get-query-results --query-id <id>
# Lambda cold starts
aws logs start-query \
--log-group-name /aws/lambda/<name> \
--start-time $(date -u -d '24 hours ago' +%s) \
--end-time $(date -u +%s) \
--query-string 'filter @type = "REPORT" | filter @initDuration > 0 | stats count() as coldStarts by bin(1h)'
# RDS Performance Insights (if enabled)
aws pi get-resource-metrics \
--service-type RDS --identifier db:<identifier> \
--metric-queries '[{"Metric":"db.load.avg"}]' \
--start-time $(date -u -d '24 hours ago' +%Y-%m-%dT%H:%M:%SZ) \
--end-time $(date -u +%Y-%m-%dT%H:%M:%SZ) \
--period-in-seconds 3600
Identify: recurring error patterns, correlation with deployments (CloudTrail), performance trends, dependency failures.
Step 5: Issue Classification & Root Cause Analysis
Severity:
- Critical: Service unavailable, data loss, security incidents
- High: Performance degradation, error rates >5%, intermittent failures
- Medium: Warnings, suboptimal configuration, minor performance issues
- Low: Informational alerts, optimization opportunities
Root Cause Categories:
- Configuration Issues: wrong settings, missing env vars, IAM permission denials
- Resource Constraints: CPU/memory/disk limits, Lambda throttling, RDS connection exhaustion
- Network Issues: security group rules, VPC routing, DNS, NACLs
- Application Issues: code bugs, memory leaks, unhandled exceptions, slow queries
- Dependency Issues: downstream timeouts, SQS/SNS failures, external API limits
- Security Issues: KMS key issues, certificate expiration
Step 6: Generate Remediation Plan
Immediate Actions (Critical):
# Lambda throttling — increase reserved concurrency
aws lambda put-reserved-concurrency \
--function-name <name> --reserved-concurrent-executions 100
# RDS connection exhaustion — reboot to reset connections
aws rds reboot-db-instance --db-instance-identifier <name>
Short-term Fixes (High/Medium): Configuration adjustments, right-sizing, CloudWatch alarm improvements, IAM corrections.
Long-term Improvements: Architectural changes for resilience, preventive monitoring, enable AWS Health Dashboard notifications via EventBridge.
Step 7: Report & User Confirmation
Present findings:
🏥 AWS Resource Health Assessment
📊 Resource Overview:
• Resource: [Name] ([Type])
• Status: [Healthy/Warning/Critical]
• Region: [Region] | Account: [Account ID]
🚨 Issues Identified:
• Critical: X | High: Y | Medium: Z | Low: N
🔍 Top Issues:
1. [Issue]: [Description] — Impact: [High/Medium/Low]
2. [Issue]: [Description] — Impact: [High/Medium/Low]
🛠️ Remediation: X immediate, Y short-term, Z long-term actions
❓ Proceed with detailed remediation plan? (y/n)
Then generate a full markdown report covering: health metrics, issues with root cause analysis, phased remediation steps with AWS CLI commands, CloudWatch alarm recommendations, and validation checklist.
Error Handling
- Resource Not Found: Ask user to clarify name/region
- Authentication Issues: Guide through
aws configure - Insufficient Permissions: List required IAM actions (
logs:*,cloudwatch:*,pi:*) - No Logs Available: Suggest enabling CloudWatch logging for the resource type
- Query Timeouts: Use shorter time windows
Success Criteria
- ✅ Resource health accurately assessed across all key metrics
- ✅ All significant issues identified and classified by severity
- ✅ Root cause analysis completed for major problems
- ✅ Actionable remediation plan with AWS CLI commands
- ✅ CloudWatch monitoring recommendations included
- ✅ Implementation steps include validation and rollback procedures
Frequently asked questions about AWS Resource Health Diagnosis
Similar skills
Turborepo
Optimized build system for JavaScript/TypeScript monorepos.
Azure Pipelines Validation
Streamline your Azure DevOps pipeline changes locally.
Azure Developer CLI
Streamline your Azure project workflows with best practices.
Azure Container Registry CLI
Manage Azure Container Registry resources with ease.
Aspire
Build and orchestrate polyglot distributed applications seamlessly.
Vercel CLI
Manage and deploy Vercel projects from the command line.
