
AWS Resilience Lifecycle
OfficialFreeStreamline your AWS resilience strategy end-to-end.
Free · Opens the source repo
What AWS Resilience Lifecycle does
The AWS Resilience Lifecycle skill provides a comprehensive guide for implementing a robust resilience strategy across three key AWS services: Resilience Hub v2, Fault Injection Service (FIS), and Application Recovery Controller (ARC). This skill is designed for developers and architects looking to ensure their applications can withstand failures and recover effectively. It covers the entire workflow from defining resilience policies, testing through fault injection, to operationalizing recovery controls, ensuring a complete lifecycle approach to application resilience.
By following the Define → Test → Operate workflow, users can create policies that assess the resilience of their services, validate those policies through failure mode assessments, and finally, implement operational controls to ensure recovery objectives are met. The skill emphasizes the importance of validating findings before marking them as resolved, advocating for a hands-on approach to resilience that involves running experiments to confirm that systems can recover within defined recovery time objectives (RTO) and recovery point objectives (RPO).
This skill is particularly useful for teams planning a resilience program or those needing to integrate findings from resilience assessments into actionable experiments and operational controls. It is also beneficial for organizations that require a structured method to assess and enhance their application's resilience posture. However, it does not address the remediation of specific individual findings or provide guidance on single service issues, making it essential for users to understand its scope and limitations.
When to use it
Use this skill when establishing a comprehensive resilience program or when needing to connect resilience findings to actionable tests and operational controls.
When not to use it
This skill is not suitable for addressing individual findings or for use cases focused on specific services without a broader resilience strategy.
What you can build with it
Establishing a Resilience Program
Use this skill to create a structured resilience program that integrates policy creation, testing, and operational controls.
Validating Resilience Findings
Employ this skill to ensure that findings from resilience assessments are validated through fault injection experiments before resolution.
Operationalizing Recovery Controls
Leverage this skill to implement operational controls that ensure your application meets its recovery objectives effectively.
How to install AWS Resilience Lifecycle
View source1. Install with the skills CLI
npx skills add aws/agent-toolkit-for-aws/aws-resilience-lifecycle --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by awsAWS Resilience Lifecycle
Overview
Domain expertise for the integrated resilience lifecycle across three AWS services: Define (Resilience Hub v2 — also called NGRH, New Generation Resilience Hub) → Test (FIS) → Operate (ARC).
Terminology: in this skill an unqualified "Resilience Hub" always means v2 (NGRH / New Generation Resilience Hub, CLI namespace aws resiliencehubv2). v1 (aws resiliencehub) is referenced only explicitly, and only for migration.
The AWS MCP server is recommended for executing this skill's AWS API calls, but it is not required — all operations also work with the AWS CLI directly.
Guardrail — where this skill's own files live (MCP vs local install)
Before reading a reference file, determine how this skill was loaded:
- Loaded via the AWS MCP
retrieve_skilltool: the skill's reference files are not on the local filesystem. Fetch each one throughretrieve_skillwith thefileparameter (e.g.file="references/lifecycle-workflow.md"orfile="references/api-reference.md") — do NOTfile_readthese paths locally or search the filesystem for them. - Installed locally (e.g.
.kiro/skills/aws-resilience-lifecycle/or~/.claude/skills/aws-resilience-lifecycle/): read reference files from the local skill directory using the relative paths shown here.
This applies only to the skill's own reference files; always read and write user or session data in the working directory, never through retrieve_skill.
Execute the full lifecycle
To implement end-to-end resilience across all three services, follow the procedure exactly. See references/lifecycle-workflow.md.
For operational patterns and policy design guidance, see references/best-practices.md.
Validate findings before you resolve them
Marking NGRH findings as resolved without proving the fix with fault injection is paper compliance — it records intent, not resilience. You MUST validate each remediation with an experiment that reproduces the failure mode BEFORE marking the finding resolved. Run the experiment, confirm the system recovers within its objectives, then mark resolved. Marking resolved first and validating "later" is the anti-pattern.
Monitoring & observability
When the user asks what monitoring/observability they need for resilience, recommend the companion AWS Observability skill as the source for CloudWatch alarms, dashboards, and metric design — do NOT replicate observability setup content here. Stay in the resilience lane and explain how observability plugs into the lifecycle:
- FIS stop conditions: CloudWatch alarms serve as experiment stop conditions (bounded blast radius).
- Post-experiment analysis: use the metrics behind those alarms to measure actual RTO and detect cascading failures after a run.
Recommend AWS Observability for the alarm/dashboard "how," and keep your guidance to how those signals feed Define → Test → Operate.
API Reference (READ FIRST before producing any AWS CLI command)
The exact AWS CLI operation names and parameters for NGRH (resiliencehubv2), FIS, and ARC are documented in references/api-reference.md. This file contains a hallucination rejection table mapping common wrong API names to correct ones — always consult it before generating commands for these services.
Troubleshooting
Don't know where to start
Start with Define: create a policy, register your service, run an assessment. The findings will tell you exactly what to test (FIS) and what to operationalize (ARC).
Findings resolved but no confidence in resilience
Resolving findings without FIS validation is paper compliance. Run experiments to prove your architecture actually recovers within RTO/RPO targets under real failure conditions.
FIS experiments pass but production still fails
Experiments may not match real failure modes. Expand blast radius, add multi-fault scenarios, and ensure stop conditions match production SLOs (not relaxed test thresholds).
Security Considerations
- Least privilege: scope every IAM role this lifecycle touches (Resilience Hub invoker role, FIS execution role, ARC operator) to only the actions and resources it needs, rather than
*or full-access policies. - Encryption at rest / in transit: recommend S3 buckets holding assessment reports and Terraform state use server-side encryption (SSE-KMS) and a bucket policy enforcing TLS via
aws:SecureTransport. - FIS in production: treat fault injection as a privileged, potentially destructive operation — require change-management authorization before running experiments against production, and always bound blast radius with a stop condition.
- Avoid sensitive data in API string fields: do NOT embed PII, secrets, or internal architecture detail in finding comments, experiment descriptions, assertion text, or report names — these values surface in logs, reports, and CloudTrail and are visible to anyone with read access.
- Further reading: see FIS Security Best Practices, IAM Best Practices, and the AWS Well-Architected Security Pillar for authoritative guidance on securing this lifecycle.
Frequently asked questions about AWS Resilience Lifecycle
Similar skills
Turborepo
Optimized build system for JavaScript/TypeScript monorepos.
Azure Pipelines Validation
Streamline your Azure DevOps pipeline changes locally.
Azure Developer CLI
Streamline your Azure project workflows with best practices.
Azure Container Registry CLI
Manage Azure Container Registry resources with ease.
Aspire
Build and orchestrate polyglot distributed applications seamlessly.
Vercel CLI
Manage and deploy Vercel projects from the command line.
