
Database Recovery Manager
FreeAutomate and optimize your database recovery processes.
Free · Opens the source repo
What Database Recovery Manager does
The Database Recovery Manager skill is designed for database administrators and developers who need to ensure robust backup and recovery strategies for PostgreSQL and MySQL databases. This skill provides a comprehensive framework for planning, executing, and verifying backup and recovery procedures, focusing on point-in-time recovery (PITR), logical and physical backups, and disaster recovery testing. By utilizing this skill, users can streamline their database management tasks while adhering to critical Recovery Point Objective (RPO) and Recovery Time Objective (RTO) targets.
With this skill, users can assess their current backup configurations and define their RPO and RTO targets based on specific business requirements. It guides users through the configuration of WAL archiving for PostgreSQL and binary logging for MySQL, ensuring that the necessary components are in place for effective data recovery. The skill also includes detailed instructions for creating both physical and logical backups, enabling users to choose the best approach based on their database size and recovery needs.
Additionally, the Database Recovery Manager skill emphasizes the importance of testing recovery processes. Users are encouraged to restore backups to separate environments, verify data integrity, and measure recovery times to ensure that they meet their RTO targets. Automated backup verification is also a key feature, allowing users to set up cron jobs that regularly test backup integrity and send notifications on success or failure. This proactive approach helps maintain confidence in the recovery processes and minimizes downtime during actual recovery scenarios.
In summary, this skill is ideal for database professionals looking to enhance their backup and recovery strategies, automate verification processes, and ensure their databases are resilient against data loss. By following the structured guidance provided, users can effectively manage their database recovery operations, making this skill a valuable addition to their toolkit.
When to use it
Use this skill when you need to establish or improve backup and recovery procedures for PostgreSQL or MySQL databases.
When not to use it
This skill may not be suitable for environments without superuser access or where database management is not a primary concern.
What you can build with it
Point-in-Time Recovery
Recover from accidental data deletion by restoring a database to a specific timestamp using WAL files.
Automated Backup to Cloud
Schedule daily backups to S3 with automated verification to ensure data integrity and compliance.
Disaster Recovery Testing
Conduct monthly drills to validate recovery procedures and ensure the team can meet RTO targets during an actual incident.
How to install Database Recovery Manager
View source1. Install with the skills CLI
npx skills add jeremylongshore/claude-code-plugins-plus-skills/managing-database-recovery --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by jeremylongshoreDatabase Recovery Manager
Overview
Plan and execute database backup and recovery procedures for PostgreSQL and MySQL, including point-in-time recovery (PITR), logical and physical backups, WAL archiving, and disaster recovery testing. This skill covers the full backup lifecycle from configuration through automated verification, ensuring Recovery Point Objective (RPO) and Recovery Time Objective (RTO) targets are met.
Prerequisites
- Database superuser or replication-role credentials
- Backup storage destination (local disk, NFS mount, S3, GCS, or Azure Blob)
pg_basebackup,pg_dump,pg_restore(PostgreSQL) ormysqldump,xtrabackup(MySQL)tar,rsync, oraws s3CLI for backup transfer and storage- WAL archiving configured for PITR (PostgreSQL:
archive_mode = on,archive_command) - Sufficient storage for backup retention (estimate 2-3x database size for full + incremental)
Instructions
-
Assess the current backup situation by checking existing backup configurations. For PostgreSQL: verify
archive_mode,archive_command, andwal_levelinpostgresql.conf. For MySQL: check if binary logging is enabled withSHOW VARIABLES LIKE 'log_bin'. -
Define RPO and RTO targets based on business requirements:
- RPO (acceptable data loss): determines backup frequency and WAL archiving interval
- RTO (acceptable downtime): determines backup type and recovery procedure complexity
- Typical targets: RPO < 1 hour (WAL archiving), RTO < 30 minutes (physical backup restore)
-
Configure WAL archiving for PostgreSQL PITR:
- Set
wal_level = replicaandarchive_mode = on - Configure
archive_command = 'test ! -f /archive/%f && cp %p /archive/%f'(or use pgBackRest/WAL-G for S3) - Verify archiving works:
SELECT * FROM pg_stat_archiver - For MySQL, enable binary logging:
log_bin = mysql-bin,binlog_format = ROW
- Set
-
Create a full physical backup using
pg_basebackup -D /backups/base -Ft -z -P(PostgreSQL) orxtrabackup --backup --target-dir=/backups/full(MySQL). Physical backups are faster to restore than logical backups for databases larger than 10GB. -
Create logical backups for portability and selective restoration:
pg_dump -Fc -f database.dump dbname(PostgreSQL) ormysqldump --single-transaction --routines --triggers dbname > database.sql(MySQL). -
Upload backups to remote storage for disaster recovery:
aws s3 cp /backups/base.tar.gz s3://backup-bucket/postgres/$(date +%Y%m%d)/with server-side encryption enabled. Implement a retention policy (e.g., daily backups for 30 days, weekly for 90 days, monthly for 1 year). -
Test recovery by restoring to a separate server or container:
- Restore physical backup:
pg_restore -d testdb database.dumpor untar base backup and start PostgreSQL - For PITR: restore base backup, copy WAL files to
pg_wal, createrecovery.signalwithrecovery_target_time = '2024-01-15 14:30:00' - Verify data integrity by running application test suite against restored database
- Measure actual recovery time to validate RTO target
- Restore physical backup:
-
Automate backup verification with a daily cron job that: takes backup, restores to a test instance, runs integrity checks (
pg_catalog.pg_classrow counts, checksum verification), and sends a success/failure notification. -
Document the recovery runbook with exact commands for each recovery scenario: full database restore, PITR to a specific timestamp, single table restore, and cross-region failover.
-
Schedule monthly disaster recovery drills to verify the runbook works and the team can execute recovery within RTO targets.
Output
- Backup configuration files for PostgreSQL (postgresql.conf changes, archive_command) or MySQL (my.cnf changes)
- Backup scripts (shell) for automated full and incremental backups with S3/GCS upload
- Recovery runbook with step-by-step commands for each recovery scenario
- Backup verification scripts that automate restore-and-check procedures
- Retention policy configuration with automated cleanup of old backups
Error Handling
| Error | Cause | Solution |
|---|---|---|
| WAL segment not found during PITR | Gap in WAL archiving due to archive_command failure | Check pg_stat_archiver for last_failed_wal; fix archive_command; consider WAL-G or pgBackRest for reliable archiving |
pg_basebackup fails with "replication connection" error | Missing replication permissions or max_wal_senders exhausted | Grant REPLICATION role; increase max_wal_senders; add entry to pg_hba.conf for replication connections |
| Backup storage full | Retention policy not enforced or backup size grew unexpectedly | Implement automated cleanup script; compress backups with gzip or zstd; monitor storage usage with alerts at 80% |
| Recovery takes longer than RTO target | Database grew since RTO was last validated, or restore is I/O-bound | Use physical backups instead of logical; restore to SSD storage; parallelize restore with pg_restore -j 4; consider standby replica for faster failover |
| Restored database has corruption | Backup taken during crash or disk error | Enable data_checksums in PostgreSQL; verify backups with pg_verifybackup; run ANALYZE and REINDEX after restore |
Examples
Point-in-time recovery after accidental table drop: At 14:30 a developer runs DROP TABLE orders on production. Recovery: restore the most recent physical backup (taken at 02:00), replay WAL files up to 14:29:59 using recovery_target_time, verify orders table is intact, then swap the recovered database into production. Total recovery time: 25 minutes for a 200GB database.
Automated daily backup to S3 with verification: A cron job runs at 02:00 UTC: takes pg_basebackup, compresses with zstd, uploads to S3 with server-side encryption, restores to a Docker container, runs row count checks against production, and sends a Slack notification with backup size and verification status.
Cross-region disaster recovery drill: Monthly exercise: restore the latest S3 backup to a different AWS region, replay WAL files to catch up, run the application test suite, measure total failover time, and document results. Target: full recovery in a different region within 1 hour.
Resources
- PostgreSQL backup and recovery: https://www.postgresql.org/docs/current/backup.html
- pgBackRest (production backup tool): https://pgbackrest.org/
- WAL-G (WAL archiving to cloud storage): https://github.com/wal-g/wal-g
- Percona XtraBackup: https://docs.percona.com/percona-xtrabackup/
- MySQL point-in-time recovery: https://dev.mysql.com/doc/refman/8.0/en/point-in-time-recovery.html
Frequently asked questions about Database Recovery Manager
Similar skills
ClickHouse Logs Queries
Efficiently manage Supabase logs with ClickHouse SQL.
EF Core D2 Database Diagram Generator
Visualize your EF Core models as D2 diagrams effortlessly.
Safe SQL Execution
Ensure secure SQL execution in Supabase applications.
Oracle to PostgreSQL Migration
Identify migration risks between Oracle and PostgreSQL.
SSMA Console
Streamline Oracle to SQL Server migrations with ease.
SQL Performance Optimization
Enhance SQL query efficiency across all databases.
