administering-linux
Manage Linux systems covering systemd services, process management, filesystems, networking, performance tuning, and troubleshooting. Use when deploying…
Design and implement disaster recovery strategies with RTO/RPO planning, database backups, Kubernetes DR, cross-region replication, and chaos engineering testing. Use when implementing backup systems, configuring point-in-time recovery, setting up multi-region failover, or
$ npx -y skills add ancoleman/ai-design-components --skill planning-disaster-recovery --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/planning-disaster-recoveryContext preview
The summary Claude sees to decide when to auto-load this skill.
Design and implement disaster recovery strategies with RTO/RPO planning, database backups, Kubernetes DR, cross-region replication, and chaos engineering testing. Use when implementing backup systems, configuring point-in-time recovery, setting up multi-region failover, or
name: planning-disaster-recovery description: Design and implement disaster recovery strategies with RTO/RPO planning, database backups, Kubernetes DR, cross-region replication, and chaos engineering testing. Use when implementing backup systems, configuring point-in-time recovery, setting up multi-region failover, or validating DR procedures.
Provide comprehensive guidance for designing disaster recovery (DR) strategies, implementing backup systems, and validating recovery procedures across databases, Kubernetes clusters, and cloud infrastructure. Enable teams to define RTO/RPO objectives, select appropriate backup tools, configure automated failover, and test DR capabilities through chaos engineering.
Invoke this skill when:
**Recovery Time Objective (RTO):** Maximum acceptable downtime after a disaster before business impact becomes unacceptable.
**Recovery Point Objective (RPO):** Maximum acceptable data loss measured in time. Defines how far back in time recovery must reach.
**Criticality Tiers:**
Maintain **3 copies** of data on **2 different media** types with **1 copy offsite**.
Example implementation:
**Full Backup:** Complete copy of all data. Slowest to create, fastest to restore.
**Incremental Backup:** Only changes since last backup. Fastest to create, requires full + all incrementals to restore.
**Differential Backup:** Changes since last full backup. Balance between storage and restore speed.
**Continuous Backup:** Real-time or near-real-time backup via WAL/binlog archiving. Lowest RPO.
RTO < 1 hour, RPO < 5 min → Active-Active replication, continuous archiving, automated failover → Tools: Aurora Global DB, GCS Multi-Region, pgBackRest PITR → Cost: Highest RTO 1-4 hours, RPO 15-60 min → Warm standby, incremental backups, automated failover → Tools: pgBackRest, WAL-G, RDS Multi-AZ → Cost: High RTO 4-24 hours, RPO 1-6 hours → Daily full + incremental, cross-region backup → Tools: pgBackRest, Velero, Restic → Cost: Medium RTO > 24 hours, RPO > 6 hours → Weekly full + daily incremental, single region → Tools: pg_dump, mysqldump, S3 versioning → Cost: Low
| Use Case | Primary Tool | Alternative | Key Feature | |----------|-------------|-------------|-------------| | PostgreSQL production | pgBackRest | WAL-G | PITR, compression, multi-repo | | MySQL production | Percona XtraBackup | WAL-G | Hot backups, incremental | | MongoDB | Atlas Backup | mongodump | Continuous backup, PITR | | Kubernetes cluster | Velero | ArgoCD + Git | PV snapshots, scheduling | | File/object backup | Restic | Duplicity | Encryption, deduplication | | Cross-region replication | Aurora Global DB | RDS Read Replica | Active-Active capable |
**Use Case:** Production PostgreSQL with < 5 minute RPO
**Quick Start:** See `examples/postgresql/pgbackrest-config/`
Configure continuous WAL archiving with full/differential/incremental backups to S3/GCS/Azure. Schedule weekly full, daily differential backups. Enable PITR with `pgbackrest --stanza=main --delta restore`.
**Detailed Guide:** `references/database-backups.md#postgresql`
**Use Case:** MySQL production requiring hot backups
**Quick Start:** See `examples/mysql/xtrabackup/`
Perform full (`xtrabackup --backup --parallel=4`) and incremental backups with binary log archiving for PITR. Restore requires decompress, prepare, apply incrementals, and copy-back steps.
**Detailed Guide:** `references/database-backups.md#mysql`
**Quick Start:** Use `mongodump --gzip --numParallelCollections=4` for logical backups or MongoDB Atlas for continuous backup with PITR.
**Detailed Guide:** `references/database-backups.md#mongodb`
**Quick Start:** `velero install --provider aws --bucket my-backups`
Configure scheduled backups (daily full, hourly production namespace) with PV snapshots. Restore with `velero restore create --from-backup <name>`. Support selective restore (namespace mappings, storage class remapping).
**Examples:** `examples/kubernetes/velero/` **Detailed Guide:** `references/kubernetes-dr.md`
**Quick Start:** `ETCDCTL_API=3 etcdctl snapshot save /backups/etcd/snapshot.db`
Create periodic etcd snapshots for control plane recovery. Restore requires cluster recreation with snapshot data.
**Examples:** `examples/kubernetes/etcd/`
**Key Services:**
**Examples:** `examples/cloud/aws/` **Detailed Guide:** `references/cloud-dr-patterns.md#aws`
**Key Services:**
Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude
Repo: ancoleman/ai-design-components
Manage Linux systems covering systemd services, process management, filesystems, networking, performance tuning, and troubleshooting. Use when deploying…
Data pipelines, feature stores, and embedding generation for AI/ML systems. Use when building RAG pipelines, ML feature serving, or data transformations.…
Strategic guidance for designing modern data platforms, covering storage paradigms (data lake, warehouse, lakehouse), modeling approaches (dimensional,…
Design cloud network architectures with VPC patterns, subnet strategies, zero trust principles, and hybrid connectivity. Use when planning VPC topology,…
Design comprehensive security architectures using defense-in-depth, zero trust principles, threat modeling (STRIDE, PASTA), and control frameworks (NIST CSF,…
Assembles component outputs from AI Design Components skills into unified, production-ready component systems with validated token integration, proper import…