Disaster recovery

High availability approaches

Network redundancy lives at three layers: devices, NICs and cables (servers run 2+ NICs for redundancy or load balancing), and router and switch paths (redundant internal and internet paths).

Active-active runs all systems live and sharing load, for maximum utilization. Active-passive keeps a standby idling until the primary fails: a reliable fallback.

Load balancers spread traffic and reroute around failed nodes with health checks. CDNs cache content on geographically distributed servers near users and reroute on failure or overload.

Designing redundant networks

Decide redundancy at the module and chassis level (power supplies, drives, whole routers), weigh cost per option, and lean on software redundancy where hardware isn’t needed. Protocol choice matters: TCP resends, UDP won’t.

Power and environmental redundancy (UPS, generators, HVAC) scale with uptime criticality. Set technical goals (uptime %) and performance standards up front. Designing redundancy in from the start is far cheaper than retrofitting.

Everything trades off time, cost, and quality.

DR metrics

Redundant sites

Platform diversity (varied OS, network vendors, cloud providers) and geographic dispersion kill single points of failure. Same site breakdown as my Security+ resilience note.

Training and exercises

A tabletop exercise is scenario discussion: cheap, theoretical. Penetration testing is live attack simulation with real tools; scope carefully, and prefer third parties or a separate internal red team.

The teams: red attacks, blue defends (sysadmins, network defenders, analysts), and white administers, referees, builds the simulated environment, and reports outcomes.