Cyber resilience and redundancy

High availability

The goal is keeping services up by minimizing downtime, via load balancing, clustering, redundancy, and multi-cloud.

Uptime is a percentage. Five nines (99.999%) is about 5 minutes of downtime a year; six nines (99.9999%) is about 31 seconds.

Load balancing spreads workload across servers so none is overloaded. Clustering ties multiple machines into one logical system for HA and scaling, surviving hardware failure. Redundancy duplicates critical components (power supplies, links, servers, services, providers) to kill single points of failure, and multi-cloud spreads across providers to avoid lock-in and single-provider failure.

RAID

The resilience categories: failure-resistant (RAID 1), fault-tolerant (1/5/6/10), disaster-tolerant (1 and 10, with data in independent zones).

Powering data centers

The power events: surge (small over-voltage), spike (short over-voltage from shorts or lightning), sag (a brief drop), brownout (prolonged undervoltage forcing shutdown), blackout (full power loss).

The protection: line conditioners (stabilize and filter, but not for outages), UPS (battery backup, typically 15-60 min, plus line conditioning), generators (grid outage backup, needing startup time), PDCs (central distribution with monitoring and load balancing).

Backups

Onsite is convenient but exposed to local disasters; offsite is geographically separate and disaster-safe.

Frequency is set by the RPO (Recovery Point Objective): the max data loss you’ll tolerate.

Encrypt backups at rest and in transit. Snapshots capture point-in-time state and store only changes. Replication is real-time copying for HA, and journaling logs changes over time for granular recovery and an audit trail.

BC/DR planning

COOP (a Continuity of Operations Plan) ensures recovery from disruptive events. The BC plan is the broad response to disruptions; the DRP is a subset focused on faster recovery after disasters (fires, floods, hurricanes).

Senior management owns the BC plan, sets goals and scope by risk appetite, and appoints a BC coordinator to lead the BC committee (IT, Legal, Security, Comms) that sets recovery priorities.

Redundant sites

Virtual sites do the same in the cloud (hot, warm, cold). Platform diversity (varied OS, network gear, cloud) and geographic dispersion reduce single-point and localized-outage risk.

Resilience and recovery testing