A multi-facility healthcare provider ran patient records, scheduling, and clinical decision-support systems on infrastructure with a documented disaster recovery plan that had never actually been tested end-to-end. When an internal audit simulated a regional outage, the honest recovery time came in at just over six hours — far outside what patient safety and regulatory requirements demanded.
The existing DR documentation described a plan that, on paper, looked reasonable. In practice, key recovery steps depended on individuals who had since left the organisation, backup restoration processes had never been timed under realistic load, and failover procedures for the clinical decision-support system were entirely undocumented.
We started by throwing out the assumption that the existing DR plan was a reliable baseline. Every recovery procedure was re-derived from a fresh business-impact assessment, defining actual tolerable downtime per system rather than relying on generic targets.
The rebuilt DR strategy reduced actual, tested recovery time for critical systems from over six hours to twelve minutes. The provider now runs quarterly live failover drills as a standing operational practice, with clear ownership and documentation that doesn't depend on any single person's knowledge.
Related Service
Platform Reliability & Disaster Recovery →We can help you find out before an outage does.