Pilot Light DR
A disaster recovery strategy keeping critical data replicated while compute resources remain inactive until failover.
Last reviewed: July 25, 2026
Pilot light disaster recovery is a low-cost DR strategy where a minimal, always-on version of critical infrastructure runs continuously in a secondary region — typically just the data layer, kept in sync via replication — while the compute resources needed to serve real traffic remain switched off until an actual disaster is declared.
How It Works
The name comes from the small pilot flame in a gas heater that stays lit at all times so the full burner can ignite quickly when needed, rather than requiring a cold start from nothing. In cloud architecture, this typically means keeping a database replica continuously updated in the DR region (so data is never stale by more than the replication lag), along with pre-configured but stopped compute resources — AMIs, container images, or infrastructure-as-code templates ready to deploy — that only get launched and scaled up when a failover is triggered.
The Tradeoff
Pilot light is one of the cheapest DR strategies available, since the DR region incurs almost no compute cost during normal operation — only storage and replication costs for the standing data layer. The cost is recovery time: because compute has to be provisioned and scaled up from scratch during an actual disaster, recovery time objectives (RTO) for pilot light are measured in tens of minutes to a few hours, rather than the near-instant failover possible with a warm standby or active-active architecture.
When It’s the Right Choice
Pilot light suits workloads where a moderate RTO is acceptable and where minimizing steady-state cost matters more than achieving the fastest possible recovery — often used for internal systems or lower-tier applications, while customer-facing critical systems tend to justify the higher cost of a warm standby or fully active-active setup instead.
Automating Pilot Light Failover
Because pilot light DR requires standing up compute resources from a cold state during an actual disaster, the speed of that failover heavily depends on how much of the process is automated in advance — infrastructure-as-code templates that can provision the necessary compute in a single command, combined with pre-tested runbooks, dramatically reduce the recovery time compared to manually working through infrastructure provisioning during an active incident. Organizations serious about pilot light DR typically conduct periodic failover drills specifically to validate that the automation actually works as expected and that the recovery time objective can genuinely be met, rather than only discovering gaps in the automation during a real disaster when there’s no room for trial and error.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.