Key Takeaways
Namecheap took more than 5,000 servers offline on August 13 after a cooling system failure at RadiusDC’s Phoenix facility put the equipment at risk of overheating.
“Continuing to operate systems in such conditions risked equipment overheating and suffering longer-term and catastrophic damage,” Namecheap later stated in an official update.
Namecheap services were down for
The outage lasted somewhere between 28 and 30 hours, according to different sources, with sites either unavailable, slow, or returning 404s. Services returned at different times on August 14.
“RadiusDC: Phoenix is a Tier 3 datacenter, designed with a high level of redundancy. That made this an extremely unusual incident, and one we had not experienced before in our 25-year history,” Namecheap wrote in its blog.
One Cooling Failure Took Down Several Services
Namecheap’s Phoenix infrastructure supported its hosting products, but also affected EasyWP, Private Email, DNS management, billing and account functions, and customer support. Getting those systems back online meant restoring core networking, virtualization, databases, load balancers, and other systems.
Namecheap has infrastructure in Phoenix, Europe, and Asia, but having infrastructure in multiple locations doesn’t necessarily mean automatic failover.
Namecheap made the decision to take down
Geographic redundancy means infrastructure exists in multiple locations, which lessens the dependence on a single physical site. Automatic workload failover is when a service can actually continue operating from a different location if its primary environment becomes unavailable.
Namecheap has not publicly shared whether all of the services affected by the Phoenix outage were capable of automatic cross-region failover.
But there is one exception to this two-day disruption: DNS resolution stayed online. Although customers couldn’t access Namecheap’s DNS management dashboards, they could still use their domains. That’s because Namecheap’s PremiumDNS service uses globally distributed Anycast infrastructure, which allows DNS requests to be routed to available servers in different locations.
Namecheap Plans for More Redundancy
To restore cooling, RadiusDC brought in temporary chillers to help lower temperatures in Namecheap’s area of the data center. Once it was safe, Namecheap began bringing its infrastructure back online in stages.
“Restoring them meant bringing these different elements back online in the right order, rather than simply switching everything back on at once,” Namecheap said.
Namecheap is also conducting a broader review of the incident to better identify where its existing safeguards and protocols failed. In the meantime, it also plans to increase redundancy across its data centers, which will include its Amsterdam and Singapore locations.
If this happens again, will services automatically fail over? Will important data be backed up in another region? Will systems be able to keep running if an entire data center goes down?
Right now, we obviously don’t know. But we do know that more data centers don’t automatically mean automatic resilience, and more servers doesn’t mean guaranteed redundancy.
