Namecheap’s Outage Shows How One Data Center Can Disrupt an Entire Hosting Platform

Namecheaps Outage Shows How One Data Center Can Disrupt An Entire Hosting Platform
Follow Us:
2.7k
1k

Namecheap took more than 5,000 servers offline on August 13 after a cooling system failure at RadiusDC’s Phoenix facility put the equipment at risk of overheating.

“Continuing to operate systems in such conditions risked equipment overheating and suffering longer-term and catastrophic damage,” Namecheap later stated in an official update.

Namecheap services were down for

~0 hours

The outage lasted somewhere between 28 and 30 hours, according to different sources, with sites either unavailable, slow, or returning 404s. Services returned at different times on August 14.

“RadiusDC: Phoenix is a Tier 3 datacenter, designed with a high level of redundancy. That made this an extremely unusual incident, and one we had not experienced before in our 25-year history,” Namecheap wrote in its blog.

One Cooling Failure Took Down Several Services

Namecheap's Phoenix location supports its shared and reseller hosting, Private Email servers, VPS, and dedicated servers, meaning the outage affected EasyWP, Private Email, DNS management, billing and account functions, and customer support. Getting those systems back online required restoring core networking, virtualization, databases, load balancers, and other systems.

Namecheap has infrastructure in Phoenix, Europe, and Asia, but having infrastructure in multiple locations doesn’t necessarily mean automatic failover.

Namecheap made the decision to take down

0 servers

Geographic redundancy means infrastructure exists in multiple locations, which lessens the dependence on a single physical site. Automatic workload failover is when a service can actually continue operating from a different location if its primary environment becomes unavailable.

Namecheap has not publicly shared whether all of the services affected by the Phoenix outage were capable of automatic cross-region failover. It also doesn't share the specifics of its architecture, though we do know some of its infrastructure is distributed across multiple regions.

Its PremiumDNS service, for example, uses globally distributed infrastructure which is why DNS resolution stayed online. Although customers couldn't access Namecheap's DNS management dashboards, they could still use their domains.

Namecheap Plans for More Redundancy

To restore cooling, RadiusDC brought in temporary chillers to help lower temperatures in Namecheap’s area of the data center. Once it was safe, Namecheap began bringing its infrastructure back online in stages.

“Restoring them meant bringing these different elements back online in the right order, rather than simply switching everything back on at once,” Namecheap said.

Namecheap is also conducting a broader review of the incident to better identify where its existing safeguards and protocols failed. In the meantime, it also plans to increase redundancy across its data centers, which will include its Amsterdam and Singapore locations.

If this happens again, will services automatically fail over? Will important data be backed up in another region? Will systems be able to keep running if an entire data center goes down?

Right now, we obviously don’t know. But we do know that more data centers don’t automatically mean automatic resilience, and more servers doesn’t mean guaranteed redundancy.