The Most Expensive Typo in History
In 2017 one mistyped command at Amazon took a huge slice of the internet offline for four hours, and the status page meant to report it was down too, because it ran on the very service that had failed.
On the afternoon of 28 February 2017, a single mistyped command took a large slice of the internet offline. Slack, Trello, Quora, Medium and thousands of other sites stuttered or went dark for around four hours. The culprit was not a hacker or a hardware failure. It was a typo, entered by an authorised engineer at Amazon, and it remains the clearest lesson the cloud has ever handed out about putting too much in one place.

A routine command on a busy afternoon
The engineer was working on the billing system for Amazon S3, the storage service that quietly holds a vast portion of the web's images, files and backups, in the US-EAST-1 region in Northern Virginia. Following an established playbook, they ran a command to take a small number of servers offline. One of the inputs was entered incorrectly, and instead of a small set, a much larger set of servers was removed.
The blast radius
That mistake knocked out two of S3's core subsystems in the region: the index, which keeps track of where every stored object actually lives, and the placement system, which decides where new data goes. Both had to be fully restarted, and because they had grown enormous and had not been restarted in years, the safety checks took hours to complete. While they did, anything depending on S3 in that region was stuck. Since US-EAST-1 is the oldest and most heavily used corner of Amazon's cloud, that turned out to be a frightening amount of the internet.

The detail everyone remembers
The part that turned a serious outage into an industry legend was the status dashboard. As services fell over, Amazon's own Service Health Dashboard kept showing reassuring green ticks. The reason was almost too perfect: the dashboard's status icons were themselves hosted on the S3 service that was down. The one tool meant to tell the world that S3 had failed could not work, because S3 had failed. Engineers ended up posting updates by hand from a bare page while they scrambled, the digital equivalent of keeping the fire extinguisher inside the burning room.

What it taught everyone
Amazon's follow-up was candid and practical. It changed the tool so it could no longer remove too many servers too quickly, audited similar tools for the same flaw, and moved its status dashboard to run across multiple regions so it could never again be taken down by the thing it was meant to monitor. The deeper lesson was about concentration. An astonishing number of companies had quietly built their entire presence on a single region of a single provider, and discovered all at once that they shared a single point of failure they had never thought about.
It is the same lesson that the OVH fire taught from the opposite direction. One was fire, one was a typo, and both ended in the same place: spread your risk. Real resilience means more than one region and more than one copy, which is part of why the geographic spread of a provider's data centres and network is worth a serious look. If even Amazon can be felled by a stray keystroke, the safe assumption is that your host can too, and the answer is to design as though it will. Our directory and plan comparison are a place to weigh providers on exactly that, and cloud hosting is only as resilient as the way you spread it out.
0 Comments
No comments yet
Be the first to share your thoughts on this article.