t was Amazon Simple Storage Service, better known as S3, in the Northern Virginia region called US-EAST-1. At 9:37 a.m. Pacific Time on 28 February 2017, an authorized S3 team member entered one command input incorrectly while investigating a slow billing process, causing more servers to be removed than intended. According to Amazon’s official post-mortem, the GET, LIST and DELETE APIs were fully restored after three hours and 41 minutes, while complete S3 operations returned after four hours and 17 minutes.
Amazon described the initiating error plainly: “Unfortunately, one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended.” The command came from an established operational playbook and was supposed to remove only a small number of servers supporting an S3 billing subsystem.
The unexpectedly large removal also took capacity away from two much more important systems. The index subsystem tracked the metadata and location of every S3 object in the region, while the placement subsystem decided where newly uploaded objects would be stored. Without enough capacity in those systems, S3 could not process its normal requests.
This shows the danger of running commands you do not know and understand.
The biggest takeaway from this for me was this paragraph:
We are making several changes as a result of this operational event. While removal of capacity is a key operational practice, in this instance, the tool used allowed too much capacity to be removed too quickly. We have modified this tool to remove capacity more slowly and added safeguards to prevent capacity from being removed when it will take any subsystem below its minimum required capacity level. This will prevent an incorrect input from triggering a similar event in the future. We are also auditing our other operational tools to ensure we have similar safety checks. We will also make changes to improve the recovery time of key S3 subsystems. We employ multiple techniques to allow our services to recover from any failure quickly. One of the most important involves breaking services into small partitions which we call cells. By factoring services into cells, engineering teams can assess and thoroughly test recovery processes of even the largest service or subsystem. As S3 has scaled, the team has done considerable work to refactor parts of the service into smaller cells to reduce blast radius and improve recovery. During this event, the recovery time of the index subsystem still took longer than we expected. The S3 team had planned further partitioning of the index subsystem later this year. We are reprioritizing that work to begin immediately.
Let’s hope we are never reminded again just how much of the web runs on Amazon AWS again like this. Usually the question is not if, but when! Especially now at this time with so much new/young-technology being integrated into production systems and processes.