news.volyx.in

Summary of the AWS Service Event in the Northern Virginia (US-East-1) Region (aws.amazon.com)

638 points by eigen-vector · 1744 days ago · 388 comments on HN

Article summary

AWS experienced a service disruption in the Northern Virginia (US-EAST-1) Region on December 7th, 2021, due to an automated activity that triggered unexpected behavior from clients inside the internal network, resulting in congestion and performance issues. The issue impacted several AWS services, including EC2, RDS, and Route 53, and caused delays and errors for customers. AWS has taken actions to prevent a recurrence of the event, including disabling the scaling activities that triggered the issue and deploying additional network configuration. The company has also acknowledged the impact on customers and apologized for the disruption.

Main themes

  • AWS infrastructure complexity
  • Service disruption and outage
  • Regional vs zonal redundancy
  • Transparency and post-mortem analysis
  • Cloud service pricing and cost
  • High availability and scalability
  • Global service management and deployment

What commenters say

  • The complexity of AWS's infrastructure can lead to difficult-to-troubleshoot issues, even with a large team of operations staff.
  • Some commenters felt that the post-mortem analysis lacked detail and was too general, while others appreciated the transparency.
  • The issue highlighted the importance of replicating services across regions, rather than relying on availability zones within a region, to ensure high availability.
  • The us-east-1 region's age and scale may contribute to its higher incidence of issues, but this is not a sufficient excuse for the problems that occur.
  • AWS's decision to host some global services, such as the management API, in us-east-1 can cause problems for customers in other regions when that region experiences issues.
  • Some commenters argued that customers who only use one availability zone or region are partly to blame for their own availability issues, while others felt that AWS should be more proactive in preventing and mitigating such problems.
  • The cost of data replication between availability zones can be significant, and some commenters felt that AWS's pricing model is unfair in this regard.