news.volyx.in

Roblox October Outage Postmortem (blog.roblox.com)

687 points by kbuck · 1701 days ago · 310 comments on HN

Article summary

Roblox experienced a 73-hour outage in October due to issues with their Consul cluster, which was caused by a combination of a new streaming feature and a pathological performance issue in BoltDB. The company has shared a postmortem of the incident, detailing the challenges they faced in diagnosing and resolving the issue. Roblox is taking steps to prevent similar outages in the future, including improving their monitoring and moving to multiple availability zones and data centers. The incident highlighted the importance of designing systems for high availability and redundancy.

Main themes

  • outage postmortem
  • high availability
  • redundancy
  • cloud infrastructure
  • service level agreements
  • system design
  • cost vs reliability tradeoff

What commenters say

  • Having multiple fully independent zones can improve reliability and failsafe, but it also introduces new modes of failure and increased costs.
  • Independent zones may not be completely independent due to common-mode failures that can cause outages to propagate across zones and regions.
  • The cost of implementing multiple availability zones may be offset by the potential loss of revenue and user trust during outages.
  • Cloud providers' service level agreements and availability guarantees may not be fully understood by users, leading to unrealistic expectations about uptime and redundancy.
  • Implementing services and architecture that can sidestep common failure modes is crucial for high availability.
  • The use of multiple cloud providers can provide an additional layer of redundancy and protection against outages.
  • The trade-off between the cost of implementing redundancy and the potential cost of outages is a key consideration for companies designing their systems.