news.volyx.in

Cloudflare outage on July 17, 2020 (blog.cloudflare.com)

522 points by tomklein · 2269 days ago · 213 comments on HN

Article summary

Cloudflare experienced a 27-minute outage on July 17, 2020, due to a configuration error in their backbone network. The error caused all traffic across the backbone to be sent to Atlanta, overwhelming the router and causing a failure in several network locations. The outage was not caused by an attack or breach, and Cloudflare has made changes to prevent it from happening again. The company has introduced a maximum-prefix limit on their backbone BGP sessions and plans to change the BGP local-preference.

Main themes

  • Cloudflare outage
  • change management
  • BGP and IGP's
  • network reliability
  • transparency and communication
  • compliance and regulation
  • error prevention and testing
  • reliability and risk management

What commenters say

  • The outage highlights the importance of proper change management and testing to prevent similar errors in the future.
  • The use of BGP and current IGP's can be problematic and may need to be replaced with more fail-safe protocols.
  • Cloudflare's transparency in communicating the cause of the outage and their efforts to mitigate it is commendable.
  • Some argue that compliance requirements, such as change management processes and incident response playbooks, are essential for preventing similar outages.
  • Others believe that having multiple people review changes and automated testing processes in place can help prevent errors.
  • The incident raises questions about the reliability of Cloudflare's services and the potential risks of relying on a single provider.
  • There are differing opinions on whether Cloudflare's post-mortem analysis was thorough and transparent enough.
  • The discussion highlights the challenges of balancing the need for rapid changes to mitigate issues with the need for careful testing and review to prevent errors.