news.volyx.in

A terrible, horrible, no-good, very bad day at Slack (slack.engineering)

631 points by ceohockey60 · 2280 days ago · 270 comments on HN

Article summary

Slack experienced a significant outage on May 12, 2020, due to a technical issue with their HAProxy server state management. The problem was caused by a bug in the program that synced the host list with the HAProxy server state, which led to a stale list of backends and ultimately resulted in the outage. The issue was resolved with a rolling restart of the HAProxy fleet. The incident highlighted the importance of monitoring and testing, particularly for systems that have been working reliably for a long time.

Main themes

  • outage postmortem
  • HAProxy configuration
  • monitoring and alerting
  • chaos engineering
  • disaster recovery
  • downtime cost
  • testing and drilling
  • production data usage

What commenters say

  • Using production data for QA is risky and can lead to data leaks and corruption.
  • Chaos engineering can help uncover hidden issues, but it is not a silver bullet and requires careful planning and execution.
  • Monitoring and alerting systems are crucial for detecting issues, but they can become stale and ineffective if not regularly tested and updated.
  • Testing and drilling for failure scenarios is essential for disaster recovery and can help prevent costly outages.
  • The cost of downtime can be significant, but it is difficult to estimate and may vary depending on the company and the nature of the outage.
  • Regularly testing and updating monitoring and alerting systems can help prevent issues like the one described in the article.
  • The effectiveness of chaos engineering depends on the specific use case and the ability to simulate realistic failure scenarios.
  • There is a trade-off between the cost of testing and drilling for failure scenarios and the potential cost of downtime and reputational damage.