news.volyx.in

Inside the longest Atlassian outage (newsletter.pragmaticengineer.com)

1232 points by andyjohnson0 · 1614 days ago · 738 comments on HN

Article summary

Atlassian experienced its longest outage, lasting over 9 days, affecting around 400 companies and 50,000 to 800,000 users. The cause of the outage was a script that accidentally deleted customer data, and restoring it has been a complex process. Atlassian's communication during the outage was criticized for being inadequate and untransparent. The company has since apologized and is working to restore service and provide a post-incident review.

Main themes

  • Atlassian outage
  • communication failures
  • data restoration challenges
  • technical debt
  • cloud services reliability
  • incident response

What commenters say

  • Atlassian's communication during the outage was inadequate and untransparent, exacerbating the issue.
  • The difficulty in selectively restoring data for affected users suggests a lack of isolation between users, which is a technical concern.
  • Sharding customer databases 1:1 could have prevented the issue, but it also has downsides such as increased complexity and potential performance issues.
  • Atlassian's tech debt and monolithic architecture choices may have contributed to the outage and restoration challenges.
  • Transparency is always better, even in bad situations, and Atlassian's lack of transparency was a major issue.
  • The outage has significant implications for Atlassian's business, particularly as it transitions customers to its cloud offering.
  • Atlassian's engineering practices and incident response process have been called into question, with some arguing that they are not up to standard.