news.volyx.in

Tarsnap outage postmortem (mail.tarsnap.com)

553 points by anderiv · 1127 days ago · 319 comments on HN

Article summary

Tarsnap, a backup service, experienced an outage due to a hardware failure in Amazon's EC2 us-east-1 region. The service was offline for approximately 26 hours and 16 minutes. The founder, Colin Percival, provided a post-mortem analysis of the incident, detailing the causes and steps taken to recover the service. The outage was resolved by spinning up a new EC2 instance and replaying log entries from Amazon S3.

Main themes

  • backup services
  • outage post-mortem
  • reliability and redundancy
  • SLAs and service guarantees
  • data importance and prioritization
  • cloud infrastructure and scalability

What commenters say

  • Some commenters appreciate the transparency and honesty of the post-mortem analysis, while others are concerned about the lack of redundancy and potential single point of failure in the service.
  • There is a debate about the importance of Service Level Agreements (SLAs) and their effectiveness in ensuring service reliability.
  • Some users prefer Tarsnap over other backup services like restic due to its granular pricing and reliability, while others are hesitant due to the lack of redundancy and potential risks.
  • The discussion also touches on the idea that all data is 'super important' to its owner, but objectively, some data may be more critical than others.
  • Redundancy and having multiple backup providers is seen as a best practice to ensure data safety.
  • SLAs are seen as more of a contractual obligation than a guarantee of service reliability, and their effectiveness is questioned.
  • The post-mortem analysis highlights the trade-offs between service reliability, cost, and complexity.