news.volyx.in

Fly.io Postgres cluster down for 3 days, no word from them about it (webcache.googleusercontent.com)

797 points by burnerbob · 1133 days ago · 477 comments on HN

Article summary

The article discusses a recent outage of Fly.io's Postgres cluster, which has been down for three days with no communication from the company. Users are expressing frustration and disappointment with the lack of transparency and support. The incident has sparked a discussion about the reliability and customer service of cloud providers. Fly.io's CEO had previously acknowledged the need to improve reliability, but the current outage has raised concerns about the company's ability to deliver on this promise.

Main themes

  • Cloud provider reliability
  • Customer support and communication
  • Outage and downtime management
  • Trade-offs between cloud providers
  • DIY and self-hosting options
  • Infrastructure and hardware management

What commenters say

  • Some users feel that Fly.io's lack of communication and transparency during the outage is unacceptable and has damaged their trust in the company.
  • Others argue that the outage is a reminder that even with cloud providers, things can go wrong, and it's essential to have a plan in place for such events.
  • There is a debate about the trade-offs between using a cloud provider like AWS, which can be more expensive but offers better support and reliability, and using a smaller provider like Fly.io, which may be more affordable but has its own set of challenges.
  • A few commenters suggest that enthusiasts and small businesses might be better off hosting their own servers in a local colo or using a DIY approach, rather than relying on cloud providers.
  • Some users appreciate the honesty of Fly.io's representative, who acknowledged the company's shortcomings and promised to improve communication and support.
  • Others are skeptical about the company's ability to follow through on these promises and are considering alternative providers.
  • There is a discussion about the importance of having a 'throat to choke' or a single point of contact when issues arise, and how this can impact the overall customer experience.
  • The outage has also raised questions about the reliability and durability of Fly.io's infrastructure and the company's ability to scale and manage its hardware.