news.volyx.in

CrowdStrike ex-employees: 'Quality control was not part of our process' (semafor.com)

565 points by everybodyknows · 694 days ago · 300 comments on HN

Article summary

Former employees of CrowdStrike, a cybersecurity firm, have come forward to describe a work environment where speed was prioritized over quality control, leading to rushed deadlines, excessive workloads, and increasing technical problems. This culture allegedly contributed to a catastrophic failure of the company's software, which paralyzed airlines and knocked banking and other services offline for hours. The company has disputed these claims, stating that it is committed to ensuring the resiliency of its products through rigorous testing and quality control. The incident has sparked a significant loss in stock-market value and multiple lawsuits.

Main themes

  • Cybersecurity
  • Quality control
  • Software development
  • Corporate culture
  • Technical failures
  • Accountability

What commenters say

  • The article's claims about CrowdStrike's prioritization of speed over quality control are not credible due to the sources being disgruntled former employees.
  • The company's lack of staged deployments and inadequate testing procedures are clear indications of poor quality control and a contributing factor to the software failure.
  • Even with the pressure to quickly respond to emerging threats, basic testing and quality control measures, such as canary releases, should always be implemented to prevent catastrophic failures.
  • The dichotomy between code updates and data updates can be dangerous if not properly understood, and CrowdStrike's alleged failure to recognize this distinction may have contributed to the incident.
  • The need for swift deployment of security updates to counter immediate threats does not justify bypassing essential testing and validation procedures.
  • CrowdStrike's actions, including the handling of the software failure, demonstrate a lack of accountability and a culture that prioritizes speed over customer safety and security.
  • The incident highlights the importance of robust error handling and the treatment of configuration data as untrusted input to prevent similar failures in the future.
  • While the pressure to quickly mitigate security threats is real, it does not excuse the failure to implement basic safeguards and testing, which could have prevented the widespread outage.