news.volyx.in

Tell NYT, Atlantic, USA Today to keep Wayback Machine (savethearchive.com)

435 points by doener · 109 days ago · 121 comments on HN

Article summary

Major media outlets such as The New York Times, The Atlantic, and USA Today are blocking the Internet Archive's Wayback Machine from preserving their content, citing concerns about AI and data harvesting. A petition is calling on these outlets to stop blocking the Wayback Machine and work with the Internet Archive to preserve their content. The Internet Archive has been preserving news articles for decades, but some outlets are now blocking it to maintain control over their content and revenue streams. This move has sparked debate about the importance of preserving online news and the role of the Wayback Machine in maintaining a historical record of the internet.

Main themes

  • News preservation
  • Internet Archive
  • Paywalls and revenue
  • AI and data harvesting
  • Journalism and historical record
  • Online archiving

What commenters say

  • Blocking the Wayback Machine will ultimately harm the news outlets themselves by reducing their online presence and accessibility.
  • The Internet Archive's preservation of news articles is crucial for maintaining a historical record of the internet and should be supported by news outlets.
  • News outlets have a right to control their content and revenue streams, and blocking the Wayback Machine is a necessary measure to protect their interests.
  • The use of paywalls and data harvesting is a flawed business model for news outlets, and they should consider alternative approaches that prioritize accessibility and preservation.
  • The debate over the Wayback Machine highlights the tension between the need for news outlets to generate revenue and the importance of preserving online content for historical and research purposes.
  • Some commenters argue that the Internet Archive's preservation of news articles is not a significant loss, as other sources can provide similar information and context.
  • Others believe that the Internet Archive plays a unique role in preserving online content and that its loss would be a significant blow to research and historical record-keeping.
  • The issue of robots.txt and the Internet Archive's respect for it is a complex one, with some arguing that the Archive should respect the directive and others arguing that it is not a reliable or effective way to control crawling and indexing.