news.volyx.in

News publishers limit Internet Archive access due to AI scraping concerns (niemanlab.org)

569 points by ninjagoo · 156 days ago · 363 comments on HN

Article summary

News publishers such as The Guardian and The New York Times are limiting the Internet Archive's access to their content due to concerns about AI companies scraping their articles. The Internet Archive operates crawlers that capture webpage snapshots, which can be accessed through the Wayback Machine, but this has raised concerns about AI companies using this data to train their models. Some publishers are blocking the Internet Archive's bots or restricting access to their content, while others are working with the Internet Archive to implement changes. This issue highlights the tension between preserving the internet's historical record and protecting intellectual property.

Main themes

  • Internet Archive
  • AI scraping
  • News publishers
  • Content preservation
  • Intellectual property
  • Web archiving

What commenters say

  • The Internet Archive's efforts to preserve the internet's historical record are being hindered by news publishers' concerns about AI scraping, which may ultimately harm the public's access to information.
  • News publishers have a right to protect their intellectual property and should be able to control how their content is used, even if it means limiting access to the Internet Archive.
  • The rise of AI-generated content may render the preservation of the internet's historical record less valuable, but it is still important to preserve accurate and trustworthy information for future generations.
  • A crowd-sourced archiving effort, potentially through a browser extension, could be a solution to preserve the internet's historical record while respecting news publishers' concerns about AI scraping.
  • Some commenters argue that serving up content to the public should imply archivability, while others believe that news publishers should have control over how their content is used and archived.
  • The use of rate limiting and other technical measures may not be sufficient to prevent AI companies from scraping content, and more robust solutions may be needed to protect intellectual property.