news.volyx.in

Meta torrented & seeded 81.7 TB dataset containing copyrighted data (arstechnica.com)

1270 points by gameshot911 · 543 days ago · 938 comments on HN

Article summary

Meta torrented and seeded a large dataset containing copyrighted data, totaling 81.7 TB. The dataset is believed to contain a large number of books, potentially in the tens of millions. The discussion revolves around the implications of this action, including potential copyright infringement and the use of the data for training large language models. The exact details of the article are not available, but the comments suggest that the dataset was obtained from a torrent and may have been used for training AI models.

Main themes

  • Copyright Infringement
  • AI Training Data
  • Corporate Accountability
  • Digital Piracy
  • Intellectual Property

What commenters say

  • The sheer size of the dataset suggests that Meta may have downloaded and seeded a massive number of copyrighted books, potentially exceeding 80 million files.
  • The use of copyrighted data for training AI models may be necessary for their development, but it raises questions about the legality and ethics of such practices.
  • Some argue that the laws surrounding copyright infringement are outdated and may need to be restructured to accommodate the growing importance of AI and its need for large amounts of training data.
  • Others believe that Meta's actions demonstrate a double standard in the justice system, where individuals are held to a different standard than corporations when it comes to copyright infringement.
  • The potential fines for Meta's actions could be substantial, potentially exceeding $6 billion, although some argue that this would not significantly impact the company's growth.
  • The case may set an important precedent for future cases involving copyright infringement and AI training data, with some arguing that it could lead to a re-evaluation of current laws and regulations.