news.volyx.in

Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge (businessinsider.com)

497 points by pyman · 387 days ago · 651 comments on HN

Article summary

Anthropic, a company backed by Amazon and Alphabet, has been training its AI chatbot Claude using millions of copyrighted books, including over 7 million pirated ones. A judge ruled that Anthropic's use of copyrighted books to train its AI models is fair use, but downloading pirated copies is not. The company had spent millions of dollars buying used print books, which were then scanned and discarded. The case highlights the ongoing debate over AI copyright and fair use.

Main themes

  • AI copyright
  • fair use
  • piracy
  • profiting from piracy
  • transformative use
  • copyright law
  • AI training data

What commenters say

  • Pirating 7 million books to train an AI model is a form of theft and profiting from piracy is unacceptable.
  • The use of copyrighted material to train AI models is fair use as it is transformative and does not harm the original creators.
  • The law on copyright and fair use is still evolving and unclear, particularly in the context of AI.
  • Individuals who pirate content for personal use are different from companies that profit from piracy at scale.
  • Copyright infringement is not the same as stealing, and the distinction is important in the context of AI training data.
  • The fact that AI models can generate content that is similar to the original copyrighted material raises questions about the boundaries of fair use.
  • The use of pirated content to train AI models undermines the ability of creators to earn a living from their work.
  • The issue of AI copyright and fair use is complex and requires a nuanced approach that balances the interests of creators and society.