news.volyx.in

OpenAI Furious DeepSeek Might Have Stolen All the Data OpenAI Stole from Us (404media.co)

1361 points by latexr · 552 days ago · 16 comments on HN

Article summary

OpenAI and Microsoft are investigating whether DeepSeek, a Chinese AI startup, improperly used OpenAI's data to train its large language model, R1. DeepSeek's model has been reported to outperform OpenAI's while using less money and older chips. The investigation centers on whether DeepSeek used a technique called distillation to learn from OpenAI's models without permission. This has raised questions about data ownership and usage in the AI industry.

Main themes

  • AI data ownership
  • Model distillation
  • Corporate espionage
  • AI industry ethics
  • Data usage agreements

What commenters say

  • The article's headline is clickbait and misrepresents OpenAI's stance on the issue, as there is no evidence they are 'furious'.
  • DeepSeek may have used corporate espionage to steal OpenAI's model weights and optimize them, which would explain the similarities in their outputs.
  • The use of distillation to learn from other models is a common technique in AI, but it raises questions about the ownership and usage of the original data.
  • OpenAI's criticism of DeepSeek is hypocritical, given their own history of collecting data without permission.
  • The article lacks evidence to support its claims and is sensationalized to provoke a reaction.
  • The investigation into DeepSeek's actions is necessary to clarify the rules around data usage and ownership in the AI industry.
  • The fact that OpenAI's model is closed source makes it difficult to determine whether distillation can be done effectively via their API.
  • The use of distillation to 'launder' ill-gotten data is a concern that extends beyond this specific incident and requires further discussion.