news.volyx.in

A small number of samples can poison LLMs of any size (anthropic.com)

1202 points by meetpateltech · 290 days ago · 439 comments on HN

Article summary

A study found that a small number of malicious documents, as few as 250, can create a backdoor vulnerability in large language models, regardless of model size or training data volume. This challenges the assumption that attackers need to control a percentage of training data to succeed. The study demonstrated that the absolute number of poisoned documents, not the percentage of training data, determines the success of the attack. The findings suggest that data-poisoning attacks may be more practical than previously believed.

Main themes

  • Data poisoning attacks
  • Large language models
  • Model security
  • Training data
  • Backdoor vulnerabilities
  • AI security

What commenters say

  • The effectiveness of poisoning attacks depends on the rarity of the trigger word in the training data, making it easier to poison models with unique tokens.
  • The study's findings are not surprising, as the poisoned documents are a significant percentage of the training data, even for the largest model.
  • The risk of poisoning attacks is limited by the need for attackers to get the rare token in front of the production LLM, which can be challenging.
  • The concept of 'reasoning' in AI models is ill-defined and may be used to anthropomorphize their capabilities for marketing purposes.
  • The use of terms like 'reasoning' and 'intelligence' to describe AI models is problematic, as they imply human-like abilities that the models do not possess.
  • The study's results are concerning, as they suggest that poisoning attacks can be successful with a relatively small number of malicious documents, regardless of model size.
  • The discussion around AI models' abilities is hindered by the lack of clear definitions and the use of terms that are misleading or confusing.
  • The focus on 'reasoning' and 'intelligence' in AI models distracts from the actual capabilities and limitations of these systems.