news.volyx.in

Universal and transferable adversarial attacks on aligned language models (llm-attacks.org)

220 points by giuliomagnifico · 1124 days ago · 157 comments on HN

The AI summary for this story hasn't been generated yet — it's produced hourly. Check back soon. Meanwhile, read the discussion on HN.