news.volyx.in

Heretic: Automatic censorship removal for language models (github.com)

745 points by melded · 250 days ago · 380 comments on HN

Article summary

Heretic is a tool that removes censorship from transformer-based language models without expensive post-training. It uses an advanced implementation of directional ablation and a TPE-based parameter optimizer to work completely automatically. Heretic can produce decensored models that rival the quality of those created manually by human experts. The tool supports most dense models and can be used with a simple command-line interface.

Main themes

  • Language model censorship
  • Directional ablation
  • Model decensoring
  • Transformer-based models
  • AI safety and ethics

What commenters say

  • The development of tools like Heretic is important in the context of increasing ideology fixation and open-sourced models.
  • Some commenters believe that language models should not be overly cautious around certain topics, while others think it's necessary to protect users.
  • There is a debate about the responsibility of LLM companies in cases where their models cause harm, with some arguing it's a matter of liability and others seeing it as a way to slow down technological progress.
  • The use of tools like Optuna for hyperparameter optimization can significantly improve the performance of language models.
  • Some argue that the concept of 'correctness' in language models is too broad to be captured by a simple mechanistic description, while others see potential in exploring this area further.
  • The alignment of language models is seen as shallow by some, making it easy to 'jailbreak' them, while others argue that newer models have stronger alignment.
  • There is a discussion about the potential for an 'arms race' in obfuscating model safety and the limitations of current approaches to censorship and alignment.