Researchers at Anthropic have developed a method called Natural Language Autoencoders (NLAs) to understand the internal workings of language models like Claude. NLAs convert the model's activations into natural-language text, allowing for a better understanding of the model's thoughts and decision-making process. The method has been applied to improve Claude's safety and reliability, and the code has been released for other researchers to build upon. The technique has shown promising results in auditing and understanding the model's hidden motivations.