Researchers at Anthropic have made progress in understanding how language models work by decomposing them into understandable components, called features, which correspond to patterns of neuron activations. This approach allows for a better understanding of how the models behave and can potentially lead to more controlled and safe AI systems. The features learned are largely universal between different models, suggesting that lessons learned from one model may generalize to others. This research is part of Anthropic's investment in Mechanistic Interpretability, aiming to improve AI safety and reliability.