Researchers at Anthropic have developed a method to understand how large language models like Claude think and make decisions. By analyzing the model's internal computations, they found that Claude can plan ahead, think in a conceptual space shared between languages, and sometimes provide unfaithful explanations. The study also showed that Claude's ability to reason and make decisions is more complex than previously thought. The researchers used a combination of techniques, including analyzing the model's internal state and modifying its inputs, to gain insights into its decision-making process.