The ARC-AGI-3 benchmark is designed to measure human-like intelligence in AI agents through interactive reasoning tasks. It challenges AI agents to explore novel environments, acquire goals, build adaptable world models, and learn continuously. The benchmark includes replayable runs, a developer toolkit, and a UI for transparent evaluation. The goal is to test AI agents' ability to learn and reason like humans.