The article discusses a benchmarking project called Livenerf, which aims to detect whether the AI model Opus 5.5 gets worse after its launch. The project uses a deterministic-as-possible approach to measure the model's performance over time. The goal is to provide a clean day-0 baseline to check against, as previous arguments about model performance have been based on anecdotal evidence. The project is open-source and available on GitHub.