The article discusses the results of benchmarking Opus 5 on SlopCodeBench, a new coding benchmark that evaluates a model's ability to maintain codebase quality over time. The benchmark consists of multiple checkpoints, where the model must evolve the codebase as new requirements are introduced. Opus 5 achieved a 24% pass rate, outperforming other models, but still failed to maintain code quality over time. The article highlights the need for better benchmarks to evaluate the maintainability of code generated by models.