The author, who has been working on a project to leverage AI models for security research, expresses skepticism about the recent progress in AI models, citing their own experience of not seeing significant improvements in their internal benchmarks despite trying out new models. The author suggests that the reported gains in AI models may not be reflective of economic usefulness or generality. The article also discusses the limitations of current benchmarks and the potential for AI labs to exaggerate their models' capabilities. The author argues that the industry needs to develop better metrics for assessing the impact of AI models.