news.volyx.in

Recent AI model progress feels mostly like bullshit (lesswrong.com)

579 points by paulpauper · 482 days ago · 458 comments on HN

Article summary

The author, who has been working on a project to leverage AI models for security research, expresses skepticism about the recent progress in AI models, citing their own experience of not seeing significant improvements in their internal benchmarks despite trying out new models. The author suggests that the reported gains in AI models may not be reflective of economic usefulness or generality. The article also discusses the limitations of current benchmarks and the potential for AI labs to exaggerate their models' capabilities. The author argues that the industry needs to develop better metrics for assessing the impact of AI models.

Main themes

  • AI model progress
  • Benchmark limitations
  • Exaggerated claims
  • Economic usefulness
  • Generality of AI models
  • Industry accountability

What commenters say

  • Recent AI model progress is mostly about reducing costs rather than improving capabilities.
  • The business model of AI companies is based on hype and selling proprietary models, rather than generating revenue through practical applications.
  • The limitations of current benchmarks make it difficult to accurately assess the capabilities of AI models.
  • Some commenters believe that newer models, such as Gemini 2.5, have shown significant improvements and are useful in real-world applications.
  • Others argue that the AI bubble will eventually burst due to the lack of viable business models and the limitations of current AI technology.
  • There is a concern that AI labs may be exaggerating their models' capabilities and that this could have negative consequences for the industry.
  • The development of better metrics for assessing the impact of AI models is necessary to ensure accountability and transparency in the industry.
  • Some commenters are skeptical about the potential for AI models to achieve general intelligence and believe that current models are not making significant progress towards this goal.