news.volyx.in

Vision language models are blind (vlmsareblind.github.io)

451 points by taesiri · 762 days ago · 191 comments on HN

Article summary

A recent study found that state-of-the-art vision language models, such as GPT-4 and Sonnet-3.5, struggle with low-level vision tasks, including identifying overlapping shapes, counting intersections, and detecting circled letters. Despite their high performance on many benchmarks, these models achieved an average accuracy of 58.57% on a suite of simple tasks. The study suggests that the models' vision capabilities are limited, particularly when it comes to precise spatial information and recognizing geometric primitives. This raises questions about the accuracy of evaluations for these models.

Main themes

  • Vision language models
  • Low-level vision tasks
  • Model limitations
  • Evaluation accuracy
  • Multimodal learning
  • AI capabilities

What commenters say

  • The poor performance of vision language models on simple tasks is embarrassing and highlights the gap between their advertised capabilities and actual performance.
  • The models' limitations are not surprising, as they are not human brains and should not be expected to perform like humans.
  • The hype surrounding AI models is misleading, and their actual capabilities are often overstated, particularly in terms of their ability to understand images.
  • The models can still be useful for certain tasks, such as assisting people with low vision, despite their limitations.
  • The approach of adding specific tasks to the training set to fix issues is not a viable path to achieving generalized problem-solving ability.
  • The study's findings are interesting but the title and claims are hyperbolic and do not accurately reflect the models' capabilities.
  • The models' performance is not a failure, but rather a reflection of their current capabilities, which are still impressive and useful in many areas.
  • The limitations of vision language models highlight the need for more accurate evaluations and a better understanding of their strengths and weaknesses.