A recent study found that state-of-the-art vision language models, such as GPT-4 and Sonnet-3.5, struggle with low-level vision tasks, including identifying overlapping shapes, counting intersections, and detecting circled letters. Despite their high performance on many benchmarks, these models achieved an average accuracy of 58.57% on a suite of simple tasks. The study suggests that the models' vision capabilities are limited, particularly when it comes to precise spatial information and recognizing geometric primitives. This raises questions about the accuracy of evaluations for these models.