The article compares the performance of two AI models, Qwen3.6-35B-A3B and Claude Opus 4.7, on a benchmark test of drawing a pelican riding a bicycle. Qwen's model, running on a laptop, produced a better result than Opus, despite being a smaller and less powerful model. The author notes that the benchmark test is meant to be humorous and not taken seriously, but has surprisingly correlated with the general usefulness of the models in the past. However, this correlation appears to have broken with the latest results.