news.volyx.in

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7 (simonwillison.net)

463 points by simonw · 136 days ago · 97 comments on HN

Article summary

The article compares the performance of two AI models, Qwen3.6-35B-A3B and Claude Opus 4.7, on a benchmark test of drawing a pelican riding a bicycle. Qwen's model, running on a laptop, produced a better result than Opus, despite being a smaller and less powerful model. The author notes that the benchmark test is meant to be humorous and not taken seriously, but has surprisingly correlated with the general usefulness of the models in the past. However, this correlation appears to have broken with the latest results.

Main themes

  • AI model comparison
  • benchmark testing
  • model performance evaluation
  • overfitting and fine-tuning
  • artistic vs. physical plausibility
  • model capabilities and pricing

What commenters say

  • Some commenters argue that Qwen's output is more artistically interesting, while others prefer Opus's more physically plausible results.
  • The benchmark test is seen as flawed and not representative of real-world performance, with some arguing that it is too narrow or easily gamed.
  • There is disagreement over whether the models are overfitting to the benchmark test, with some pointing out that the results are still surprisingly good despite the test's limitations.
  • Others argue that the comparison is unfair, as Qwen's model is a smaller and less powerful model than Opus, and that the results should be considered in the context of the models' respective capabilities and pricing.
  • Some commenters suggest that the benchmark test has outlived its usefulness and is no longer a meaningful way to evaluate model performance.
  • The results of the benchmark test are seen as having implications for the broader evaluation of AI models, with some arguing that they highlight the need for more robust and realistic testing methods.
  • There is also discussion of the potential for models to be fine-tuned for specific benchmarks, and the impact this could have on the validity of the results.
  • The speed and efficiency of the models are also considered, with Qwen's model being noted as faster than Opus's.
  • themes": [