news.volyx.in

Kimi K3, and what we can still learn from the pelican benchmark (simonwillison.net)

405 points by droidjj · 41 days ago · 222 comments on HN

Article summary

The article discusses the release of Kimi K3, a new AI model with 2.8 trillion parameters, and its performance on the 'pelican benchmark', a test that generates an SVG of a pelican riding a bicycle. The author notes that while the benchmark is not a perfect measure of a model's capabilities, it can still provide valuable insights into its performance and characteristics. The article also touches on the limitations of the benchmark and the potential for models to be trained on similar tasks. The author concludes that the benchmark is still useful as a 'hello world' exercise for prompting a model and for comparing performance between different models.

Main themes

  • AI model releases
  • Benchmarking and evaluation
  • Model capabilities and limitations
  • Training data and bias
  • Cost and efficiency
  • Creative tasks and generation

What commenters say

  • The pelican benchmark is a useful tool for evaluating AI models, despite its limitations, as it can provide insights into a model's ability to generate novel and coherent images.
  • The benchmark is not a reliable measure of a model's overall capabilities, as it may be influenced by the model's training data and may not reflect its performance on more complex tasks.
  • The cost of generating a pelican SVG is relatively cheap compared to human labor, and the focus on cost efficiency is misguided when considering the capabilities of AI models.
  • The presence of pelicans in the training set may not necessarily mean that models are deliberately trained on this specific benchmark, but rather that they are exposed to a wide range of images and concepts.
  • The benchmark is still useful as a 'hello world' exercise for prompting a model and for comparing performance between different models, even if it is not a perfect measure of a model's capabilities.
  • The ability of a model to generate a pelican SVG is not a guarantee of its overall quality or usefulness, and other factors such as its ability to reason and operate tools should also be considered.
  • The use of benchmarks like the pelican test can lead to overfitting and may not provide a comprehensive picture of a model's abilities, and therefore should be used in conjunction with other evaluation methods.
  • The comparison between human and AI costs is not relevant, as AI models and humans have different strengths and weaknesses, and the focus should be on the capabilities and limitations of each.