news.volyx.in

GPT-6 Astra (openai.com)

2270 points by kibae · 5 days ago · 2072 comments on HN

Article summary

The article discusses GPT-6 Astra, a new model that has achieved high scores on the ARC-AGI-3 benchmark, with reports of 98.6% to 100% scores. The model's performance is attributed to its ability to use a custom compaction algorithm and a continuous conversation harness. However, some have raised questions about the fairness of allowing the model to use these custom settings. The model's efficiency and cost are also discussed, with some users reporting high token usage and costs.

Main themes

  • GPT-6 Astra performance
  • ARC-AGI-3 benchmark
  • custom settings and fairness
  • model efficiency and cost
  • true intelligence and understanding
  • benchmark limitations and revisions

What commenters say

  • The high scores achieved by GPT-6 Astra on the ARC-AGI-3 benchmark are impressive, but may not be directly comparable to human scores due to the scaling of the benchmark.
  • Allowing GPT-6 Astra to use a custom compaction algorithm and harness may be unfair to other models, as it gives it an advantage in the benchmark.
  • The use of custom settings to achieve high scores on the ARC-AGI-3 benchmark may be a sign of the model's true capabilities, rather than a flaw in the benchmark.
  • The high cost and token usage of GPT-6 Astra may be a significant limitation to its adoption, despite its impressive performance.
  • The ARC-AGI-3 benchmark may need to be revised or updated to account for the capabilities of models like GPT-6 Astra.
  • The performance of GPT-6 Astra on the ARC-AGI-3 benchmark is not necessarily a sign of true intelligence or understanding, but rather a demonstration of its ability to optimize for a specific task.
  • The efficiency gains claimed by the developers of GPT-6 Astra may not be realized in practice, and the model may be more expensive to use than expected.