news.volyx.in

Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference (cerebras.ai)

427 points by benchmarkist · 626 days ago · 156 comments on HN

Article summary

Cerebras has achieved a record-breaking performance of 969 tokens per second on the Llama 3.1 405B model using their Inference platform, outperforming other solutions by a significant margin. This breakthrough enables frontier AI models to run at instant speed, allowing for real-time interaction and improved user experience. The Cerebras Inference platform is set to become available in Q1 2025, with pricing starting at $6 per million input tokens and $12 per million output tokens. The company's wafer-scale chip technology is credited for this achievement, offering a unique approach to AI computing.

Main themes

  • AI computing performance
  • Wafer-scale chip technology
  • Cerebras Inference platform
  • Llama 3.1 405B model
  • GPU vs CPU performance
  • Memory bandwidth and size

What commenters say

  • The Cerebras Inference platform's performance is astonishingly fast, but its high cost and limited availability make it inaccessible to most users.
  • The use of wafer-scale chip technology is a key factor in Cerebras' achievement, but its power consumption and cooling requirements are significant challenges.
  • The comparison between Cerebras and GPU-based solutions is not straightforward, as memory bandwidth and size are major limiting factors for GPUs in AI computing.
  • The cost of Cerebras systems is prohibitively expensive, with estimates ranging from $1.36M to $2.5M per system, making it unlikely to be widely adopted in the near future.
  • The performance of Cerebras Inference platform is expected to become more affordable and widely available in the next 3-5 years, following the historical trend of cost reduction in the semiconductor industry.
  • The Llama 3.1 405B model will become less interesting in the future as newer models are developed, but smaller models will still be useful for practical applications where compute budget is limited.
  • The math used to estimate the required memory bandwidth for GPU-based solutions is only valid for batch size = 1, and large enough batches can saturate compute throughput instead of bottlenecking on memory bandwidth.
  • The Cerebras patent on wafer-scale chip technology may be challenged by prior work, but its commercial success is largely due to the existence of market demand for AI computing solutions.