news.volyx.in

Qwen 3.8 27B available on Cerebras at 1500 tokens/s (inference-docs.cerebras.ai)

690 points by altertable · 5 days ago · 228 comments on HN

Article summary

Cerebras has made the Qwen 3.8 27B model available on their platform, with a speed of 1500 tokens per second. The model is part of their public endpoint offerings. The company provides various tools and features, including prompt caching, to support model usage. The documentation index is available for users to explore and discover more about the models and capabilities.

Main themes

  • AI model performance
  • Prompt caching and pricing
  • Context size limitations
  • Speculative decoding and draft models
  • Customer support and billing issues
  • Model usage and implementation

What commenters say

  • The Qwen 3.8 27B model is one of the strongest models hosted on the Cerebras public endpoint, but its high speed can make it expensive to use without prompt caching.
  • The introduction of prompt caching is seen as a positive development, but some users are skeptical about its actual benefits and pricing implications.
  • The limited context size of 128k for the Qwen model is a disappointment for some users, who find it insufficient for certain tasks.
  • The speed and performance of the Qwen model vary depending on the specific use case and implementation, with some users reporting impressive results and others experiencing slower speeds.
  • The use of speculative decoding and draft models can impact the throughput and performance of the Qwen model, depending on the output token distribution and training data.
  • Some users have experienced issues with the billing system and customer support, including problems with credits and account access.