news.volyx.in

The Llama 4 herd (ai.meta.com)

1235 points by georgehill · 484 days ago · 658 comments on HN

Article summary

Meta AI has released two new models, Llama 4 Scout and Llama 4 Maverick, which are part of the Llama 4 series and offer advanced multimodal intelligence. These models are designed to enable more personalized experiences and are available for download on llama.com and Hugging Face. Llama 4 Scout has 17 billion active parameters and a 10M context window, while Llama 4 Maverick has 17 billion active parameters and 128 experts. The models were trained using a mixture-of-experts architecture and have achieved state-of-the-art performance on various benchmarks.

Main themes

  • Llama 4 models
  • Multimodal intelligence
  • Model architecture
  • Context length
  • AI applications
  • Model training

What commenters say

  • The release of Llama 4 models may have been accidental, as there was no accompanying press release or blog post at first.
  • The 10M context window of Llama 4 Scout enables new use cases such as long chats and video processing.
  • The use of mixture-of-experts architecture in Llama 4 models allows for more efficient training and inference, but may require more VRAM.
  • Some commenters believe that LLMs should not use first-person pronouns or phrases that imply moral superiority, while others argue that this is a necessary aspect of generating human-like text.
  • The Llama 4 models may have picked up habits from consuming GPT output, such as using phrases that imply moral superiority.
  • The release of Llama 4 models is seen as a significant step forward in AI development, with potential applications in various fields.
  • Some commenters are skeptical about the ability to run Llama 4 models on local hardware, due to the large VRAM requirements.
  • The use of a 10M context window in Llama 4 Scout is a notable achievement, and may be the first time a model with such a large context window has been released.