news.volyx.in

ETH Zurich and EPFL to release a LLM developed on public infrastructure (ethz.ch)

716 points by andy99 · 383 days ago · 101 comments on HN

Article summary

ETH Zurich and EPFL are releasing a large language model (LLM) developed on public infrastructure, trained on a supercomputer and open to the public under the Apache 2.0 License. The model is multilingual, capable of understanding over 1,000 languages, and will be released in two sizes. The development of this model is part of the Swiss AI Initiative, a collaboration between multiple Swiss universities and institutions. The model's training data will be transparent and reproducible, but not entirely publicly available due to copyright and distribution limitations.

Main themes

  • open-source AI
  • language models
  • multilingual capability
  • public infrastructure
  • data transparency
  • AI innovation
  • copyright and distribution
  • LLM performance and limitations
  • trustworthy AI development

What commenters say

  • Respecting web crawling opt-outs during data acquisition may not significantly impact the performance of LLMs.
  • Open-source models may not be able to surpass proprietary models due to limited access to certain data.
  • The quality gap between compliant and non-compliant LLM training data may be smaller than expected.
  • The release of open-source LLMs could drive innovation and accountability in the field of AI.
  • The use of open-source LLMs may be hindered by the lack of access to certain data, such as documentation blocked by robots.txt.
  • The architecture of the model may be more important than the data used to train it in determining its overall performance.
  • The transparency and reproducibility of the training data are crucial for the development of trustworthy AI models.
  • The announcement of the LLM's release without a concrete release date may be seen as premature or strategic.