Researchers introduce a 1-bit Large Language Model (LLM) variant, BitNet b1.58, which uses ternary parameters {-1, 0, 1} and achieves performance equivalent to full-precision models. This model is more cost-effective in terms of latency, memory, throughput, and energy consumption. The 1.58-bit LLM defines a new scaling law and recipe for training new generations of LLMs. The model's performance is demonstrated through experiments and comparisons with full-precision models.