The article introduces Consistency Large Language Models (CLLMs), a family of efficient parallel decoders that can accelerate inference by 3.5x. CLLMs are trained with a consistency loss and an autoregressive loss, allowing them to predict multiple tokens at once and maintain generation quality. The authors demonstrate the effectiveness of CLLMs on various datasets, including specialized domains and open-domain conversational challenges. CLLMs require moderate fine-tuning costs and can achieve significant speedup with a relatively small amount of training data.