StyleTTS2 is an open-source text-to-speech model that achieves human-level synthesis using style diffusion and adversarial training with large speech language models. The model surpasses human recordings on single-speaker datasets and matches them on multi-speaker datasets. The project provides pre-trained models, demo code, and instructions for training and fine-tuning. The model's performance and quality are demonstrated through audio samples and online demos.