Stable Diffusion 3 is a new text-to-image generation model that outperforms state-of-the-art models in typography and prompt adherence. The model uses a novel Multimodal Diffusion Transformer architecture that improves text understanding and spelling capabilities. The research paper provides technical details of the model and its performance, which will be accessible on arXiv soon. The model will be downloadable, with multiple variations ranging from 800m to 8B parameter models.