The article introduces Dia, a 1.6B parameter open-weights text-to-speech model that generates realistic dialogue directly from a transcript. Dia can condition the output on audio, enabling emotion and tone control, and also produce non-verbal communications like laughter and coughing. The model is available through Hugging Face Transformers and has been tested on GPUs, with plans for future optimization and quantization. The creators provide guidelines for using the model and demonstrate its capabilities through examples and a demo page.