Mercury is a new generation of large language models based on diffusion, designed for coding applications. It comes in two sizes, Mini and Small, and has achieved state-of-the-art throughputs on NVIDIA H100 GPUs. The models are trained to predict multiple tokens in parallel and have been evaluated on various code benchmarks. A public API and free playground are available for testing.