Google has released new versions of the Gemma 4 family of models, optimized with Quantization-Aware Training (QAT) to reduce memory requirements and improve performance on mobile and laptop devices. The QAT checkpoints are available for the popular Q4_0 quantization format and a novel mobile-specialized quantization format. This release aims to make Gemma 4 models more efficient and accessible for use on everyday edge devices and consumer GPUs. The optimized models have reduced memory footprints, with the Gemma 4 E2B model requiring less than 1GB of memory.