Google has introduced the Universal Speech Model (USM), a state-of-the-art speech recognition model trained on 12 million hours of speech and 28 billion sentences of text, spanning over 300 languages. The model is designed to perform automatic speech recognition on widely-spoken languages as well as under-represented languages. USM has achieved a word error rate (WER) of less than 30% on average across 73 languages and has outperformed the Whisper model in certain scenarios. The model is intended for use in YouTube, such as for closed captions.