Google Research's Brain Team has introduced Imagen Video, a text-conditional video generation system that uses a cascade of video diffusion models to generate high-definition videos. The system consists of a base video generation model and multiple spatial and temporal super-resolution models. Imagen Video can generate videos with high fidelity, controllability, and world knowledge, including diverse videos and text animations in various artistic styles. The model's capabilities and limitations are discussed, along with potential applications and safety concerns.