The authors of Infinity AI have trained a video diffusion transformer model that can generate realistic AI characters that speak, allowing for expressive and realistic-looking characters. The model takes in a single image, audio, and other conditioning signals and outputs video. The authors claim this is the first time someone has trained a video diffusion transformer driven by audio input. The model has been trained for about 11 GPU years and is still actively being trained.