FLUX 3 is a new multimodal foundation model that jointly learns from images, videos, and audio within a unified architecture. It is designed to develop real-world visual intelligence and can generate images and videos with audio, as well as predict actions. The model is currently available in early access, with plans to release more capabilities and technical details in the coming weeks and months. FLUX 3 is intended to be a versatile model that can be used for various applications, including content creation and physical AI.