AI
Flux 3 X Mimic: The Next Generation of Video-Action Models
Black Forest Labs announced an early version of its new multimodal foundation model Flux 3 is running on robots through a collaboration with Mimic robotics. The companies said the combination of Mimic's strength in robot learning and deployment with the model's world knowledge and BFL's foundation model expertise produced Flux-mimic, described as the next generation of video-action models.
Flux 1 and Flux 2 generate images while Flux 3 expands into multimodality and generates audio-visual content jointly. The model provides the foundation of Flux-mimic, a video-action model running robots that have been tested and deployed at Audi. BFL said Flux 3 is one model jointly trained across images, video and audio from the beginning, with video prediction accounting for over 95% of total compute costs.
The company said generating realistic videos requires learning contact, motion, weight, cause and effect. Audio makes up less than 0.5% of tokens in a 720p video with audio. Actions follow the same shape as a low dimensional representation of a robot's state tightly coupled to visual observations. BFL said teaching Flux 3 to predict actions should not incur lasting costs, noting human ratings on text-to-video and image-to-video initially fell by up to 10% before regaining full previous quality after 3,500 steps while also predicting actions.
Flux-mimic is built on the Flux 3 backbone and decodes actions from the learned world representation. Mimic builds robots and brings expertise in robot learning, dexterous manipulation and production deployment while BFL brings multimodal training and modeling expertise. The companies said they built a next-generation model for general-purpose manipulation adapted to industry requirements and integrated into Mimic's full-stack deployment system.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from bfl.ai and reviewed by the T&B editorial agent team.