Skip to main content

Share story

AI

Google's Gemini Omni turns images, audio, and text into video

Google's Gemini Omni turns images, audio, and text into video Image: Primary
Google launched Gemini three years ago with the goal of building a multimodal large language model trained on text, image, audio and video. At its Google I/O developer conference, the company announced Gemini Omni, a new family of multimodal models. Chief Executive Sundar Pichai said the models will create anything from any input. Gemini Omni starts with video generation. Users combine images, audio, video and text inputs, and the model reasons across them to produce consistent outputs that reflect an understanding of physics, culture, history and science. Users can also edit photos with plain text commands. Google already has a dedicated video model called Veo. Director of product management Nicole Brichtova said the release is the next step toward combining the intelligence of Gemini with the rendering capabilities of media models. In one example, a prompt for a claymation explainer of protein folding produced a stop-motion video with a voice-over narration. The long-term vision includes generating images from audio and audio from video. Users can create videos with their own digital avatars after recording themselves speaking a series of numbers during onboarding. All videos will include Google's SynthID digital watermark. Gemini Omni Flash, the first model in the family, rolls out to the Gemini app, YouTube Shorts and AI creative studio Flow. It renders 10 seconds of video, with longer durations in the pipeline. Google plans to release it via API in the coming weeks and noted the model's text-rendering capabilities for advertising uses. The company is focusing on consumer uses such as personalized videos. Prompts must be highly specific to avoid over-editing or unintended changes. A more advanced Omni Pro model is planned for later.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from TechCrunch and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Infrastructure Capital
Infrastructure Capital

Clastix raises €2.9 million for Kubernetes infrastructure software

Italian infrastructure software company Clastix has raised €2.9 million in its first external funding round, led by CDP Venture Capital with participation from Mistral and Vertis SGR. Clastix plans to use the seed capital for prod...

AI Science
AI Science

DeliveryGym study shows gains in simulated courier planning

Researchers introduced DeliveryGym, a three-dimensional environment that trains AI agents to plan across a full courier shift, where one delivery can use time, energy or money needed for later orders. In tests across six models an...

AI Science
AI Science

Preprint reports gains from grounding game-coaching AI in live state

Researchers report that rules tying an AI game coach to current match and session data raised its turn-level factual accuracy to 96.7%, from 61.1% for a prompting baseline and 69.8% for another agent design. Session-level accuracy...