AI summary
- Google announced Gemini Omni at I/O 2026 on May 19. The first model, Gemini Omni Flash, started rolling out the same day.
- Any-to-any multimodal. Inputs: any combination of text, images, audio, and video. Outputs: high-quality video with synchronized audio. Image and audio output are planned to follow.
- Conversational editing. You prompt, the model generates, you give a follow-up instruction in plain English, the model edits while keeping characters consistent, physics holding, and the scene remembering what came before.
- Clips capped at 10 seconds in the rollout (a deployment decision, not a model constraint). Higher-end Omni Pro is planned with no announced date.
- Availability: Gemini app and Google Flow for AI Plus, Pro, and Ultra subscribers globally. Free for YouTube Shorts and YouTube Create App users. API access for developers in the coming weeks.
- Every video carries an imperceptible SynthID watermark. Verifiable through the Gemini app, Gemini in Chrome, and Google Search.
- This is the most significant Google AI video announcement since Veo 3.1, and one of the most consequential video-generation releases of 2026 from any lab.
Google announced Gemini Omni at I/O 2026 this morning. The first model in the family, Gemini Omni Flash, started rolling out the same day. This is the most ambitious AI video release Google has shipped since Veo 3.1, and the pitch is bigger than “another video model.” Omni is positioned as a single model that takes any combination of text, images, audio, and video as input and produces edited video with synchronized sound as output. Koray Kavukcuoglu, Google DeepMind’s CTO, posted the launch under the headline that the company has been telegraphing for months: a model that can “create anything from any input.”
Here is what was actually shown, how it compares to the rest of the field documented in our Real State of AI Video in 2026 Special Report, and what we think it means for everyone making short-form video.
Above: the Google I/O 2026 keynote replay. Scrub to the Gemini Omni segment for the live demos.
What is Gemini Omni in plain English?
Gemini Omni is a single multimodal AI model that takes mixed inputs and produces video. You can hand it a sentence, an image, a clip of audio, a short reference video, or any combination of those, and it generates a new video that follows your instruction. You can then talk to it through the same interface to refine the result. Change the lighting. Swap the character’s outfit. Add a sound effect. Re-time the scene.
The reason this matters more than “another video model” is the architecture choice. Prior video systems usually relayed work between specialized components: one model writes the prompt expansion, a second generates the visual frames, a third dubs audio, a fourth handles editing operations. Each handoff loses information. Omni is presented as a single model handling input parsing, generation, editing, and audio synthesis end to end. According to Google, that single-model approach is what makes the conversational editing work. The model knows what it generated, so it can change it coherently when you ask.
The other clean way to think about it comes from Google itself. Elias Roman, VP of Product at Google Labs, framed it as “Omni is like Nano Banana, but for video.” Nano Banana is Google’s image model, and the analogy is useful: same any-to-any input philosophy, same conversational editing surface, same level of grounding in Gemini’s world knowledge, but generating moving pictures with sound instead of still images.
What can it actually do?
The launch demos covered three capability areas the team wanted to highlight.
Physics. The rolling marble demo became the unofficial signature of the launch. A small ball rolls across a table, bounces, strikes a tiny bell, and continues. The bounce trajectories are right. The bell rings at the moment of impact with audio that matches the strike force. Subsequent bounces have decaying audio amplitudes that feel correct. Google’s own framing is that Omni has “improved intuitive understanding of forces like gravity, kinetic energy and fluid dynamics.” We watched the demo. The framing is fair.
Style transfer at scale. The second highlighted demo was a claymation-style explainer of how protein folding works. A complex scientific concept rendered in a coherent, consistent visual style across a 10-second clip with appropriate background sound. Style coherence across a long-form video has been the failure mode of every earlier video model. This is the area where Omni shows the most visible advance.
Character and scene memory. Google’s claim is that “your characters stay consistent, the physics hold up and the scene remembers what came before.” In practice this means you can prompt for an initial scene, then iterate. Change the camera angle. Re-cast the lighting. Add a second character. The original character’s face, clothing, and movement style all persist through the edits. This is the practical breakthrough most short-form video creators have been waiting for, because the inability to iterate without losing continuity was the single biggest workflow blocker in prior models.
Synchronized audio. Veo 3 already produced video with synchronized sound, and Omni continues that work. Voice, ambient, music, and effect tracks all generate in alignment with the visual content. The implication is that the same model handles audio generation, which is a meaningful unification.
How does it compare to Veo 3.1, Sora, Runway, Kling, and Seedance?
The 2026 AI-video field has roughly seven players that matter, covered in depth in our Special Report. Here is the comparison as of launch day.
- Veo 3.1 (Google’s previous flagship). Strong at cinematic shots and synchronized audio. Less flexible in conversational editing. Omni is the meaningful step up from Veo on input flexibility and iteration speed.
- Sora 2 (OpenAI, discontinued). Strong physics in early demos. The discontinuation story is its own piece. Omni now occupies the “best physics demos” position in the public conversation.
- Runway Gen-4.5. The professional creator workflow leader, with the deepest editing surface, the strongest team support, and the largest enterprise install base. Runway’s moat is the editor itself, not the model. Omni does not directly threaten that yet.
- Kling 3.0 (Kuaishou). The leading Chinese model, very strong on cinematic shots, strong physics. Geopolitically a question mark for many Western enterprise buyers. Omni’s broad availability matters competitively.
- Seedance 2.0 (ByteDance). Specialized in short-form vertical video for TikTok-adjacent workflows. Different niche than Omni.
- Pika 2.5. Independent, strong on stylization, weaker on physics. Niche but defensible.
- Adobe Firefly Video. The integration play. Lives inside Premiere, Photoshop, and the Creative Cloud. Strongest in the hands of professional editors already on Adobe.
The fair read is that Omni is the most capable launch-day video model we have seen in 2026. The conversational editing surface is the meaningful differentiator. The physics quality is at or above Sora 2’s discontinued peak. The 10-second clip cap will frustrate longer-form creators, but the cap is a deployment decision and presumably lifts when Google sees the compute economics.
Who can use it today?
Availability is unusually broad for a launch-day rollout.
- Google AI Plus, Pro, and Ultra subscribers get access globally through the Gemini app and Google Flow starting today.
- YouTube Shorts and YouTube Create App users get access at no cost starting this week. This is the surprise distribution move. Hundreds of millions of YouTube Shorts creators can now generate Omni clips inside the YouTube app itself.
- Developers and enterprise customers get API access in the coming weeks, on terms that have not been published yet.
The YouTube distribution is the strategic decision here. Most prior AI video tools have lived inside specialist editors that creators have to leave their normal workflow to enter. Omni lives inside the app YouTube creators are already using to publish. The friction curve drops to near-zero for a meaningful slice of the world’s short-form creators. Watch the YouTube Shorts feed in the next 60 days. Omni-generated content density will rise visibly.
What about the avatar feature?
Google previewed a personal-avatar capability inside Omni. Users record a short reference of their own face making a set of expressions and reading aloud random numbers. Once verified, they can generate video featuring themselves doing things they did not actually do. The verification step is the safeguard against using someone else’s face.
This is the feature Google has been most cautious about, and the launch coverage made clear that some of the most powerful versions of it are being held back from general release. The fair read: avatar generation done well is one of the highest-leverage creator features anyone could ship, and avatar generation done badly is the single fastest path to deepfake harm. Google is staging the rollout carefully. SynthID watermarking on every generated video is the most important corresponding safeguard.
What about SynthID watermarks?
Every video generated by Omni carries an imperceptible SynthID watermark. This is Google’s cryptographic-style provenance marker that survives normal editing, re-encoding, and platform reuploads. Three places will detect it for you in 2026: the Gemini app, Gemini in Chrome (via right-click), and Google Search results that include the verification metadata.
SynthID does not solve the deepfake problem. It does meaningfully raise the cost of plausible denial, because anyone can verify whether a clip was generated by a Google model. This is the most concrete provenance infrastructure any major lab has shipped at this scale. The remaining challenge is interoperability: SynthID detection works for Google-generated content, not for content from Sora, Kling, Runway, or any other model. Industry-wide provenance is still the open problem.
What does this mean for AI video as a category?
Three things change with this launch.
First, conversational editing becomes the new table-stakes feature. Every other model in our Special Report will need to ship its own version within a year, or accept that the workflow gap becomes structural. Runway and Adobe will both move fast here, because their editor moats depend on staying ahead on iteration speed.
Second, distribution-via-the-app is now a competitive vector. Google can put Omni inside YouTube the way no other lab can reach a billion-creator pool. OpenAI’s equivalent move would be ChatGPT-as-creator-app. Meta’s equivalent would be Reels with their own model. The integration shipped today raises the bar for what “shipping a video model” actually requires to compete.
Third, the 10-second cap is the next benchmark. The model can probably do longer. Compute economics keep it capped. Whichever lab is willing to absorb the cost of longer-form generation first will own the long-form-video creator market that does not yet exist. Helion-style power-purchase agreements may be the unspoken precondition for the next leap.
When does Omni Pro arrive?
Google’s official answer is “when we see a step change above Flash.” No date. The naming convention (Flash now, Pro later) parallels the Gemini 2.0 release pattern, where Flash shipped months ahead of Pro. Watch for Omni Pro in late 2026 or early 2027 on that historical pattern. The Pro tier will probably remove the 10-second cap, raise the resolution, and improve the physics edge cases that Flash still misses.
Frequently asked questions
How do I try Gemini Omni right now?
Three paths in May 2026: the Gemini app (any AI Plus, Pro, or Ultra subscriber), Google Flow (same subscription tiers), or YouTube Shorts and YouTube Create App (free tier, rolling out this week). Developer API access is “in the coming weeks.”
How long are the clips?
10 seconds maximum in the current rollout. Google has been explicit that this is a deployment decision driven by compute demand, not a model constraint. Expect the cap to lift over time.
Can I generate video from just text?
Yes. Text is one of the four input modalities (text, image, audio, video). The any-to-any framing means any subset of those works as input.
Does it work for commercial use?
The Gemini app and Google Flow versions are commercial-use-permitted under standard Google AI subscriber terms. YouTube Shorts and YouTube Create generation falls under YouTube’s existing creator terms. Enterprise API terms are pending.
How is Omni different from Veo?
Veo is the predecessor video model. Omni is the next-generation any-to-any model that can also do what Veo did but adds conversational editing, multimodal input, and integrated audio in a single unified model. Existing Veo workflows continue to work; Omni is the upgrade path.
The Beginners in AI position on Gemini Omni
This is one of the most consequential AI launches of 2026, and it is also one of the cleanest examples of the dual-use problem that defines the next decade of media. Conversational video editing is real progress. The creator economy gains real new capability. The protein-folding claymation demo is a glimpse of how science communication can be done by anyone with a curious mind and a phone.
It is also a model that, in less careful hands, makes deepfake-grade video accessible to billions of YouTube users. SynthID is a serious safeguard. The avatar verification flow is a serious safeguard. The held-back features are evidence that Google has thought about this. None of those measures will be a sufficient safeguard if industry-wide provenance does not arrive soon. The right policy answer involves cross-lab watermarking standards, browser-level verification by default, and a more thoughtful set of takedown norms for platforms. None of those are in place today.
We are pro-technology and we are excited by what Omni unlocks. We are also pro-human first, and that means being clear-eyed about what the world looks like in six months when AI-generated short-form video saturates the YouTube Shorts feed. Watch the platform-level adaptations as much as you watch the model launches. The model is impressive. The infrastructure around it is the part that decides whether the next year goes well.
If you are reading this on a Sunday morning and you are an AI-curious creator, do this today: install the Gemini app or open YouTube Create, run the first Omni clip you can think of, and notice what surprises you. The technology is real. The conversation about how to use it well starts the moment you have your own example to think with.
Sources
- Introducing Gemini Omni, the official Google blog announcement by Koray Kavukcuoglu
- New agents, mobile apps and Gemini Omni for Google Flow, by Elias Roman, VP of Product, Google Labs
- Google I/O 2026 Keynote, the full Shoreline Amphitheatre replay
- Everything new in Google AI subscriptions, fresh from I/O 2026, the Plus / Pro / Ultra tier breakdown
- 9to5Google’s coverage for additional demo details
- VentureBeat’s enterprise-angle reporting
- CNBC on Google’s broader I/O 2026 AI roadmap
- SynthID overview, the Google DeepMind provenance system
Get Smarter About AI Every Morning
Free daily newsletter. One story, one tool, one tip. Plain English, no jargon.
Free forever. Unsubscribe anytime.
You might also like
- The Real State of AI Video in 2026, our Special Report on every video model that matters
- Google Veo 3.1 Explained, the predecessor flagship
- Nano Banana 2 Explained, Google’s image-generation cousin to Omni
- Runway Gen-4.5 Explained, the leading professional creator workflow
- Kling 3.0 Explained, the Chinese flagship
- Seedance 2.0 Explained, ByteDance’s short-form specialist
- What Happened to Sora 2, the OpenAI retrospective
- All Beginners in AI Special Reports
Two ways to go further
The AI Prompt Library
1,000+ ready-to-use prompts for Claude, ChatGPT, and Gemini. Stop staring at a blank box.
Get it for $39 →2-Hour Live AI Crash Course
A private, beginner-friendly session across Claude, ChatGPT, Gemini, and the wider landscape.
Book for $125 →