Omni fits short-form creators and marketers who want to describe an edit in plain language rather than open a traditional editor, not a production pipeline: there is still no developer API, so teams needing programmatic video generation today should use Veo 3.1 instead. Casual creators can try Omni free through YouTube Shorts and YouTube Create.
Gemini Omni is Google DeepMind's any-to-any video world model, generating a preview clip in under 15 seconds and then editing it through follow-up prompts that swap the weather, an outfit, or the location. It replaces separate Veo, Imagen and Lyria pipelines with a single generation pass, announced alongside the Gemini 3.5 line in May 2026.
Provider: Google DeepMind · Family: Gemini Omni
More about Google DeepMind on HokAI
Input modalities: text, image, audio, video · Output: video, audio, image, text
About Gemini Omni
Gemini Omni is a multimodal "world model" that Google DeepMind introduced at Google I/O in May 2026, alongside Gemini 3.5 Flash and the Antigravity 2.0 developer platform. Unlike the numbered Gemini reasoning line, Omni is a separate generative system that folds capabilities previously split across Veo, Imagen, Lyria and Nano Banana into a single any-to-any reasoning pass. The first publicly available variant is Gemini Omni Flash, a faster, lower-cost tier aimed at consumer-scale rollout; Google has signaled a higher-quality "Omni Pro" tier will follow, though no release window has been given. Because Omni's outputs are video and audio rather than text completions, Google has not run it through the standard LLM benchmark suite (SWE-bench Verified, GPQA Diamond, AIME, MMLU-Pro, ARC-AGI 2) and has not published comparable scores. Early hands-on testers instead graded it on generative-video quality, reporting that it handles "object permanence" (a character walking behind an obstacle re-emerging with the same face, clothing and proportions) more reliably than earlier diffusion-based video generators from Google and rivals like Seedance, though no independently verified numeric score for that test has been published. Omni accepts text, static images, audio clips and existing video as input, and can combine several of these in one prompt, for example a photo plus a voice note plus a reference clip, to produce a new or edited clip. Its defining feature is conversational editing: after a clip is generated, a follow-up request can change the setting, the weather or a character's outfit without re-describing the whole scene, and the model is built to keep that character's identity and the scene's context stable across the edit chain. Google DeepMind researcher Nicole Brichtova told TechCrunch that Omni Flash's clip-length ceiling is a deployment choice to manage compute demand during rollout, not a fixed architectural limit, implying longer clips become possible once capacity allows. Omni does not expose function calling, tool use, or code execution; it is a generation and editing model, not an agentic reasoning model, and those capabilities remain with the Gemini 3.5 line. The only confirmed access paths are the Gemini consumer app, Google Flow (Google's AI filmmaking tool), and the YouTube Shorts Remix and YouTube Create apps. There is no self-hosting option: Omni is closed-weight and API-only, consistent with the rest of the proprietary Gemini line (Google's open-weight releases are branded separately as Gemma). No Omni-specific system card, red-teaming partner list, refusal-rate benchmark, or training-data cutoff had surfaced by launch, beyond the SynthID watermarking on its outputs. Gemini Omni Flash is explicitly the first step in a longer rollout: Google has signaled a longer-clip, higher-resolution "Omni Pro" tier is coming, and a Vertex AI developer API is expected to follow, though neither had shipped as of this writing.
Pricing
There's no standalone API price for Omni yet: it's bundled into Google's AI Plus, Pro and Ultra subscriptions (see the pricing FAQ for exact tier amounts) and offered free, with daily limits, through YouTube Shorts and YouTube Create. Any per-token or per-second developer rate quoted online today is a third-party estimate anchored to Veo 3.1 and Gemini 3.5 Flash pricing, not a confirmed Google price.
Key Features
- Any-to-any multimodal generation: Takes text, image, audio and video as input in a single prompt and outputs a new or edited video, folding what used to be four separate models, Veo, Imagen, Lyria and Nano Banana, into one pass.
- Conversational video editing: A follow-up prompt like changing the weather or a character's outfit edits the existing clip instead of generating a new one, and identity stays consistent across the edit chain.
- World-model physical reasoning: Simulates gravity and spatial consistency well enough that characters who walk off-screen and back keep the same face, clothing and proportions, the 'object permanence' test reviewers use to grade video generators.
- Fast Omni Flash generation: A 5-second preview clip renders in under 15 seconds, quick enough for iterative back-and-forth editing sessions.
- SynthID watermarking: Every generated video and audio file carries Google's invisible SynthID watermark, the same provenance system used across Veo and Imagen outputs.
Pros
- Folds four separate generative systems (Veo, Imagen, Lyria, Nano Banana) into one prompt instead of chaining tools by hand.
- Conversational editing keeps a character's identity and scene context stable across multiple revision turns, a common failure point for other video generators.
- Free access through YouTube Shorts and YouTube Create lowers the barrier for casual creators to try it.
Cons
- Omni Flash's clip length is hard-capped at 10 seconds, ruling out longer-form content for now.
- No developer API or Vertex AI access at launch; it can't be wired into an automated pipeline as of June 2026.
- Official developer API pricing is unannounced, so the per-token and per-second figures circulating online are unconfirmed estimates, not real Google prices.
Frequently Asked Questions
What are Gemini Omni's pricing plans in 2026?
Gemini Omni Flash has no standalone developer price. Google folds it into three consumer subscriptions, AI Plus at $20 a month, AI Pro at $30, and AI Ultra at $100, all reachable through the Gemini app and Google Flow. A Vertex AI developer API was promised at Google I/O 2026 but had not shipped as of this writing, so the $1.50-$2.50 per 1M input token and $0.20-$0.60 per second of video figures floating online are unconfirmed third-party estimates, not official Google pricing.
Does Gemini Omni have a free plan?
Yes: it's free on YouTube Shorts Remix and the YouTube Create app, rationed to roughly 50 Flow credits a day, enough for one or two short generations. Getting it through the Gemini app itself requires an AI Plus, Pro or Ultra subscription; there is no free tier there.
What are the best alternatives to Gemini Omni?
On hokai, the closest video-generation alternative is Seedance 2.5, ByteDance's model built around one continuous take per generation rather than short conversational edits. Outside hokai, OpenAI's Sora 2 competes in the same any-to-any video space, though neither vendor has published benchmark numbers that allow a direct comparison. Anyone who needs text or code output instead of video should look at Gemini 3.5 Flash, the reasoning line Omni sits alongside.
Gemini Omni or Seedance 2.5: which should you pick?
Both generate video from several input types but for different workflows. Seedance 2.5 favors one continuous take per generation, while Omni's strength is revising a clip you already generated: changing the setting, an outfit or the time of day across several turns without losing the character's identity. Neither vendor has published head-to-head benchmark scores, so the practical choice comes down to long-form generation versus short-form iterative editing.
How do you get started with Gemini Omni?
There's nothing to install: open the Gemini app or Google Flow with a Google AI Plus, Pro or Ultra account and enter a prompt combining text, an image, or an audio clip. YouTube creators can reach the same tools for free through YouTube Shorts Remix or the YouTube Create app, limited by the daily Flow-credit ration. Once a clip exists, follow-up prompts like 'make it raining' or 'swap the jacket for leather' apply as conversational edits directly on that clip.
Top Alternatives
- Seedance 2.5: Pick Seedance 2.5 for one continuous take per generation; pick Gemini Omni to conversationally revise a clip you already made.
- Gemini 3.5 Flash: Pick Gemini 3.5 Flash for text, coding or agentic tasks; pick Gemini Omni only when the output itself needs to be video, audio or image.