Gemini Omni 1.1 Flash pricing, plans and limits

Google's fast, conversational video model for iterative editing, positioned alongside the cinema-focused Veo line rather than replacing it.

  • ga
  • proprietary
  • multimodal
  • Gemini Omni family
checked

Gemini Omni 1.1 Flash shipped as a direct upgrade to Google's original May 2026 preview, replacing single-shot clips with chat-driven, chained scene extension. It suits teams doing fast, conversational video iteration rather than single-take cinematic production, where Google's own Veo line is the better fit.

Gemini Omni 1.1 Flash is Google's multimodal video generation and editing model, generally available since August 27, 2026, and it currently ranks first on two Artificial Analysis Video Arena leaderboards when audio isn't scored. It takes text, image, audio, and video input together and returns video with synchronized audio.

Provider: Google · Family: Gemini Omni

More about Google on HokAI

Input modalities: text, image, audio, video · Output: video, audio

About Gemini Omni 1.1 Flash

Gemini Omni 1.1 Flash is a multimodal video generation and editing model built by Google, developed within Google DeepMind. The model reached general availability on August 27, 2026, replacing the gemini-omni-flash-preview endpoint that Google is retiring on September 30, 2026. It belongs to the Gemini Omni family, first previewed in May 2026 at Google I/O, and sits alongside Google's separate Veo line rather than replacing it: Omni Flash targets fast, conversational iteration, while Veo targets cinematic, single-take production. Unlike Google's text-first Gemini models, Omni Flash accepts text, image, audio, and video together in one request and returns video with synchronized audio. Because the model generates video rather than text, it is not scored on benchmarks like SWE-bench or MMLU. Google instead submits it to the crowd-sourced Artificial Analysis Video Arena, where human raters compare blind clip pairs. As of early September 2026, Gemini Omni 1.1 Flash leads the Text-to-Video Arena in the no-audio category and the Image-to-Video Arena in the no-audio category. Once audio quality is scored, its lead narrows: Alibaba's Wan 3.0 edges it out on Text-to-Video with audio, and Minimax's H3 Max leads Image-to-Video with audio, an honest gap the 1.1 update has not closed. Google has not published a text-style context window for this model, because its native unit is seconds of video rather than tokens. A single generation pass runs 3 to 10 seconds. Extension passes chain additional 10-second segments up to a 40-second cumulative scene, and each extension pass can analyze up to 10 seconds of the prior clip for continuity, a sharp jump from the 1-second final-frame reference the original May 2026 preview supported. Google has not disclosed a native resolution above 720p; higher-resolution output options in the API are documented as upscaling passes over that base, not separate native renders. Inputs span text, still images, audio clips, and short reference videos limited to three seconds apiece, up to three per request, so a single prompt can combine a product shot, a voice line, and a style reference. Outputs are video with audio generated in sync with the visuals. The model supports first-and-last-frame interpolation, generating the connecting motion between two supplied keyframes, plus the scene-extension mode described above. One capability the architecture supports is deliberately switched off in this release: altering a real person's recorded speech in an uploaded video is technically possible but disabled while Google works through the safety review for that specific misuse surface. Gemini API pricing for this model follows Google's standard per-token structure: one rate for input across every modality, and separate output rates for generated video and generated text. Video output cost scales directly with resolution and duration rather than a flat per-clip fee, since Google converts rendered seconds into a fixed number of output tokens before billing. There is no free tier for this model; access is paid from the first generation, unlike some Gemini text models that offer a limited free quota in AI Studio. The exact current rates are in the pricing FAQ below, since Google revises API pricing on its own schedule. Access runs through Google AI Studio and the Gemini API directly, with the model also available through Vertex AI, the Gemini Enterprise Agent Platform, and Flow, Google's video-editing product. For the scene-extension feature specifically, Google AI Plus, Pro, and Ultra subscribers can use it inside the consumer Gemini app without touching the API. Google has not published a fixed requests-per-minute or tokens-per-minute limit for this model; Vertex AI serves it under Dynamic Shared Quota, allocating available capacity across all customers in a project's region rather than guaranteeing a fixed rate. Users in the European Economic Area, Switzerland, and the United Kingdom face additional restrictions specifically on editing and extending uploaded videos, separate from generating a clip from scratch. Google trained the model on audio, video, image, and text data, with the video and audio portions annotated with text captions at multiple levels of detail and filtered for policy compliance, safety, quality, and duplication before training. Safety mitigations run at two stages: diverse synthetic captioning during pre-training to broaden concept coverage, and production content filters plus mandatory SynthID watermarking after generation, applied to every output with no opt-out and checkable through the Gemini app, Chrome, and Search. Google's model card cites internal automated red-teaming and external human red-teaming by unnamed specialist teams, plus an ethics and safety review against its AI Principles, ahead of the August 2026 release. The model fits teams doing fast, iterative video work: social clips, ad drafts, and workflows built around refining a result through chat rather than one perfect upfront prompt. Its closest published rival is Google's own Veo 3.1, which renders natively at a higher resolution with longer single takes and spatial audio mapped to scene geometry, aimed at broadcast and cinematic delivery; teams needing a genuine native high-resolution deliverable should reach for Veo 3.1 instead. Teams chasing the top score in every leaderboard category should also weigh Alibaba's Wan 3.0 and Minimax's H3 Max, which each beat Gemini Omni 1.1 Flash once audio quality is part of the judging. Google has not published a training-data cutoff date for this model, unlike the system cards it publishes for its text models, and the model card does not spell out an API input-retention window or a zero-retention enterprise option for this specific endpoint. English is the only prompt language Google's own documentation lists as fully supported; other languages may produce usable results but are not guaranteed.

Pricing

Input (text, image, video, audio) is $1.50 per 1M tokens. Video output is billed at $17.50 per 1M tokens, converted at a fixed 5,792 tokens per second of standard-resolution video, about $0.10 per second. Text output during conversational editing bills separately at $9.00 per 1M tokens. No free tier; paid-tier only.

Key Features

  • Scene Extension to 40 Seconds: Chains 10-second generation passes up to a 40-second cumulative scene, analyzing up to 10 seconds of prior footage per extension.
  • First-and-Last-Frame Interpolation: Generates a continuous transition clip between two supplied keyframes, giving direct control over how a shot starts and ends.
  • Native Synchronized Audio: Generates audio in sync with the video from the same prompt, billed separately from video tokens.
  • Resolution Control, 360p Drafts to 4K Upscale: Renders fast 360p previews for iteration, then upscales the same generation to 1080p or 4K; native output tops out at 720p.
  • Reference Video Input: Accepts as many as three short reference clips, capped at 3 seconds each, so a scene keeps consistent characters, style, or camera movement across edits.
  • Conversational Multi-Turn Editing: Refines a generated clip through follow-up chat instructions in AI Studio or the Gemini app instead of one single upfront prompt.

Pros

  • Chat-based re-editing removes the need to rewrite one perfect prompt: users nudge a generated clip through follow-up turns instead of restarting from scratch.
  • The chained scene-extension feature reaches a longer cumulative clip than a single take from Google's own Veo line, useful for anything needing more than one short shot.
  • Mandatory SynthID watermarking on every clip gives compliance teams a built-in provenance check with no extra integration work.
  • Accepts image, audio, and reference videos alongside text in one call, more input flexibility than a text-and-image-only rival.

Cons

  • No free tier: every generation costs money from the first clip, with no AI Studio quota to test on before committing budget.
  • Higher-resolution outputs are upscales of a lower native render rather than true higher-resolution generations, a gap Google's own Veo line does not have.
  • Speech editing on uploaded video is disabled, limiting one class of conversational edit some users will expect from a video-editing model.
  • Google has not published a training-data cutoff, an architecture description, or a fixed rate limit for this model, less transparency than its text-model system cards provide.

Benchmarks

  • artificial analysis text to video arena elo no audio: 1324
  • artificial analysis image to video arena elo no audio: 1364
  • artificial analysis text to video arena rank no audio: 1
  • artificial analysis image to video arena rank no audio: 1
  • artificial analysis text to video arena elo with audio: 1237
  • artificial analysis image to video arena elo with audio: 1180
  • artificial analysis text to video arena rank with audio: 2
  • artificial analysis image to video arena rank with audio: 4

Frequently Asked Questions

What does Gemini Omni 1.1 Flash actually cost?

Gemini Omni 1.1 Flash runs on one paid tier: $1.50 per 1 million input tokens across text, image, audio, and video, and $17.50 per 1 million tokens for generated video, converted from a fixed 5,792-token-per-second rate at standard resolution (about $0.10 per second of finished clip). Replies generated during conversational editing are text output, billed separately at $9.00 per 1 million tokens. Google has not published a batch or provisioned-throughput discount for this model.

Can you use Gemini Omni 1.1 Flash without paying?

No. Google ships it on the paid tier only, so every clip costs money from the first generation, with no free quota in AI Studio the way some Gemini text models offer. The one no-cost path is scene extension inside the consumer Gemini app, bundled into an existing Google AI Plus, Pro, or Ultra subscription rather than billed per API call.

What are Gemini Omni 1.1 Flash's closest competitors?

Google's own Veo line is the closest in-house alternative, trading Omni Flash's conversational editing for native high-resolution rendering and spatial audio aimed at broadcast and cinematic work. Alibaba's Wan 3.0 currently edges out Gemini Omni 1.1 Flash on the Artificial Analysis Text-to-Video Arena once audio is scored, and Minimax's H3 Max leads the Image-to-Video Arena once audio is scored. None of the three publish an equivalent chained scene-extension feature.

What separates Gemini Omni 1.1 Flash from Veo 3.1?

Pick Gemini Omni 1.1 Flash for fast, conversational iteration and social clips built around chat-based re-editing rather than a single polished render. Pick Veo 3.1 when the deliverable needs a native high-resolution render, a single longer take, or spatial audio tied to camera geometry, since Omni Flash's higher-resolution option is an upscale of a lower native render, not a native render. Both apply Google's mandatory SynthID watermark to every output.

How do you get started with Gemini Omni 1.1 Flash?

Sign in to Google AI Studio or call the Gemini API directly with model ID gemini-omni-1.1-flash and a billing account attached, since there is no free-tier key for this model. Submit a text prompt alone, or add a reference image, audio clip, or a couple of short reference videos for more control, then use follow-up chat turns to extend the scene or adjust the result. Consumer users can skip the API and use scene extension inside the Gemini app if they already hold a Google AI Plus, Pro, or Ultra subscription.

More AI Models on HokAI

Visit Gemini Omni 1.1 Flash Official Page