Javlon Baxtiyorov

Gemini Omni Flash: Google's $0.10-a-Second AI Video Model

Google's Gemini Omni Flash generates 720p video with native audio and edits conversationally for ~$0.10 per second. The specs, pricing, benchmarks, and where AI video actually fits.

Gemini Omni Flash: Google's $0.10-a-Second AI Video Model

Google just turned video generation into an API call you can afford. Gemini Omni Flash takes text, images, and video in — and produces 720p clips with synchronized native audio, edited conversationally, for about ten cents a second. It's the same "fast, cheap, good-enough" playbook that made Gemini 3.6 Flash the workhorse of the text world, aimed squarely at video. Here's what it does, what it costs, and where it actually fits.

$0.10per second of video
720p+ native synced audio
3 / 5 / 10sclip durations
text·image·videoinputs → video

What it is

Gemini Omni Flash is Google's cost-efficient video generation and conversational-editing model. You describe a shot, or hand it an image or a reference clip, and it returns a 10-second 720p video with audio baked in — in a single pass, no separate sound step. Because it's a Gemini model, it draws on the family's world knowledge ("history, biology, narrative logic," per Google), so prompts like "show the reaction a chemist would have to this result" land closer to intent than a pure pixel model.

The headline feature is conversational editing: you don't re-prompt from scratch, you talk to the clip — "make it night," "add a voiceover," "keep the character, change the room" — and it iterates. That's the workflow shift, not just the fidelity.

The specs that matter

Property Gemini Omni Flash
Output 720p video, native synchronized audio
Durations 3s, 5s, or 10s (enum — 4/6/7/8s error, they don't round)
Inputs text, image, video references
Price $0.10 / second → $0.31 (3s), $0.51 (5s), $1.02 (10s)
Where Google AI Studio, Gemini API (public preview), Gemini Enterprise, Gemini app, Google Flow
Companion Nano Banana 2 Lite — text-to-image in ~4s at $0.034 / image

A few honest limitations up front: audio-reference upload and scene-extension aren't wired into the API yet, video reference inputs over ~3s "aren't correctly processed," and character consistency drifts across scene changes and big camera moves. This is a Flash model — built for speed and price, not flawless long-form cinematography.

The benchmarks

Google reports Gemini Omni Flash scoring strongly across the four evals that matter for a video model — and early public arenas back it up:

Text-to-video (arena)top tier
Image-to-video (arena)top tier
Video editing (preference)strong
Reference-to-videostrong

Those are Google's human-rated internal benchmarks (MovieGenBench prompts for text-to-video, VBench pairs for image-to-video, a 500-prompt fast-motion suite for sports/action) plus early public arena rankings that placed it at or near the top of text-to-video and image-to-video. Treat "top tier" as directional until neutral boards settle — but the direction is clearly competitive with the best dedicated video models, at a fraction of studio cost.

The economics change what you build

At $0.10/second, a 10-second clip is ~$1. That's not "generate a film" money — it's "generate a thousand product variations, ad cuts, or explainer snippets in a nightly batch" money. The same repricing that made text-tier Flash a routing default now applies to video:

  • Marketing / social — per-SKU product videos, A/B ad variants, localized cuts at scale.
  • Product & docs — auto-generated feature demos and onboarding clips from a script + a screenshot.
  • Prototyping — storyboard-to-motion in seconds inside Google Flow, before committing to expensive production.

Pair it with Nano Banana 2 Lite ($0.034/image, ~4s) for the stills pipeline and you have a full cheap-and-fast generative-media stack behind one API. For the text side of that stack, the Gemini 3.6 Flash cost math applies the same way — cache the stable prefix, pour volume through Flash.

The honest caveats

  • 10 seconds is the ceiling today. "Longer durations coming soon," but right now this is a short-clip tool, not a long-form one.
  • API ≠ full app. Several features (audio refs, scene extension) exist in the consumer apps before the API. Check the schema before you architect around them.
  • Consistency is the weak spot. For character-driven, multi-shot narratives, budget for retries or a reference-locking workflow.

My take

Gemini Omni Flash won't replace a production house — and it isn't trying to. It replaces the "we can't justify video for this" decision. When a usable 10-second clip with audio costs a dollar and edits by conversation, video stops being a project and becomes a feature you sprinkle across a product. That's the same quiet revolution Flash pulled in text: not the smartest model in the room, the one cheap and fast enough to use everywhere.


Sources

Specs and pricing reflect Google's launch details as of 29 July 2026 and will evolve as the preview matures. Spot a corrected figure? Tell me and I'll update.

Read next All writing →
← All writing Get in touch →