All News DISPATCH AI VIDEO

Gemini Omni 1.1 Flash Upgrade Adds Precise Control to Veo 3

Google updated its underlying architecture with Gemini Omni 1.1 Flash, bringing enhanced control and faster inference to Veo 3 video generation. Directors and visual effects artists can now steer composition and subject consistency with greater precision.

Google Veo 3

Google Veo 3, Google's flagship AI video generation model, now integrates Gemini Omni 1.1 Flash to give creators finer control over prompt adherence and spatial direction. The architecture update reduces generation latency while giving users tighter command over camera movement, subject framing, and temporal consistency. This release addresses previous limitations where complex visual prompts resulted in unpredictable output drift.

What's new

Google's Gemini Omni 1.1 Flash update introduces low-latency multimodal steering for video generation pipelines. In Google Veo 3, this manifests as direct controls for frame-by-frame structural guidance and improved multi-turn editing capabilities as of May 2025.

Key additions to the generation engine include:

  • Tighter prompt adherence for multi-subject scenes, reducing visual artifacts when managing background and foreground elements simultaneously.
  • Improved camera path conditioning, allowing directors to specify pans, tilts, and tracking shots without sacrificing subject fidelity.
  • Faster generation speeds via the Flash architecture, reducing clip render times compared to standard Gemini Omni models.
  • Enhanced native audio matching, synchronizing generated sound effects and ambient dialogue directly with visual motion cues.

How it fits your workflow

For directors, visual effects artists, and video editors, Google Veo 3 with Gemini Omni 1.1 Flash shortens the iteration cycle during pre-visualization and shot prototyping. Where previous iterations required repeated prompting to achieve specific framing, the updated system respects precise layout instructions on the initial attempt.

In comparison to Runway Gen-4 and Kling 2.0 Master, Google Veo 3 utilizes Google's native multimodal understanding to interpret multi-image reference inputs alongside detailed text prompts. Animators and storyboard artists can feed keyframe roughs or character turnaround sheets into Google Veo 3, maintaining visual style across sequential shots without relying on third-party control nets or external image editors.

This control improvement makes Google Veo 3 a practical alternative to Sora 2 Pro for sequence generation, particularly when producing complex B-roll or rapid scene transitions. Visual effects teams can use the higher prompt fidelity to match lighting setups and focal lengths across multiple generated assets before dropping them into compositing software like Nuke or After Effects.

What it costs / how to try it

Google Veo 3 capabilities powered by Gemini Omni 1.1 Flash are rolling out to registered developers and enterprise users via Google AI Studio and Vertex AI, with expanded availability planned for VideoFX. Pricing follows standard Google Cloud API and token-based pricing schedules.

Read the original announcement on Google Veo 3 ↗

Powered by ReelStack

Help keep this running

Your tip funds servers, models, and the time it takes to ship new tools faster. Set any amount below — every bit helps.