All News DISPATCH WORKFLOW

Gemini Omni 1.1 Flash Brings 4K Output and Video Editing to ComfyUI

ComfyUI added native support for Gemini Omni 1.1 Flash, bringing rapid multimodal video generation, conversational editing, and 4K upscaling into custom node graphs. Technical directors and VFX artists gain fine-grained control over prompt-guided video pipelines without leaving their local setups.

ComfyUI

ComfyUI, the open-source node-based AI generation interface, integrated Google's Gemini Omni 1.1 Flash model to support faster multimodal video generation, conversational editing, and 4K resolution processing directly inside local and cloud graph setups. The update enables creators to chain Google's high-speed multimodal model with existing ComfyUI nodes for automated upscaling, keyframing, and video-to-video modifications.

What's new

  • Conversational video editing: Creators can feed existing video sequences into Gemini Omni 1.1 Flash nodes and apply text-guided editorial changes, style adjustments, or localized modifications across frame batches.
  • 4K output support: The model integration supports upscaled generation and keyframe rendering up to 4K resolution, reducing the reliance on secondary external upscalers like Topaz Video AI or standalone ESRGAN passes.
  • Keyframe interpolation: ComfyUI nodes can now leverage Gemini Omni 1.1 Flash to generate smooth transition frames between disparate keyframes, maintaining temporal stability across complex camera movements.
  • Lower latency processing: As a Flash-tier multimodal model, Gemini Omni 1.1 Flash provides significantly faster frame-processing times compared to heavier models like Gemini 1.5 Pro or Veo 2, making iterative workflow building feasible on local graphs.

How it fits your workflow

For motion designers and VFX artists using ComfyUI, this integration bridges the gap between fast multimodal API calls and complex diffusion pipelines. Instead of hopping out to closed web UIs like Runway Gen-3 Alpha, Kling 2.0, or Pika 2.1 to generate source plates, artists can now direct video generation and revision passes right alongside their FLUX, SDXL, and ControlNet nodes.

The inclusion of conversational editing inside a node interface allows iterative revisions: artists can pipe output clips into a Gemini Omni node, supply targeted prompts ("adjust lighting to dusk", "add cinematic grain"), and feed the resulting video directly into custom composite pipelines or AnimateDiff stages. This makes ComfyUI a viable competitor to all-in-one studio platforms while maintaining complete node-level customization.

What it costs / how to try it

ComfyUI users can access the new Gemini Omni 1.1 Flash nodes by updating their ComfyUI core installation or pulling the custom node pack from the ComfyUI Manager. Running the node requires a valid Google AI Studio or Google Cloud Vertex AI API key, with usage billed at standard Google Gemini API rates.

Read the original announcement on ComfyUI ↗

Powered by ReelStack

Help keep this running

Your tip funds servers, models, and the time it takes to ship new tools faster. Set any amount below — every bit helps.