All News DISPATCH WORKFLOW

Wan 3.0 Integration Brings 30-Second Video Generation and 20-Asset Reference Control

ComfyUI added native nodes for Alibaba's Wan 3.0 model, enabling single-pass 30-second video generation and up to 20 reference input assets. Indie filmmakers and VFX artists can now maintain subject consistency and edit existing footage without chaining short clips.

ComfyUI

ComfyUI, the node-based open-source generative AI interface, integrated native nodes for Alibaba’s Wan 3.0 video generation model as of March 2025. The update allows users to generate continuous 30-second clips in a single pass, incorporate up to 20 reference assets for character and environment consistency, and perform direct text-guided video editing within node graphs. This release removes the need to stitch short 4-second snippets together or rely on frame-interpolation chains.

What's new

The native Wan 3.0 implementation brings three primary capabilities directly into the ComfyUI workspace:

  • 30-second single-pass video generation: Instead of rendering brief 4- to 5-second segments and using frame extension tricks, Wan 3.0 outputs sustained 30-second video renders directly from a prompt and base seed.
  • Multi-asset reference control: Users can plug up to 20 image or video assets into a single node setup. These inputs act as visual, character, or lighting references, maintaining visual fidelity across long takes.
  • Instruction-driven video editing: The new setup accepts raw video footage alongside text instructions, modifying specific elements—such as altering wardrobe, changing lighting conditions, or swapping backgrounds—without breaking motion stability.

The model implementation operates natively inside standard ComfyUI workflow pipelines, allowing creators to chain Wan 3.0 outputs into existing upscalers, ControlNets, and audio-synchronization nodes.

How it fits your workflow

For technical directors, animators, and VFX artists, this update simplifies long-form AI video production. Previously, generating a 30-second sequence required cascading multiple generation steps, often leading to noticeable visual drift, shifting character details, and temporal flickering. By handling 30 seconds natively, Wan 3.0 inside ComfyUI reduces generation overhead and minimizes manual cleanup in post-production.

The 20-asset reference system addresses one of the persistent limitations in open-source video tools: subject consistency. Where earlier workflows relied on stacking multiple IP-Adapter nodes or custom LoRAs, Wan 3.0 allows editors to pass comprehensive style guides and multi-angle character sheets into one generation pass. This makes the setup a direct open-source alternative to proprietary web platforms like Runway Gen-3 Alpha, Luma Dream Machine, or Kling 2.0 Master, giving artists node-level precision over frame processing and memory management.

VFX artists can also use the instruction-driven editing nodes for targeted plate modification. Rather than re-rendering an entire shot, users can input an existing clip and instruct the model to alter specific visual components while retaining the camera movement and subject motion of the source footage.

What it costs / how to try it

ComfyUI is free and open-source software. The Wan 3.0 model weights are publicly available for local installation through Hugging Face, provided your local GPU meets the high VRAM requirements for long-context video generation. Creators without high-end local GPUs can run these updated workflows through hosted cloud instances that run ComfyUI environments.

Read the original announcement on ComfyUI ↗

Powered by ReelStack

Help keep this running

Your tip funds servers, models, and the time it takes to ship new tools faster. Set any amount below — every bit helps.