ComfyUI: Setting up Wan 2.1 and InfiniteTalk models

Demo summary
A walkthrough of the model group in ComfyUI, showing the configuration of Wan Video Block Swap for VRAM management, the Light X2V LoRA for faster generation, and the InfiniteTalk GGUF model for audio conditioning.
Step-by-step
- Configure the Wan video torch compile settings to enable faster subsequent runs.
- Set the Wan Video Block Swap 'blocks to swap' setting to 15 to manage VRAM via CPU offloading.
- Load the Light X2V LoRA to reduce generation steps from 30-50 down to 4-10.
- Load the InfiniteTalk GGUF model using the Wan multi-talk model loader.
- Select the single person variant for audio conditioning unless syncing multiple speakers.
- Load the UMT5 double XL text encoder for prompt processing.
- Use the Clip Vision Loader to encode the reference image for identity preservation.
- Configure the Wave 2 vector model for acoustic feature encoding.
Options
- Increase 'blocks to swap' if you encounter CUDA out of memory errors.
- Use the multi-person variant of InfiniteTalk if you need to sync multiple speakers.
Watch out for
- The first run after enabling torch compile will be slow.
- Do not set the Light X2V LoRA strength to 1.0 as it causes color shifts and reduces identity preservation.
- InfiniteTalk is not a standalone generator; it requires Wan 2.1 video to function.
- You must use a video VAE rather than a standard image VAE to preserve motion coherence.
Tips
- Set the Light X2V LoRA strength to 0.8 for the best balance of speed and facial consistency.
- Use the Light X2V LoRA to reduce generation time from several hours to a fraction of that for long videos.
Highlights
“This is where things get technically dense. We have six separate models loading here.”
All demos from “AI Talking Head Videos With Perfect Lip Sync (ComfyUI + InfiniteTalk)”
2:560:54Image preparation for InfiniteTalkThe video shows the process of loading a portrait image and resizing it to the specific 384x640 resolution required by the Wan video model using standard ComfyUI nodes.ComfyUI· AI Crop Image
3:503:32Setting up Wan 2.1 and InfiniteTalk modelsCurrentA walkthrough of the model group in ComfyUI, showing the configuration of Wan Video Block Swap for VRAM management, the Light X2V LoRA for faster generation, and the InfiniteTalk GGUF model for audio conditioning.ComfyUI· AI Animation Generator
9:081:03Sampling and video output generationThe demonstration shows the final sampling process using the Lightning LoRA settings (CFG 1, 7-10 steps) and combining the decoded frames with audio for the final video file.ComfyUI· AI Animation Generator- Watch “AI Talking Head Videos With Perfect Lip Sync (ComfyUI + InfiniteTalk)” →
AI Animation Generator
1:420:59Seed hunting with a multi-stage LTX 2.3 workflowThe creator demonstrates his custom ComfyUI workflow that generates four low-resolution LTX 2.3 samples simultaneously to find a 'golden seed' before upscaling to 1080p.Fox•Fur•Essence Films
16:361:56Setting up HunyuanVideo 1.5 in ComfyUIThe video demonstrates how to update ComfyUI and import the HunyuanVideo 1.5 JSON workflow files to create a node-based generation environment.AI Search
20:151:53Text-to-Video generation in ComfyUIA step-by-step demo of configuring the Hunyuan nodes in ComfyUI, entering a prompt for a 'giant cat', and rendering the final 720p video.AI Search
29:151:33Running HunyuanVideo with GGUF (Low VRAM)The video shows how to use the GGUF loader node to run a compressed version of HunyuanVideo 1.5, enabling video generation on GPUs with as little as 6GB of VRAM.AI Search
0:510:36Configure Infinite Talk and Wan 2.1 models in ComfyUIThe user demonstrates loading the Infinite Talk model alongside the Wan 2.1 I2V 14B model within ComfyUI, including enabling block swap and torch compile for VRAM optimization.Olares
1:551:13Configure sampling and window settings for long video generationThe user walks through the Wan Video Wrapper sampling node, explaining how to set frame window size, motion frame overlap, and start steps for consistent video generation.Olares
0:310:46Generate cinematic video with LTX MSR workflowThe creator demonstrates using the LTX MSR workflow in ComfyUI to generate a video from multiple reference images and a prompt, highlighting the 3D camera movement and character consistency.Apex Artist
0:290:24Load source footage and models in ComfyUIThe user demonstrates importing source video footage and loading the necessary model nodes including WAN Video, VAE, and Clip Vision within the ComfyUI interface.ComfyUI
0:431:12Configure LTX-2.3 MSR LoRA and Prompt Relay in ComfyUIThe creator demonstrates setting up a ComfyUI workflow using the LTX-2.3 model with the MSR LoRA and Prompt Relay nodes to manage video generation on 8GB of VRAM.bigboss97
4:261:28Overview of the TensNodes Consistent Character Workflow in ComfyUIThe creator walks through a four-stage ComfyUI workflow utilizing TensNodes to process reference images and text for stable LTX video generation.SOTAI
3:411:25Configure LTX 2.3 foundational parametersA walkthrough of setting dimensions, frame counts, and loading core models including the LTX 2.3 distill model and audio VAE within a ComfyUI workflow.SOTAI
5:060:37Process visual and audio inputs in ComfyUIThe demo shows how to use the LTX V image-to-video condition node and empty latent audio node to prepare data for the generation engine.SOTAI
ComfyUI