demobook

ComfyUI: Combining audio, images, and prompts

Demo summary

A demonstration of layering a specific action prompt (patting stomach) over a specific audio timestamp to create a fully directed AI scene.

Step-by-step

  1. Upload your custom audio file to the timeline
  2. Add your base images to the project
  3. Identify specific timestamps in the audio for character actions
  4. Insert a text prompt at the desired timestamp to trigger a specific movement

Options

  • Combine custom audio, images, and prompts to direct the scene

Tips

  • Sync specific action prompts (like 'patting stomach') with relevant dialogue keywords (like 'fat') for more realistic character animation

Highlights

something cool you can do... to really fully direct the scene

All demos from “How to use LTX Director - A Free Open Source Tool for Creating LTX 2.3 AI Videos Locally in ComfyUI

  1. 3:070:38Image to Video with LTX DirectorThe creator demonstrates dragging an image into the LTX Director timeline and entering a text prompt to generate a video of a woman waving her hand.ComfyUIImage to Video
  2. 8:150:26Using Keyframes to guide video generationThe demo shows how to slide an image to a later point in the timeline to use it as a target keyframe, allowing the model to generate the action leading up to that specific visual.ComfyUIVideo to Video
  3. 10:070:59Lip-syncing with custom audioThe creator demonstrates importing an audio file into the timeline and using a specific prompt structure to synchronize the character's mouth movements with the audio track.ComfyUIAI Lip Sync Generator
  4. 11:060:26Combining audio, images, and promptsCurrentA demonstration of layering a specific action prompt (patting stomach) over a specific audio timestamp to create a fully directed AI scene.ComfyUIAI Animation Generator
  5. Watch “How to use LTX Director - A Free Open Source Tool for Creating LTX 2.3 AI Videos Locally in ComfyUI” →

AI Animation Generator

  1. 1:420:59Seed hunting with a multi-stage LTX 2.3 workflowThe creator demonstrates his custom ComfyUI workflow that generates four low-resolution LTX 2.3 samples simultaneously to find a 'golden seed' before upscaling to 1080p.Fox•Fur•Essence Films
  2. 16:361:56Setting up HunyuanVideo 1.5 in ComfyUIThe video demonstrates how to update ComfyUI and import the HunyuanVideo 1.5 JSON workflow files to create a node-based generation environment.AI Search
  3. 20:151:53Text-to-Video generation in ComfyUIA step-by-step demo of configuring the Hunyuan nodes in ComfyUI, entering a prompt for a 'giant cat', and rendering the final 720p video.AI Search
  4. 29:151:33Running HunyuanVideo with GGUF (Low VRAM)The video shows how to use the GGUF loader node to run a compressed version of HunyuanVideo 1.5, enabling video generation on GPUs with as little as 6GB of VRAM.AI Search
  5. 0:510:36Configure Infinite Talk and Wan 2.1 models in ComfyUIThe user demonstrates loading the Infinite Talk model alongside the Wan 2.1 I2V 14B model within ComfyUI, including enabling block swap and torch compile for VRAM optimization.Olares
  6. 1:551:13Configure sampling and window settings for long video generationThe user walks through the Wan Video Wrapper sampling node, explaining how to set frame window size, motion frame overlap, and start steps for consistent video generation.Olares
  7. 0:310:46Generate cinematic video with LTX MSR workflowThe creator demonstrates using the LTX MSR workflow in ComfyUI to generate a video from multiple reference images and a prompt, highlighting the 3D camera movement and character consistency.Apex Artist
  8. 0:290:24Load source footage and models in ComfyUIThe user demonstrates importing source video footage and loading the necessary model nodes including WAN Video, VAE, and Clip Vision within the ComfyUI interface.ComfyUI
  9. 0:431:12Configure LTX-2.3 MSR LoRA and Prompt Relay in ComfyUIThe creator demonstrates setting up a ComfyUI workflow using the LTX-2.3 model with the MSR LoRA and Prompt Relay nodes to manage video generation on 8GB of VRAM.bigboss97
  10. 4:261:28Overview of the TensNodes Consistent Character Workflow in ComfyUIThe creator walks through a four-stage ComfyUI workflow utilizing TensNodes to process reference images and text for stable LTX video generation.SOTAI
  11. 3:411:25Configure LTX 2.3 foundational parametersA walkthrough of setting dimensions, frame counts, and loading core models including the LTX 2.3 distill model and audio VAE within a ComfyUI workflow.SOTAI
  12. 5:060:37Process visual and audio inputs in ComfyUIThe demo shows how to use the LTX V image-to-video condition node and empty latent audio node to prepare data for the generation engine.SOTAI