demobook

Stable Audio 3 in ComfyUI: Create AI Music and Sound Effects (Ep19)

YouTube ↗
pixaromapixaroma·11.5K views · May 2026

Clips

  1. 0:441:57Generate audio from text in ComfyUI with Stable Audio 3The creator demonstrates setting up a ComfyUI workflow using the Stable Audio 3 Medium model and T5 Gemma text encoder to generate a 30-second instrumental audio clip from a text prompt.ComfyUIAI Music Generator
  2. 3:340:28Using Tiled VAE Decode for low VRAM audio generationThe narrator shows how to replace the standard VAE decode node with a tiled version in ComfyUI to handle longer audio generations (120 seconds) on hardware with lower VRAM.ComfyUIAI Music Generator
  3. 4:270:30Copying and replacing prompts in ComfyUIThe user demonstrates a custom node feature that allows hovering over prompt examples to copy and then replace the current prompt in the workflow with a single click.ComfyUIAI Image Generator
  4. 8:420:58Enhance audio prompts with Gemma 4A workflow is shown where a simple piano prompt is processed by a Gemma 4 model to generate a more detailed audio prompt before being passed to the Stable Audio sampler.ComfyUIAI Music Generator
  5. 10:550:28Image-to-Music generation in ComfyUIThe creator demonstrates an experimental workflow that loads an image of a bunny, uses Gemma 4 to describe it as a music prompt, and then generates corresponding audio.ComfyUIAI Music Generator
  6. 13:030:51Managing node colors with PixaRoma nodesThe video demonstrates UI updates for PixaRoma nodes in ComfyUI, including selecting color swatches, copying/pasting colors between nodes, and setting favorite colors.ComfyUIAI Image Generator
  7. 14:161:11Advanced image loading and padding in ComfyUIThe narrator demonstrates the updated Load Image node featuring folder filtering, image previews, and manual padding controls for outpainting tasks.ComfyUIAI Outpainting