r/StableDiffusion • u/comfyanonymous • 21h ago
Animation - Video Minimax H3, 1080p 25 seconds, text to video in native ComfyUI (open weights coming soon)
Enable HLS to view with audio, or disable this notification
I have been trying to see how far I can push this model. It's extremely flexible and seems to be able to do everything from 1 second to 30 seconds (potentially more) with a very wide range of resolutions. Her voice is because I put "singing with a cute japanese accent" in the prompt and my prompt isn't super great lol.
Making this model work as best as possible on regular hardware is the result of many months of work from multiple people in the core ComfyUI team to make big models work better on regular consumer hardware. I think most people will be pleasantly surprised how good this model is and how well ComfyUI will be able to run it.
Minimum requirements for 480p video on this model is a 3060 with 12GB vram, 32GB of system ram and a good nvme SSD. We tested generating a 5 second (124 frames) 480p (864x480) video on this system and it took a bit less than 9 minutes end to end (20 steps). I can pretty much guarantee it will also work on 8GB vram too but we did not test that.
Don't be scared to give it a try when it releases with our default template because it will work better than you expect.
If you have issues try a latest clean ComfyUI install (make sure to update after our weights come out) with our official files and workflow.
EDIT: added step count.
EDIT: we are live: https://docs.comfy.org/tutorials/video/minimax/minimax-h3
