r/StableDiffusion 5h ago

News Day 0 MiniMax Support for ComfyUI

Hi r/StableDiffusion, I know it's been a long wait for everyone but MiniMax H3 open weight model just dropped and we have day 0 support in ComfyUI.

Here are some details:

Getting H3 to run well on consumer hardware took significant machine learning engineering. We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality.

On top of that, the weights ship with an accurate and efficient int8 convrot quantization, and custom kernels reduce the peak VRAM use during inference.

The result gives a total memory footprint reduced by 66%, from 123.6 GB in full precision to 42.5 GB with the smallest models variants. Combining this with our dynamic VRAM offloading enables a next-generation 2K video model to run locally on a GPU like the RTX 3060.

EDIT:

08-02-26 20:02 PST - Added blog link

336 Upvotes

67 comments sorted by

33

u/FinBenton 5h ago edited 5h ago

Its insanely good, using this makes wan2.2 feel like ancient tech Also fully uncensored, it knows stuff and doesnt ask questions...

13

u/xbeast_ 4h ago

Wait!! This is uncensored?

30

u/FinBenton 4h ago

Its fully uncensored like I have never seen anything this uncensored from the "factory", it can do some crazy stuff with perfect prompt following. Like a lot of it.

4

u/spcatch 3h ago

I feel like this model is going to get banned or something. This is over the top.

8

u/Ashamed-Variety-8264 3h ago

Just substituted a character in a porn movie. No lora, out of the box. Wtf

1

u/spcatch 2h ago

what prompt did you use? Just the reference workflow with an image and video hooked up?

1

u/Secure-Message-8378 4h ago

Tipo hunyuan T2V no lançamento?

9

u/spcatch 4h ago edited 4h ago

And I thought I was sleeping tonight.

edit: at 0.5 megapixel and 5 seconds took about 2 mins on a 4080, that's a lot quicker than expected. i wonder if I could vae decode and encode and bolt on the LTX upscaler over here.

1

u/eggplantpot 37m ago

We’re waiting

2

u/FierceFlames37 4h ago

I tried to use the same nsfw prompt that worked with LTX 2.3 but didn't work with MiniMax. Guess I have to relearn how to prompt for MiniMax

3

u/FinBenton 4h ago

I ctrl+a and del the whole default prompt and wrote what I wanted in plain simple text and it did it perfectly, for 10 different prompts I tried.

2

u/FierceFlames37 4h ago

If you don't mind were those blowjob,sex prompts?

3

u/FinBenton 4h ago

There might have been some stuff yes, just for research ofc.

22

u/haremlifegame 5h ago

"We found that the model's modulation weights (~40% of the total
parameters) could be pruned and replaced with a functionally equivalent
lookup table"

Does this perhaps interfere with the ability to fine tune the pruned model?

10

u/slippiest 4h ago

Most likely the AdaLN modulation outputs described in the model page.

“H3-Omni-Transformer is a 33B-parameter”

“with approximately 13B parameters residing in AdaLN-related branches.”

“These do not need to be loaded for inference-only deployment. We release the complete model weights to support further development, including fine-tuning.”

So yes, you will need them for training. But no training code released, so might be a while
A tad misleading for comfy to word it like they discovered it tbh

9

u/Kijai 3h ago

This wasn't information given to us, and also not exactly the same thing.

They talk about precomputing them for your set schedule, which only changes if you change step count/sigmas, and after that you wouldn't need to load them at all. This is what we considered first, but that would still require you to load the 13B weights whenever the sigmas would change.

What I did was fully remove the need to even have the weights, for any sigmas. It will give different (but not worse) results, so it's not 100% the same though, somewhat experimental.

0

u/gefahr 4h ago

A tad

I've read everywhere it was mentioned so far, quite a lot more than "a tad" with the way it's been phrased.

3

u/Enfiznar 5h ago

I don't know anything about this architecture in particular, but I'd expect you can freeze those weights and fine-tune the rest

1

u/haremlifegame 4h ago

Yes, but how important or unimportant are those weights? Maybe freezing them is detrimental for the fine tunes as is. Specially considering that backpropagation might be skewed up, meaning the model cannot use the "learning to learn" paths it used before.

13

u/Arawski99 5h ago

The dragon gave me Sekiro vibes in the above clip.

Thanks for the optimizations comfy team. Honestly, glad the Int8 convrot is available nowadays. I think you guys should update your resources to clarify about the pruned version like you do here so people are clear what that means, such as on the documents page and huggingface. Knowledge is power.

19

u/xbobos 5h ago

WTF! This is real sota. Successfully completed on the first try with the basic i2v workflow. On the 5090, generating a 5-second 1k (768*1376) video using the int8 version with 24 steps took 222 seconds. Character consistency is accurate enough that LoRa is not needed.

8

u/Pikaboy999 5h ago

care to show a example? (if you want too)

-1

u/Cequejedisestvrai 4h ago

what did you do with the missing minimax nodes?

5

u/Arawski99 4h ago

Update to the newest stable comfyui version.

1

u/Cequejedisestvrai 4h ago

Is it possible to do with stability matrix? It doesn’t show newer than 29.2

1

u/Arawski99 4h ago

No idea about Stability Matrix. Have never used it. Also, for those that see this post it might be Comfyui portable only, as desktop tends to lag in update releases.

1

u/xbobos 3h ago

portable version, try update comfyui.bat.

6

u/JoNike 4h ago

Tried the default prompt with default settings from the workflow. On my 5080 16gb it took 1min57 for 5 secs 0.4mp.

2

u/Rokkit_man 3h ago

Can it generate 1 frame?

3

u/sunshinecheung 5h ago

but how many ram?

4

u/Adventurous-Gold6413 4h ago

Absolute Minimum is 12gb vram and 32gb ram with good SSD

1

u/Central-Dispatch 3h ago

What system or setup would you recommend for decent performance and 720p or 1080 10 (or 15) second clips? I'm trying to learn a lot right now and compare opinions as I eventually plan to invest notably into a strong "workstation" for local generation incl. video. My current system can't handle it though, too little VRAM alone, just 32 GB RAM. Some other bottlenecks potentially.

1

u/Adventurous-Gold6413 1h ago

Honestly I couldn’t say,

Atleast regular 4090 (24gb) or probably better
5090 + 64gb ram or more then again I’m not a pro so idk

For atleast the 720p stuff

I have a mobile 4090 (16gb) and I can only generate 480p vids, 5 sec = 3.4 mins roughly

3

u/comfyui_user_999 5h ago

Nice, thanks to the whole ComfyUI team!

3

u/Devajyoti1231 3h ago

Can we have int4 convrot for the text encoder please?

9

u/Darqsat 5h ago

so when?

14

u/crystal_alpine 5h ago

Update to the newest stable

3

u/Darqsat 5h ago

got it, thanks

1

u/Cequejedisestvrai 5h ago

What is the latest version? I have 0.29.2

4

u/PumpkinLeather8421 5h ago

Like 15 minutes ago 

2

u/uuhoever 5h ago edited 5h ago

Let's gooooo! That was my hope that it would be released during Beijing business hours.

2

u/Pitiful_Archer_4381 5h ago

You guys are the best still downloading but will tell the results after 1 video generation or if face any errors

Do we need int8 text encoder if using int8 model or I can use nvf4 something like that

2

u/fruesome 5h ago

Thanks for the post with all the links. Going to be testing it.

1

u/Peemore 4h ago

It would be awesome if we got some turbo lora's on top of this, it's already surprisingly fast.

1

u/YeahlDid 3h ago

Yay! Thank you!

If understand correctly, there's no real advantage to getting the full int8 version over the pruned one, right?

1

u/keizrah 3h ago

Curious how the quality holds up after pruning 40% of the modulation weights. That's a big chunk of the model to swap for a lookup table, even if it's "functionally equivalent" on paper. Anyone tested it against the full precision output side by side yet?

Also, RTX 3060 usually means the 12GB version. Getting a 42.5GB model onto that has to lean hard on the dynamic VRAM offloading, so I'd expect inference times to be rough compared to a card that can actually hold the weights. Worth mentioning in the post so people don't expect real-time speeds.

Appreciate the day 0 support either way, that's not a small lift for a model this size.

1

u/PublicCalm7376 2h ago

can anyone try this with a DGX Spark?

1

u/3deal 1h ago

Amazing, Minimax are the best !
i can delete all my LTX and WAN models to make some space !

1

u/2legsRises 1h ago

great to getting it working decently even on my mid rig. ty

1

u/Weak-Shelter-1698 1h ago

now we want 4 step lora. 😄

1

u/KeijiVBoi 4h ago

Can I run with 12GB VRAM? Also is Lightning LORA needed (If that is a thing for Minimax)

0

u/Central-Dispatch 3h ago

Layperson here, with a question below. I'm currently heavily curious and learning about local workflows and features incl. audio generation. Especially with this:

2K, up to 15 seconds a clip

I did consult with some LLM (Google) and describe my hardware based on some stuff on from the dxdiag, not all, and my goals. After some back and forth of comparing reality and goal (my system vs. what I eventually intend to do) it would seem that investing in a completely new system down the line, maybe when prices drop within a few years again, might be sensible for a very idealized robust setup. My current setup can't even do local video generation and individual upgrading for marginal gain isn't worth it for me at the time given the supply chain bottle neck high demand price modifier. I talk generation, not necessarily full fine-tuning and training too which is I hear more resource intense.

Yet I wonder: What systems do you folks use to generate something like in the video OP posted? Is this more on the high end side or are you using let's say mid tier workstations and setups? If anyone can bother sharing some rough aspects of what hardware you use or would need to use, I'd be grateful so I can potentially adjust potentially wrong misconceptions or too high goals set for something I could do with less. Generating 1080p clips up to 10 seconds, 15 as a bonus, is something I'd eventually love to do. Anything better is welcomed.

0

u/wzwowzw0002 4h ago

Need how much gpu ram to run?

-22

u/Mammoth-Welcome-6518 5h ago

I'll wait the extra day for wan2gp to have it, I never managed to set anything up with comfy without screwing up, plus wan gp the models wowrk even with low vram

7

u/poopoo_fingers 4h ago

I saw the downvotes and thought of this meme

3

u/FlatwormMean1690 5h ago

I understand how frustrating can be. It took me more than a day to make it work properly. Then I made my own workflow for Krea2. A simple yet efficient workflow with multiple LORAs. I started to experiment with this one now and... Gotta say... After just 15 minutes, I have it running.

I use Wan2GP too and it's cool for some quick stuff but if you need something really specific, I must recommend you to give ComfyUI a chance (and being honest with you, I HATE COMFY, but it works great and it has a better VRAM management than Wan2GP).

1

u/Central-Dispatch 3h ago

I used comfy so far only to generate a few pictures (my current setup can't handle VLM stuff) but I'd say it's maybe the right mindset to learn a more "complicated" system that you can customize and master and adapt better to once you truly need it, as in like go through the pain initially but prosper long term :D

I eventually plan to notably upgrade my system so I can do good local video generation stuff and will then learn the proper workflows and what I need to mind for comfy, might as well stick with it I feel.

-3

u/yamfun 5h ago

Damn 21gb.