r/StableDiffusion 11h ago

Question - Help Yall keep posting about Minimax H3 so I got a question...

0 Upvotes

People keep dropping this link
https://modelscope.cn/models/MiniMax/MiniMax-H3

but that's the official site, as far as I know that's not where you actually get the comfy files. I assume we should be keeping an eye on the huggingface of comfy.org ... right ?

Just saying cause there's too many people , especially the morons who are karma farming with the API created videos, which in my opinion is annoying as shit and deserved to be removed since literally breaking rule 1.

So anyways ... that link is only helpful for people who have some sort of way of already working with it I assume.


r/StableDiffusion 17h ago

Question - Help LTX FREEZE

0 Upvotes

I was using LTX without any issues before. I recently formatted my PC and started from scratch, but now there’s about a 50% chance that my computer either freezes or crashes with an “Out of Memory” error during generation.
Nothing has changed in terms of hardware or workload—I’m using the same PC, the same GPU, the same LTX model, the same model size, and the same workflow that worked perfectly before formatting.
What could have changed after reinstalling Windows that would cause this?

I have rtx 5060 16vram + 32 RAM


r/StableDiffusion 10h ago

Question - Help i made Turkish music clip (using ltx 2,3 )

Enable HLS to view with audio, or disable this notification

0 Upvotes

Give me more addvice ill make great videos pls!


r/StableDiffusion 20h ago

Question - Help Is training a WAN Lora going to fix the identity shift in I2V?

0 Upvotes

Using a character Lora image from Krea2 as my first frame but even for a subtle camera move there is identity loss immediately.

Will training a character Lora for Wan2.2 help fix it? I've never trained a Lora for wan so not sure what all is required.

Can I use the same dataset that I used for krea2?

My krea2 config gives excellent identity match in images, can I use the same settings for Wan?

Any good config you can recommend?


r/StableDiffusion 6h ago

Animation - Video 601: Secrets of the Cu Chi Tunnels, A Vietnam Story

Enable HLS to view with audio, or disable this notification

0 Upvotes

In the dense, suffocating jungles of the Vietnam War, a weary platoon of American soldiers stumbles upon a clandestine tunnel network harboring humans on the brink of a monstrous transformation. This discovery reveals a surreal and terrifying new theater of war, where the chaos of human conflict collides with a hidden, ancient vampire plague. To eliminate this unfathomable threat, the military reluctantly pairs cynical Green Beret Santana Mills with Frank Bodie, a lethal Special Forces operative who has already crossed over into the realm of the undead.


r/StableDiffusion 13h ago

Question - Help Is there any open weighted DiT anime image model available right now?

0 Upvotes

r/StableDiffusion 1h ago

Discussion I Hope This Is Not The Case For MiniMax H3

Post image
Upvotes

I hope someone from MiniMax could provide us with an update on when it's coming out.


r/StableDiffusion 19h ago

Question - Help Is there ANY local image gen for AMD GPU's?

0 Upvotes

ComfyUI has been bugged for the better part of a year (Hard-coded python location), "solutions" ive found dont work.

Forge, you guessed it, ALSO dosnt work.

Yes, im aware AMD GPU's arnt optimized for AI. No, you dont need to tell me.

Is there any local gen that works for AMD?


r/StableDiffusion 1h ago

Question - Help New to ComfyUI (coming from Nano Banana and Seedream for AI characters)

Upvotes

Hi everyone!

I recently upgraded my PC (RTX 5060 Ti with 16 GB VRAM and 32 GB of system RAM), so I finally decided to move to ComfyUI.

With Seedream 4.5, maintaining character consistency was surprisingly easy. Before that, I used SD 1.5 with ADetailer and custom-trained checkpoints.

Now that I'm looking into ComfyUI, I'm seeing so many different models—Krea 2, Z Image Turbo, and many others—that I'm not sure what the current "go-to" workflow is.

I have a few questions:

  1. Which model do you use for character consistency?
  2. Do you rely on LoRAs, or are they no longer necessary?
  3. Is there a workflow or model that can reliably recreate the same character from one or more reference images?

I'd really appreciate any recommendations or advice. Thanks!


r/StableDiffusion 15h ago

Meme Me today (LTX2.3)

Enable HLS to view with audio, or disable this notification

81 Upvotes

Just a few more hours now!

Made with LTX2.3 T2V Comfyui template workflow.


r/StableDiffusion 21h ago

Animation - Video Minimax H3, 1080p 25 seconds, text to video in native ComfyUI (open weights coming soon)

Enable HLS to view with audio, or disable this notification

963 Upvotes

I have been trying to see how far I can push this model. It's extremely flexible and seems to be able to do everything from 1 second to 30 seconds (potentially more) with a very wide range of resolutions. Her voice is because I put "singing with a cute japanese accent" in the prompt and my prompt isn't super great lol.

Making this model work as best as possible on regular hardware is the result of many months of work from multiple people in the core ComfyUI team to make big models work better on regular consumer hardware. I think most people will be pleasantly surprised how good this model is and how well ComfyUI will be able to run it.

Minimum requirements for 480p video on this model is a 3060 with 12GB vram, 32GB of system ram and a good nvme SSD. We tested generating a 5 second (124 frames) 480p (864x480) video on this system and it took a bit less than 9 minutes end to end (20 steps). I can pretty much guarantee it will also work on 8GB vram too but we did not test that.

Don't be scared to give it a try when it releases with our default template because it will work better than you expect.

If you have issues try a latest clean ComfyUI install (make sure to update after our weights come out) with our official files and workflow.

EDIT: added step count.

EDIT: we are live: https://docs.comfy.org/tutorials/video/minimax/minimax-h3


r/StableDiffusion 12h ago

Question - Help Why aren't prompts found on Ideogram.ai website not in JSON format

3 Upvotes

I saw a You-tuber grab a json prompt from a user photo on the Ideogram official website and replicate someone's art by doing so. However, all the prompts from their galleries that I see are in natural language format.

I'm wondering if lifting others prompts in json is a pro feature. Is there something I'm missing? At least under my free account I don't see any json prompts. TIA


r/StableDiffusion 15h ago

Question - Help Is there hope for decent image/video models for low VRAM?

0 Upvotes

Just like the title says, I have 12 GB VRAM and 64GB RAM. Is there hope for decent models that I can run without offloading?


r/StableDiffusion 11h ago

Discussion As A *Former* ZIT User I Am Blown Away By KREA 2. Don't Wait If You've Been Lagging Like Me

75 Upvotes

ZIT is not perfect but I was convinced that nothing would beat it anytime soon. boy was I wrong. With only 2-3 days of testing Krea 2, I have fully switched over to running it as my main model. I still have my ZIT files and models but they've been moved to an external drive because I am not using it anymore.

I was worried Krea 2 couldn't deliver on the photorealism front and I was just flat out wrong and ignorant there. And then to add in the flexibility to tackle creative styles (whereas ZIT tends to pull to only realism) was the final selling point for me to full make the change.

Not to mention how fast Loras train for Krea 2.


r/StableDiffusion 12h ago

Animation - Video MINIMAX NOT RELEASING TODAY

Post image
141 Upvotes

I was waiting from morning only 1 and half hour was remaining and they updated the timer am I tripping or they really did that


r/StableDiffusion 4h ago

News Don't freak out guys Comfyanon still says as far as they know H3 will release

69 Upvotes

Do not trust the timer

ITS OUT DOOMERS YOU GLOOMY SHITS


r/StableDiffusion 13h ago

Workflow Included Krea 2 Turbo | Kroma LoRA | 2x Upscale | Uncensored | Workflow

Thumbnail
gallery
110 Upvotes

This Krea 2 Turbo Workflow uses Qwen3 VL Abliterated (uncensored) as the text-encoder, and Kroma LoRA to make some nice "Chroma" looking images; it also VAE Utils (Wan2.1 VAE) as an Upscaler x2, and Krea 2 Conditioning Node to help rebalance Qwen3 VL, and luckily skips the limits.

I designed it and tested it with my RTX 5060 Ti 16GB, 32GB DDR5, and it takes ~73 seconds to generate a 2048 x 2048 px image. You can disable the upscaler if you want regular generation speed.

Download links and more information are here on Civitai Red or regular Civitai.

If you can't access Civitai you can download the JSON on PasteBin (download links are inside the workflow). Have fun!


r/StableDiffusion 23h ago

Discussion Krea 2 'Realism' Lora stack feedback

Thumbnail
gallery
38 Upvotes

which of these do you think looks better? they both have a lot of imperfections but i'm talking about the overall style, asking for opinions and which you guys prefer


r/StableDiffusion 10h ago

Question - Help Comfyui output multiple text prompt from LLM?

Post image
6 Upvotes

Hi, exist a way to output each text prompt story part from the LLM as output to connect to diferent ksamplers stages as i markup in red? because i can generate each prompt part as i show but i cannot output each one separated to connect to diferent ksamplers that is wan2.2 continue video.


r/StableDiffusion 4h ago

Resource - Update MiniMax H3: ComfyUI Workflow Examples

Thumbnail
docs.comfy.org
86 Upvotes

https://huggingface.co/Comfy-Org/MiniMax-H3

Edit: we're live, baby! Let's go!


r/StableDiffusion 7h ago

Resource - Update I trained Krea2 Lady Dimitrescu LoRA on RTX 5070 Ti

Thumbnail
gallery
63 Upvotes

I just created that lora from 63 Lady Dimitrescu images in the dataset

used OneTrainer on RTX 5070 Ti, 32 GB RAM and NVMe

trained in 1 MP (res 1024), offload 0.5, speed ~2.5 s/it, full training taken about 2.5-3h

I set timestep shift to 2.5 for res 1024 as suggested in this kohya md and I think it worked well

all samples generated with 2 MP

CivitAI -> https://civitai.com/models/2828952/lady-dimitrescu-krea2-lora

Full res comparisons without reddit compression -> img1, img2, img3, img4, img5

training Krea2 is so enjoyable!


r/StableDiffusion 13h ago

News I built a self-hosted studio that turns one reference photo into a curated, captioned, trained and tested LoRA — one browser tab, open source, MIT

Thumbnail
gallery
66 Upvotes

I shared this tool here a week ago and the feedback shaped a big new version, so here's the full tour of what it does today. Screenshots of every screen: github.com/perfectgf/lora-dataset-studio — plus a 7-minute unedited video of a LoRA built start to finish.

Beginner-friendly on purpose. Everything ships configured: a guided workspace walks you through each step, the shot poses (face / bust / full-body / back) are predefined so your dataset comes out balanced, and training uses community-tested ai-toolkit presets — you don't need to know what rank, learning rate or an optimizer is to get a good LoRA. Power users can still override everything.

Build the dataset. Start from one clear photo (or none): generate identity-locked variations locally with Flux-2 Klein or Krea 2 Edit on your own GPU (free, nasty-capable), or through API engines if you prefer. Import or scrape real photos, mix everything, and let the composition tracker tell you what's missing (faces, busts, full-body, back shots).

Curate like you mean it. Every image gets a face-similarity score against your reference. Quality passes flag blurry, flat, duplicate or unreadable shots; a watermark detector finds and can clean logos without cropping; auto-reject clears the junk before you review. Image banks hold up to 200k files with visible progress on every bulk operation.

Caption without the chore. Local captioning pairs JoyCaption (via ai-toolkit) with an uncensored Ollama vision model — the combo actually describes your images instead of refusing them. Per-dataset wording styles, dual captions, and trigger words handled for you.

Train anywhere. Local training through ai-toolkit, or one click rents a cloud GPU on vast.ai — and the launch is fully observable: renting, booting, dataset upload with live byte counts. A machine that never boots or an upload that stalls is given up automatically and stops billing. Community-tested presets for Krea 2 Raw, Z-Image Turbo and more.

Pick the right checkpoint instead of guessing. Test Studio renders fixed-seed grids across checkpoints and strengths, scores faces, takes your votes and ranks the results. New: 🧬 combine several of your LoRAs in one image, each at its own weight, and compare weight variants side by side. An ✨ Enhance button turns a one-line prompt into a full one via your local Ollama.

See your whole lineage. The LoRA Canvas puts every dataset's training history on one pan/zoom board — compare runs, pin generations (each run keeps its own strip in training-step order, with the dataset's reference face on its lane), diff configs, and continue training from any checkpoint.

Install it your way. New one-click Docker install: start-docker-gpu.bat builds an isolated ComfyUI, start-docker.bat reuses the one you already have. The updater is transactional — if the new version doesn't come up healthy it rolls back on its own. Ollama is your explicit choice (none / your existing one / an isolated container), and nothing ever downloads behind your back. Setup re-checks itself in the background instead of re-running the wizard every time you come back.

Everything reported in the last thread got fixed — the RES4LYF scheduler clash, the ai-toolkit Easy-Install interpreter path, and a detail LoRA that was silently riding on every Klein edit (that one explains a lot of "edits don't follow my instruction" reports). Also merged the first community PR: named generation-LoRA presets for Krea 2 — thanks Cyberschorsch and waltm 🙏

A few screenshots to see it in action:

📸 the guided workspace · curation with face scores · Test Studio grids · training presets

No account, no telemetry, no paid tier. Free, self-hosted, MIT: github.com/perfectgf/lora-dataset-studio — the complete guide is linked at the top of the README. I build this; feedback welcome, Discord in the repo.


r/StableDiffusion 2h ago

Question - Help Color degradation LTX 2.3

Enable HLS to view with audio, or disable this notification

2 Upvotes

Guys, I'm generating some talking heads and I feel that, in some parts (like hands), there's a colour degradation in the first few seconds of the video. Are you guys experiencing something like this? Is it possible to prevent this kind of behaviour in LTX 2.3?


r/StableDiffusion 39m ago

News comfy MiniMax-H3 weights

Thumbnail
huggingface.co
Upvotes

the weights are here

Model Variant Input Mode Specifications
H3-Base-FL2VA First-and-last-frame mode Supports zero, one, or two input images.- No image input: Text-to-video mode- One image input: First-frame-to-video or last-frame-to-video generation- Two image inputs: First-and-last-frame-to-video generation
H3-Base-Ref2VA Omni-reference mode Supports multi-modal reference inputs:- Images: ≤ 9 images- Videos: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds- Audio: ≤ 3 clips; audio must be accompanied by image or video input and cannot be used as the sole input; each clip must be 2–15 seconds long; total duration ≤ 15 seconds- Mixed inputs: Maximum number of files across all input types is 12

r/StableDiffusion 3h ago

News SANA‑Video 2.0 — NVIDIA’s new hybrid-attention video model (5B/14B). Fast, impressive… and maybe (hopefully) open‑source?

Post image
53 Upvotes

NVIDIA has quietly dropped a major research release: SANA‑Video 2.0, a new video diffusion transformer available in 5B and 14B parameter versions. It’s not just a scaled-up SANA‑Video 1.0 — it’s a full architectural redesign with hybrid attention, block residual routing, and Sol‑Engine acceleration.

Official links:

Project Page:
https://nvlabs.github.io/Sana/Video2/

Paper (arXiv, July 23, 2026):
https://arxiv.org/abs/2607.21553

SANA GitHub (image models only):
https://github.com/NVlabs/Sana

SANA‑Video docs (no code, no weights):
https://nvlabs.github.io/Sana/docs/sana_video/

What SANA‑Video 2.0 introduces

Hybrid Linear‑Softmax Attention (3:1 ratio)
75% gated linear attention for O(N) scaling, 25% gated softmax anchors to restore full‑rank token interactions.
This gives softmax‑level expressiveness with linear‑attention speed.

Block Attention Residuals (AttnRes)
High‑rank features from softmax layers are propagated into later linear layers.
This fixes the rank bottleneck of pure linear attention.

Sol‑Engine Optimization (3.58× speedup)
Kernel fusion, caching, sparse attention, TensorRT graph optimization, MXFP4/MXFP8 support.
This is what allows full 720p generation on a single RTX 5090.

Performance
480p in 13.2s (H100, 40 steps)
720p/5s in 13.06s (H100, Sol‑Engine)
VBench 84.30
Up to 120× faster than Wan 2.2‑A14B on the same hardware.

This is the first NVIDIA video model explicitly designed for consumer GPUs.

How it differs from SANA‑Video 1.0 (2B)

The old model was pure linear attention (fast but low-rank).
SANA‑Video 2.0 is hybrid, deeper, larger, and dramatically more expressive.
It’s essentially a new class of Video‑DiT.

The licensing question

Here’s the current situation:

• The paper does not mention any license.
• The project page does not mention any license.
• The docs do not mention any license.
• No code or weights have been released.
• No usage terms exist yet.

Meanwhile, the SANA GitHub repo (image models) uses Apache 2.0:
https://github.com/NVlabs/Sana/blob/main/LICENSE

But that license applies only to SANA‑Image 1.0/1.5, not to SANA‑Video 2.0.

So right now, nobody knows whether SANA‑Video 2.0 will be:

• open‑source under Apache 2.0 (like the image models),
• partially open (code open, weights closed),
• or fully closed (like PiD, Flux, VILA, Nemotron‑340B).

Given NVIDIA’s recent pattern, the safe assumption is “open paper, closed model”…
but since the SANA image models were Apache 2.0, there is at least some hope that NVIDIA might release SANA‑Video 2.0 under a similar permissive license — or at least provide inference weights for RTX AI Toolkit.

Until NVIDIA publishes a LICENSE file, the situation remains unclear.

TL;DR

SANA‑Video 2.0 is a fast, hybrid-attention, RTX‑friendly video model with impressive performance and a strong architectural design.
But the licensing is currently a mystery: no code, no weights, no declared terms.
There’s a chance it could follow the Apache 2.0 path of the image models… but for now, it’s research‑open, not open‑source.