Z-Image's eye for texture on MiniMax-H3's engine. Drop-in replacements for the standard H3 checkpoints: same identity, same voices, same speed, same VRAM, same workflows - but sets and surfaces render noticeably richer. Peeling paint peels harder, rust bleeds further, water carries more light.
And the extra detail stays flat across chained shots - measured on the same seed, the graft ran a 0.99 texture ratio over three joins where stock drifted to 1.11. More detail, no per-shot sharpening creep from the graft.
What is this, actually
Z-Image is a 6B image model with exceptional fine-texture rendering. Both it and MiniMax-H3 normalise attention queries per head - so the shape of how sharply each attention head commits to fine detail can be transplanted between them by rescaling those normalisation weights. That's the whole trick: no retraining, no new knowledge, H3 keeps everything it knows and attends to texture the way Z-Image does. Early blocks are deliberately left untouched (grafting them creates a weave artifact in tile-like textures - we measured, so you don't have to). This is the second marriage in this line - Joy-LTX 2.5 put JoyAI-Echo's performance on LTX-2.5's engine the same way, donor to engine. The statistics-transplant mechanism here is our own; a nod to TenStrip, whose H3 attention experiments sparked the question of what an image model could donate.
Which file
Two versions on this page - fl2va (default) chains and lands on supplied frames; ref2va adds reference images, voice anchoring and the identity bank for characters that must persist. Every format is attached to its version:
RTX 30 / 40 - take a GGUF (needs ComfyUI-GGUF; the Multishot pack's loader takes both formats):
curve-zs05-Q8_0- 21.5 GB - 32 GB cards, closest to full precision.curve-zs05-Q5_1- 15.2 GB - the 24 GB pick.curve-zs05-Q4_0- 11.5 GB - the 16 GB pick.plain (non-curve)
Q5_1/Q4_0(andQ3mixon fl2va) - for workflows built on the original bakes.
RTX 50 - take a comfy-native (stock Load Diffusion Model, ComfyUI 0.32+):
comfy-int8- 21 GB - the fastest file on Blackwell, 32 GB cards.comfy-fp8- 21 GB - the fp8 twin (ref2va also shipsfp8e5m2).comfy-w4a8/nvfp4/mxfp8- ~11-12 GB - the 16 GB family (ref2va also shipsw4a4).int8_convrot(ref2va) - the Lightricks-style convrot build.
The master: ref2va pruned zs05 bf16 - 40.2 GB - quantise your own cuts from it.
Install
Put the file where your H3 checkpoints live. Pick it in your loader. Done - every H3 workflow works unchanged, including the MiniMax-H3 Multishot seamless-chain canvases.
Links
All formats: the GGUF and comfy-native repos on Hugging Face (joeygambino). Workflows: the MiniMax-H3 Multishot page. Questions: comment here - I answer.
Description
The default variant. Chains seamlessly, lands on supplied first/last frames.
MiniMax-H3-fl2va-curve-zs05-Q8_0.gguf- 21.5 GB - 32 GB cards.MiniMax-H3-fl2va-curve-zs05-Q5_1.gguf- 15.2 GB - 24 GB cards.MiniMax-H3-fl2va-curve-zs05-Q4_0.gguf- 11.5 GB - 16 GB cards.
Plain (non-curve) bakes and comfy-native formats (bf16 / fp8 / int8 / w4a8 / nvfp4) are on Hugging Face.
FAQ
Comments (54)
Would be nice to have some actual examples.
I agree. You should make some!
@joeygambino You went through the effort to make and upload this... and yet you wont post a single example? I assume you tested this before you uploaded it, why upload one of those samples?
@q5sys I did post an example... it's literally in the header with the title image. I am not going to post a library of examples, though, no. There are 25 checkpoints in this group, only three of which are uploaded so far, this repo is far from complete. If you want examples, you have the technology.
I agree with the comment. I'm not downloading any 20Gb files without a convincing example. I'll wait till somebody else wastes their time instead of me.
@randombrowser1234 So.. don't? What do you think I lose when you don't use my completely free model, exactly? What is the threat? I genuinely don't understand.
i think this are the demos https://huggingface.co/joeygambino/MiniMax-H3-x-Z-Image-GGUF/tree/main/demo
@attc123 oh thanx !
@joeygambino fair enough, thanx
Are you still uploading models on HF? Can't see them there.
Five minutes, about to flip public.
@joeygambino Good model so far but I think paired with SLA and "low resolution" (running at 544p) texture thing is not as vivid. Maybe need some more tests with regular dareties turbo and in higher res. But so far model doesn't ruin anything on its own so I'd say it's a good job.
@DigitalGarbage Yeah, these are meant to generate higher quality renders, and I haven't even tested Turbos LoRAs with them at all - I've yet to test turbo loras on any of my models, to be honest. I generally run my tests at 1280x736 - most of my real content is for social media and vertical aspect ratio.
managed to find your HF page. nice you're using same username everywhere,
wondering why you don't bother with the int8 models and such? are you finding the ggufs are better quality and speed wise?
@QualityControl Thank you! Should be INT8 models in all of the repos as well, unless I missed one somewhere? I try to get up every format I can think of so everyone has something they can use.
@joeygambino my mistake. I was in the gguf H3 Zimage repo. not this : joeygambino/MiniMax-H3-x-Z-Image-native
which is the same as this civit model right?
btw the int8 isnt posted on this ref2va civit page. only the gguf. The fl2va has an int8 version.
@QualityControl Civ is extremely slow to upload, so I get things up on HF first. I am working on those now. Yes, same models - sometimes HF has more because I get annoyed with uploads here ;) Also, I meant to tell you before that I'm a big fan of your Amateur Hour lora - most of my non-demo content is vertical found footage content for socials, and it was one a I used all the time! You should totally make one for LTX2.5 and H3.
Currently having decent outputs with LTXV 2.5 on RTX5070 12GB. Could I run minimax h3 or is it a lot heavier than LTXV 2.5 ?
It's quite a big heavier, but my Q4 GGUF is 10.68GB for 12-16GB cards. You'll still go over when you throw in the TE, VAE's, etc., but it will work, just kind of slow. LTX is much more forgiving where VRAM and speed is concerned, but H3 beats it in quality all day.
That said, I do have some LTX merges that you might want to check out as well:
Joy-LTX 2.5 Distilled GGUF / INT8 / NVFP4 / W4A8 / MIX4x8 - 16GB | LTX Video Checkpoint | Civitai
i run it on same card, no problem, just a little slower
@joeygambino 10gb on SSD, but in VRAM more than 14gb. You need Q3 or Q2
@gambikules858 There is a Q3 at that link. I built a Q2, but didn't upload it, because it's awful.
Unfortunately, on a 12GB card, you'll never avoid offloading.
@joeygambino Just found your nvfp4
I got a Blackwell so I figured I'll try this one :
MiniMax-H3-fl2va-pruned-zs05-comfy-nvfp4.safetensors
NVFP4 4-bit, Blackwell-optimized
But upon downloading it, the file name is different : minimaxH3XZImageRicherSetsAnd_fl2va.safetensors.
So, are they the same one ?
@randombrowser1234 You should be able to just add a GGUF loader. I'd try Q3 or Q4. There's a few GB difference, and if the quality is good with Q3, no reason not to save the VRAM for a little speed boost. I test these at 736x1280 and upscale at 1.5 from there usually and the quality is pretty good, even at Q3.
That said, I have never tested turbo loras on any of these, so I can't account for quality if you do.
@randombrowser1234 Full disclosure on the nvfp4 - I am not a big fan and I'd take the slower renders to get the better quality and still stick with GGUF over it. Or maybe even try the w4x8.
@randombrowser1234 You may need to update your nodes/ComfyUI. It doesn't support Minimax GGUF's natively. You can also add support manually:
Open:ComfyUI/custom_nodes/ComfyUI-GGUF/loader.py
Search for the error string or locate the architecture check block (around line 90).
Modify the code block to bypass the error for minimax_h3 by adding and arch_str != "minimax_h3": [1, 2]
# Before
elif arch_str not in IMG_ARCH_LIST and not is_text_model:
# After
elif arch_str not in IMG_ARCH_LIST and not is_text_model and arch_str != "minimax_h3":@joeygambino Thnx fo your efforts, I'm not going to fight uncooperating GGUF. Going to try w4a8. Isnt it annoying that Civitai keeps renaming all your files to the same generic name ? lol. I am renaming them manually to avoid confusion in the folders.
for me 4070 laptop 8gb vram minimax is lighter than ltx2.5
Can you create an example with human skin?
Yes.
It looks like the bf16 file was uploaded incorrectly.Damn, Civ mislabeled it. I'll upload the real bf16 today.
int8 convrot plz
It's on his huggingface, sadly HF has been very slow for me, have to DL overnight.
I'm working on the IN8 upload here today
Z IS GOOD~
YOU! is GOOD!
NEED fl2vapruned int8_convrot.safetensors ~
Uploading now. ;)
You may want to download from my HF repo, uploading to Civ is ridiculously slow:
https://huggingface.co/joeygambino/MiniMax-H3-x-Z-Image-native/blob/main/minimax_h3_fl2va_zs05_int8_convrot.safetensors
@joeygambino Thanks for sharing. You can simply upload one model here, and then specify the HF download links for the rest in the description so they can download them from the HF website.
@wyxzddsjj919 I was adding HF links for other models and then I forgot and stopped. But I will update with links. I don't only upload on HF, simply because some new kids seem to be confused by it, so I try to keep them in both places.
Since the Z-Image DiT consists of a set of values mediated by a refiner, extracting the DiT in isolation is largely meaningless. While altering the attention gain—regardless of whether Q-norm is used—would likely change the resulting image, there is no guarantee of improvement. Furthermore, regarding textures, tweaking the higher-level layers has virtually no effect.
Nothing is extracted from the Z-Image DiT, and no Z-Image weights run at inference. Z-Image only donates a statistics profile, measured offline: per-block row norms of the fused-QKV q-slice, mean-normalized into a depth curve. That curve rescales H3's own per-block q_norm weights by small factors (roughly ±10% at the shipped dose). So it's a per-depth attention-temperature adjustment of H3, shaped by where Z-Image concentrates its query energy - much closer to a hyperparameter transfer than a weight transplant. (Also, Z-Image Turbo is single-stage - there's no refiner in its pipeline - and either way none of its weights are involved here.)
On "no guarantee of improvement" - which is why the claim is empirical rather than theoretical. These bakes only shipped after same-seed A/Bs (stock vs graft, identical prompt/seed/settings). The dose is deliberately 0.5 for the same reason.
Where the gain lands matters more than that it lands. Early-block boosts produced checkerboard artifacts on H3, which is why the shipped bake is late-block only. That late-only profile is what produced the visible set-dressing and texture differences in tested pairs.
bro can you make an int8_convront version too?
not pruned
It's there, but Civ likes to name things badly.
https://civitai.com/api/download/models/3256384?fileId=3141163
I am confused by what this is doing. can you show a side by side comparison of native model vs your version?
Added some A/Bs. It's not terribly noticeable for most, but I was getting frustrated with certain missing fine details, especially where foliage is concerned. It's also a much bigger problem on LTX than it is on H3, but I figured while I was building the LTX merges (which you really do notice the difference on), I figured I'd do these as well.
@joeygambino thank you, can you extract a lora from it and release the lora as well?
You may need to update your nodes/ComfyUI. It doesn't support Minimax GGUF's natively. You can also add support manually:
Open:
ComfyUI/custom_nodes/ComfyUI-GGUF/loader.py
Search for the error string or locate the architecture check block (around line 90).
Modify the code block to bypass the error for by adding : [1, 2minimax_h3and arch_str != "minimax_h3"]
# Before
elif arch_str not in IMG_ARCH_LIST and not is_text_model:
# After
elif arch_str not in IMG_ARCH_LIST and not is_text_model and arch_str != "minimax_h3":
You should do that, yes.
I fail to see the improvement in the provided comparisons. They look a bit different, but not necessarily better in terms of detail rendering.
They're better, but you also don't have to use them. We all have free will, different visual processing capabilities, and are allowed to make mistakes. It isn't my job to convince you; I'm not getting paid for this.
Which sampler and scheduler do you use in your workflow? I don't see any advantages to a standard fp8 model—the sound is like it's coming from a toilet, and the image quality is worse?Euler/beta57 for faster renders, res_2s for better motion. There are seven models in this repo though - which did you use, and what are your hardware specs?
@joeygambino I tested the fp8 SafeTensor (Best match, ref2va) with 32 GB and 64 GB of RAM.
Details
Files
minimaxH3XZImageRicherSetsAnd_fl2va_3140100.safetensors
Mirrors
Available On (1 platform)
Same model published on other platforms. May have additional downloads or version variants.
