10Eros_Max beta2 × MiniMax H3 ref2va — INT8 ConvRot
An int8_convrot quantization of TenStrip's 10Eros_Max work on the MiniMax H3 ref2va (reference-to-video-audio) variant, at less than a third of the bf16 size. Drop-in for standard H3 ref2va workflows: same loaders, same native int8 kernels as the official Comfy-Org int8 files, no LoRA node, no custom nodes.
What it is (per TenStrip, the author). This build is Wan VACE layers on H3 out_proj, plus Krea2 turbo weights across the mid-late blocks, on top of the beta2 Wan attention graft — then merged with Silver's turbo and the timetable layer restoration. It is an early experimental turbo merge (a first test build at the time of quantizing), not a finished release.
Intended use. Primarily for video reference and fast multi-video reference — this is a turbo/video-oriented model. For pure image reference, TenStrip's FL model performs better and is the recommended choice. Match the model to the task: this one for video/multi-video reference scenarios, the FL model for image reference.
Quantization fidelity. Quantized with ComfyUI's own TensorWiseINT8Layout convrot quantizer (group-wise Hadamard rotation, groupsize 256, deterministic rounding) — the same code path behind the official int8_convrot releases — using the official ref2va int8 file as a structural template. Calibration against that official file: 100% of elements within ±1 quantization step; all comfy_quant configs and unchanged tensors byte-identical to the official release. Expect ordinary int8-vs-bf16 differences, nothing more.
Requirements. Standard H3 stack: H3-truncated Qwen3-VL text encoder, MiniMax video + audio VAEs, and a ComfyUI recent enough for native comfy_quant int8_convrot loading (if the official H3 int8 files load for you, this loads). Ref2va prompting uses the official six-section Full-Reference format — see the MiniMax-H3 repo's ref-mode prompt guide.
Credits & license. Base model by MiniMax; all 10Eros_Max fine-tune and graft work by TenStrip — all creative credit is theirs. This is a community-made quantization of their released weights, not an official MiniMax or TenStrip release. Use is governed by the MiniMax H3 Community License Agreement (same terms as the source checkpoints — license linked from the base model repo).
Description
This is the official ref2va release by Tenstrip. I simply quant the model into int8.
I'm currently testing this model vs my previous merge and will upload a new improved model if possible. I'm thinking of merging some anatomy lora's and specific action later on as the models seem to really struggle with anatomy.
Link to his HF: https://huggingface.co/TenStrip/10Eros-Max/tree/main
FAQ
Comments (29)
Love this checkpoint, thanks a lot. I compared to your I2V checkpoint and it's pretty fun to use even as I2V++ the ref2v prompts are fastidious, but when they work, they work.
when animating one illustrated image, I used:
1MP, 8 seconds, 12 steps, res_multistep/simple and fl2v turbo 8 step and it was actually my favorite result, movement a tiny bit faster than demanded, but overall stable and fluid, and at 12 steps the sound has zero turbo-related glitches.
@MysteriousString420508 What turbo lora are you using exactly? All i have tested with 8steps and a strenght of 1 were shit. The only one reliable is the 600 ema pruned for me.
things change fast, I saw they released a 4-step ref2va lora 0.1 , so now I'm running that with 8 steps
we're all poking in the dark
Your model descriptions kind of wrong. This is wan vace layers on H3 out_proj and Krea2 turbo weights across the mid-late blocks on top of the beta2 wan attention graft, then merged in with silver's turbo and the timetable layer restoration. It's mainly only meant for video reference and fast multi-video reference. The FL model is actually somewhat better at pure image reference. Also I can post my own models when they're actually tested this was the first test of a turbo merge is only 10 hours old.
Thanks for the correction — I was inferring the layer makeup from the diff rather than the actual lineage, so that's really useful; I'll update the description with what it actually is (Wan VACE on out_proj, Krea2 turbo mid-late, Silver's turbo + timetable restoration) and note it's video-reference focused with FL being better for pure image ref. And fair point that this one was a fresh untested turbo test — I'll hold off on mirroring your experimental builds and check with you first before reposting new ones.
I appreciate the feedback :)))
p.s. currently testing this with comfy kitchen
Amazing model not only for ref2v but for I2V aswell!
Are there any recommendations on which sampler and sheduler are best for this model? I rock with either euler/beta or res_multistep/simple. But there is a secret recommendation from the community aswell. It is the combination of "dpmpp_sde_gpu/beta". The prompt adherence with this combination is just amazing.
Ive used this combo for some time and i also like it best, but generation time is nearly twice as long. I also like er_sde and sgm_uniform, not as good but faster.
You should link the model mentioned here "Match the model to the task: this one for video/multi-video reference scenarios, the FL model for image reference." I had to look around for awhile to understand what exactly this means. It sounds like one of the two models on this page is intended for FL2VA - but clearly not, they are both for Ref2VA. I realized what you are referencing is a model authored/published by tenstrip - I'm assuming it's this one here: https://civitai.red/models/2851079/h3-eros-max
Everything works great. Please develop the project, don't abandon it. And don't break what's already working well.
One question...do we need an additional penis lora for this?
What I do is just use a picture of what ever body part is not generating properly and inject that ref into the prompt. Works well. I tried with loras, but the outcomes were not that good, so I opted for a 90% accuracy with ref image. even better if you create a sheet :D... but ain't nobody got time to create a body part sheet lol
@Stuubzzz I am using penis loras and I dont like it. I will try your suggestion, thanks!
I hope it's possible to crack that anatomy in a fine tune at some point. That would just open up so much.
The whole point of the ref model is showing it images/videos of what you need, bypassing need for loras. Loras will also have little effect on most ref outputs if the motion or concept is tied to a video reference input. The ref model will constrain to the motion in the video at a much larger stricter influence that most motion loras have by themselves and they get overpowered.
@tenstrip Thanks for the info!
Reference images result in basically great looking penises throughout, this applies to FL2VA as well. H3 is very good at maintaining the likeness of whatever it is animating, and I think people are still having a hard time understanding this. They're used to other video models that are unable to achieve this without loras.
That said, it hasn't even been a full 2 weeks yet. People have very short memories and just assume WAN2.2 was perfect from day one and all the LoRAs that we have now that just work were perfect. They forget it took months to get to where it is now. LTX2.3 is still shitty, and LTX2.5 ain't any better it's just faster. And LTX2.5 is HORRIBLE at character identity retention and following prompts.
@DaddyWolfgang I don't have a hard time understanding it. But I don't want to increase my gen time, and I want to be able to simply type what I want. And yeah, I am used to other video models that can do this without loras or start frames. Having access to these concepts in text means you can do things like use wildcards and do all sorts of things that maybe you don't want to do too much setup for, nor, again, use more gen time for.
I'm SO HAPPY to finally be past LTX. This is clearly the next Wan, so I'm looking forward to it handling what Wan could handle. I know it can.
Every checkpoint I've used so far struggles with penis and vagina concepts. Dicks turn to flesh lumps and vaginas turn into weird doughy tacos. I get trying to focus on the movement, but not having anatomy concepts just break the movement anyways because things get deformed and the movement tracks badly.
@tenstrip Tied to video reference? In other word you can prompt H3 to copy the camera motion and/or exactly the same pose/movement/motions or interaction and it can handle it?
@Insistent yes, but you have to prompt it corectly following the official guide. I have a workflow on my page that helps with that. For exampl, instead of using a character lora, you can simply input an image and the model will keep it extremely consistent. or you can ask it to keep the motion from a video and replace the character and artstylee with your audio references. Same with audio reference; keep full track or use voice/ sound as a reference for a new sound.
@tenstrip i consider myself a man of culture and yet i hadn't fully unlocked H3's capabilities until i saw this comment thread. it's absolutely insane how effective feeding it additional anatomy reference photos is
@BBBAAA2 Yes it's very simple too "<picture 2> is an erect male penis." done.
@Insistent You need to look at the developer reference prompting document to understand formatting. When you reference a video you describe the video in the first part of prompt, then you describe what in the video is actually referenced in another block that describes what parts of reference to use and how, like "all motion strictly referenced", loose motion reference, the person in the video, the things in the video. Then in the actual main output prompt you call <video x> when you need that reference and what the reference does. It helps to set an agent or LLM like grok up with the md first then ask them to construct it as you give them all the references and edit and work it from there.
@Stuubzzz Nice, gonna try that. because i was getting some realy nightmare inducing penisses yesterday.
int8 pruned possible ?
it looks to me that the v1.1 is a pruned int8
Hello, is there no corresponding accelerated LoRa for this model? Is it that the 10Eros model cannot be directly used with the LightX2V 4-step or 8-step LoRa? Can it only run the full 20 steps?
its minimax h3 model. use minimax ref2va lora
