beta 5:
Use TURBO-hybrid_int8 by default. Other options are for experimenting. Full versions here.
beta_5 is built with a different normalization technique. It's also built off 7 different concept-grouped grafts and consensus merges from over 20 Loras, no full direct lora merges. It allows for a non-turbo version which is available. The files with TURBO have a hybrid turbo-delta fusion baked in saving 4.2 gigabytes of memory instead of having to load both ref/fl turbos. Other concept Loras will also load on top more readily, usually needing lower strength 0.2-0.6. For t2v and even some i2v you should 100% be loading mystic_v4, anatomy enhancer, or a related concept lora on top.
Also consider using the non-turbo with PDD, 6 warmup and 4 PDD turbo steps.
Prompt is the entire key to quality and success. Any issues you'd want to blame on a model or workflow, you can go ahead and take a look at the prompt instead. Describe things cleanly and literally:
❌ He puts his penis in her pussy, make hot sex
✅ Live-action pornographic explicit sexual and sensual intimate POV recording: The man slowly moves his lower body forward with his finger on the base of his penis shaft. The tip of the penis slowly disappears into her pussy hole and her wet labia open allowing it entry. He keeps swinging his pelvis forward until his crotch touches hers. Then he backstrokes and starts a repeated thrusting sex motion. The sex act makes her whole body recoil into the couch, bouncing her breasts wildly. She stares lustfully into the camera.
Prompt in temporal sequence on a linear time flow. The longer and more embellished the prompt, the better the output will be usually. 100% use some kind of LLM enhancer; Grok is the one I use since he has a skill preset for reference prompting.
The version just labeled 'hybrid' is non-turbo. Non-turbo full step audio will always be better than audio on turbo versions.
Some working sampling setups:
er_sde/beta57 4-6 steps (seems to be best preservation of style and reduced drift)
res_multistep/simple 6-9 steps (good motion quality, beta will be the sharpest fast motion at 8-9 steps but will give a plastic/burned look)
LCM/simple 6-8 steps (best turbo audio)
Euler/simple 4-8 steps
Next version should be mainly powered by the first round of Sulphur training. From what I've seen it will fix most missing audio and missing concept issues.
Updated - Full credits to these lora makers for having deltas involved in a consensus merge in beta_5 version:
alcaitiff, MisticRain69, diogod, FourBunny, HearmemanAI, tazmannner, simonishere, QualityControl, blo01, ComfyTinker, kermitfrog1202, qdr1en, coachbate, salttaro, alternative_penguin, misterxrex, freek22, definitelynotadog
beta 4:
rebuilt on beta3 config with some loras changed out for newer versions. Turbo was consensus merged to construct a ref/t2va hybrid turbo lora. That's the key element to the merge, there is no non-turbo version. The non-turbo version is bad and doesn't work since the turbo weight is normalizing the merge by it's strength. That custom turbo merge will need further improvement. This one preforms video, reference, and motion well in 6-8 steps without the issues from beta3, but audio needs shift configuration.
Use sampling like Euler/simple 8 steps with sampling shift - 12 video/ 7+ audio. LCM/simple or beta with 6-8 steps and no shift can also be better for audio and drawn styles.
Audio is lackluster and it's becoming somewhat apparent that H3's integrated audio is not good, like terrible actually. Any future multimodal models should avoid integrated audio if they intend to open source. I have almost no control over how the audio works inside the model. Don't post about it. I focused on motion and prompt response and of course when I get those working well, the audio ends up bad, go figure. I'll look at what kind of different turbo configs can enable more audio crispness to come back or likely will have to wait for Sulphur to replace MysticXXX which is contributing to the audio quality drop.
beta 3:
Rebuilt on the delta1024 reference hybrid model. Does t2va and reference. Treat i2v as a single image reference, don't do i2v prompting. Silver's merged turbo is integrated and tuned for 6 steps with simple schedule with the extra _emb layers tacked on to the model, not sure if they're needed.
Samplers like er_sde or multires or other turbo sampling setups work. Best imo is just er_sde/simple 6 steps, no shift, no spectrum, no cache, only a comfy_kitchen backend selection node. 3 video references in a 15+ second outputs can be done in under 10-15 minutes now on larger cards with no extra cache or quality hit needed.
Many (like a lot, all the good ones) on-site loras were fully combined into a consensus weighted merge with ranked drop-out to form an initial part. That merge is put against new, more powerful Wan and LTX grafts as a blend/reshape that uses the loras to consensus shape the grafts, but it also allowed some of the better loras clean pass-through. This is not a linear list of loras just merged. The main element that can present most is probably MysticXXX which was given the most pass-through weight since it's just good--and all 3 release steps of that are inside it. However, they're all-combined with a ton of other loras with agreement and consensus of shape and then it's only reshaping the grafts parts. The results of that are the actual weighted loras that are used to make the model.
Due to drop-out and consensus merge, pretty much all of the loras can all still be used easily on top if needed, and might work better even. All of this was only done to create a large rank dummy lora similar to what sulphur data will look like as a lora or extracted lora so I can start looking at how to apply it cleanly.
It's definitely not a few on-site loras that are linear merged, uncredited, and then renamed with some emojis. I'll only do this until sulphur tuning steps are in my hands and I can work with more targeted and shifted stuff, plus I was tired of waiting and I wanted fast easy i2v.
There is one quirk of the hybrid h3 usage: don't use it i2v. It should either be always used in reference prompting mode, or t2va prompting mode. Even if there is just one single image input it needs to be used as reference and prompted in the ref2va format. If you run an underdeveloped or manually written prompt you will get odd outputs, random camera changes, and blue lighting color shifts when you use the i2v prompt style.
Full credits to these lora makers for being involved somewhat in beta3 version:
alcaitiff, MisticRain69, diogod, FourBunny, HearmemanAI, tazmannner, simonishere, QualityControl, blo01, ComfyTinker, kermitfrog1202
Beta2 and previous:
This started a finetune-by-graft. Or maybe a GST - grafted shift of transformer (cross-architecture). I made both up, because there aren't any projects that have done it that I know, except one reddit post that made me look into it. I experimented with Wan and LTX on the side which led to the initial LTX Eros scripts that became what powered this, all before H3 ever came out. It seems like unified unbiased models like MMH3 can technically take attention influence from any other DiT without breaking if done correctly. Anima, Krea2, LTX, Wan2.2, Flux1 were all tried out, configs tested, about ~40 hours maybe of working in the dark without any paper or technical documents from Minimax. Eventually I developed linear-magnitude blend application and specific block and head gate targets allowing for a smoother graft on an attn-triplet-unfused version of H3 output as a patch file. That sent to lora extraction, then merged to checkpoint at taste. This is a merge but a merge of LoRas I extracted that interact to produce this current shift. I saved 5 ponds of water by recycling data in a few minutes on a single card instead of toasting a server up.
Turbo not recommended yet for i2v, especially when used with other LoRas. T2V use with turbo is better. Use 20-25 steps normal sampling with no dialogue, 25 steps with dialogue along with cache nodes and attn modes. More steps over 25 are not neccessarily better, and can be worse. Use full int8: int8 model, int8 VAE (if it doesn't crash comfy), int8 qwen3vl along with current cache or attn mode nodes. For smaller cards: quants, macOS ports, and Wan2gp support will likely appear on huggingface but not from me.
Known quirks:
Audio difference v.s. Base - This model's audio changes come from attention shifts seeking alternate audio pairing. Attn triplets were unfused before graft, both standard and triplet q_attn was grafted holding about maybe 10-15% audio influence, attn_k was frozen and MLP fc2 layers were untouched resulting in minimal audio interference. This was the main issue with the entire transformer graft and protecting audio. However this version is slightly louder overall than the base model.
Low resolution detail smearing - Some finger digits and fine motion will smear more at low resolution, also a problem in base model. As memory use gets more efficient increase resolution or work on the composition to get around it.
Odd outputs - This can attempt certain concepts more liberally than base model, but that can lead to some undesirable outputs in bad prompting and certain contexts. Data shift comes from completely different transformers and architecture. This shouldn't even work, so it is what it is.
This model is not dedicated to NSFW as that would violate community license agreement. Sure it can do it, just like base. Any NSFW generations are purely the result of advanced reasoning and tokenization resulting from experimental changes. All terms from the H3 community license also still apply to the users of this version. Don't be a dumbass.
H3 usage still requires very intense prompting for maximum effect. Every motion, every interaction, every sound plainly and fully described. Not with slang terms; with proper actionable words that can be tokenized. Refer to the h3 developer prompting guide, hand that .md file to an LLM or Chat agent and have them enhance or refine prompts along the released H3 developer prompt guide styles using the model's tag system. Certain concepts can be made from pure token reasoning. Consult the prompts in my previews to see certain physical descriptions that I use for some things. When using enhancement give the agent feedback about any issues in the generation and get them to describe motions in alternate fashion, or manually edit it yourself adding a negative like "no X, no Y". Still requires prompt refinement and trial/error for best outcomes.
Sulphur Project has 10k banked to attempt actual tuning. Right now training pipelines are sub-optimal. As always Eros is my personal side project, and this beta was also essentially a speed-run of finetuning, figuring out exactly in what configurations and target areas do you get helpful/harmful changes in the model. This is also a proof-of-concept of what and where to target while leaving the reinforcement quality of base unharmed by being additive.
Description
normalized, base model shape regression, and maybe an actual working hybrid turbo for now.
FAQ
Comments (169)
Does beta5 uses Shift? Is not mentioned in the description
Default 12/3, but like I said higher audio shift up to 6-10 on low steps on some samplers isn't gonna hurt.
*Sigh*, here we go again
What are you sighing about? Just use the default MiniMax H3 if you don't appreciate people trying to improve it. You aren't being inconvenienced at ll so no need to sigh about anything.
That's literally how I feel every version too. A handful will always find one niche A/B v.s. something else and not do any prompt or seed rolls and post about something on the page like it's a version-wide issue that isn't fixed with one prompt change.
Guess you people consumed to much porn and not enough memes
why is the generation time even on turbo 2 times longer?
If youre vram is full thats maybe the issue on youre end.
Model must fit insode the vram otherwise rendering time increases maybe 1.5-3x times.
With rtx 5090 i have to use int8 version and not the bf16 version otherwise rendering takes longer but it still works.
First of all, there's no way in hell you're claiming the model is twice as slow when it's 90% faster than base 20 steps without turbo. If you download the 'fp8' which is actually a w4a8 you need to check this out too for the correct use and optimization files: https://huggingface.co/LokkenJP/10Eros_Max_optimized_w4a8_exp_learned/tree/main
@tenstrip I don't have an rtx 6000 pro, the eros 4 turbo at the same 7 steps and 0.74 mp was generated in 120 seconds, the eros 5 turbo geenrit in almost 200 seconds, there are no problems with vram, I rather want to understand why this is so and what can be done about it
@baamad9400 Many things.but if u are using sla attention it somehow turned off sometimes(check terminal log).
Some setups compile on first try then render faster.
The issue is definitely on youre end.
beta 5 Turbo int8 is pretty good. definitely an improvement from beta 4. The audio is improving as well. Thanks for your efforts.
If you have audio issues they're most likely caused by the flow or other loras. You might need to add a lora or two at low strengths. For example, the blowjob V3 lora at around 0.25 - 0.35 str completely eliminates the crunching sounds. And as tenstrip said in the description, you'll want to use soemthing like MysticXXX V4 at low strength to improve things even more.
Just don't use high lora weights or you will get crazy stuff happening. Start low, like real low, and work up until you find a setting that works. I always, always start at 0.10 with H3 for every lora. For Turbo loras I start at 0.5 and go up and down from there until the quality is good and nothing looks too wonky. I don't care what the creator of the loras say, starting at 1.0 is rarely a good idea, only one Turbo lora has ever worked at or over 1.0 and that is Plaguekind Parasyte H3 Turbo lora.
The real way to fix the audio is to run the non-turbo full steps. Audio and the actual explicit sounds can come out correctly that way.
@DaddyWolfgang oh I'm talking about celebrity voices which I don't use as often. Even then, it's a step in the right direction. The audio in general is fine. I do use MysticXXX around .4 for adult stuff. If I'm making a dwarf tossing scene with Tyrion Lanister I may need to use the base MiniMax model but otherwise beta 5 gets the job done.
@iodrg244 This should still be 1:1 with base I've done House MD t2v gens and alot of normal reference outputs. The advantage of hybrid is you also have ref t2v where you can make whole t2v scenes with just referenced characters and it's more loose and willing than the strict reference model.
@tenstrip I grabbed the non Turbo and used 25 steps and (Tyrion's voice in this instance) does sound good. Good reminder for me to take the time for full/steps for certain videos. Thanks.
Thank you for great job! is it possible to train lora on your model?
There is a bf16 non-turbo version on the Huggingface that can technically be trained to output a hybrid fl/ref lora. But it's not base and a specific merge that will probably be unique since I'm not gonna build off of it going forward, it's a hybrid, and it's also pruned so not sure if the outcome will be usable with anything but this version.
Finally, Beta 5 is here. I’ve been using Beta 4 for a while—it’s the best version I’ve used so far, and I haven’t encountered any complex or annoying issues with it. That said, I’m going to try Beta 5 and see what’s changed. thnx a lot bro
Beta4 has been amazing (40GB) and I can't wait to try Beta5 (40GB)!
All the models I've tried the Ero versions have been the most consistent across the board.
And I can attest to the prompt being super duper important. H3 understands direction well, but not nuance, "Make sex" is gonna just make it render an approximation of it. Describe things more literally and it'll absolutely be able to render more and add in details that accentuate the prompt. But it needs more of a launch pad. This is good.
Use Qwen-VL prompter and you can load up as many reference images and video as you want (You need to use a multi image (Image Batch Load by KJNodes is perfect) loading node, select the amount of images you want to use, say 3, and then attach its output to the "image 2" input on the Qwen VL node. Then, change the "frame count" number to the number of images you have attached to image_2.
Then use the Qwen 3.5 4b Heretic V2 model and then select what type of prompt you want from the drop down menu.
You may need to guide it a bit further, so in the prompt area you could write:
The man <Subject 1> is in <Picture 1>.
The entire video will take place in <Picture 2> and only <Subject 1> will be transferred to <Picture 2>.
The man is dancing on top of a giant balloon. He is jovial and has a kind face; there are people below cheering him on. The camera pans around him as he dances and zooms out slowly showing the audience below.
Click run, and it'll generate a much better prompt for you than that. YOu might need to edit it, but after awhile you'll find a combo that works for you. Save those instructions before editing them so you can reuse them again.
For anime I made these rough extractions to be able to stack Singularity on top of this. They don't need a very high strength, something like 0.2-0.5. https://huggingface.co/TenStrip/Minimax-h3_Singularity-Lora
Testing beta 5. So this model isn't designed for any speed patches just the built in turbo lora?
Nothing but comfy_kitchen backend, or maybe SLA. The other speed ups aren't worth using with their warmup being half the gen anyways and they reduce a lot of quality.
I thought I had damaged my Comfy installation, because the new Eros looked awful. What happened was that I downloaded the non-turbo version and tried to adjust the settings from the turbo version. Then, looking at the links again, I saw that there were two different versions, but they had the same name when downloaded. A wonderful model.
Yeah I didn't want to write NONTURBO in caps, all the previous versions were named like that.
在原本就领先的finetune上又做出了卓越的工作
It's not really a finetune since it's not bound to a dataset. It's a revolving merge mix and config that I refine more and more to reach the function that I want from it.
Phenomenal work on Beta 5. Sound quality is VASTLY improved over Beta 4 and all-around quality and prompt adherence is fantastic. Good call on using the Mystic v4 LoRA too, definitely helps. As does the VBVR LoRA, IMO
You can run any lora stack or loras even the ones used. All the loras were combined with each other on a consensus drop out before grouped by concept, anatomy, style, gender. It cut out most of their individual strength so they can be doubled up on just fine.
@tenstrip Eros 5 is 🔥 Thank you 🙏🏻 for Eros 6 can we get more intimacy and passion?
There is a big powerful lora/model coming soon that's a sulphur h3 preview. It's only 90k clips but it should fill a ton of t2v, motion, and audio gaps. That may help make a fully generalized 1.0 version.
@tenstrip - Thanks for the update, do you know when that is going to launch?
When do we have an LTX2.5 Eros?
Sulphur2 and Eros loras all work on 2.5. Any forward merging of them just makes a worse version of the 2.3 models.
Now that minimax has clarified their license;
What was the minimax license clarification?
when you generate 18+ sex content you lose 2000 of social score
I could've been arrested for using this model in the US and making it do this, but I have the pass now.
Thank you so much for your hard work! I'm seeing noticeable improvements over Beta 4, especially with prompt adherence in my tests. I also found the w4a8 version pretty good for anime style, showing better understating of anime pussies and anuses.
As far as checkpoints go this is easily the most capable for uh, stuff. It sticks it in deep enough without going like, really overboard like Motion Booster does.
It does degrade reference capability and ups their cup sizes by about two letters, but there's actually a solution for that - go back to the original for the last step or two, ideally the upscaled pass. Then you'll still have the motion and it can slap the references (mostly) correctly back onto it.
It also goes overboard removing clothes and strips them off in even more of a hurry than the base model does, often just disappearing rather than ripping them off. I have an unsuccessful LoRA for that, may try a higher rank.
Also struggles with actually adhering to a first frame, but that is a general problem with the model atm.
I use Kijai rank 31 resize turbo, 12 steps 0.15 MP, 1 step 2x LBM 123 AI latent upscale, er_sde/sgm_uniform with Spectrum.
How do I re-apply a reference image during the upscaling step? In LTX 2.3, I knew it was possible to use shader nodes during upscaling to match and correct the character and tone; I’m not sure if this method works in H3, and I assume it wouldn't be feasible with multiple reference images—it likely only works with a single image. Do you have a workflow I could study? Thanks.
Have you tried ever fully elaborating on the undressing motions in detail. Also the pace of things is generally set by how many tokens need to be realized in the gen. If its a short prompt and long gen it will be relaxed and slow usually. Jamming like 10 motions into 5 seconds will cause it to rush especially if they're not accurately described in detail.
@Silicon_Mirage In R2V you don't need to. You can use the same conditioning for both passes.
For I2V, you will need to downscale the latent and the text encoding for the first pass. h3-latent-upscaler Miinmax H3 Conditioning Upscale, Element_easy Minimax_H3-Latent Upscaler. Then you'll only need to run the conditioning one time. Run the conditioning node at the 2nd pass resolution.
beta 5 so far is so damn good! keep up the great work!!!
Great model, thank you. But could somebody explain to me how can I make the character stop talking ? Everybody without a cock into the mouth is speaking constantly even if I described the sound et explicitly says that I don't want any dialogue...
Are you using the proper prompt layout with overall_soundscape? "no dialogue" doesn't work that usually means it's looking for dialogue. I haven't had that issue in a long time, but even putting "she moans" without putting it into context can cause random speech. The first sentence in detailed description also frames the entire generation: "live-action erotic and sensual," in those scenarios you usually just do "breathless exerted breathing" and that alone keeps anyone from speaking.
beta5 is amazing work! rooting for you!!! do you have plans for nvfp4?
dude this is amazing! blows pinkcherry out of the water!
Not really a fair comparison. That one is actually training on a dataset. Instead I leveraged what the model was missing from loras to just gap fill until it worked like I want it to. Although the approach to training like cherry is doing is basically slow destructive crawl that will eventually unwind the base model's reinforcement more and more and make it worse and focused entirely on nsfw, creating a lot of other issues. But it would be way better at nsfw by then but not as versatile. This one is normalized to be additive and not tweak out too much since I have way more control than just training.
Hi creator, I've been using your 10Eros fine-tuned model and I think it works very well. I have a question about the latest beta 5 version. Previously, I only used Turbo models, which are models that incorporate accelerated LoRa. This time, I want to try a non-Turbo version of the beta 5 model. You mentioned that setting it between 6-8 steps is suitable. I can use Kitchen Attention for acceleration. Besides that, if I want to use accelerated LoRa, should I use the fl2v version or the ref2v version?
Can the 8-step LoRa in the lightx2v repository at this URL (https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main) be directly used for 10Eros?
Piggybacking on this comment: Is there a specific Turbo LoRA that you are merging to make the "TURBO" version? Or is it a more complex things such as a combination of various Turbos?
Personally, I'm downloading the HYBRID model but would like the flexibility to test out Turbo results via applying the Turbo LoRA directly.
It's this 4.2gig combined hybrid lora that I used, but in the merge it is layer-scaled across the model starting low and ending higher strength. At inference the match would be 0.7 strength or around there. https://huggingface.co/TenStrip/MinimaxH3-Turbo_Shenanigans/tree/main
The real point of the non-turbo model is to run it with PDD or Fast accelerations, or mainly with full steps. If you want turbo I used the best turbo configuration I found through testing already inside the TURBO labeled ones.
I think beta4 is the best so far despite the camera moving issues. - it produces very lively gens even for SFW. Thank you again and again for your amazing work.
a version of beta5 that will fit into my VRAM? awesome, gotta test that out latet tonight.
In beta5 some sort of issue with vaginal sex. Even when you spell out in detail that it has to be vaginal sex only—and that anal sex is forbidden—the penis suddenly shifts from the vagina to the anus (even with the "correct" prompt from the beta5 example). This happened occasionally in beta4 too, but if you explicitly specified "vaginal sex" in the prompt, the problem would be resolved. Is the anal LoRa too intense in beta5?
What "anal lora"? I havent found any such (hetero) loras. I've sometimes got the opposite problem starting as anal sex and then morphing into vaginal. Prompting and reference images don't help as well. an anal lora would help.
@QuietSpark I meant that this model leans towards anal sex. And in beta5, even a well-crafted, detailed prompt doesn't help (unlike in beta4). I used the term "LoRA" for the sake of general understanding. But overall, the beta5 is certainly of higher quality than the previous model, especially in terms of audio.
I'd wonder if it's i2v if the input image just leans more towards an anal shot. The model will take shortcuts to whatever sex type it sees. Also it's "pussy" and not "vagina", you can embellish vaginal with a paragraph of prompting to force it, mention "labia opening wrapping around the shaft" "Wet pussy sliding up and down the shaft" "missionary sex" "cowgirl riding sex" etc
@QuietSpark The homosexual loras are all present as well which is why anal is even in the model. As a last resort just add a random reference of a close up of whatever type of sex you want and reference that. Reference will work better than i2v or t2v even with just one image.
@d8167895 What’s wrong with anal sex? Especially if the woman herself doesn't mind... ;)
@marsele117268 Nothing wrong. But the video becomes glitchy when the penis teleports from the pussy to the anus.
@tenstrip I checked - this happens in reference (ref2va) too - not only i2v.
@d8167895 just use this lora or mysticXXX, the model is base not one-all can do everything and it's still a beta version. https://civitai.red/models/2923850/anal-doggystyle-by-funphantom?modelVersionId=3308371
@tenstrip OK. Thanks. What strength should I use for these LoRAs? You specify a range of 0.2 to 0.6 in the description, right?
@d8167895 For Loras that are brand new and not in the model you can go 0.5-1.0, or what the maker specifies, or whatever strength starts making it work.
sorry if dumb question, but this is a unet model right? for a workflow i found, is this suppose to be used in the unet_fl2va, unet_ref2va or both?
beta5 is awesome ,dude.
CAN you make it Like GROK please i know model has potential
Right now in a lot of cases it's better than Grok private on Venice was, just more involved to prompt. Especially if you do reference and some loras.
Which of the two fp8 20GB model to download for RTX4060 8GB VRAM, 40GBB DDR5 RAM? Is Turbo one better than the other one? I usually find w4a8 models slowers. Using Minimax H3 extender workflow without any loras.
this model always produces big silicone tits other than that its good model. but please remove the big tits lora or whatever you merged to always get big silicone tits
There's no specific one. They're all consensus merged. The loras all democratically choose big breasts, but that's if you're very simply prompting and letting the model choose. I've stated many times the model is almost entirely prompt driven, you want something keep writing for it, you can prompt "tiny breast, petite skinny female" etc. If that doesn't work there are loras for it as well.
Much appreciated that you didn't paywall it! Also very dumb question, is this fl2va or ref2va? And can someone plz share their working workflow for this model, I couldn't find any in the gallery.
I wouldn't paywall something built off community contributions that relies on using credited loras. I think it's biggest strength is as a ref model and H3 is best with reference, but it can do all modes.
@tenstrip oh thats cool that it can do everything! You have the thanks of the community sir 🫡
I really like Beta 4. It’s my main model now, and it’s amazing. But I’ll also try Beta 5. Thank you for all your hard work—that’s all I can say.
Just want to give a huge thank you for making and posting these models. I do not belive you hear that enough
Could you tell me what I need to make it work apart from the file? I need the Clip, Vaes and that's it? Or do I need something else? Do you have a workflow? Why is the fp8 version so economical in terms of its weight? The fp8 version is First / Last frame?
The 'fp8' is mislabeled, its w4a8 and there's no tag for it on civit. That is just an option only for 12-16g cards, the int8 is gonna be the most standard one to use. They all drop straight into the normal minimax H3 templates and need qwen VL and VAEs. But the Turbo models sample in 4-8 steps instead of 25.
@tenstrip Thank you😌
I'm going to test it, I've got the components, it looks good.
A doubt, w4a8 Is it Fl2vid or ref2vid?
God bless you, you do an excellent job
@chamo9009 They are all hybrids that can do all modes.
@tenstrip Wow, that's even cooler
Hey dude. I just want to say, thank you very much for this model. The prompt adherence is top notch. I do not know what sorcery you did. haha!
Hello, in your experience, do we need to prompt it for I2VA using the whole REF2VA heavy structure or does it understand simple enriched linear prompting with a starting frame ?
You use ref as i2v. Just prompt ref which gives you full subject level controls and much better control over output. In ref prompt just lock the input image as the start frame, mid frame, or end frame.
using fp8, int8. keeps crashing my gpu. But also thank you for having a bf16 version in your huggingface page. prev was using beta4 bf16 and had no issues.
need to wait for comfyui to patch this or something.
ComfyUi error TLDR from chatgpt:
W4A8 quantization/dequantization in comfy-kitchen triggers a CUDA “unknown error” during the INT4 → INT8 weight conversion.
Yeah use int8 turbo then, I can't make civitai recommend the right version idk by it chooses the w4a8 as the recommended by default.
@tenstrip even the int8 also crashes... lol also great thanks to you that you have the b16 version. working without issues.
not the model issues, just comfyui problem.
The model is excellent, but there is one issue: the character finishes its own lines and then goes on to say the prompt that follows. I’ve tried the official model with the same prompt, and this doesn’t happen. However, your model is truly powerful and excellent.
<d>[voice description] speech. </d> you might've missed the closing tag or you're not using them. In overall_soundscape: her/his one line of speech, or multiple lines of dialogue.
@tenstrip Thank you for your reply; I’ll give it another go.
@784559563390 如何?
@aa1803989464704 I haven’t tried it yet.
For me, beta2 still has the best visual quality and consistency compared to later hybrid version. The lora compatibility of beta2 is also better than the latter version. With the help of realism lora, you can turn beta2 into stunning image edit model. The beta5 have better prompt adherence, but the visual quality is still no way near what beta 2 offers.
also added the gguf version
https://civitai.red/models/2902153/h3-eros-max-gguf?modelVersionId=3308580
Hi, will test it tonight!
Do we need an extra penis lora?
No, but the underneath angles might not look right all the time unless you convince it it's the underside shape of it.
@tenstrip Thanks!
i was curious for some time, what is "consensus merging"? your v5 is the best minimax and i'm using every version so far, compared to v4 there's no longer a quality drop with anime eyes from afar, really thanks >__<
modified TIES merging based on shared signal. One lora says this is a penis, the other lora says no this is a penis, the merge forces them to agree and combine on all that shared stuff.
Does V5 need audio/video shift setting?
i use 12 video and 3 audio shift. works fine with turbo version
@hoummel745 That's default, so you could remove the node and get the same. That's how I test and make the model without shift. But increasing audio shift to double or triple like 9 can improve audio in some low step situations.
@tenstrip i didn't know that, thanks ^^
v5 is a clear improvement in many ways for me, but the main issue with turbo-enabled stuff and H3 is the over amplified thumping and knocking sounds in NSFW scenes.
The non-turbo version doesn't have this issue but gives me fingers from horror-hell and mangled body parts at random with the same prompts.
These knocking sounds really are a killer for NSFW-H3 atm for me. They kill the immersion of a scene.
Didn't find a reliable way to control this with prompting yet.
Replacing any objects that can move (wooden table with a sturdy marble table) can help slightly.
It seems to me, that the model normalizes certain sfx to ~0dB that normally are very low volume. The model can control these levels properly as the non-turbo variants show, but the turbo lora messes up the sound level balancing of the different audio elements somehow.
Exactly!
That and the chomping sounds for blowjobs means that LTX 2.3/2.5 is what I'll be using.
I hope someone comes up with a fix, but I simply can't use it as it is. Other people don't seem to notice, or care.
I swear up and down, in the base FL2V INT8 Convrot, with one of the earlier Turbo LoRAs (It's either Larry EMA or Lightx2v's), specific noises like the slaps were retained. Lightx2v's recent LoRA 4step 1.2 allegedly helps with sound, though I haven't tried it yet.
It seems that beta 5 still sometimes suffers from color drift. Not as often as beta 3, but still, some generations tend to turn blue-ish.
I only saw it once on one seed in i2v node. It was fixed by prompting "warm orange lighting" in the prompt which had no mention of lighting.
beta5 make euler/simple cant work. It brings many noise in vedio
Yeah it's er_sde/beta57(res4lyf) for 4 steps to denoise most of it, more steps for audio though. I like LCM/simple 6-8 steps for better audio, multires/beta 8-9 steps for dark, cinematic, and high motion.
not only euler/simple, maybe something wrong in turbo int8 version. Anybody know which version is as stable as v4, I like v5 how to describe enveronment, but every move is blur and break stability.
@tenstrip I will try again. And I want to make sure turbo int8 is best choice for now.
@tenstrip 请问Multires在哪儿啊,我在sampler里面没有找到啊,是不是要安装什么新的节点?
@353865690492 I always keep mistyping the nickname I have for it. It's called res_multistep and default comfyui selected one in the H3 template.
@tenstrip thx
Try euler / beta, 8 steps
It's a great model, but for some reason it's in the doggy pose, no matter how hard I try to be a woman static, that is, not to move my hips towards the man, it's impossible to do this, in general, in the doggy pose, the model is most unstable, otherwise everything is fine, but as if some variation is missing, after 100 generations, no matter how you change the promptness, plus or minus you get the same results, am I doing something wrong or is this the maximum of the model so far? well, there is also one gap, in some poses in i2v, in the picture the penis is fixed in one position, but in the video for some reason it changes position quite strongly, thereby deforming the labia, I do not criticize if anything, I just give feedback, I hope it helps, I would like to get the most dynamic results from the model, lively and diverse results
alot of NSFW models in general are bad with having just the man thrust ive noticed this with LTX as well where if you want the man thrusting for doggy or even missionary it takes several trys because it will by default make the woman be thrusting into the man. I think its a training data issue/caption issue in the datasets since even the best LLM's require alot if tweaking to make good NSFW captions.
1 additional lora at 0.5 would've saved you "100 generations". Don't use i2v, use reference. Use additional reference images to control anatomy look and behavior. Use a better start image or make image adjustments so that composition is closer up and able to be more detailed. Fully change the prompt with an enhancer because your wording doesn't seem to be correct. Don't manually prompt from scratch. Increase the resolution and decrease length.
@tenstrip Low megapixels reduces prompt adherence then? I've had back and fourth as to wether that's the case or the reverse.
@tenstrip Thanks for the response. I just want to make sure I understood your recommendations correctly:
Which specific additional LoRA do you mean at 0.5 strength? Are you referring to one of the LoRAs mentioned in the model description, such as Mystic V4, Anatomy Enhancer, or another specific concept LoRA?
When you say “don’t use I2V, use reference,” do you mean that I should use the REF2VA workflow and provide my starting image as a reference rather than as the actual first frame? Can REF2VA work properly with only one reference image, or do you recommend using two or more images? If multiple images are recommended, what should each one control—for example, identity, pose, anatomy, or composition?
My prompts are already written and refined by Grok specifically for NSFW MiniMax H3 prompting. Is that still not enough? Are you recommending that I use the LLM Prompt Enhancer built into your workflow and generate the complete six-section full-reference prompt format instead? If so, what exactly should be changed in the prompt structure or wording?
@baamad9400 Mystic v3 might be better, but v4 wasn't merged in this time so it actually adds something. The Lora should just be on the concept like normal. For anal there are loras, poses like doggystyle there are doggystyle, cowgirl, standing pose loras. If there's a lora for the concept you probably want to use it.
REF2VA works with up to 9 images. If you see something like the female backside doesn't look right, you find an image that looks like how you want it and add that as a reference un-cropped. partially_preserved as the anatomy of <Subject X>'s rear ass/nudity. I use Grok but on a skill where he is doing the REF2VA prompting, directed to create subjects out of every notable thing in the scene beyond just standard. <Subject> is the main reason reference goes totally beyond i2v because you create hard tags for the actual things inside the images that the model will strictly follow and you can control if it loosely or strictly references them.
@DarkEngine2024 It's just about how much detail something has. If it's too small it won't look right. Like the blurry faces at distance.
@tenstrip Understood, thanks.
ah man, can anyone make one that fits for 12gb vram? need space for loras too
Sorry for the newcomer question but, what's PDD in "Also consider using the non-turbo with PDD, 6 warmup and 4 PDD turbo steps."?
Thank you 🪵🐸🪵
It's a whole separate turbo pipeline that needs all this to set up. https://huggingface.co/aptech0081/MiniMax-H3-Acc-LoRAs-ComfyUI
@tenstrip 🪵🐸🪵 Thank you, will definitely check that ❤️
Tested beta5 in I2V/REF2VA workflows - Turbo merge vs non-Turbo.
The turbo merge is giving me significantly better results [better adherence/quality/reference] than using the regular version + my own turbo lora. So yeah I highly recommend using the Turbo merge.
Generally speaking while the model has better understanding of NSFW concepts and NSFW prompt adherence it seems to be biased towards female anatomy. In some of my generations it would add female parts to males [especially stockier guys] even if the prompt explicitly states males as the subject. This was especially a problem on the non-Turbo version. Massive breasts are one of the issues some other people already noted.
Overall a great model and direct upgrade over native H3 for NSFW but keep in mind the female bias.
Hey, first of all, great models! They beat any other H3 tune I've tested so far, even for SFW purposes.
For my specific use case, I've noticed that beta2 has the best visual look for I2V/FL2V, whereas beta4 and beta5 have the best motion. The issue I have is beta4 and 5 also alter and bias the look a lot more than beta2.
From what I've looked up, I assume there's no way to actually merge beta4's 0-19 blocks and beta2's 20-49 blocks to get a combination of visuals + movement from both?
What should I do to ensure that the cucumber is inserted into the anus instead of the vagina in the final close-up
subject_definitions:
<Subject 1> The man: An amateur male protagonist, seen from a POV perspective.
<Subject 2> The female corpse: A woman with large breasts and large hips, based on <Picture 1>.She's wearing a semi‑transparent skirt,She is in a sculpture-like state, completely still, no breathing, and no autonomous movement. Her arms hang loosely, and her body is slightly collapsed.
<Subject 3> The wardrobe: A wooden bedroom wardrobe.
<Subject 4> The breasts: Full and voluminous, revealed after the bra is pulled off.
<Subject 5> The underwear: A pair of panties that reveals the pubic hair and vagina.
<Picture 1> The reference image of the beautiful woman with large breasts and hips.
summary: [reference generation]
retention_analysis:
<Subject 2> (出现在 [Shot 1]): fully_preserved - The female corpse is the central focus, maintaining its sculpture-like state throughout.
<Picture 1> (出现在 [Shot 1]): fully_preserved - Defines the appearance of the female corpse.
<Subject 4> (出现在 [Shot 1]): attribute_transfer - The breasts are revealed and massaged.
<Subject 5> (出现在 [Shot 1]): attribute_transfer - The underwear is pulled down to reveal the vagina.
The target video is in a raw, amateur smartphone cinematography style, characterized by low image quality, significant digital noise, heavy grain, and a soft, slightly blurred focus to create a realistic snapshot texture.
[Shot 1]
POV perspective. The scene begins as a man enters a bedroom and reaches out to open a wardrobe (<Subject 3>). Inside, <Subject 2> is curled up with bent legs, her Her eyes rolled back, showing only the whites, red bloodshot eyes,and remained completely motionless,in a sculpture-like, motionless state.The man reached out, stroked the female corpse's head first, and then pinched her cheek,then The man reaches in and grabs her bra, pulling it down to reveal her full breasts (<Subject 4>). He then uses both hands to firmly massage the breasts for a few seconds. Next, he pulls down her underwear (<Subject 5>), exposing the pubic hair and vagina. He inserts his fingers into the vagina, moving them back and forth in a rhythmic motion. The man then grabs one of her feet, lifting it up to bring the sole of the foot into a close-up shot directly in front of the lens. Finally, he lifts the heavy, limp corpse. The corpse remains passive, with arms hanging loosely and the body slightly collapsed. The man then flips the corpse over so her hips face upward. He grabs her waist with both hands and lifts her hips up into a kneeling position, with the camera moving into a close-up of her buttocks.exposing the anus and vagina ,Her anus is clearly visible and distinct from the vagina,Then he took a cucumber and put it into her anus ,not the vagina
overall_soundscape: The muffled sound of footsteps on a bedroom floor, the creak of the wardrobe door opening, the fabric rustling as the bra and underwear are pulled, and the wet sounds of the fingers massaging and moving in the vagina.
non_diegetic_music: N/A
You have no reference summary. "summary: [reference generation] Amateur phone POV. He opens the wardrobe, finds her still, strips her,then lifts her hips and inserts a cucumber in her butt with anal."
References should be "partially_preserved" if only certain parts referenced or "fully_preserved" to use it exactly and fully.
You're already mixing vaginal fingering and then something else it is probably too confused by that.
"Close up on on the buttocks: He seats a cucumber in her butt hole. She never moves on her own. Only he moves her." Then probably add any kind of anal lora, like this one at mid strength or until it works https://civitai.red/models/2923850/anal-doggystyle-by-funphantom?modelVersionId=3308371
@tenstrip Thank you very much for your plan, but it basically doesn't work. The medium and long shots are okay, but even if the lora intensity is set to 1, the anus in close-up shots is basically impossible to be recognized and interact with. Moreover, I found that, for example, krea2, it's also very difficult to achieve interaction with the anus in close-up shots. Maybe it's because there are too few training sets for this
@louqwer It's like a bit of a necrophile the prompt, isn't it?
Should I use Anime Style Concept LoRAs for your betas? I know one of your videos was generated with an anime art style, but by and large it's determined to generate in live-action, even when following your prompt in said example.
It is inconsistent with the half frame rate by-two anime motions, and seems to be tied to image input style, which needs to be extremely manga looking and the prompt needs to reference a Studio Ghibli or anime show to boost it. That's a default H3 ref/i2v turbo issue. For anything you'd do https://civitai.red/models/1952560/anime-flat-style?modelVersionId=3225946 or https://civitai.red/models/2861135/2d-anime-style-nsfw-lora-h3?modelVersionId=3286171 even though it's nsfw just a small amount should effect the motion style.
It would be very nice if there were gguf version of this model. Please do it :)
Please!
The w4a8 version is practically a gguf, for 12gb and 16 cards with 12 ram or more
@chamo9009 You're right, it is close. But just a little off the mark for my 8GB Vram and 32 GB Ram build... gguf of normal H3 runs way faster than w4a8 version because of a tiny difference in model size that makes the whole process more than double slower..
You are evoking an anatomy_helper lora in your description, but I couldn't find any on civitai. Any recommendation ?
MysticXXX V4 is all you need... It fills in MOST of the missing anatomy knowledge (if you need something VERY specific, perhaps look up keywords NSFW and AOI (All-in-one) to see if it covers your niche)
@valentinkognito365 You can also try https://civitai.red/models/2834417/hmnsfw-aio-sex-lora
For me MysticXXX V4 tends to give bigger boobs than intended, otherwise is a great model.
I waiting for DR34ML4Y to come out I expecting big boom from that one!
@HugMeIntoFace it`s a pic lora?I think we are talking about minimax-h3
@mistania Sorry I posted the wrong link. It is fixed now.
Does Beta 5’s sexual dynamic performance seem like a regression compared to Beta 4? Under the exact same prompt, Beta 4 could easily render oral sex penetrating from the tip to halfway down the shaft, with great persistence and overall continuity.
In Beta 5, however, it is extremely difficult. Even with constant prompt adjustments, oral sex still only inserts a tiny fraction including just the glans. In the same video, there are even camera angle inconsistencies—where each thrust goes deeper in the POV shot, but the side view shows it barely inserted at all.
Subsequent deepthroat dynamics are also quite difficult for Beta 5. It almost never performs a deepthroat blowjob with continuous penetration; instead, it pulls the dick completely out of the mouth before re-entering for each thrust.
The target video is shot in a clean, realistic photographic style with soft warm interior lighting, maintaining correct adult human proportions and highly detailed, natural facial micro-expressions at all times. Both <Subject 1> and <Subject 2> remain on the carpet at the foot of the bed inside <Subject 3> for the entire duration; no background elements change.
[Shot 1] The video opens exactly and locked on the first frame of <Picture 1> inside <Subject 3> as the sole opening frame. <Subject 1> kneels on the carpet at the foot of the bed directly in front of <Subject 2>, her entire body and face oriented toward the left side of the screen so that she faces <Subject 2>. The large circumcised penis of <Subject 2> is fully and deeply inserted into the mouth of <Subject 1> and remains completely inside her mouth at all times; the penis never appears outside the mouth of <Subject 1>. <Subject 1> tightly wraps her lips around the fully inserted thick shaft of <Subject 2> and continuously performs sucking motions; every forward thrust of the penis forces her tightly wrapped lips to slide farther along the shaft together with it. Both hands of <Subject 2> firmly and forcefully grip the hair and head of <Subject 1>, with all fingers of both hands pressing hard against her scalp and hair. Both hands of <Subject 1> are raised in front of her own chest forming clear V-sign gestures (index and middle fingers extended) and remain in this position with slight continuous movement. Thin saliva strands stretch from the corners of her mouth along the shaft, and small droplets of saliva drip from her mouth onto the shaft and down onto the large testicles of <Subject 2>, making the penis wet and shiny. Her facial expression is vivid: brows softly furrowed showing mild discomfort, cheeks subtly tensed, and eyes directed upward toward <Subject 2>.
From the locked opening frame of <Picture 1>, the camera begins a slow continuous push-in (dolly-in) that gradually moves closer to the face of <Subject 1> without any hard cut. As the slow push-in progresses, both hands of <Subject 2> remain firmly and forcefully gripping the hair and head of <Subject 1> with all fingers pressing hard against her scalp and hair. <Subject 2> swings his hips forward with large amplitude while simultaneously pulling the head of <Subject 1> toward his crotch; a substantial portion of the thick heavily veined shaft remains continuously inside her mouth while each thrust drives the meat-pink glans and more of the shaft even deeper into her throat for continuous deepthroat. He then pulls her head back only slightly so that the penis stays partially buried and immediately drives forward again, repeating the motion in a rapid rough rhythm so that the shaft repeatedly plunges deeper into her throat without ever fully leaving her mouth. Each powerful forward thrust forces the tightly wrapped lips of <Subject 1> to stretch and slide farther down the wet shiny shaft toward the base while her continuous sucking motions create visible hollows in her cheeks. Both hands of <Subject 1> stay raised in front of her own chest performing continuous V-sign gestures with slight ongoing movement and repeated slight finger curling and recovery. The facial expression of <Subject 1> grows more intense with each deepthroat thrust: brows deeply furrowed, eyes looking upward at <Subject 2> with a clear strained and uncomfortable look, cheeks strongly hollowed and tensed, and visible muscle tension around her eyes and mouth reflecting the discomfort of the rapid deepthroat. Thin saliva strands continue to stretch from the corners of her mouth, and small droplets of saliva drip onto the wet, shiny shaft and testicles of <Subject 2>.
By approximately 00:03.000 the slow continuous push-in has fully arrived at a tight close-up focused on the face of <Subject 1>. From 00:03.000 to 00:14.000 under this continuous facial close-up, the exact same actions are maintained without interruption or simplification: both hands of <Subject 2> remain firmly and forcefully gripping the hair and head of <Subject 1> with all fingers pressing hard against her scalp and hair; <Subject 2> continues to swing his hips forward with very large amplitude while pulling the head of <Subject 1> toward him so that a substantial portion of the thick shaft remains continuously inside her mouth and each thrust drives the penis even deeper into her throat for continuous deepthroat, then only partially withdraws and plunges forward again in a rapid rough rhythm without the penis ever fully leaving her mouth; the tightly wrapped lips of <Subject 1> are continuously forced to slide farther along the shaft toward the base with every deepthroat thrust while she keeps performing sucking motions; both hands of <Subject 1> stay raised in front of her own chest performing continuous V-sign gestures with slight ongoing movement and repeated slight finger curling and recovery; the facial expression of <Subject 1> remains vivid with deeply furrowed brows, strained upward gaze, and strongly hollowed tensed cheeks showing clear discomfort from the rapid deepthroat thrusting; thin saliva strands and small droplets continue to appear, keeping the penis wet and shiny.
At 00:14.000 the camera pulls back and switches to a side close-up centered on the upper body of <Subject 1> as the main subject. In this final side close-up <Subject 1> remains kneeling on the carpet at the foot of the bed with her entire body and face oriented toward <Subject 2>. Both hands of <Subject 2> continue to firmly and forcefully grip the hair and head of <Subject 1>, with all fingers of both hands pressing hard against her scalp and hair. <Subject 2> continues to swing his hips forward with very large amplitude while pulling the head of <Subject 1> toward his crotch so that a substantial portion of the thick shaft remains continuously inside her mouth and each thrust drives the penis even deeper into her throat for continuous deepthroat, then only partially withdraws and plunges forward again in a rapid rough rhythm without the penis ever fully leaving her mouth. The tightly wrapped lips of <Subject 1> continue to be forced farther along the shaft toward the base with every deepthroat thrust while she keeps performing sucking motions. Both hands of <Subject 1> remain held in front of her own chest in continuous V-sign gestures with slight ongoing movement and slight finger curling and recovery. The facial expression of <Subject 1> stays vivid: brows deeply furrowed, eyes looking upward at <Subject 2> with a strained uncomfortable look, cheeks strongly hollowed and tensed, reflecting the clear discomfort caused by the rapid deepthroat thrusting of the rough oral, with thin saliva strands and small droplets continuing to drip onto the wet, shiny shaft and testicles of <Subject 2>. The side close-up centered on the upper body of <Subject 1> holds until the end of the video. Background elements of <Subject 3> remain completely unchanged throughout all camera transitions. The opening frame of <Picture 1> is never revisited or used as any ending composition.
These are some of my key words.
I watched some other videos related to oral sex demonstrations. When I was using beta5, did I need to use an external Lora in addition to the main model in order to achieve the same effect as beta4?
@anyezhixie Yes it's normalized instead of serving the loras at more strength. But the loras around aren't enough to generalize it well and just add more dataset skew issues like b4 is full of. So waiting on a sulphur training run to have an actual improved version.
@tenstrip OK, it's worth it. Currently, the non-turbo version of Beta 5 can effectively support complex changes in character poses and camera movements in longer videos. I just need to figure out why high-intensity LORA causes the entire video's initial scene to be redrawn, and why the original initial frame reference is used as the end frame reference instead.
@anyezhixie Reference stuff would be prompting. Most Loras aren't trained on reference and they will impact it negatively. That's one reason to make this since when they're merged into the model the way I do it they kind of become reference able data.
It also work slower and eat more memory
Will there be a ref 2 video model for EROs Minimax?
I think its already a hybrid ref2va and fl2va for eros max so same model does both. Works very well at both ref2va and i2va for me. Worst side effects I've seen on this model is extra saturation and a tendency to make the titties larger than the input image. Neither are the end of the world lol.
I literally only use this version with reference now, it's mostly been made for reference.
I'd like to suggest you to base beta 6 on beta 4, rather then 5. v4 gives way better results, IMO.
Oh yeah, and THANK YOU for your checkpoint. It's really good.
It's the wrong way to go around doing it. I'm done polluting with undercooked loras. Right now I'm making refmods for every missing major concept and just fixing it with 5mb files and workflow.
First of all, thank you for all your hard work. Beta 5 has a huge improved in terms of prompt adherence, but there are still some lingering issues: The video quality is somewhat worse compared to the original official model. (I'm not sure if this is caused by conflicts between certain LoRAs), there are still problems with the depiction of genitals, with the pussy looking like a swollen lump, the penis appearing blurry, and the male and female genitals interfering with each other. Women also tend to have large breasts even when the prompt specifies a flat chest. The liquid effects have also gotten worse (possibly due to the same issue causing the drop in image quality).
All of the issues mentioned above are specific to T2V, not I2V. MiniMax H3 has consistently performed very well with I2V, but it runs into problems when T2V has to create things from scratch.
personally beta2 gives most realistic skin color, afterwards the skin are very greasy
For the turbo look mitigation you can use a sampler like LCM/simple 6-8 steps or use the latent contrast node from https://github.com/Tr1dae/ComfyUI-MiniMaxH3_LatentUpscaler and reduce to 0.85 to preserve the norms from the turbo which eliminates that.
Thank you for another update! I might even post a part 2 to my "Don't sleep on 10Eros" post on reddit lol. When it comes to sampler/scheduler combos these are my top picks using the TURBO version. And please just try them (especially if u get plastic skin):
heunpp2 x ddim_uniform @6 steps (Amazing - a little "slow")
res_multistep x simple @6 steps (Great - Fast)
ddim x ddim_uniform @6steps (Great - Fast - But Different)
I usually run either of the bottom 2 for quick gens, just depends on which scene lighting and vibe I like most.
thank you for your masterpiece of work. Actually I'm completely new to wan2gp and want to try this model. Is there any config or guide to add and use this model?