CivArchive
    4K 4-Step MiniMax H3 - Text-to-Single-Image | Multi-Reference Image-to-Image Workflow - v4.0
    NSFW
    Preview 142171543

    Workflow is able to +4K Text-to-Image and also Multi-Reference Image-to-Image.

    UPDATES: v4.0 Put page on pause as I sorted out the new turbo 4step_v1.2_768p_comfyui_bf16 as it has solved many of the problems in skin rendering and a slew of other things, this one is great at 4 steps.

    Workflow v4.0:

    It'’s got two sampler groups: the regular Sampler Custom, and the Clown Shark Sampler from RES4LYF, you can swap between the working groups with the Fast Groups Muter.

    Plus several of its optional nodes — Detail Boost, Implicit Steps, Momentum and Sigma Scaling. I left notes from the official RES4LYF example workflow for explanations of what they do, along with some quick notes on my testing of them.

    I tried most of the samplers and landed on what feels like the sweet spot for Clown Shark: diag_implicit/pareschi_russo_2s with Beta57. It’s a little slower, but totally worth it IMO.

    Got rid of the text enhancer and replaced it with the Video/Audio Shift node as it had a better effect on prompt adherence.

    This workflow is pretty much a set-and-forget workflow, I found settings that give a nice balanced image so nothing "should" need adjusted, still tweak to your hearts content as there is plenty to adjust if you want.

    The DasiwaMinimaxH3_dasiwaREF2VAHybridV1 model Darksidewalker pulled is a must, every setting in the workflow is tuned for this exact model, if you use anything else you're going to have to find your own settings.

    Quick notes:

    • The Beta57 scheduler by RES4LYF is a must for image quality in this workflow.

    • For extra detail, run a SeedVR detail pass with the same resolution as the image (no upscaling).

    • I’d keep the LoRA strength light as these were made for video.

    • Eyes at a distance is still a problem as of v4.0 but it's much better this round and prompt adherence is still spotty, let's hope the image model fixes it.

    • The model I used to create the images was DaSiWa MiniMax H3 and the CLIP was Huihui-Qwen3-8B-abliterated-v2-FP8-Comfy as these both gave much better results.

    • Pretty much everything you can tweak is sitting on the subgraphs. Fair warning: cranking the resolution too high will probably throw errors.

    • My pin naming is a bit weird, but that’s on purpose. It helps me keep track of what goes where when I’m unplugging and plugging a bunch of stuff.

    • I put the Save Image node on a switch, I only save the ones I actually like. Leave the seed on fixed, generate, then flip the switch and hit run to save (or just right-click the preview and save it manually).

    • Reference Image input size has a great deal to do with rendering time, the bigger the image the slower the renders. I set the resize @2.0, if you have a slower card set this lower.

    About lighting, it is very hard to tame as MM always wants to shine a spotlight on the characters. Want a dark dimly lit atmosphere, forget about it, at least I haven't found a way yet after many prompts of lighting techniques with very few exceptions. The closest I can get to good lighting is using this phrase at the beginning and adapt the second half per scene: "Ray tracing. Make no studio, set, self emissive character, spot, vanity or hero key lights. Soft, diffused, bounced, natural daylight window key lighting from all sides." If anyone finds a way to get pinpointed lighting that you want let me know, I will be grateful for it. If you want better examples drag the v4 example images into Comfy as they nearly all have the current workflow embedded.

    The Big Three:

    The three biggest contributors to image quality is ETA in Sampler Group and Implicit Steps, Lying Strength with Lying Start Step in the Options Group

    Adjusting the two down in Lying will produce a more grainy but sharper image, up will make them smooth and blurrier. ETA in Sampler Group does the same thing.

    Implicit Steps can have multiple effects, sometimes 2 looks better sometimes 1 will, varies from prompt to prompt and seed to seed. Switch to 2 when you get a image you like to save time. 2 is also good for busy scenery.

    Some things I learned on this journey. MM in this format loves, loves and loves some more the details. If you don't explain exactly what you want in a detailed way you will get bland output. For example just saying a woman is moaning won't get you much, putting things in like the state of her face such as open mouth, furrowed eyebrows and facial tension will get you much further, this applies to everything, even more when it comes to lighting.

    Nodes and models you will need to have installed, the list below is the bare minimum.

    custom_nodes:

    ## Model Links

    diffusion_models:

    text_encoders:

    vae (HARD REQUIREMENT):

    vae_approx (TAE):

    loras/h3:

    ## Highly Recommended

    diffusion_models:

    text_encoders:

    Description

    FAQ

    Workflows
    MiniMax H3

    Details

    Downloads
    37
    Platform
    CivitAI
    Platform Status
    Available
    Created
    9/8/2026
    Updated
    9/9/2026
    Deleted
    -