CivArchive
    MiniMax H3 - Text-to-Single-Image | Multi-Reference Image-to-Image Workflow - v3.0
    NSFW
    Preview 141755442
    Preview 141752563
    Preview 141750853
    Preview 141750879
    Preview 141750881
    Preview 141751230

    Workflow is able to +4K Text-to-Image and also Multi-Reference Image-to-Image.

    UPDATES: v3.0 - Removed SDE node and replaced with 3 new nodes. Removing SDE and adding the others gives a much more detailed natural look to skin, while a bit softer the tradeoff was worth it IMO, also seems to have better prompt adherence.

    v2.0 on left, v3.0 on right

    Clown-Shark Workflow v3.0:

    It'’s got two sampler groups: the regular Sampler Custom, and the Clown Shark Sampler from RES4LYF, you can swap between the working groups with the Fast Groups Muter.

    Plus several of its optional nodes — Detail Boost, Implicit Steps, Momentum and Sigma Scaling. I left notes from the official RES4LYF example workflow for explanations of what they do, along with some quick notes on my testing of them.

    I tried most of the samplers and landed on what feels like the sweet spot for Clown Shark: diag_implicit/pareschi_russo_2s with Beta57. It’s a little slower, but totally worth it IMO.

    The Klein node for prompting (in System Settings Subgraph) you can just delete from the workflow or disable if you don't want it, but I find it does help.

    Quick notes:

    • If you have the DasiwaMinimaxH3_dasiwaREF2VAHybridV1 Darksidewalker pulled you are in for a treat when it comes to up close portraits and nearly everything else except for eyes at a distance where it lacks.

    • The Beta57 scheduler by RES4LYF is a must for image quality in this workflow.

    • For extra detail, run a SeedVR detail pass with the same resolution as the image (no upscaling).

    • I’d keep the LoRA strength light as these were made for video.

    • Eyes are a problem when distance is involved and prompt adherence spotty, let's hope the image model fixes it.

    • If you use JonXL's Generator LORA with text-to-image, I found a strength setting of .20 @8 steps is best, with reference images you can go higher.

    • The model I used to create the images was DaSiWa MiniMax H3 and the CLIP was Huihui-Qwen3-8B-abliterated-v2-FP8-Comfy as these both gave much better results.

    • Pretty much everything you can tweak is sitting on the subgraphs. Fair warning: cranking the resolution too high will probably throw errors.

    • My pin naming is a bit weird, but that’s on purpose. It helps me keep track of what goes where when I’m unplugging and plugging a bunch of stuff.

    • I put the Save Image node on a switch, I only save the ones I actually like. Leave the seed on fixed, generate, then flip the switch and hit run to save (or just right-click the preview and save it manually).

    • Reference Image input size has a great deal to do with rendering time, the bigger the image the slower the renders. I set the resize @2.0, if you have a slower card set this lower.

    Nodes and models you will need to have installed, the list below is the bare minimum.

    custom_nodes:

    ## Model Links

    diffusion_models:

    text_encoders:

    vae (HARD REQUIREMENT):

    vae_approx (TAE):

    loras/h3:

    ## Highly Recommended

    diffusion_models:

    text_encoders:

    Description

    • Removed SDE node, added 3 more. Split Clown Options into their own subgraph. Added notes from the official RES4LYF workflow, added personal notes.

    Workflows
    MiniMax H3

    Details

    Downloads
    73
    Platform
    CivitAI
    Platform Status
    Available
    Created
    9/4/2026
    Updated
    9/4/2026
    Deleted
    -

    Files

    minimaxH3TextToSingleImage_v30.json

    Mirrors