JonXL's Minimax H3 Photorealistic Image Generator (LORA + Workflow)
Consider to download the workflow (from 'optional downloads') for easier image extraction.
A conceptual lora and workflow to create high resolution photorealistic images with Minimax H3.
Remember image generation on WAN? Now you can generate images with H3 too.
Photorealistic datasets and special training has made it possible to create a photorealistic image generation lora for Minimax-H3, with output up to 3840x2176.
With two types of available workflows, the photorealism image lora is used to either capture the first and last frame out of a 22 frame video (Default) or generates 1 frame (Cicaloo Node+VAE).
Pick either the 3780 or 3990 step variant. Lora 3990 seems to be more vibrant / contrasty and sharper detailed.
Text 2 Image
This lora has been originally developed with the intent to generate photorealistic images from text.

Example from Text to Image
Image 2 Video
You can also use the "Image2Video" mode to inject an input image into a new output. Results may vary, lora strength may play a factor.

All images above were made using the "Default workflow".
Reference Model:
With support for reference model in the default workflow (not the 1 Frame workflow), you can also use the ph0t0r34l lora with a references.

You can try either the FL2VA model or REF2FLVA model, both may yield different results.

Cicaloo 1 Frame node support
As of 26.08, two additional workflows using "Cicaloo 1 Frame node" have been added.
Speed up times on my 5090 are about 5 times faster for a 3840x2176 image generation at 55 seconds vs 5 minutes+.
My recommendation is to use the "Default workflow" up to 2560x1440 for best quality and the "Cicaloo 1 Frame workflow" for renders up to 3840x2176 higher resolution.
Comparison 4K:
JXL_IMGEN_DEF_V4b_single_lora workflow

0.9 lora strength - cfg 1 - 5 minutes 53 seconds
JXL_IMGEN_1FRAME_T2I_v4 workflow (CICALOO 1 Frame node - based)

1.5 lora str / cfg 1.5 - 55 seconds
Note: Adjusting lora strength and CFG can have an effect on sharpness versus stability.
Prompt:
ph0t0r34l, A twenty four year old blonde female looking out a rain-streaked train window, blurred landscape passing by, soft interior lighting, introspective and melancholic , depth of field
Pro tip: The 1 frame gen workflow works quite well at 3840x2176. If you can afford it, try running it at 40 steps for best results.
Authors comment
Feel free to post your safe for work results on this LoRa's page too!
Please enjoy making content and experimenting with this photorealistic lora and workflow!
Usage:
Preferably, use the following trigger word to activate the image generation
ph0t0r34l
Seems to work very well on resolutions up to 2560x1440 resolution (Default workflow) and 4K when using the Cicaloo workflow.
The lora when run at lower lora strength can also help bring out any of your character loras n a more static position.
As an experiment, you could try increasing/decreasing lora strength and maybe combining both loras with different strengths. 1.0 lora strength is standard for "default" ans 1.5 for the Cicaloo workflow.
Let me know your feedback, as I believe this is the only photorealistic image generator that exists for H3 so far :)
Known issues:
The workflow might complain about an image needed to be loaded, just feed it an image (ignore that it will have any effect though)
You may need to install some frequently used nodes, mostly Kijai's nodes and VideoHelperSuite
When using the Cicaloo workflows, make sure to clone the custom node and download the special VAE, no turbo lora required.
https://github.com/cicalooo/ComfyUI-MM-1Frame
Comparisons:
Left is with lora, right is without lora.
Right click, open image in new tab, notice the difference in sharpness and contrast.





Description
FAQ
Comments (51)
The question is
Can it goon
Officialy, the Minimax H3 developers dont't appreciate NSFW loras and finetunes. Therefore, I cannot comment on it specifically. However, as it is trained on realistic-like datasets it may or may not improve texture details of a certain gender ahem. I mean, you can always try and download right?
@JonXL What an interesting answer, seems like i should download and try for the good of this investigation, for the culture... it is very important, the culture
Winnie the Pooh approves
@JonXL its hilarious because H3 was trained on porn too. I guess they just follow party guidelines to avoid trouble.
Technically H3 should not be used outside of China and few approved countries but they pretty much said they don't care and whole USA/EU ban is just to avoid legal troubles.
@bitzupa Better safe than sorry I guess xD
It's both sad and funny that in this day and age we have to rely on the "People's Republic" for creative FREEDOM :)
@Postaldude1985 If I were to hazard a guess with a tinfoil hat, the technopurity apparatus of the PRC is some kind of immune system for completely uncensored gen AI and dumping it into Western civ may be a play at destabilizing things.
But, I doubt that seriously, they just have a more research open mentality since they aren't as badly corrupted by capitalism and it eats at centralized Western corporate power more than anything.
The workflow seems to require an input image? Even when text to video or manually bypassed the loan image node throws and error?
Thanks for your comment.
You can load an image and press run, it's explained in the description. No need to manually bypass things. If you want to know why, it's because the base workflow is multi purpose. It won't use an image it just needs one to be loaded. Why multi-purpose? Because I don't want to maintain multiple separate workflows.
Short answer: load image and run.
If you somehow still have an issue after you re-opened the workflow as it was downloaded, loaded an image and pressed run, then let me know.
By the way, I just added an extra disclaimer on top to let people know they should load an image.
I'll also consider either rewriting the workflow to 100% account for no image loading in text to video mode, or uploading a 100% text to video workflow.
Fixed by the way in workflow version V2. No longer requires an arbitrary image to be loaded.
what about image edit? :)
Thanks for your question.
I haven't played with editing in H3, but you can use this lora and workflow to inject an input image into a new scene. Not intended, but seems to work (may need some help describing the character still)
Minor update to workflow 24.08
Workflow version V2 no longer needs an arbitrary image to be loaded for the image generation to take place.
What about ref2va?
You tell me, what about it? :)
The lora seems to work fine with the reference model by the way. You can easily use any reference image with the ref model together with a text prompt to generate a photo/image.
I'll try including a reference portion in the workflow this week.
I've seen recently on github something named "MiniMax H3 Combined Image And Reference to Video" node that makes life easier, grab it put in custom_nodes, restart, then in workflow swap the "MiniMax H3 First Image to Video" to "MiniMax H3 Combined Image And Reference to Video" it and you can plug references, even with FL2VA
@AmplitudeH3 Thanks for the tip! I'll take a look at your recommendation soon.
By the way, if you download the image of the blonde woman on the beach near a blue car, that image already has a beta workflow using the (default) REF node with a few bypassable inputs (also using the FL2A model).
(Haven't had time to test it thoroughly, but should work, hence why I haven't uploaded it officially yet)
You can download image throw the image in ComfyUI for the REF workflow in the meantime.
@AmplitudeH3 I updated the workflows by the way :) So if you'd like to play with reference now you can :)
tried the lora with cicaloo's 1frame node, it helps giving the image some extra details that don't look forced
Interesting, I"ll try to see how they compare.
Thanks for your feedback!
Also tried the 1Frame. My view the 1F is quicker but skin has that plastic look and finer details are lost, this workflow is slower and produces more detail to it's detriment. For example skin is rough looking, find a in between and you're golden.
Guys,
I uploaded two additional workflows (zip archive) that now support the Cicalooo 1 FrameNode.
I made sure to tweak the settings as best as possible, output is now a lot closer to the default workflow. Also 4K resolution is possible within a minute on 5090, so that's cool, since the default workflow would take about 5x longer.
If you want to experiment, try changing the lora strength and cfg values.
(Especially for I2I this seems to matter a lot more)
Update 26.08
Now included a set of workflows V4:
https://civitai.com/api/download/models/3260165?fileId=3151042
Contains the default workflow and two additional workflows for use with the cicalooo/ 1 Frame node.
https://github.com/cicalooo/ComfyUI-MM-1Frame
(Faster generations, even at 4K resolution, at the cost of some stability / unsharpness, although I tweaked it pretty good)
tried JXL_IMGEN_1FRAME_T2I_v4 wf , it will not run. these nodes missing:
MM1FrameMiniMaxH3ReferenceToImage
MM1FrameEmptyMiniMaxH3LatentAV
using comfyUI 0.33.4
Hey, thanks for your message!
For that worfklow you are required to install the following node, and download the VAE
https://github.com/cicalooo/ComfyUI-MM-1Frame
minimax_h3_t1_image_vae_step1597.safetensors
Can we use a ref model, with more reference images, to generate images?
Hey there!
Yes, while not intended, it's possible:
I just uploaded a new set of workflows that you can download:
https://civitai.com/api/download/models/3260165?fileId=3179647
Note that only the regular workflow got support for References.
You can either try using the FL2VA or REF safetensors model for different results.
(The 1 Frame node does not have support for reference)
Cool! Although there are only 5 pictures for reference, it's generally enough for most users.
@ccbeauti5965954 Whenever there is another need for an update to the workflow then I will nclude some more referemce images. Five for now should be plenty I'd assume.
@ccbeauti5965954 Updated the workflow to now support 9 (max?) image references
https://civitai.com/api/download/models/3260165?fileId=3179647
Update 02.09
Now included a workflow that has support for reference as well.
https://civitai.com/api/download/models/3260165?fileId=3179647
Officially not supported but seems to work :)
How did you write the reference model prompts? The official prompt structure doesn't seem to work; is a natural description sufficient?
ph0t0r34l,
<Subject 1> is the character represented in <Picture 1>. In the <Picture 2>, <Subject 1> is the tattooed woman from <Picture 1>.
If you havent switched the model loader to load the ref model yet, consider doing so. They can lead to different results.
Prompt structure like worked for me:
... ... ... <Picture 1> ... .... ... <Picture 2> .... ..... ....
@JonXL I've tried something similar to what you've done, but the images often don't blend together properly, failing to merge the head reference onto the train figure as you recently showed on the page. I'm using 20 steps: cfg 1 lora 1.2 and 22frame
@Jasonchow1022 Thanks for letting me know. Did you give loading the reference model safetensors a try? Running the lora at 1 strength?
Have you tried running your prompt on a different reference workflow that doesn't use my lora?
For me I had some good resiults even though the lora was never trained on referece data, I can try posting my prompt here later for you to try:
ph0t0r34l, replace the redhead female from <Picture 2> with the blonde female of <Picture 1>. Keep the subject likeness of <Picture 1>. Keep the train setting of <Picture 2>. Change the outside lighting to <Picture 3>. Static, no movement..
@Jasonchow1022
I just uploaded a new workflow zip file, including the reference images, and including a pre-filled reference workflow that you can load as is.
Please see if you can run that workflow to see if the reference output image gets generated.
https://civitai.com/api/download/models/3260165?fileId=3179647
@JonXL Thank you so much for your workflow and the suggested keywords. I've run it and it's working. I hope you continue to optimize this workflow; the only downside right now is that it's too slow.
@Jasonchow1022 Hey there! Thanks for your feedback. There is no way to optimize the workflow, the slowness that you mention is mostly likely the generation time it will take regardless of the workflow running on your GPU. You are free to change the resolution of course for faster inference.
How is this regarding render time versus rendering lets say a 3 sec video?
Also since it supports references, how does it handle that weird distance face thing of h3 minimax face reference where distant faces become pixelated / weird?
Depending on which workflow you use, it either uses 5 internal frames for the "1 Frame" workflow or 22 frames for the standard workflow. So its 5/72th or 22/72th the generation time compared to a 3 second video.
Regarding the faces in the distance, this should be less of an issue as long as you generate on high resolution, which you can and should. After all, since you only need 22 frames to generate. you can use way higher resolution in the same VRAM budget as a 3 second video or longer would take. A resolution of 2560x1440 seems optimal for the default workflow. The lora was also trained on static still images / videos, meaning less motion to jitter things up as well.
Of course, you can download the lora and workfows and try out for yourself :)
Amazing work as always.
Thank you AiInfluence :)
great work! sadly, it takes almost 1min to generate a qwen equivalent image, you work is great but just minimax face consistency is really bad so I can't use it to edit people :/
Hey marin, thanks for yor comment!
Can you tell me more about what you intended to do?
sorry, I was thinking that if minimax could be a great image edit for face identitiy but its just not there yet. your work is very good but justminimax destroys face identitiy . I asked a woman to move around into another psoition and she ended up with a differente face
@marin73tomas840
I haven't personally done much with editing H3, but you did use the reference model with <Picture X> tags?
Also have you tried Image 2 Video to make the character transition and grab the last frame?
So technically, if it were possible, your case would also be solved if a lora existed that could generate a still image with a different pose with the same character based on an input image?
I might be able to try training a first frame last frame lora regarding this pose editing at some point :D







