More examples on HuggingFace
Find the Native Workflow as a .zip in the optional files
- This LoRA applies on MiniMaxh3 ref2va model
Changes the Reference video identity face with the Reference(s) one(s)
Uses a single Trigger word (Faceswap)
Uses at strenght 1
The LoRA is pruned from blocks under the set fro treshold (less heavy, less lora artifacts, identical strength to trained base)
Workflow is shared, (you can use your own Ref2va workflow)
(CRT-Nodes current version is 2.19.0, fetch it from github if you want to use these nodes, ComfyUI registry has it stuck to 2.17.0, i'm working on it to pass checks and make current version available on the manager), you can use your own ref2va workflow, the LoRA will work fine too.
If you want to support this model, you can do it on https://buymeacoffee.com/designedbycrt
Or https://runpod.io?ref=u7b2habt if you are using RUNPOD, so we both benefits with free credits, everything I earn from these is destined to train and share here and HF https://huggingface.co/UntMods anyway. Appreciated
Description
FAQ
Comments (61)
"make sure your input video respect the MH3 format"
I'm getting error "The value Ref Video 1 for CRT_MiniMaxH3UnifiedSampler's aspect_ratio is not available."
I assume my video is the wrong aspect ratio? Any tips on what are valid? I've been trying a wide variety and have always had this error.
Did you had CRT-Nodes before? If so, you should update it in the manager. You can use any aspect ratio, the node will automatically use the ref video aspect ratio and resize it to your megapixels value (in the settings tab), the length is also automatic based on your video set length (video loader).
But yeah you should try to update the nodes, restart comfy and reload the workflow
@pgc Same error here. Using the linked workflow. Everything updated.
@deestovel send your console trace from "[INFO] got prompt". I can't tell what could be the issue without that
@pgc [INFO] got prompt
[ERROR] Failed to validate prompt for output 149:
[ERROR] * CRT_MiniMaxH3UnifiedSampler 135:
[ERROR] - Value not in list: aspect_ratio: 'Ref Video 1' not in ['1:1 (Square)', '2:3 (Portrait)', '3:4 (Portrait)', '4:5 (Portrait)', '5:7 (Portrait)', '5:8 (Portrait)', '7:9 (Portrait)', '9:16 (Portrait)', '9:19 (Portrait)', '9:21 (Portrait)', '3:2 (Landscape)', '4:3 (Landscape)', '5:3 (Landscape)', '5:4 (Landscape)', '7:5 (Landscape)', '8:5 (Landscape)', '9:7 (Landscape)', '16:9 (Landscape)', '19:9 (Landscape)', '21:9 (Landscape)']
[ERROR] Output will be ignored
@deestovel So you have outdated nodes, current CRT-Nodes is "(2.19.0)"
https://github.com/PGCRT/CRT-Nodes/commit/d1bbdcb1d49cc5bdc0c43db75ba44d649f10b611
Commit d1bbdcb
"- Add "ControlNet" I2V aspect mode and "Ref Image 1" / "Ref Video 1" canvas aspect sources; rework frame-count precedence; default video_frames_override to True; log per-run reference content fingerprints."
I think there is an issue with the comfyUI registry that shows 2.17.0 as current, I will investigate, for now you can grab it from github with git clone into your custom nodes folder
git clone https://github.com/PGCRT/CRT-Nodes.git
The reason is a feature in the "Unsloth Studio Bridge (CRT)" that was flagged and didn't pass the checks, i'm glad you guys reported this because otherwise I wouldn't know,
@pgc Ahhh..Thanks. Yeah, Comfy registry has 2.16 latest. I installed nightly version and that works.
@deestovel If you go to "custom_nodes\crt-nodes\pyproject.toml" you will see the version you have, last is 2.19.
I was about to go sleep but fuck ^^
@pgc I have 2.19, same error (edit, my error was my own dumbass fault, video I chose had no audio lol), it was the same overall generic error but once I checked the logs I realized my error.
I had the same issue even with the newest version.
You have to choose the aspect ratio in the Minimax H3 Node -> click on R2V, choose aspect ratio -> run
@Lora_Addict of you are on prior 2.19 yes probably, otherwise set r2v aspect to "ref video" for automatic aspect ratio
changing video upload node type from none to animatediff allowed me to input custom height and width for aspect ratio alignment.
Fantastic work! did a quick test, The Fugitive scene starring a certain famous person. worked really well in wan2gp
how did you get this to work in wan2gp? Di you simply use the prompt "Faceswap"?
@ltcdrbroccoli yes. use reference image. reference video. worked for me. i didn't use any turbos. so if it fails, try once like that
请问我是否能不使用CRT节点直接让LORA在我自己的工作流中运行
It's a Ducking Great! Thanks a lot. And what keywords can I use instead of "faceswap"?
also wondering if we can have two subjects in a video, and only swap one face
hell yeah. using my own ref2va workflow and a few tweaks, it just works.
Edit: any way to stop getting combined faces? most gens combine the reference face with the original face...
i'm using the same turbo lora, 8 steps, etc.
Man, this looks so dam cool, but I'm struggling to find a reason to use it, when and how woudl it work. Man, this looks so good. But...when woudl be a good time to use this? It's great, really good work. I wish you could at least apply a differetn diolouge, but I know that's not how it works.
Thanks. Working well and it's easy to use.
As Minimax could do this on it's own already with reference stuff and the correct prompt i wonder if you could do the same but with swapping the whole subject including body and or clothes? It would be just so much easier with just one prompt word instead of these long prompts you need for a subject swap.
Yes, this is why I trained it, it can natively do it with a fancy prompt but it's less reliable and often doesn't work, you can probably use both to enforce
I noticed that cuts can be delayed on both native / with lora, since a latent frame is a small batch or images, it wait the next latent frame to perform the cut.
A single trigger word is easy but you can experiment and see what works best, i knew that civitai will bitch on me with the examples videos so I didn't generated/experiment much
@pgc i see! Thanks for your response!
Same request, I find it really challenging to prompt Minimax correctly for that kind of things (and I know the prompting guide and I've been experimenting with many system prompts with a 26b VLM + LLM). It's definitely not obvious and if you found it easier, lucky you because a considerable number of us are struggling. I haven't tested this face swap lora yet but I concur, clearly some loras helping swapping a full character would be great
@valentinkognito365 There is alot of unexpected uses cases that works, the black chad cat is an example (I didn't train anything that images and videos A/B of face swaps, no animals and surely no cats ^^, so maybe it's also possible to "faceswap" an old bue car by a yellow porsche using the faceswap trigger, if it works then I think that ≈anything could be replaced, including full character, but I didn't test edge cases tbh (other than this cat swap)
@pgc if that's true then you are a hero of mine for having done this LoRa. Would you (please please) do the extra step of posting in your description a few (bulletproof as much as possible) prompts that you used to do face swap and character swap ? That would really help making the best out of your LoRa, and i'm sure people here could contribute in the comments
Thx. Worth to test. Where can i get this workflow? https://blobs-b2.civitai.com/file/blobs-managed-public/9eb1502189dbe0fc7cd017763529a462.png
Clone this repo into you custom nodes directory
git clone https://github.com/PGCRT/CRT-Nodes.git
You will find the workflow on the optional files in this model page on CivitAI as a .json
It's crazy. You can use old blurry and super dark videos, use reference pics of the character seen and prompt it to be a good quality and bright video and you get like the same video, just in good quality. This works without this lora, but at -5 it adds even more bright and realistic visuals.
Why/how would this work? It's a LoRA, not a change to the underlying structure of generation. Swaps can happen if you coax the the inputs/prompt into doing so but I can't see a LoRA doing anything like that.
A LoRA isn't an instruction beyond what is on top of the generation. It can't know what a faceswap is since it doesn't actively compare two different inputs.
you know nothing john snow. look at what has been done with the identity lora on krea2...
(btw joke aside, I'd be interested as well to know about the magic :-)
The LoRA works because Ref2VA is a multimodal model, meaning all the things the Ref2VA model can understand, it can be conditioned to enhance those or adapt new behavior.
works like a charm
Works great, I use my usual REF2VA workflow.
Nice, thanks! Would you mind sharing a couple of specifics from your setup? Mainly:
- denoise value on the sampler
- steps / steps_turbo
- which enable_ toggles you have on/off (sol_attn, chunk_ff, spectrum)
- how many reference face images you used, and the mode (Faceswap?)
I'm getting basically the untouched source video with artifacts, face isn't swapping at all, so trying to figure out which setting is off on my end.
@madara_01012000 1.0 denoise (with DMD, Taomate has manual sigmas for denoise). I use the the DMD pruned turbo lora from Drbaphs repo at 8steps, euler simple. or the Taomate with the 4-step manual sigmas that Drbaph has there, euler. I use Spectrum default settings, KJ Low VRAM attention, Comfy-kitchen attention. I used 3 reference images.
Also I RIFE interpolate the video to 24 frames, and slice 73/etc (where applicable) and make sure the vid length/frames are the same.
Could you link to this "usual REF2VA workflow" please? Because with this workflow i get OOM.
@voayer2003587 literally just the ref2va comfyui template with your node additions, like CK attention, spectrum, sla, vram attention, etc.
@madara_01012000 Lol...this is exactly the same workflow i used. But thx for the info
Running this workflow for faceswap (3 reference face images + a reference video, R2V, Turbo LoRA + FaceSwap LoRA, 8 steps). Generation completes with no errors, but the output is basically the original source video with artifacts — the face isn't being swapped at all. It looks like the input video, just noisier/distorted.
I've tried adjusting denoise without much visible change. Could you share what denoise / reference-strength value is expected for a proper faceswap effect in this workflow, or whether I might be missing a setup step for the references?
Are you using native nodes or the workflow I shared?
If you are using native nodes then you need to pre-process the video using this https://github.com/ostris/ComfyUI-AIToolkit-MiniMaxH3
@pgc Oh, thank you very much for the reply, I’ll try it now.
@pgc I’m using your workflow.
@madara_01012000 I was asking because the inference node doesn't provide denoise or reference strenght values so i wonder what did you adjust, it's all automatic and ready to use after you load your reference video and reference (face) image.
@pgc I’m launching it like this, but the face doesn’t change. https://drive.google.com/file/d/1IeBkN-ZHcEFFypB_ogOgG6gIA4xchxEn/view?usp=sharing
For me, the face swap results are infinitely better WITH this LORA than just a prompt with a reference or even references. If you need to swap face (or in my case, repair/restore a face lost by too much compression and low resolution) it's a no brainer to at least try it. It's on sale for exactly $0 today.
You don't need a special workflow, just load the LORA into your Ref2VA WF, include the word 'faceswap' in your prompt along with a basic R2VA wf, load a video, load a <picture 1> for reference (you can do more than one) and voila. Disable the lora and try it again without it.
good work ;)
Thanks for this LoRA! If you have time and resources please also train a character swap. I know Ref2VA can do this out of the box, but this seems to be easier and improves/enforces the success rate.
Any idea why the video would come out completely unchanged?
Nice works really good, one thing though, how do i use the background from another image? i always add this prompt "Faceswap, Use background from <Picture 3>." with 0.7 weight but its always only the face and outfit but using the same background as the video source?
I have 16 gb GPU and 32 gb sist. Ram and got OOM all the time. Even when i changed resolution from default 0,5 to 0,4. Maybe author should put a warning, that this workflow is not for 16 gb or less GPU's.
im using it for 1.0 pixels resolution while using the normal minimax h3 on a 5060 Ti 16gb works fine just takes time, try changing your workflow or maybe its the resolution of the video is too high.
When you open the task manager during inference, in the "Performance" tab, do you see the "shared GPU memory" getting filled before OOM?
If not, make sure that your comfyUI .bat launcher doesn't have this argument "--disable-dynamic-vram"
You can also add "--vram-headroom 2.0"
Make sure you do not use the int8 convrot model but a smaller quant like W4A8.
I don't recommend GGUF for any DIT models.
@pgc Except for the text encoder (mine is the "Heretic" version), which is the same size as the one in the workflow, I used the exact same models listed in the workflow. I gave Grok AI to analyze the error logs, and it told me that at one point the system tried to load 19 GB into my GPU, triggering an "out of GPU memory" error.
@CallMeEviL I have the same GPU and 32 GB system ram. Which main model did you used in this workflow? The one in the workflow? 20 GB?
does this work on several characters?
Hi, hoping someone can help, I keep getting the original video back with no face swap, every single run.
CRT-Nodes version: 2.20.0
Models used (all downloaded from these exact sources):
- UNET: MiniMax-H3-Ref2VA-pruned_int8_convrot.safetensors — https://huggingface.co/DeepBeepMeep/MiniMax-H3/blob/main/MiniMax-H3-Ref2VA-pruned_int8_convrot.safetensors
- FaceSwap LoRA: SS_FaceSwap_MiniMax_H3_REF2VA.safetensors — https://huggingface.co/UntMods/FaceSwap_MiniMaxH3_REF2VA/tree/main
- Video VAE: minimax_h3_video_vae_int8_convrot.safetensors — https://huggingface.co/Kijai/MiniMax-H3-experimental/blob/main/minimax_h3_video_vae_int8_convrot.safetensors
- CLIP/text encoder: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors — https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
- Audio VAE: minimax_h3_audio_vae_fp32.safetensors — https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/vae/minimax_h3_audio_vae_fp32.safetensors
- Turbo LoRA: minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors — https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors
Workflow setup:
- workflow_mode: R2V, mode: Faceswap
- Both LoRAs at strength 1.0 (also tested FaceSwap LoRA at 2.0, no visible change)
- 3 real reference face images + 1 reference video, preprocessed through AI-Toolkit H3 Reference Video (resampled to 24fps, frame count auto-adjusted)
- Sampler runs clean, no errors, full "R2V complete" in the log (confirmed: models_pipe correctly wired, ref2va_turbo_model ✓, 3 images + 1 video refs loaded with real, non-empty content)
Result: every time, the output is basically the original reference video, face unchanged. No errors anywhere in the console.
Already tried:
- Raising FaceSwap LoRA strength to 2.0
- Reducing to a single reference image instead of 3
Neither changed the result. Since the pipeline runs cleanly end-to-end with no exceptions, wondering if these exact model versions are compatible with each other, or if there's a setting on the model/LoRA side I'm missing, or specific reference-image prep (crop/alignment/resolution) needed for the FaceSwap LoRA to actually kick in. Any ideas welcome.
WorkFlow: https://huggingface.co/UntMods/FaceSwap_MiniMaxH3_REF2VA/blob/main/SS_FaceSwap_MiniMax%20REF2V.json
Your workflow seems correct, thanks for providing details.
You said "preprocessed through AI-Toolkit H3 Reference Video", but if you are using the workflow I provide, the pre-processing is already handled by the nodes, so this is not the cause of your problem.
both lora should be set to 1
make sure that the reference images are focusing on the head as a tight portrait, (I suggest you to use https://nomacs.org/ as your main windows image viewer it's great, you can use the "c" key to crop the displayed image).
You can send me your video and reference images on discord if you want, I can test and tell you what is wrong