27/05 : Updated with SVD INT4 Quant for Flux https://github.com/mit-han-lab/ComfyUI-nunchaku
Special thanks @theunlikely for the work done on quantification (took 6 hours on a H100 GPU), and thanks to JIB (J1B Creator Profile | Civitai) for his WF and clarifications on using this type of model:
For using this model you need to install the nunchaku project following the instructions there. and the nunchaku ComfyUI custom nodes to get it to work. Download and unzip the archive from Civitai to: \comfyui.git\app\models\diffusion_models\svdq-int4-flux-dev-de-distill - A Nunchaku workflow you can use : https://civarchive.com/models/617562
This is a repost from hugging faces, I did not play any role at all in the creation of this epic, groundbreaking model from nyanko7. I just had to post it asap because it is F A N T A S T I C
Flux.dev as we have known until now was a distilled model, meaning it was trained by flux.pro as its teacher. These new models change everything ! This is the first experimental, effectively de-distilled version of Flux.dev, meaning it is much closer to what flux.pro is capable of. And it's just the start !
(NB : This is not stating it was trained with flux.pro - I don't know the exact method)
/!\ READ ALL IMPORTANT STUFF BELOW (AFTER THE EXAMPLES) IN THIS DESCRIPTION OR IT WILL BE A FAIL FOR YOU /!\
EXAMPLES MADE MYSELF
Distilled CFG 8 VS Real CFG 8 - Fixed seed, no prompt change except for the text
Each picture is the first shot on each model : no cherrypicking / cheating
Ear gauge LoRa : "Cute blonde raver young girl smiling, facing viewer, with emo makeup and puffy emo hairstyle, green shiny colored hairstyle. 3argauge, both ears have Ear gauge plug with very large circular hole in her lobe. She is wearing a black hoodie with text "DEDISTILLED MAKES MY LORAS WORK" in golden letters
Cum on face Lora (with dedistilled it actually works anywhere) : "COF, Young woman with white sticky cum on face,white sticky semen on face,white sticky sperm on face,face covered with white sticky cum, face covered with white sticky semen, face covered with white sticky sperm. She is wearing a black hoodie with text letters "DISTILLED" in silver font on it"
No LoRa : "Dutch shot view of a black car speeding away from a massive explosion in the streets of a futuristic city with large buildings, towards the viewer, escaping the blast of the explosion. The motion lines around the car give the impression of speed. The front plate of the car has text "DISTILLED" on it."
No Lora : "A small kitten playing with a ball of yarn is seen through an old wooden window of a rustic house. The scene is cozy, with weathered wooden furniture inside the house and soft afternoon light streaming through the window, casting gentle shadows. Outside, in the distance, a photographer is approaching, camera in hand, ready to capture the playful moment of the kitten. The photographer is wearing a brown jacket and is framed by the soft glow of the golden hour light, adding a sense of warmth and tranquility to the scene. The overall atmosphere is peaceful, with a touch of nostalgia from the vintage setting."
SUMMARY OF THE IMPORTANT STUFF (aka current knowledge about it)
Disclaimer : These models are very new. So, just gathering here what is known about them for now. Please share your experiences in the comment section so we can update this together.
PARAMETERS
You can now forget about Distilled CFG and use real CFG (I've tried up to like 14)
NEVER, ever use it with CFG = 1 - It will automatically be a complete disaster, and most of the time the reason why you don't get results
You SHOULD use at least 40 - 60 steps, depending on the CFG you use.
It will be much longer but SO worth it
Unfortunately the current hyperdev 8-steps Lora doesn't seem to work with it to reduce steps
Dedistilled allows NEGATIVE PROMPTS
BENEFITS
Prompt adherence will be EXCEPTIONAL, even with Loras.
Faces Loras will work better, details will be better, text will be WAY better...
Everything from the prompt will be better basically
/!\ IF YOU DON'T SEE ANY IMPROVEMENT FROM DISTILLED MODELS : VERIFY YOU ARE NOT ACTUALLY USING REAL CFG = 1 WITHOUT KNOWING. NOT FLUX GUIDANCE. IT IS TRICKY /!\
AS ALL WORKFLOWS WERE TAILORED FOR DISTILLED, CONSIDER TRYING WITH FORGE IF IT'S NOT WORKING FOR YOU IN COMFY ?
GUIDELINES FOR USE IN FORGE UI
Works in Forge without any change. Will be loaded as if it was Schnell model, disabling Distilled CFG (cool).
EDIT : uploaded all the new Quants, you should find at least one working for you
If you are new to Forge, make sure you use similar settings :Flux workflow
DeDistilled as the checkpoint
In VAE / Text encoder files, provide the vae (ae.sft / ae.safetensors) + clip_l (or a modified clip) + t5xxl (whatever the quant you are using, fp16, fp8, etc).
Otherwise it won't work as they are NOT bundled in the model files herePlease also set Diffusion in Low Bits = Automatic (FP16 Lora), otherwise you might be in trouble with LoRas. This applies to any checkpoint in Forge.
I then recommend these settings for Dedistilled :
GUIDELINES FOR USE IN COMFY UI
Works in ComfyUI using a pretty standard workflow, the one cited below uses the GGUF Loader, Dual CLIP Loader for t5xxl and clip_l prompts, and KSampler Efficient
Recommended settings for Comfy :
Dual CLIP Loader guidance: 3.5 KSampler cfg: 2 to 10 Steps: 50 to 60 Negative Prompt: Can be left blank or can be provided if needed, will affect the image if provided
Sampler: DDIM or euler Scheduler: beta or exponential
WORFKLOW HERE : https://gist.github.com/dasilva333/87bdd5b5b8ebba5515a9919ede0e3c05
Found this one also on reddit (drag & drop it into Comfy) : https://files.catbox.moe/y99yl7.png
TRAINING AND LORAS
I have just trained myself my first LoRa using De-distilled and guidance = 6, after failing hard with Distilled and guidance = 1. Results are awesome, it basically saved my LoRa. Works great with both De-distilled and distilled (but better with De-distilled).
I will be using it to train from now on.
The first checkpoint fined tune with dedistilled has been posted on civitai here : https://civarchive.com/models/690991/sapianf-nude-men-and-women-for-flux-now-de-distilled
Awaiting answers from the author to update here
Sources
Dedistilled model FP16 : https://huggingface.co/nyanko7/flux-dev-de-distill
Dedistilled model FP8 : https://huggingface.co/MinusZoneAI/flux-dev-de-distill-fp8/tree/main
Dedistilled FP8 GGUG & Q4_ K_M GGUG : https://huggingface.co/TheYuriLover/flux-dev-de-distill-GGUF/tree/main
Description
FAQ
Comments (26)
Can you share your workflow here? The discord link doesn't go anywhere for me. Just trying to make an account took long enough.
Every workflow I've tried to make comes out either very blurry (like upscaling in photoshop from 256p to 4K) or blocky (like upscaling in MS Paint).
Can you try drag & dropping any of @DaSilva or @azimuthalobserver pics directly in ComfyUI ? I think this should import his workflow (take recent ones, he initially was stuck with real CFG 1)
Sorry, my bad, I've edited a lot the description page and the workflow link was supposed to be github, not discord. This is fixed, thanks.
@valentinkognito365 I think I figured it out, when I replaced the Flux sampler node, I also disabled the model sampler node (the max_shift/base_shift one). With that back in place I'm back up and running.
@beitris I just found another workflow on reddit, i will post it. Are you having good results now ?
@valentinkognito365 The prompt adherance is definitely better, though I haven't played around enough to see how far that goes. This seems to lead to much better quality, I'm getting 2nd pass quality with just one ksampler (a glass of water being clear instead of cloudy for instance), but I am having more trouble getting loras up to the level of quality I had with the distilled model. I'm sure it's just a matter of fiddling though.
@beitris yes probably, LoRas become much more powerful from my experience, to the point I feel they barely work on Distilled
@valentinkognito365 Well I'm about ready to give up. The GGUG version works great, but I can't for the life of me get FP8 or FP16 versions to work without gridding artifacts (faint lines that make it look like a bad photocopy). I'm rendering at 1024x1024, I've tried steps 15, 30, 50, 70 & 90, I've move the max/base shift parameters till the sliders broke off. I just cannot get any workflow to work. It's such a shame because even through the artifacts the clarity and detail is obvious, much better than the GGUG versions.
@beitris I'm sorry not be of much help with ComfyUI, that I stopped using a while ago in favor of Forge; Maybe check the sampler and scheduler you are using, and also are u using Ksampler ? If everything fails, why not trying Forge ?
Updated the description with examples and the correct link for the ComfyUI workflow
maybe looks nice,
but why should i use this for 5x longer inage generation time
i can use a FLUX REAL lora and a normal FLUX checkpoint 25steps and CFG=1
up to you, depends on what you are making and if you're using LoRa. I prefer having what I want first shot and perfect than make 10 lousy pictures instead and loose time as well and/or give up. I think just try it and you will see.
This is better mostly for those who train. For the end user: its iight.
I feel like saying "This is better and if you dont think so its probably your fault" is not the right way to go about it lol. "Better" is as subjective as art.
This is not what I intended to say at all, so I revised my wording, thanks. It's more that it seems there are people that have results and think it's awesome, and some others that don't and think it's meh.
I think a great part of this comes from the fact setting it up is a bit new and CFG settings can be very misleading. So I just wish for people to be able to try it with the proper settings before rightfully deciding if it fits them or not.
There's nothing wrong about choosing speed over quality if that's your priority. Personally I think the difference is not light at all, it is massive, but then that's for my use.
I can tell most the commenters are reddit cultists... anyway, thanks for posting. Let's give it time and see the masters in here use this new model to unlock new possibilities.
@ishadowxx The very first checkpoint finetuned with it has been uploaded to civitai : https://civitai.com/models/690991/sapianf-nude-men-and-women-for-flux-now-de-distilled
Unfortunately for me, it crashes on Forge with the dreaded mat1 mat2 exception...
I hope we'll soon have proper anime finetuned checkpoints
Civit gooners never cease to amaze me lol.
Actually there is the Flux Booru checkpoint that is also dedistilled and anime finetuned with 3.5 M pictures
I was very sceptical, but it really works. Unfortunatel at the cost of 5 times more render-time (on a 4090). Example: "four persons, three of them are female, one is male. Females doing this, male is doing that.." With the dedistilled FP8 NOT a single error like 3 males, 4 woman, etc. Quality was good as expected. Nevertheless, I am back to FP8 for speed. Very interesting!
I also have a rtx 4090 (16 gb laptop version). I do 512 x 768 in 1 min 30 (CFG = 8) where as with Distilled FP8 30 steps it must take around 30 - 45 s and with Hyperdev FP8 10 steps around 15 - 20 s. I think its interesting to use low resolution + high res fix/upscale with dedistilled so you have reasonable speed and good res.
@valentinkognito365 I have a full-size 4090. Normally I do 1152x896 20-30 steps FP8 images in 14-24s (no Lora). With this dedistilled FP8 version it was around 1m 30s for same size, same prompt, CFG 2 (great prompt-following) 60 Steps. OK, it is 3-5 times.
@TToby interesting thx. I haven't tried CFG 2. Maybe you can do less steps with CFG 2, according to some test someone has done
@valentinkognito365 https://civitai.com/models/686704?modelVersionId=768547
Works in around 30 steps for me
@rag37735300 Didn't know this one. If using distilled I use hyperdev Q8_0 or hyperdev acorn Q8_0
Very interesting, do things normal flux won't do. It feels like superior technology, like a 'sentient' flux, with deeper understanding of the image and more creative. it is still a work in progress though, slow and buggy. Anyway if you want more speed and sharper images, you can always render with this at lesser resolution (0'5 megapixel for instance) and then upscale using normal flux.
I've just tried like 10+ pictures of my model having a drunk expression with distilled, each time adding more keywords because it wouldn't do it. First try on dedistilled : pretty good. I don't get why people think they save time with distilled because it's faster. You only save time if you're really making basic stuff, otherwise you're WASTING your time making crap when you could have it done first shot.
