Source https://huggingface.co/city96/FLUX.1-schnell-gguf/tree/main by city96
This is a direct GGUF conversion of Flux.1-schnell.
The model files can be used with the ComfyUI-GGUF custom node.
Also working with Forge since the latest commit!
☕ Buy me a coffee: https://ko-fi.com/ralfingerai
🍺 Join my discord: https://discord.com/invite/pAz4Bt3rqb
Description
FAQ
Comments (18)
what is this ks stuff why is there no info :/
honestly, the quality of gguf quantized models from schnell is MEH, it is better to use quantization of dev model, and the fastest one if NF4 - and it gives way better quality than Schnell ... ;)
I think FS actually produces better looking images. Yes, I know, opinions are subjective. Dev's output has this overly smooth corporate look that I don't like. I have confirmed this with an on-site generator many times. And I have actually tested some dev versions but ended up removing them due to license inconsistencies. And I'm glad I did. I wouldn't want to use dev versions anymore anyways. For my rtx 3060, FS Q5_K_S is just the perfect combination of quality and speed.
Sorry, Q4 is still ahead of all your gguf. It is evident that nobody could post any good images created with your other models.
Please explain in detail about the vae, clip, text encoder and sampler along with cfg and generation steps for Forge and exclusive workflow files for each model to use in ComfyUI.
Regards
What all? What other Qs are there than gguf's? Check under Q5.KS. I argue that Schnell makes better images than Dev. And it has become very evident to me.
@PirateGirl Yes, it does great. If you have not replied, I would have missed using Q5.KS. It is one of your best works. Actually, my comment was, 'your earlier model does far better than the latest one'. Still, I expect wonderful models from you.
Regards
My breakdown of all Flux Schnell GGUF Q models that I've tested with my RTX 3060 12Gb:
Q2_K - I can't recommend this to anyone. It's on the edge of coherency and pixelated mess. Images done with this have all the elements of better versions in it, but lacks coherent details and images don't look any good at all. I think you could run this with 4Gb vram, but it's really not worth it. I suggest you to try Q3_K_S instead.
edit: Bernoulli Q2_K is a lot better than city96's version. It actually makes coherent images unlike this version here. Probably runs well on 4Gb vram and 16Gb system just fine.
Q3_K_S - It's on edge of acceptable quality and coherency. Absolute minimum version I would run. It's much better than Q2_K and produces acceptable images, but lack some contrast in some cases compared to better versions, but loses it's coherency at higher resolutions. I guess you could run this confortably with 4Gb vram with a bit of swapping with the system ram. If you have 6Gb vram, I believe you could run this quite comfortably.
Q4, Q4_1, Q4_K_S - Things get better here. I have not tested these, but from what I've seen of the images posted, the quality is somewhere around NF4 and FP8 versions. I guess you could run these versions with at least 6Gb vram with some swapping with system ram. If you have 8Gb vram, I believe one of these versions is optimal for your gpu.
Q5, Q5_1, Q5_K_S - All of these produce excellent quality images. Easily surpasses FP8 quality. If you have 12Gb vram then one of these is optimal for your gpu. My personal favorite is Q5_K_S that offers optimal quality and speed. And it leaves some vram for the background processes for you to browse youtube or watch HD videos at the same time without losing generation speed. Like someone said in reddit post, it's Kood enough and Small.
O6_K - Excellent quality. However, not much visible difference to Q5 versions imho. And it's a bit slower than Q5 versions. It's great for 12Gb vram for sure. I just choose to use Q5_K_S instead due to personal preferences.
Q8 - edit: The Ultimate version. I just tested this and it still fits on 12Gb vram and it runs nicely without having to rely on swapping with system ram.
edit:
In summary: If you have a GPU with -
4Gb - Q3_K_S, or maybe some of the Q4 versions if you don't mind swapping.
6Gb - Q4 should fit in 6Gb vram. Q4_1, Q4_K_S if you don't mind some swapping. Feel free to experiment how Q5 perform in your system.
8Gb - Q5, Q5_K_S. You're golden with either version. Q5_1 does not fit in 8Gb, you may have to rely on swapping.
10Gb - Is there even 10Gb Nvidia gpus? Maybe laptop versions? If you have 10Gb, then you can run Q6_K.
12Gb or more - Q8. You can run any GGUF version comfortable. If you want to experiment and you have enough ssd space, then you could also keep some of the Q5 versions or Q6K. They all have a slightly different looking output, and depending on the prompt, any version of them is good enough and can randomly produce the best looking result for a given prompt with the same exact seed.
If you disagree something with my take, please reply. All opinions are appreciated. We all just want to create better looking images.
I have 16Gb and failed to run Q8 consistently without maxing out VRAM. Perhaps some Loras take up some of it. I'm giving Q5KS a try now.
@Phraxas How much system ram you have? I have 64Gb and Q8 run just well. I guess if you have 32Gb or less, you could run in to troubles with Q8. Q5KS is a great choice as well. Excellent quality, and very consistent. Hardly any visible difference compared to Q8.
"10Gb - Is there even 10Gb Nvidia gpus? Maybe laptop versions? If you have 10Gb, then you can run Q6_K."
The RTX 3080 (non Ti) has 10GB of VRAM.
one doubt, i can train loras from this gguf with ai toolkit?
Thanks for the write, good job! Would love to sticky that comment, but there is no option :(
To get the loras to work for Q6k and Q8, you need to have the t5xxl_fp16.safetensors vae. and to set your Diffusion in Low Bits option to - Automatic (fp16 LORA)
upload with zip? cause extracting gets error every time no matter what. and yes it was downloaded correctly
What about licence? Is Author blocked selling images? Schnell usually open for selling
Here is the license file (Apache 2.0): https://huggingface.co/datasets/choosealicense/licenses/blob/main/markdown/apache-2.0.md
Details
Available On (1 platform)
Same model published on other platforms. May have additional downloads or version variants.
