f2512_fp8 — Cinematic Realism & Emotional Depth
Description: This model is a high-fidelity FP8 merge optimized for 24GB VRAM cards (RTX 3090/4090), designed to deliver cinematic photorealism without compromising on emotional storytelling. Tested across diverse scenarios—from intimate portraits and complex lighting setups to humorous and narrative-driven scenes—f2512_fp8 consistently produces images with lifelike textures, accurate anatomy, and a distinct "fine art" aesthetic.
✨ Key Features:
Masterful Lighting: excels in natural light simulation, including Golden Hour glow, dramatic Chiaroscuro, and complex rim/backlighting. Handles subsurface scattering (SSS) on skin with high accuracy.
Rich Textures: Superior detail rendering in skin pores, vellus hair, wet/dry hair transitions, and diverse materials (fabric, fur, metal, wood).
Anatomy & Proportions: Reliable hand generation, expressive facial micro-expressions, and accurate handling of body diversity (contrast in age, size, and physique).
Versatility: seamlessly adapts between high-end fashion, realistic documentary style, and narrative humor/cinematic storytelling.
Optimized for 3090: The FP8 compression maintains quality while fitting comfortably into 24GB VRAM, allowing for 1024x1536 resolution generation without OOM errors.
⚙️ Recommended Settings:
Sampler:
Euler a(for soft, artistic results) orDPM++ 2M Karras(for sharper details).Steps:
30–40.CFG Scale:
5.5–6.5(Best balance of creativity and adherence).Scheduler:
KL_optimal.Resolution:
1024x1536or1024x1024.Clip Skip:
2(if applicable to the base architecture).
📸 What It Does Best:
Cinematic Portraits: Dramatic lighting, emotional depth, and photorealistic skin.
Texture Studies: Water droplets, wet hair, fabric folds, and organic surfaces.
Narrative Scenes: Complex interactions between subjects (e.g., human-animal bonding) and humorous/situational contexts.
📝 Notes:
For best results, use descriptive prompts focusing on lighting conditions and textures.
The model favors a "fine art photography" aesthetic; adding "photorealistic" or "cinematic" helps anchor the style.
FP8 version is a compressed variant that saves VRAM and generation time with negligible quality loss compared to FP16.
🔗 Merging Method: Merged using [Method, e.g., TIES-Merge / Linear Merge] with FP8 post-processing for efficiency.
Description
FAQ
Comments (5)
love these series of qwen image 2512 models. Glad your still giving this model some love. photorealistic results are getting better.
Is it require the Qwen 2.5 text encoder as well? If so, It would definitely would not fit in VRAM of 4090 RTX without offloading to RAM which would cause painfully slow generation, since the text encoder is almost 8Gb itself, plus you need at least couple GB for the rest stuff.
Everything fits comfortably in a 4080 with 16GB and 64GB of RAM (32 shared with GPU). 2-4s/it in QWEN. Adding RAM is easier. And the latest versions of Comfy handle dynamic memory allocation extremely well. Works with 40GB models, like LTX2.3. I also recommend adding sageattention. A 10-second video in LTX 2.3 with audio can easily be made in 1+ minute.
It produces the best results among other QWEN-based models even on standard settings with 4-step Lora. By the way, the latest versions of Comfy allow you to easily run even 40GB models on 16GB video cards, if the RAM allows it.
By the way, in the 4-step lighting, there are no artifacts. There is a reasonable opinion that the problem of squares in QWEN is due to VAE. Not because of the models. More steps - more artifacts. But with larger steps comes greater precision. Alas. It's one or the other. Waiting for the new QWEN VAE \0/ :)



