Krea 2 Turbo INT8 TensorWise — AMD ROCm Optimized for ComfyUI
⚠️ Running the models alone is not enough to unlock full AMD optimization. Follow all implementation instructions below to build a fully AMD ROCm-optimized ComfyUI stack.
⚡ Krea 2 Turbo optimized for AMD ROCm, ComfyUI, FlashAttention, Triton/AITER, and low-memory inference.
Release video
This release is the result of a full Krea 2 optimization project focused on making Krea 2 Turbo and Style Reference substantially more practical on AMD GPUs.
It combines:
Krea 2 Turbo INT8 TensorWise diffusion
Krea 2 Turbo Style Reference with the LoRA pre-fused before INT8 quantization
Krea Qwen3-VL-4B INT8 TensorWise text encoder
ROCm FlashAttention using AMD Triton/AITER
Qwen VAE Triton W8A8 acceleration
optimized Text-to-Image and Reference-to-Image ComfyUI workflows
The diffusion checkpoints use native ComfyUI int8_tensorwise quantization and do not depend on ConvRot.
🚀 Performance
On the same Ubuntu AMD ROCm host, using a 1368×768 / 8-step reference-to-image workflow:
ConfigurationFirst workflowWarm workflowWarm samplingStock-style ComfyUI INT8 + runtime Style Reference LoRA + PyTorch cross-attention242.4s81.7s71.0sFully optimized project stack85.0s75.0s65.0s
Measured gains
First workflow: 242.4s → 85.0s
64.9% lower end-to-end time
Warm workflow: 81.7s → 75.0s
8.2% faster
Warm sampling: 71.0s → 65.0s
8.4% faster
Iteration time: 8.87s/it → 8.16s/it
8.0% lower
Cold-start results include initialization and compilation overhead. Warm sampling and sec/it are better indicators of steady-state inference performance.
FlashAttention matters
On the tested system, switching from PyTorch cross-attention to ROCm FlashAttention with AMD Triton/AITER reduced:
warm generation: 82.5s → 75.0s
warm sampling: 71.0s → 65.0s
warm sec/it: 8.915 → 8.160
ComfyUI's separate --enable-triton-backend flag produced no meaningful additional gain in this particular A/B test, so it is not required for this release.
📦 Which files do I need?
The full release contains three optimized model components.
1. Krea 2 Turbo INT8 TensorWise
krea2_turbo_int8_tensorwise_v1.safetensorsUse this for normal Krea 2 Turbo text-to-image generation.
all 224 main transformer GEMM weights converted to INT8 TensorWise
no ConvRot
approximately 13.16 GiB
approximately 46% smaller than the original BF16 transformer
Place it in:
ComfyUI/models/diffusion_models/2. Krea 2 Turbo Style Reference — Fused INT8 TensorWise
krea2_turbo_style_ref_fused_int8_tensorwise_v1.safetensorsUse this for Krea 2 Style Reference / reference-to-image workflows.
The Ostris Style Reference LoRA has already been fused into the BF16 Krea 2 Turbo model before quantization.
That means:
✅ No runtime Style Reference LoRA loading
✅ 224 fused transformer weights quantized to INT8 TensorWise
✅ 32 fused txtfusion weights retained in BF16
✅ Reference conditioning remains available
Important
Do not load the Style Reference LoRA again.
The LoRA is already fused into this checkpoint.
You still need the Krea 2 Ostris Edit/reference-conditioning path:
https://github.com/ostris/ComfyUI-Krea2-Ostris-Edit
Original Style Reference project:
https://huggingface.co/ostris/krea2_turbo_style_reference
Place the fused model in:
ComfyUI/models/diffusion_models/3. Krea Qwen3-VL-4B INT8 TensorWise Text Encoder
krea2_qwen3vl_4b_int8_tensorwise_v1.safetensorsThis converts 356 attention, MLP, and vision projection weights to native ComfyUI INT8 TensorWise.
Final size is approximately 4.50 GiB, around 45% smaller than the verified BF16 source.
Place it in:
ComfyUI/models/text_encoders/Load it using ComfyUI CLIPLoader with:
type: krea2
device: default⚡ Recommended AMD ROCm setup
For the best performance measured during this project, use:
AMD ROCm-capable GPU
current compatible ROCm + PyTorch environment
ROCm/flash-attention
AMD Triton/AITER FlashAttention backend
Qwen VAE Triton W8A8
the optimized models from this release
Set:
export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUELaunch ComfyUI with:
python main.py \
--use-flash-attention \
--disable-xformersDo not blindly copy ROCm, PyTorch, Triton, GFX override, or environment-variable settings from another AMD machine.
GPU architecture and ROCm support should be detected first.
🧠 Copy/Paste AMD ROCm Installation Prompt
I created a complete installation prompt that can be pasted into ChatGPT, Claude, Gemini, or another capable LLM with web access.
It guides you interactively through:
GPU and GFX architecture discovery
AMD driver validation
ROCm/HIP setup
Python 3.12 environment creation
compatible PyTorch + TorchVision selection
ROCm FlashAttention installation
AITER and Triton validation
ComfyUI installation
environment-variable tuning
final GPU and generation validation
ROCm FlashAttention + AMD Triton/AITER installation prompt:
⚡ Qwen VAE Triton W8A8
Krea 2's Qwen Image VAE can be a major cold-start and resolution-change bottleneck.
My Qwen VAE Triton W8A8 custom node accelerates validated VAE decoder layers while keeping quality-sensitive layers in native precision.
Docker isolation benchmarks showed approximately 20–21% lower VAE execution cost in the tested first-run and resolution-change cases.
GitHub:
https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton
ComfyUI Manager search:
Qwen VAE Triton W8A8Recommended preset:
AggressiveConnect it as:
VAELoader
↓
Patch Qwen VAE Triton W8A8
↓
VAEDecodeUse the normal:
qwen_image_vae.safetensorsThe VAE itself is not included in this model release.
🧩 Optimized ComfyUI Workflows
Validated Text-to-Image and Style-Reference workflows are available here:
https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton/workflows
Full Hugging Face release:
https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton
The Hugging Face repository also contains:
benchmarks
validation evidence
model provenance
tensor policies
checksums
licenses
optimized workflows
detailed AMD ROCm installation guidance
🛠️ Basic Installation
Place the optimized models here:
ComfyUI/
└── models/
├── diffusion_models/
│ ├── krea2_turbo_int8_tensorwise_v1.safetensors
│ └── krea2_turbo_style_ref_fused_int8_tensorwise_v1.safetensors
│
└── text_encoders/
└── krea2_qwen3vl_4b_int8_tensorwise_v1.safetensorsFor the optimized VAE path, also install:
https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton
For Style Reference, install:
https://github.com/ostris/ComfyUI-Krea2-Ostris-Edit
🔬 Technical Notes
Diffusion quantization
The primary Krea 2 Turbo model contains 28 main transformer blocks.
Eight large GEMM families per block are quantized:
attn.wq
attn.wk
attn.wv
attn.gate
attn.wo
mlp.gate
mlp.up
mlp.downTotal:
28 × 8 = 224 INT8 TensorWise weightsEach quantized layer uses native ComfyUI TensorWise metadata:
.weight INT8
.weight_scale FP32
.comfy_quant {"format":"int8_tensorwise"}Style Reference fusion
The Style Reference LoRA was fused at multiplier 1.0 while the base model was still BF16:
W_fused = W_base + (B @ A)The fused transformer weights were then quantized.
This avoids:
INT8 → dequantize → merge LoRA → requantize⚠️ Compatibility
This project was developed and benchmarked primarily for AMD ROCm + ComfyUI.
Performance varies according to:
GPU architecture
ROCm version
PyTorch version
FlashAttention/AITER/Triton versions
ComfyUI revision
resolution
workflow
memory configuration
kernel compilation state
Benchmark numbers are measurements from the tested environments, not universal guarantees.
📜 License & Attribution
This is an independent community derivative release.
It is not an official Krea, Qwen/Alibaba, Ostris, ComfyUI, AMD, or KJnodes product, and it is not endorsed by those parties.
Krea 2 Turbo / Style Reference
Krea 2 Turbo and its derivatives are governed by the Krea 2 Community License Agreement.
Please review the license before redistribution or commercial deployment:
https://www.krea.ai/krea-2-licensing
Official Krea 2 Turbo:
https://huggingface.co/krea/Krea-2-Turbo
Qwen3-VL
The upstream Qwen3-VL-4B-Instruct text encoder is published under Apache License 2.0:
https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct
The Hugging Face release contains the complete license copies and modification/attribution notices.
🚀 More Projects / Contact
If this project saved you VRAM, inference time, or days of debugging an AMD Krea 2 setup, check out my other optimization work.
LinkedIn — Allen Barnard
https://www.linkedin.com/in/allen-b-3a35505a/
YouTube — PuppetVisionAI
https://www.youtube.com/@PuppetVisionAI
Website
https://puppetvision.nl
GitHub / More Projects
https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton#work-with-me--more-projects
I’m interested in AI Systems Engineering opportunities involving model optimization, inference systems, GPU acceleration, quantization, ROCm/CUDA, Triton, and generative-AI infrastructure.
Links
Full model release:
https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton
Optimized workflows:
https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton/workflows
ROCm FlashAttention / AMD Triton installation prompt:
https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton/docs/ROCM_FLASH_ATTENTION_INSTALLATION_PROMPT.md
Qwen VAE Triton W8A8:
https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton
Release video:
https://youtu.be/fD-J2CYetIg
Description
Comments (2)
Something in here optimized for AMD! Damn thats rare! Thanks!
You are welcome, I got Minimax-H3 release ready but unlike the comfy org and Civitai's official models, my models are not linked to exemption under the community licence agreement. So there is a very rare licensing blocker for that release, but I am actively attempting to kiss Minimax's ass while asking them exemption.




