CivArchive
    Krea 2 Turbo optimized for AMD ROCm - INT8 Style Ref 2 img v1.0
    NSFW
    Preview 144317103
    Preview 144317189
    Preview 144317190
    Preview 144317194
    Preview 144317195

    Krea 2 Turbo INT8 TensorWise — AMD ROCm Optimized for ComfyUI

    ⚠️ Running the models alone is not enough to unlock full AMD optimization. Follow all implementation instructions below to build a fully AMD ROCm-optimized ComfyUI stack.

    ⚡ Krea 2 Turbo optimized for AMD ROCm, ComfyUI, FlashAttention, Triton/AITER, and low-memory inference.

    Release video


    This release is the result of a full Krea 2 optimization project focused on making Krea 2 Turbo and Style Reference substantially more practical on AMD GPUs.

    It combines:

    • Krea 2 Turbo INT8 TensorWise diffusion

    • Krea 2 Turbo Style Reference with the LoRA pre-fused before INT8 quantization

    • Krea Qwen3-VL-4B INT8 TensorWise text encoder

    • ROCm FlashAttention using AMD Triton/AITER

    • Qwen VAE Triton W8A8 acceleration

    • optimized Text-to-Image and Reference-to-Image ComfyUI workflows

    The diffusion checkpoints use native ComfyUI int8_tensorwise quantization and do not depend on ConvRot.


    🚀 Performance

    On the same Ubuntu AMD ROCm host, using a 1368×768 / 8-step reference-to-image workflow:

    ConfigurationFirst workflowWarm workflowWarm samplingStock-style ComfyUI INT8 + runtime Style Reference LoRA + PyTorch cross-attention242.4s81.7s71.0sFully optimized project stack85.0s75.0s65.0s

    Measured gains

    • First workflow: 242.4s → 85.0s

      • 64.9% lower end-to-end time

    • Warm workflow: 81.7s → 75.0s

      • 8.2% faster

    • Warm sampling: 71.0s → 65.0s

      • 8.4% faster

    • Iteration time: 8.87s/it → 8.16s/it

      • 8.0% lower

    Cold-start results include initialization and compilation overhead. Warm sampling and sec/it are better indicators of steady-state inference performance.

    FlashAttention matters

    On the tested system, switching from PyTorch cross-attention to ROCm FlashAttention with AMD Triton/AITER reduced:

    • warm generation: 82.5s → 75.0s

    • warm sampling: 71.0s → 65.0s

    • warm sec/it: 8.915 → 8.160

    ComfyUI's separate --enable-triton-backend flag produced no meaningful additional gain in this particular A/B test, so it is not required for this release.


    📦 Which files do I need?

    The full release contains three optimized model components.

    1. Krea 2 Turbo INT8 TensorWise

    krea2_turbo_int8_tensorwise_v1.safetensors

    Use this for normal Krea 2 Turbo text-to-image generation.

    • all 224 main transformer GEMM weights converted to INT8 TensorWise

    • no ConvRot

    • approximately 13.16 GiB

    • approximately 46% smaller than the original BF16 transformer

    Place it in:

    ComfyUI/models/diffusion_models/

    2. Krea 2 Turbo Style Reference — Fused INT8 TensorWise

    krea2_turbo_style_ref_fused_int8_tensorwise_v1.safetensors

    Use this for Krea 2 Style Reference / reference-to-image workflows.

    The Ostris Style Reference LoRA has already been fused into the BF16 Krea 2 Turbo model before quantization.

    That means:

    ✅ No runtime Style Reference LoRA loading
    ✅ 224 fused transformer weights quantized to INT8 TensorWise
    ✅ 32 fused txtfusion weights retained in BF16
    ✅ Reference conditioning remains available

    Important

    Do not load the Style Reference LoRA again.

    The LoRA is already fused into this checkpoint.

    You still need the Krea 2 Ostris Edit/reference-conditioning path:

    https://github.com/ostris/ComfyUI-Krea2-Ostris-Edit

    Original Style Reference project:

    https://huggingface.co/ostris/krea2_turbo_style_reference

    Place the fused model in:

    ComfyUI/models/diffusion_models/

    3. Krea Qwen3-VL-4B INT8 TensorWise Text Encoder

    krea2_qwen3vl_4b_int8_tensorwise_v1.safetensors

    This converts 356 attention, MLP, and vision projection weights to native ComfyUI INT8 TensorWise.

    Final size is approximately 4.50 GiB, around 45% smaller than the verified BF16 source.

    Place it in:

    ComfyUI/models/text_encoders/

    Load it using ComfyUI CLIPLoader with:

    type: krea2
    device: default

    ⚡ Recommended AMD ROCm setup

    For the best performance measured during this project, use:

    • AMD ROCm-capable GPU

    • current compatible ROCm + PyTorch environment

    • ROCm/flash-attention

    • AMD Triton/AITER FlashAttention backend

    • Qwen VAE Triton W8A8

    • the optimized models from this release

    Set:

    export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE

    Launch ComfyUI with:

    python main.py \
      --use-flash-attention \
      --disable-xformers

    Do not blindly copy ROCm, PyTorch, Triton, GFX override, or environment-variable settings from another AMD machine.

    GPU architecture and ROCm support should be detected first.


    🧠 Copy/Paste AMD ROCm Installation Prompt

    I created a complete installation prompt that can be pasted into ChatGPT, Claude, Gemini, or another capable LLM with web access.

    It guides you interactively through:

    • GPU and GFX architecture discovery

    • AMD driver validation

    • ROCm/HIP setup

    • Python 3.12 environment creation

    • compatible PyTorch + TorchVision selection

    • ROCm FlashAttention installation

    • AITER and Triton validation

    • ComfyUI installation

    • environment-variable tuning

    • final GPU and generation validation

    ROCm FlashAttention + AMD Triton/AITER installation prompt:

    https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton#copy-and-paste-implementation-instructions-for-chatgpt-or-other-flagship-llm


    ⚡ Qwen VAE Triton W8A8

    Krea 2's Qwen Image VAE can be a major cold-start and resolution-change bottleneck.

    My Qwen VAE Triton W8A8 custom node accelerates validated VAE decoder layers while keeping quality-sensitive layers in native precision.

    Docker isolation benchmarks showed approximately 20–21% lower VAE execution cost in the tested first-run and resolution-change cases.

    GitHub:

    https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton

    ComfyUI Manager search:

    Qwen VAE Triton W8A8

    Recommended preset:

    Aggressive

    Connect it as:

    VAELoader
       ↓
    Patch Qwen VAE Triton W8A8
       ↓
    VAEDecode

    Use the normal:

    qwen_image_vae.safetensors

    The VAE itself is not included in this model release.


    🧩 Optimized ComfyUI Workflows

    Validated Text-to-Image and Style-Reference workflows are available here:

    https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton/workflows

    Full Hugging Face release:

    https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton

    The Hugging Face repository also contains:

    • benchmarks

    • validation evidence

    • model provenance

    • tensor policies

    • checksums

    • licenses

    • optimized workflows

    • detailed AMD ROCm installation guidance


    🛠️ Basic Installation

    Place the optimized models here:

    ComfyUI/
    └── models/
        ├── diffusion_models/
        │   ├── krea2_turbo_int8_tensorwise_v1.safetensors
        │   └── krea2_turbo_style_ref_fused_int8_tensorwise_v1.safetensors
        │
        └── text_encoders/
            └── krea2_qwen3vl_4b_int8_tensorwise_v1.safetensors

    For the optimized VAE path, also install:

    https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton

    For Style Reference, install:

    https://github.com/ostris/ComfyUI-Krea2-Ostris-Edit


    🔬 Technical Notes

    Diffusion quantization

    The primary Krea 2 Turbo model contains 28 main transformer blocks.

    Eight large GEMM families per block are quantized:

    attn.wq
    attn.wk
    attn.wv
    attn.gate
    attn.wo
    mlp.gate
    mlp.up
    mlp.down

    Total:

    28 × 8 = 224 INT8 TensorWise weights

    Each quantized layer uses native ComfyUI TensorWise metadata:

    .weight        INT8
    .weight_scale  FP32
    .comfy_quant   {"format":"int8_tensorwise"}

    Style Reference fusion

    The Style Reference LoRA was fused at multiplier 1.0 while the base model was still BF16:

    W_fused = W_base + (B @ A)

    The fused transformer weights were then quantized.

    This avoids:

    INT8 → dequantize → merge LoRA → requantize

    ⚠️ Compatibility

    This project was developed and benchmarked primarily for AMD ROCm + ComfyUI.

    Performance varies according to:

    • GPU architecture

    • ROCm version

    • PyTorch version

    • FlashAttention/AITER/Triton versions

    • ComfyUI revision

    • resolution

    • workflow

    • memory configuration

    • kernel compilation state

    Benchmark numbers are measurements from the tested environments, not universal guarantees.


    📜 License & Attribution

    This is an independent community derivative release.

    It is not an official Krea, Qwen/Alibaba, Ostris, ComfyUI, AMD, or KJnodes product, and it is not endorsed by those parties.

    Krea 2 Turbo / Style Reference

    Krea 2 Turbo and its derivatives are governed by the Krea 2 Community License Agreement.

    Please review the license before redistribution or commercial deployment:

    https://www.krea.ai/krea-2-licensing

    Official Krea 2 Turbo:

    https://huggingface.co/krea/Krea-2-Turbo

    Qwen3-VL

    The upstream Qwen3-VL-4B-Instruct text encoder is published under Apache License 2.0:

    https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct

    The Hugging Face release contains the complete license copies and modification/attribution notices.


    🚀 More Projects / Contact

    If this project saved you VRAM, inference time, or days of debugging an AMD Krea 2 setup, check out my other optimization work.

    LinkedIn — Allen Barnard
    https://www.linkedin.com/in/allen-b-3a35505a/

    YouTube — PuppetVisionAI
    https://www.youtube.com/@PuppetVisionAI

    Website
    https://puppetvision.nl

    GitHub / More Projects
    https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton#work-with-me--more-projects

    I’m interested in AI Systems Engineering opportunities involving model optimization, inference systems, GPU acceleration, quantization, ROCm/CUDA, Triton, and generative-AI infrastructure.


    Full model release:
    https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton

    Optimized workflows:
    https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton/workflows

    ROCm FlashAttention / AMD Triton installation prompt:
    https://huggingface.co/PuppetVision/krea-2-amd-rocm-optimized-comfy-triton/docs/ROCM_FLASH_ATTENTION_INSTALLATION_PROMPT.md

    Qwen VAE Triton W8A8:
    https://github.com/AllenCraigBarnard/ComfyUI-Qwen-VAE-Triton

    Release video:
    https://youtu.be/fD-J2CYetIg

    Description

    Checkpoint
    Krea 2

    Details

    Downloads
    32
    Platform
    CivitAI
    Platform Status
    Available
    Created
    9/30/2026
    Updated
    10/11/2026
    Deleted
    -

    Files

    krea2TurboOptimizedFor_int8StyleRef2ImgV10.safetensors