For the speed!
“好就是快,快就是好 —— 沃兹基・硕德”
==================================
Anima Turbo INT8ConvRot Quantized | Low VRAM | Minimal visual loss | Fast Anime Generation
中文介绍
INT8ConvRot 为 AI 绘画推理加速带来全新方案。本次我将该量化方案与高速 Anima Turbo 模型进行整合,借助 INT8ConvRot 量化实现推理速度再度提升,整体生成性能优于常规 SDXL 模型。
量化范围不仅包含 UNet,qwen_3_06_base 同样采用 INT8 量化。需要说明:该优化存在小幅取舍,生成图像会轻微柔和发虚。为此我搭配了同经 INT8 量化的 hdrVAEAnimaKrea2QWN VAE,尽可能弥补画质损耗。
最终成品实现了均衡表现:版本新颖、推理迅速、文件体积更小,相较于原版权重,画质损失控制在极低范围。
搭载 RTX 20 系及以上 NVIDIA 显卡,开启 INT8ConvRot 可获得50% 以上推理提速。
RTX 10 系显卡用户注意:受硬件限制无法启用 INT8ConvRot 加速,但 Anima Turbo 原生速度依旧出色。同时 INT8 量化大幅降低显存占用,让低显存显卡也能够运行当前主流前沿二次元模型。
English Description
INT8ConvRot unlocks new performance possibilities for AI image generation. I've integrated this quantization scheme with the ultra-fast Anima Turbo model for additional inference acceleration, outperforming standard SDXL models in generation speed.
Quantization covers not only the UNet module but also qwen_3_06_base with INT8 precision. There is a minor tradeoff: outputs may appear slightly softened. To offset this visual degradation, the VAE has been swapped for the INT8-quantized hdrVAEAnimaKrea2QWN.
This build achieves a well-rounded balance: cutting-edge architecture, blazing-fast inference, reduced file size, with only minimal quality loss compared to the original weights.
RTX 20-series and newer NVIDIA GPUs can gain over 50% faster inference speed with INT8ConvRot enabled.
For RTX 10-series users: INT8ConvRot acceleration is unavailable due to hardware limitations. Fortunately, Anima Turbo retains its native fast inference. Furthermore, INT8 quantization lowers VRAM usage, allowing low-VRAM GPUs to run state-of-the-art anime models.
Description
FAQ
Comments (2)
Some generation suggestions would be appreciated, cause using the usual Anime/Turbo setting results in nothing but static.
Could you make one for this model?
https://civitai.red/models/2855007/anima-29b?modelVersionId=3224434

