Milor123 commited on
Commit
8e38bdc
·
verified ·
1 Parent(s): 62be072

Upload Flux2-Klein-9B-True-V3-INT8-ConvRot.safetensors

Browse files

INT8 quantized model with ConvRot (requires the nodes https://github.com/BobJohnson24/ComfyUI-INT8-Fast/) is much faster than the FP8 model on my RTX 4070 with 12GB VRAM; I believe it is approximately twice as fast.

The model was quantized using the same tools provided by the node developer on GitHub
con

| Configuration | Time per Iteration | Speedup |
| --- | --- | --- |
| INT8 ConvRot + SageAttention | 2.26 s/it | 🚀 2.48x faster |
| FP8 + SageAttention | 5.60 s/it | 1.00x (baseline) |

Note that it does not work well with torch compile (if you are using it, it might be better to disable it or check the author's GitHub)
As for whether it loses quality or not, I believe I understood that it even has better quality as long as it is ConvRot when compared to FP8. I don't know, I don't understand any of this technically; I compared it visually in ComfyUI and didn't see any appreciable loss or notable gain, things look very similar.

Could you add the file or convert it yourself? And also add it to the `readme.md` so people know they can use it; it supposedly works very well on older GPUs like the 3000 and 2000 series. Thank you very much for everything!!

Flux2-Klein-9B-True-V3-INT8-ConvRot.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ff2f21d41dbfde4376658a6c38ac58562d2755137bfd802c9f46b40e899ad32d
3
+ size 9405843376