cs.CL 2410.09426

FlatQuant: Flatness Matters for LLM Quantization

FlatQuant employs learnable affine transformations with Kronecker decomposition to achieve sub-1% accuracy loss in W4A4 quantization of LLaMA-3-70B, boosting speed by 2.3x.

Yuxuan Sun, Ruikang Liu, Haoli Bai et al.

2024-10-12 29