Training-Free Hashing-Based Attention via Binary Principal Components
BinaryPC employs data-aware binary principal components for training-free sparse attention, maintaining accuracy and boosting decoding throughput by 3.56×.
Daohai Yu, Zhanpeng Zeng, Keyu Chen et al.