What Matters in Transformers? Not All Attention is Needed
This study introduces a similarity-based pruning method to remove 50% of Attention layers in Llama-2-70B, achieving 48.4% speedup with only 2.4% performance loss.
Shwai He, Guoheng Sun, Zheyu Shen et al.