A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
A unified pizza–sandwich–lasagna framework shows that 2× KV-cache reduction usually preserves quality, while aggressive compression favors sandwich-middle.
You Wu, Haoyi Wu, Kewei Tu