What's in the Image? A Deep-Dive into the Vision of Vision Language Models
This study conducts a hierarchical attention analysis in VLMs, revealing how middle layers facilitate cross-modal information flow and global feature storage.
Omri Kaduri, Shai Bagon, Tali Dekel