Making Sense of Vision and Touch: Learning Multimodal Representations for Contact-Rich Tasks
Proposes a variational multimodal representation learning framework integrating RGB-D, force, and proprioception, boosting contact-rich task sample efficiency by 20%.
Michelle A. Lee, Yuke Zhu, Peter Zachares et al.