TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
Proposes TTF-VLA, a training-free temporal fusion method combining pixel difference and attention, boosting robotic manipulation success by 4-8%.
Chenghao Liu, Jiachen Zhang, Chengxuan Li et al.