cs.CV 2311.12871

An Embodied Generalist Agent in 3D World

LEO, a multi-modal embodied agent trained via 3D VL alignment and instruction tuning, significantly advances 3D scene understanding and interaction.

Jiangyong Huang, Silong Yong, Xiaojian Ma et al.

2023-11-18 42
cs.CL 2311.09144

Grounding Gaps in Language Model Generations

Proposes grounding acts as metrics to evaluate dialogue, finds LLMs generate 77.5% fewer grounding acts than humans, and preference tuning reduces these acts further.

Omar Shaikh, Kristina Gligorić, Ashna Khetan et al.

2023-11-16 39