cs.CV 2401.09413

POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images

POP-3D introduces a zero-shot open-vocabulary 3D occupancy prediction model leveraging multimodal self-supervised learning, achieving 78% mIoU on nuScenes without manual annotations.

Antonin Vobecky, Oriane Siméoni, David Hurych et al.

2024-01-18 39
cs.CV 2312.14150

DriveLM: Driving with Graph Visual Question Answering

DriveLM integrates graph-structured visual QA with pre-trained VLMs for end-to-end autonomous driving, achieving superior zero-shot generalization on nuScenes and CARLA.

Chonghao Sima, Katrin Renz, Kashyap Chitta et al.

2023-12-22 35