cs.CV 2310.08587

Pseudo-Generalized Dynamic View Synthesis from a Video

Proposes a pseudo-generalized dynamic view synthesis method from monocular videos relying on geometry and temporal depth consistency, outperforming scene-specific approaches.

Xiaoming Zhao, Alex Colburn, Fangchang Ma et al.

2023-10-13 41
cs.CV 2310.06773

Uni3D: Exploring Unified 3D Representation at Scale

Uni3D employs end-to-end pretraining of a 2D ViT, scaled to 1 billion parameters, aligning 3D point clouds with image-text features, achieving state-of-the-art results.

Junsheng Zhou, Jinsheng Wang, Baorui Ma et al.

2023-10-11 37
cs.CV 2310.01415

GPT-Driver: Learning to Drive with GPT

GPT-Driver transforms motion planning into a language modeling task using GPT-3.5, achieving 0.84m average L2 error on nuScenes, with high interpretability.

Jiageng Mao, Yuxi Qian, Junjie Ye et al.

2023-10-03 467 citations 37
cs.CV 2309.17444

LLM-grounded Video Diffusion Models

LVD leverages large language models to generate scene layouts guiding diffusion-based video synthesis, achieving high fidelity without training.

Long Lian, Baifeng Shi, Adam Yala et al.

2023-09-30 33