HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
HOLODECK 2.0 integrates vision-language models with 3D generative models for diverse, styled scene creation and interactive editing.
Zixuan Bian, Ruohan Ren, Yue Yang et al.
HOLODECK 2.0 integrates vision-language models with 3D generative models for diverse, styled scene creation and interactive editing.
Zixuan Bian, Ruohan Ren, Yue Yang et al.
Proposed a method using Calibration Tokens to extend monocular depth estimators to fisheye cameras, significantly improving depth estimation accuracy.
Rit Gangopadhyay, Jung-Hee Kim, Xien Chen et al.
BridgeDepth uses bidirectional latent alignment with cross-attention transformers to fuse monocular and stereo depth, reducing zero-shot error by >40%.
Tongfan Guan, Jiaxin Guo, Chen Wang et al.
Event camera-based UAV detection leveraging sparse asynchronous data, achieving low latency and high robustness in challenging conditions.
Gabriele Magrini, Lorenzo Berlincioni, Luca Cultrera et al.
VITAL integrates visual tools with multimodal chain-of-thought for long video reasoning, achieving state-of-the-art results in QA and temporal grounding.
Haoji Zhang, Xin Gu, Jiawen Li et al.
HPSv3 uses a wide-spectrum dataset and VLM-based model with Bayesian ranking to improve human preference evaluation accuracy.
Yuhang Ma, Yunhao Shui, Xiaoshi Wu et al.
StreamAgent combines future event prediction and hierarchical KV-cache to enable proactive streaming video understanding with 15% accuracy boost.
Haolin Yang, Feilong Tang, Lingxiao Zhao et al.
SPFSplat achieves SOTA performance in 3D Gaussian splatting without pose supervision using sparse views.
Ranran Huang, Krystian Mikolajczyk
Proposes a Transformer-based unified framework achieving 85% mIoU in multimodal referring segmentation.
Henghui Ding, Song Tang, Shuting He et al.
PixNerd introduces an end-to-end pixel diffusion model using neural fields, achieving 2.15 FID on ImageNet 256×256 without VAE or cascades.
Shuai Wang, Ziteng Gao, Chenhui Zhu et al.
Introduced Alpha-CLIP and SMS score to improve OV-3DIS, achieving 32.7% mAP on ScanNet200.
Sanghun Jung, Jingjing Zheng, Ke Zhang et al.
X-Omni integrates reinforcement learning into discrete autoregressive models, significantly improving image quality and instruction-following capabilities, enabling unified multimodal generation.
Zigang Geng, Yibing Wang, Yeyao Ma et al.
Deep learning-based endoscopic depth estimation employs monocular and stereo networks, utilizing synthetic and real datasets, achieving sub-millimeter accuracy.
Ke Niu, Zeyun Liu, Xue Feng et al.
FROSS leverages 2D scene graphs and Gaussian models for real-time 3D semantic scene graph generation, achieving significant speedup over traditional methods.
Hao-Yu Hou, Chun-Yi Lee, Motoharu Sonogashira et al.
RGA3 integrates STOM and SAM2 for object-centric video reasoning, achieving state-of-the-art results in video QA and segmentation benchmarks.
Haochen Wang, Qirui Chen, Cilin Yan et al.
Hierarchical multi-platform GUI benchmark MMBench-GUI evaluates content understanding, element grounding, task automation, and collaboration, introducing EQA for efficiency.
Xuehui Wang, Zhenyu Wu, JingJing Xie et al.
Proposes DINO-world, a latent space video predictor trained on 60M videos, outperforming SOTA in prediction and physics understanding.
Federico Baldassarre, Marc Szafraniec, Basile Terver et al.
TAGC automatically tunes gamma from image color statistics to enhance low-light photos without manual settings.
Ghufran Abualhail Alhamzawi, Ali Saeed Alfoudi, Ali Hakem Alsaeedi et al.
NEAT adapts only normalization layers, resolving VLM negation shifts with under 0.01% trainable parameters.
Haochen Han, Alex Jinpeng Wang, Fangming Liu et al.
Proposes a scalable 2D-to-3D data lifting pipeline combining scale-invariant and scale-aware depth estimation, generating ~2 million realistic 3D scenes, boosting spatial understanding.
Xingyu Miao, Haoran Duan, Quanhao Qian et al.