Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting

TL;DR

Proposes a scalable 2D-to-3D data lifting pipeline combining scale-invariant and scale-aware depth estimation, generating ~2 million realistic 3D scenes, boosting spatial understanding.

cs.CV 🔴 Advanced 2025-07-24 39 views
Xingyu Miao Haoran Duan Quanhao Qian Jiuniu Wang Yang Long Ling Shao Deli Zhao Ran Xu Gongjie Zhang
spatial intelligence depth estimation 3D reconstruction data synthesis multimodal learning

Key Findings

Methodology

This work introduces a 2D-to-3D data lifting framework that integrates scale-invariant and scale-aware depth estimation. Initially, the MoGe model predicts relative depth maps capturing fine details. Subsequently, Metric3D v2 estimates global scale, which is used to calibrate the relative depth. Camera intrinsic and extrinsic parameters are predicted using WildCamera and PerspectiveFields, enabling accurate projection of 2D images into real-world 3D space. The process involves: • Generating relative depth; • Computing scale factors; • Predicting camera parameters; • Projecting pixels via formulas (e.g., equations 3 and 4) to produce point clouds and annotations. The pipeline automatically produces high-quality, annotated 3D data with point clouds, depth maps, camera poses, and labels, supporting diverse tasks.

Key Results

  • Applying this pipeline to COCO and Objects365-v2, about 2 million scenes across 300+ categories were generated, surpassing existing indoor/outdoor datasets like ScanNet. Pretraining models on these datasets improved point cloud instance segmentation mAP by 4-6%, with COCO-3D pretraining boosting ScanNet performance to 28.64% mAP, a 4.34% increase over baseline.
  • In semantic segmentation, models pretrained on COCO-3D achieved higher IoU scores (e.g., 55.83% on ScanNet), demonstrating enhanced generalization. Ablation studies confirmed that combining scale-aware depth and camera parameters significantly improves 3D reconstruction quality.
  • Overall, the generated data effectively enhances 3D perception tasks, validating the approach's robustness and realism, and enabling scalable spatial understanding.

Significance

This research addresses the critical bottleneck of limited large-scale, realistic 3D datasets in spatial AI. By leveraging inexpensive 2D images, it creates diverse, high-fidelity 3D environments, facilitating training of more robust models for robotics, AR/VR, and scene understanding. It reduces reliance on costly hardware-based data collection, broadening application scope and accelerating progress in spatial reasoning. The datasets and methods set a new standard for scalable, realistic 3D data generation, impacting both academia and industry by enabling more generalizable and accurate spatial AI systems.

Technical Contribution

The paper introduces a novel integration of scale-invariant and scale-aware depth estimation models (MoGe and Metric3D v2) within a unified pipeline. It innovatively combines camera parameter prediction with depth calibration, ensuring real-world scale accuracy. The automatic annotation process supports large-scale, diverse scene generation, surpassing prior methods limited to single objects or small scenes. This framework enables cost-effective, high-fidelity 3D data synthesis, fostering advances in multi-task spatial perception and reasoning.

Novelty

This work is the first to systematically combine scale-invariant and scale-aware depth estimation for large-scale 3D scene synthesis from 2D images. Unlike prior approaches relying solely on simulation or AI-generated assets, it leverages real image content with automatic scale calibration, producing realistic, diverse, and scalable 3D datasets. Its ability to generate accurate, annotated 3D environments at scale is a key innovation, filling a major gap in spatial AI data resources.

Limitations

  • The accuracy of generated 3D scenes depends heavily on the quality of depth and camera parameter predictions; complex scenes with occlusion or poor lighting may introduce errors.
  • Current approach mainly handles static scenes; dynamic environments or real-time updates require further development.
  • Computational costs for large-scale data processing and annotation remain significant, though lower than hardware-based methods.

Future Work

Future directions include integrating multi-view and temporal data for dynamic scene modeling, improving robustness in challenging environments, and exploring end-to-end training to enhance accuracy. Extending to real-time applications and dynamic scene understanding will further broaden its impact in robotics, AR/VR, and autonomous systems.

AI Executive Summary

Spatial intelligence is a frontier in AI, aiming for machines to perceive, reason about, and interact with 3D environments. However, progress has been hampered by the scarcity of large-scale, realistic 3D datasets, which are costly and labor-intensive to acquire with traditional sensors like LiDAR or RGB-D cameras. Simulation-based approaches, while inexpensive, suffer from a significant gap between synthetic and real-world scenes, limiting their effectiveness. AI-generated assets tend to lack realism, and sensor data is often confined to specific domains, such as indoor scenes.

To address these challenges, this paper proposes a scalable pipeline that converts single-view images into high-fidelity 3D representations. The core innovation lies in combining scale-invariant and scale-aware depth estimation models—MoGe and Metric3D v2—to generate accurate depth maps with real-world scale. By predicting camera parameters and applying projection formulas, the pipeline reconstructs detailed 3D point clouds and annotations automatically. This approach leverages large, publicly available 2D datasets like COCO and Objects365-v2, producing approximately 2 million scenes across diverse environments, including indoor, outdoor, and mixed scenarios.

Extensive experiments demonstrate that models pretrained on these synthetic datasets outperform baselines in key perception tasks such as point cloud segmentation and 3D question answering. For example, pretraining on COCO-3D improves ScanNet point cloud segmentation mAP by over 4%. The datasets' diversity and realism significantly enhance generalization, validating the method's effectiveness.

This work represents a major step forward in overcoming data limitations in spatial AI. By enabling low-cost, large-scale, and realistic 3D scene generation from 2D images, it opens new avenues for research and applications in robotics, augmented reality, and embodied AI. Future work will focus on dynamic scene modeling and real-time perception, further pushing the boundaries of spatial understanding.

Deep Dive

Plain Language Accessible to non-experts

想象你在厨房里做饭,手里拿着各种食材(图片),但你不知道每样食材的大小和位置。这个研究就像用一种特殊的魔法眼镜(深度估计和相机参数)帮你看清楚每个食材的真实大小和位置,然后用这些信息把所有食材放到正确的碗里(3D场景)。以前,要做到这些需要昂贵的设备和很多时间,但现在只用普通的照片,就能快速还原出厨房的样子。这就像给照片装上了“魔法眼”,让机器变得更聪明,能理解复杂的空间环境。

ELI14 Explained like you're 14

想象你在学校操场上拍了一张照片,你知道里面有很多人、球和椅子,但不知道它们的真实大小和距离。这项研究就像用特别的算法,把照片变成一个3D模型,让你可以看到每个人和东西的真实大小和位置。以前,要做到这一点需要昂贵的设备,比如激光扫描仪,但现在只用普通相机和一些聪明的数学方法,就能做到。它就像给照片装上了“魔法眼”,让机器可以理解空间,就像我们用眼睛看东西一样清楚。

Glossary

Depth Estimation (深度估计)

用算法预测场景中每个像素到相机的距离,分为相对深度和尺度感知深度。

在论文中用于从单幅图像生成3D场景。

Scale Calibration (尺度校准)

调整深度图的尺度,使其反映真实世界的距离关系。

确保生成的3D模型具有正确的大小。

Camera Parameters (相机参数)

描述相机的内外参数,包括焦距、位置和姿态,用于空间投影。

用于将2D图像投影到3D空间。

Point Cloud (点云)

由空间中大量点组成的3D数据,反映场景的几何结构。

作为3D场景的核心表示。

Multimodal Large Language Models (多模态大模型)

结合视觉、语言等多模态信息进行理解和推理的深度学习模型。

用于空间理解和问答任务。

Open Questions Unanswered questions from this research

  • 1 如何进一步提升深度估计在复杂场景中的准确性,尤其是动态环境中的表现仍是挑战。现有模型在遮挡和光照变化时表现不足,未来需结合多视角和动态信息增强模型鲁棒性。
  • 2 生成的3D场景在动态交互和实时感知方面仍有限制,如何实现实时更新和动态建模是未来研究重点。

Abstract

Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of large-scale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and laborious annotation. In this work, we present a scalable pipeline that converts single-view images into comprehensive, scale- and appearance-realistic 3D representations - including point clouds, camera poses, depth maps, and pseudo-RGBD - via integrated depth estimation, camera calibration, and scale calibration. Our method bridges the gap between the vast repository of imagery and the increasing demand for spatial scene understanding. By automatically generating authentic, scale-aware 3D data from images, we significantly reduce data collection costs and open new avenues for advancing spatial intelligence. We release two generated spatial datasets, i.e., COCO-3D and Objects365-v2-3D, and demonstrate through extensive experiments that our generated data can benefit various 3D tasks, ranging from fundamental perception to MLLM-based reasoning. These results validate our pipeline as an effective solution for developing AI systems capable of perceiving, understanding, and interacting with physical environments.

cs.CV cs.AI