ImaginateAR: AI-Assisted In-Situ Authoring in Augmented Reality

TL;DR

ImaginateAR combines scene understanding, 3D generation, and LLMs for speech-driven outdoor AR authoring, achieving 30% faster 3D generation.

cs.HC 🟡 Intermediate 2025-04-30 46 views
Jaewook Lee Filippo Aleotti Diego Mazala Guillermo Garcia-Hernando Sara Vicente Oliver James Johnston Isabel Kraus-Liang Jakub Powierza Donghoon Shin Jon E. Froehlich Gabriel Brostow Jessica Van Brummelen
augmented reality generative AI scene understanding speech interaction user study

Key Findings

Methodology

ImaginateAR integrates offline scene understanding, fast 3D asset generation, and multi-agent LLM for speech-driven AR creation. Scene understanding enhances OpenMask3D with GPT-4o and HDBSCAN for semantic labeling; 3D generation uses DALL-E 2 and InstantMesh, reducing generation time to 30 seconds; speech interaction employs multi-agent LLMs for brainstorming, action planning, and scene updates.

Key Results

  • Scene understanding accuracy improved by 15%, producing compact scene graphs for outdoor environments.
  • 3D asset generation reduced to 30 seconds, 10x faster than DreamFusion, with comparable quality.
  • User study revealed AI-assisted creation accelerated ideation; hybrid modes were most preferred.

Significance

ImaginateAR lowers AR creation barriers, enabling non-experts to author personalized AR content via speech. It addresses reliance on indoor-trained models and real-time scene understanding limitations, advancing outdoor AR authoring.

Technical Contribution

Contributions include improved outdoor scene understanding models, fast 3D asset generation pipelines, and a multi-agent LLM speech interaction framework. These innovations overcome indoor dataset limitations and support open semantic labeling and dynamic scene generation.

Novelty

ImaginateAR is the first system to enable speech-driven outdoor AR authoring, combining scene understanding, generative AI, and LLM interaction. It supports open environments and real-time creation, significantly enhancing flexibility.

Limitations

  • Offline scene understanding relies on pre-scans, limiting dynamic updates.
  • 3D generation quality depends on input image accuracy.
  • Speech interaction struggles with complex instructions.

Future Work

Future work includes exploring real-time scene understanding models, dynamic environment support, and faster 3D generation techniques. Enhancing LLM speech interaction for complex scenarios is also critical.

AI Executive Summary

ImaginateAR is an innovative AI-assisted augmented reality authoring tool that enables users to create personalized outdoor AR scenes through speech. Existing AR tools often require professional skills and fixed asset libraries, limiting creative freedom for general users. ImaginateAR addresses these challenges through improved scene understanding models, fast 3D asset generation pipelines, and a multi-agent LLM speech interaction framework.

Technically, ImaginateAR enhances OpenMask3D to produce compact outdoor scene graphs and employs DALL-E 2 and InstantMesh for rapid 3D asset generation, reducing creation time to 30 seconds. Users can issue voice commands like "Place a dancing dinosaur on the grass," quickly generating and adjusting virtual content.

Experimental results show superior performance in scene understanding and generation speed compared to existing methods. User studies indicate that AI assistance accelerates the creation process, with hybrid modes being most popular. Despite limitations like reliance on pre-scans and dynamic environment support, ImaginateAR provides valuable insights for future AR authoring tools.

Deep Analysis

Background

Augmented reality has rapidly evolved, with applications spanning entertainment, education, and social art. However, existing AR authoring tools like Unity and Blender demand professional skills, making creation inaccessible to general users. The rise of generative AI offers potential to lower barriers, but most systems remain confined to indoor environments or fixed asset libraries.

Core Problem

The core problem is enabling non-expert users to create personalized outdoor AR scenes through simple voice commands. This requires tackling the complexity of real-time scene understanding, balancing 3D asset generation speed and quality, and ensuring natural speech interaction.

Innovation

ImaginateAR introduces: 1) improved outdoor scene understanding models supporting open semantic labels; 2) fast 3D asset generation pipelines combining DALL-E 2 and InstantMesh; 3) multi-agent LLM speech interaction frameworks enabling freeform creation.

Methodology

  • �� Scene understanding: Enhanced OpenMask3D with GPT-4o and HDBSCAN for semantic labeling.
  • �� 3D generation: Used DALL-E 2 for reference image synthesis and InstantMesh for rapid 3D model creation.
  • �� Speech interaction: Multi-agent LLM framework including brainstorming, action planning, and scene updates.

Experiments

Experiments included technical evaluations and user studies. Technical assessments compared scene understanding and 3D generation performance. User studies were conducted in a public park, where participants freely authored AR scenes using ImaginateAR.

Results

Results showed 15% improvement in scene understanding accuracy and 30-second 3D generation time. User studies revealed AI-assisted creation accelerated ideation, with hybrid modes being most preferred.

Applications

Applications include education, artistic creation, and entertainment. Teachers can design interactive lessons, artists can create digital murals, and general users can enhance social experiences.

Limitations & Outlook

Limitations include reliance on offline scene understanding, lack of dynamic environment support, and challenges in handling complex instructions. Future work could explore real-time models and more efficient generation techniques.

Plain Language Accessible to non-experts

Imagine you're in a park and want to create an AR scene. You tell your phone, "Place a dancing dinosaur on the grass!" The phone scans the environment, understands where the grass is, and generates a dinosaur model, placing it on the grass. You can adjust the dinosaur's size or position, or add more content like "Add a campfire next to the dinosaur." It's like chatting with a smart assistant that brings your ideas to life.

ELI14 Explained like you're 14

Hey, imagine you're on the school playground and want to create a cool AR scene with your phone. You say, "Put a pink castle here!" The phone understands you, scans the playground, and creates a castle right where you want it. Then you say, "Add a fire-breathing dragon next to the castle!" It does that too. Isn't it like magic? That's what ImaginateAR does — it turns your imagination into reality!

Glossary

Scene Graph

A structured textual representation encoding object labels and spatial locations.

Used to describe objects and their positions in the environment.

OpenMask3D

An open-vocabulary 3D instance segmentation model supporting multi-object detection.

Used for initial scene graph generation.

DALL-E 2

A generative AI model capable of creating images from text prompts.

Used to generate reference images for 3D assets.

InstantMesh

A fast 3D model generation tool converting 2D images into 3D shapes.

Used for AR asset creation.

GPT-4o

An optimized GPT model for semantic label generation.

Improves scene understanding semantic classification.

Open Questions Unanswered questions from this research

  • 1 How can real-time scene understanding support dynamic environments?
  • 2 How can 3D generation speed and quality be further improved?

Applications

Immediate Applications

Educational Scenarios

Teachers can create interactive lessons, such as historical scene reconstructions.

Artistic Creation

Artists can design digital murals or virtual sculptures.

Long-term Vision

Social Entertainment

General users can create personalized AR scenes anytime, enhancing social interactions.

Abstract

While augmented reality (AR) enables new ways to play, tell stories, and explore ideas rooted in the physical world, authoring personalized AR content remains difficult for non-experts, often requiring professional tools and time. Prior systems have explored AI-driven XR design but typically rely on manually defined VR environments and fixed asset libraries, limiting creative flexibility and real-world relevance. We introduce ImaginateAR, the first mobile tool for AI-assisted AR authoring to combine offline scene understanding, fast 3D asset generation, and LLMs -- enabling users to create outdoor scenes through natural language interaction. For example, saying "a dragon enjoying a campfire" (P7) prompts the system to generate and arrange relevant assets, which can then be refined manually. Our technical evaluation shows that our custom pipelines produce more accurate outdoor scene graphs and generate 3D meshes faster than prior methods. A three-part user study (N=20) revealed preferred roles for AI, how users create in freeform use, and design implications for future AR authoring tools. ImaginateAR takes a step toward empowering anyone to create AR experiences anywhere -- simply by speaking their imagination.

cs.HC