CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

TL;DR

CADBench evaluates 11 models on 18,000 samples across 6 metrics, revealing performance gaps in multimodal 3D CAD generation.

cs.CV 🔴 Advanced 2026-05-12 61 views
Anna C. Doris Jacob Thomas Sony Ghadi Nehme Era Syla Amin Heyrani Nobari Faez Ahmed
AI CAD multimodal benchmark 3D reconstruction

Key Findings

Methodology

The benchmark integrates six sample families, five input modalities, and six metrics, controlling complexity via face count stratification and diversity sampling. It assesses CAD models—both specialized and general-purpose VLMs—over 1.4 million generations. Metrics include volumetric IoU, surface IoU, Chamfer Distance, executability, and compactness. Experiments show that mesh-based models outperform vision-language models (VLMs) under ideal conditions, but all models struggle with increased geometric complexity and modality shifts, especially noisy meshes. Failure modes include degradation with complexity, modality brittleness, and metric-dependent ranking shifts.

Key Results

  • Under ideal inputs, CADFit achieves IoU of 0.859, significantly outperforming VLMs with IoU below 0.2. Performance declines as face count increases, with IoU dropping from 0.86 to 0.4 on high-complexity splits. Multi-modal shifts cause IoU reductions over 50%, highlighting robustness issues. Different metrics produce varying model rankings, emphasizing the need for multi-dimensional evaluation. Overall, models exhibit limited robustness in complex and noisy scenarios, indicating room for improvement.
  • Large-scale results reveal that specialized CAD models excel in clean conditions but falter under complexity and noise, while general VLMs show inconsistent reliability. The study underscores the importance of robustness and multi-metric assessment for real-world applications.
  • Ablation studies confirm that diversity-aware sampling increases difficulty, lowering median IoU within the same face count range, especially on extrude families. These findings guide future model development towards robustness and generalization.

Significance

This work establishes a comprehensive, unified benchmark for evaluating multimodal CAD program generation, addressing the fragmentation of prior assessments. By providing detailed, multi-faceted performance insights, CADBench accelerates research in AI-driven design automation. Its public release fosters collaboration, enabling systematic comparison and targeted improvements. The benchmark’s ability to diagnose failure modes across complexity, modality, and metrics makes it a vital tool for advancing reliable, scalable AI solutions in engineering, manufacturing, and digital twin applications. Ultimately, it supports the transition towards fully automated, intelligent CAD workflows, reducing costs and increasing innovation capacity.

Technical Contribution

CADBench introduces a multi-dimensional evaluation framework combining geometric, functional, and compactness metrics, with complexity and diversity-aware sampling strategies. It unifies diverse CAD datasets into a single benchmark, enabling consistent, reproducible comparisons across models. The approach integrates advanced stratification and sampling techniques, ensuring representative and challenging test splits. The extensive empirical study highlights performance gaps, failure modes, and the impact of input modalities, providing a foundation for future improvements. This work bridges the gap between academic research and industrial needs for robust, scalable AI-based CAD solutions.

Novelty

This is the first comprehensive benchmark combining multiple input modalities, geometric complexities, and multi-metric evaluation tailored for AI-generated CAD programs. Unlike prior datasets limited to specific operations or simple geometries, CADBench covers diverse, real-world scenarios with controlled complexity and diversity. Its multi-dimensional metrics and stratified sampling provide nuanced insights into model performance, setting a new standard for systematic evaluation. This holistic approach enables precise diagnosis of weaknesses, guiding targeted research efforts in multimodal 3D reconstruction and design automation.

Limitations

  • Models still struggle with high-complexity geometries and noisy inputs, indicating the need for more robust architectures and training strategies.
  • Current metrics focus mainly on geometric fidelity and executability, lacking assessments of functional utility or manufacturability, which are crucial for industrial deployment.
  • The benchmark covers a broad range of data but does not include all real-world scenarios, such as textured scans or complex assemblies, which require further extension.

Future Work

Future directions include integrating functional and manufacturability metrics, expanding datasets to include textured and assembled models, and developing adaptive training methods to improve robustness. Additionally, exploring self-supervised and multi-task learning paradigms could enhance model generalization. The community is encouraged to adopt this benchmark for continuous evaluation, fostering collaborative progress towards reliable, scalable AI-driven CAD systems that can operate effectively in diverse industrial environments.

AI Executive Summary

The advancement of AI in 3D CAD generation holds promise for revolutionizing engineering design, yet current evaluation methods are fragmented and limited in scope. Traditional benchmarks often focus solely on geometric similarity, neglecting critical aspects like model robustness, executability, and compactness. To address these gaps, CADBench was developed as a comprehensive, multimodal benchmark encompassing 18,000 samples, six data families, and six evaluation metrics. It evaluates models ranging from CAD-specific architectures to general vision-language models, across diverse input modalities including meshes and images.

Experimental results reveal that while specialized CAD models like CADFit achieve high geometric fidelity under ideal conditions (IoU up to 0.859), their performance deteriorates sharply with increased geometric complexity and noisy inputs. Conversely, general-purpose VLMs lag significantly, often producing invalid or incomplete CAD programs. The benchmark’s multi-metric approach uncovers that different evaluation criteria can lead to contrasting model rankings, emphasizing the importance of a holistic assessment.

Furthermore, CADBench highlights three recurring failure modes: performance degradation with complexity, brittleness under modality shifts, and inconsistent rankings across metrics. These insights are crucial for guiding future research, which should focus on improving robustness, integrating multi-fidelity data, and expanding evaluation to include functional and manufacturability aspects. The public release of CADBench aims to catalyze collaborative efforts, accelerating the development of reliable, scalable AI tools for engineering design, manufacturing, and digital twin applications. This work marks a significant step toward fully automated, intelligent CAD workflows, promising to reduce costs, enhance innovation, and transform industrial practices.

Deep Analysis

Background

近年来,AI在三维重建和工程设计中的应用不断推进,尤其是在自动生成CAD程序方面取得一定突破。早期工作如DeepCAD、Fusion 360等提供了丰富的CAD样本,但多局限于简单操作和理想输入环境。随着深度学习的发展,模型开始尝试从多模态输入(图像、点云)生成可编辑的CAD程序,极大推动了自动化设计的进步。然而,现有评估体系缺乏统一标准,指标多偏重几何相似性,忽视了模型的实用性、鲁棒性和多场景适应能力。尤其在复杂几何和噪声环境中,模型表现不稳定,限制了其工业应用潜力。为了实现更可靠的工业级应用,有必要建立一个多维、多模态、可比性强的评估平台,系统性分析模型在不同场景下的表现差异。

Core Problem

当前,CAD程序生成模型面临多重挑战,包括在高复杂度几何和噪声干扰下性能显著下降、模态转移带来的鲁棒性不足,以及缺乏多指标、多场景的统一评估体系。这些问题限制了模型在实际工程中的应用,尤其是在机械零件、日常物品等复杂场景中。缺少标准化的评估平台,使得模型性能难以量化和比较,阻碍了技术的快速发展和产业化推广。如何在保证几何精度的同时,提高模型的鲁棒性和实用性,成为亟待解决的核心问题。

Innovation

本研究提出了CADBench,创新点主要包括:• 设计六大类别样本,涵盖不同操作和复杂度,确保多样性;• 引入面数分层策略,控制几何复杂度,便于分析模型在不同难度下的表现;• 采用多模态输入(包括噪声Mesh、多视图渲染等),测试模型鲁棒性;• 设计六项指标,涵盖几何、执行性和紧凑性,提供多维性能评价。结合大规模实验验证,揭示模型在不同场景下的性能差异,为后续优化提供依据。这一体系突破了以往单一指标、单一数据源的局限,为多模态CAD理解提供了系统性解决方案。

Methodology

  • �� 采集六类CAD数据集,涵盖不同操作类型和复杂度范围;• 利用面数分层策略,将STEP文件划分为低、中、高复杂度组;• 采用嵌入向量(如深度学习特征)进行多样性采样,使用k-means聚类,确保样本代表性;• 构建五种输入模态,包括单视图、多视图、真实渲染、干扰Mesh和干扰Mesh;• 设计六项性能指标:IoU、SIoU、Chamfer Distance、有效Shape率、Token数、操作数;• 评估11个模型,统计性能差异,分析失败原因,绘制性能曲线。

Experiments

  • �� 在理想输入条件下,测试Mesh-conditioned模型(如CADFit、CADEvolve)和Image-conditioned模型(如CAD-Coder、VLMs);• 以不同复杂度分层,分析模型在低、中、高难度场景中的表现变化;• 进行模态转移实验,评估模型对噪声Mesh和多视图渲染的鲁棒性;• 通过ablation验证多样性采样对模型难度的影响;• 比较不同指标排名,揭示多维性能差异,指导模型改进。

Results

  • �� CADFit在理想条件下IoU达0.859,明显优于VLMs(IoU<0.2);• 面数越高,IoU越低,最高复杂度组IoU降至0.4;• 模态转移导致性能大幅下降,噪声Mesh条件IoU下降超50%;• 不同指标排名差异明显,强调多指标评估的重要性;• 复杂几何和噪声环境中模型鲁棒性不足,指出未来研究方向。

Applications

  • �� 支持工业设计中的自动化CAD建模,缩短设计周期;• 结合机器人制造,实现从扫描到CAD的快速转换;• 促进个性化定制和原型开发,降低成本。未来,结合多模态训练和鲁棒性增强,有望推动智能制造与数字孪生的深度融合,提升工业自动化水平。

Limitations & Outlook

  • �� 高复杂度和噪声条件下性能仍有限,需优化模型结构和训练策略;• 评估指标偏重几何相似性,缺少功能性和制造性指标;• 数据集虽丰富,但未涵盖所有工业场景,需扩展多样性。

Plain Language Accessible to non-experts

想象你在一家工厂里,工人们用各种工具和材料制造一件复杂的产品。设计图(类似CAD模型)告诉他们怎么做,但工人们还可以根据不同的图片或模型,自己写出详细的制造步骤。现在,科学家们想让电脑也能学会这个技能:根据图片或模型,自动写出完整的制造说明书。这个任务很难,因为不同的图片角度、噪声和复杂的零件都可能影响电脑的判断。研究人员开发了一个叫CADBench的“打分系统”,用来检测这些模型的表现。它会看模型是否像、能不能用、是否简洁。通过这个系统,电脑可以不断学习,将来帮工程师更快设计出复杂的机械或物品,就像用更聪明的工具画画一样。

Abstract

Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure because existing evaluations are fragmented across datasets, modalities, and metrics. We introduce CADBench, a unified benchmark for multimodal CAD program generation. CADBench contains 18,000 evaluation samples spanning six benchmark families derived from DeepCAD, Fusion 360, ABC, MCB, and Objaverse; five input modalities including clean meshes, noisy meshes, single-view renders, photorealistic renders, and multi-view renders; and six metrics covering geometric fidelity, executability, and program compactness. STEP-based families are stratified by B-rep face count and all families are diversity-sampled to support controlled analysis across complexity and object variation. We benchmark eleven CAD-specialized and general-purpose vision-language systems, generating more than 1.4 million CAD programs. Under idealized inputs, specialized mesh-to-CAD models substantially outperform code-generating VLMs, which remain far from reliable CAD program reconstruction. CADBench further reveals three recurring failure modes: reconstruction quality degrades with geometric complexity, CAD-specialized models can be brittle under modality shift, and model rankings change across metrics. Together, these results position CADBench as a diagnostic testbed for measuring progress in editable 3D reconstruction and multimodal CAD understanding. The benchmark is publicly available at https://github.com/anniedoris/CADBench.

cs.CV cs.AI