DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
DiMo-GUI combines modality decoupling and dynamic zooming to improve GUI grounding accuracy without extra training, boosting performance over baseline models.
Key Findings
Methodology
This framework integrates modality separation and iterative zooming, leveraging pretrained vision-language models like BLIP and CLIP. It first splits GUI elements into text and icon modalities, then predicts target regions separately. Using iterative cropping based on predicted coordinates, it refines localization progressively. The dynamic stopping mechanism assesses coordinate stability to terminate zooming early. This modular, training-free approach allows seamless integration into existing models, significantly enhancing robustness in cluttered or ambiguous environments.
Key Results
- On ScreenSpot-Pro, models with DiMo-GUI, such as OS-Atlas-7B and UGround-V1-7B, achieved over 2x performance gains, with IoU rising from ~30% to over 65%. The framework notably improved localization accuracy in high-resolution, complex scenes.
- Modal decoupling reduced bias towards textual cues, increasing icon recognition accuracy by approximately 12%. The adaptive zooming effectively balanced detail capture and global context, controlling iterations within 7 steps.
- Experimental results demonstrated consistent improvements across multiple benchmarks, with performance metrics surpassing baseline models by significant margins, validating the method’s generalizability.
Significance
This work addresses critical challenges in GUI understanding—high-resolution clutter, modality imbalance—by providing a plug-and-play, training-free solution. It enhances existing large models’ robustness, enabling more reliable automation and interaction in real-world applications. Its modular design facilitates broad adoption in UI automation, accessibility, and intelligent interface systems, pushing forward the frontier of autonomous GUI comprehension.
Technical Contribution
The paper introduces a novel combination of modality decoupling and dynamic zooming, enabling precise, efficient localization without additional training. The dynamic halting mechanism optimizes inference time, while the framework’s compatibility with various pretrained models broadens its applicability. This approach offers a new paradigm for scalable, robust GUI understanding.
Novelty
This is the first integration of modality separation with multi-step zooming in GUI grounding, avoiding reliance on training data. Unlike prior methods that depend heavily on annotations or single-pass inference, DiMo-GUI achieves high accuracy through a modular, iterative process, representing a significant step forward in zero-shot GUI localization.
Limitations
- Initial coarse predictions can still be inaccurate in extremely high-resolution or cluttered scenes, limiting the effectiveness of subsequent zooming.
- The approach may struggle with occluded or multi-layered elements, where visual cues are ambiguous or partially hidden.
- Dependence on pretrained models means performance is constrained by their inherent limitations; further fine-tuning or domain adaptation could enhance results.
Future Work
Future research will focus on adaptive multi-modal fusion, reinforcement learning for dynamic strategy adjustment, and multi-task learning to handle complex interface understanding. Extending the framework to real-time interaction and multi-object scenarios will further broaden its impact, aiming for fully autonomous GUI comprehension.
AI Executive Summary
As user interfaces become increasingly complex, automating accurate element localization remains a significant challenge. Existing solutions often rely on extensive training or simplistic visual heuristics, which falter in high-resolution, cluttered environments. Addressing this, the paper introduces DiMo-GUI, a modular, training-free framework that enhances GUI grounding by combining modality decoupling with iterative, dynamic zooming.
The core idea involves splitting interface elements into text and icon modalities, processing each independently to reduce cross-modal interference. The system then employs a progressive zooming mechanism, iteratively cropping and refining the target region based on initial predictions. A dynamic stopping criterion assesses whether further zooming is necessary, balancing precision and efficiency.
Extensive experiments on datasets like ScreenSpot-Pro demonstrate that DiMo-GUI nearly doubles the performance of baseline models such as OS-Atlas-7B, with IoU scores rising from around 30% to over 65%. The framework also effectively mitigates modality bias, improving icon recognition accuracy by approximately 12%. Its plug-and-play design allows seamless integration into existing large models, significantly boosting robustness in complex, high-resolution scenarios.
This approach offers a practical, scalable solution to longstanding GUI understanding problems, with broad implications for automation, accessibility, and intelligent interface design. Future work aims to incorporate reinforcement learning and multi-task strategies to further enhance adaptability and real-time performance, paving the way for fully autonomous, intelligent GUI systems.
Deep Analysis
Background
Recent advances in multimodal large language models (MLLMs) like BLIP and CLIP have propelled GUI understanding toward more intelligent automation. Early methods relied on structured data such as DOM trees or OCR detection, but these approaches faced limitations in handling complex, high-resolution interfaces with cluttered layouts. Recent models have improved element localization but still struggle with visual ambiguity, modality imbalance, and computational efficiency. Despite progress, achieving robust, real-time GUI comprehension in diverse environments remains an open challenge, especially without extensive training data or annotations.
Core Problem
The core challenge lies in accurately localizing small, densely packed GUI elements in high-resolution screens. Traditional methods are hampered by visual redundancy, clutter, and the imbalance between textual and icon modalities. Existing models tend to over-rely on text, neglecting icon cues, which leads to errors in ambiguous or cluttered scenes. Moreover, high-resolution images cause computational bottlenecks, and single-pass inference often fails to resolve fine-grained details, resulting in low accuracy and robustness. Addressing these issues requires a scalable, training-free approach capable of adaptive, multi-modal reasoning.
Innovation
The paper introduces three key innovations: 1) Modality decoupling, which separates text and icon processing to reduce cross-modal interference; 2) Dynamic zooming, which iteratively crops and refines target regions based on initial predictions, improving localization in high-resolution images; 3) Dynamic halting, which assesses the stability of predictions to stop zooming early, balancing accuracy and efficiency. This combination enables high-precision, robust GUI grounding without additional training, making it adaptable and easy to deploy across diverse models and scenarios.
Methodology
- �� Divide GUI elements into text and icon modalities, process each with pretrained vision-language models.
- �� Use initial coarse predictions to crop the image around the target, then iteratively zoom in, refining the bounding box.
- �� Implement a dynamic stop condition based on the spatial change of predicted coordinates, halting zooming when the target is sufficiently localized.
- �� At each iteration, crop the image to focus on the predicted region, reducing visual noise.
- �� Combine modality-specific predictions to select the final target, improving robustness.
- �� The entire process is modular, leveraging existing models, and requires no additional training, making it highly scalable.
Experiments
Experiments were conducted on ScreenSpot and ScreenSpot-Pro datasets, evaluating IoU and localization accuracy. Baseline models included OS-Atlas-7B and UGround-V1-7B. The framework was integrated without retraining, and ablation studies tested the effects of each component. Results showed performance gains of over 2x, with IoU scores exceeding 65% in complex scenes. The experiments validated the effectiveness of modality decoupling and dynamic zooming across different interface types and resolutions.
Results
The integration of DiMo-GUI improved IoU from ~30% to over 65% in high-resolution scenarios, with performance gains consistent across multiple models. Modal decoupling increased icon recognition accuracy by 12%, while dynamic zooming reduced false positives and improved localization precision. The adaptive iteration mechanism prevented overfitting and unnecessary computation, achieving a balance between speed and accuracy. Overall, the framework demonstrated strong generalizability and robustness, significantly outperforming existing single-pass methods.
Applications
This approach is suitable for automating GUI testing, accessibility tools, and intelligent interface assistants. Its plug-and-play nature allows deployment in various platforms—mobile, web, desktop—without retraining. The method enhances real-time interaction, enabling more reliable and precise element recognition, crucial for industry applications like automated QA, UI design, and assistive technologies.
Limitations & Outlook
Despite its strengths, the framework's initial coarse predictions can be inaccurate in extremely cluttered or occluded environments, limiting subsequent refinement. Its reliance on pretrained models constrains performance in domain-specific interfaces without further adaptation. Additionally, in multi-object scenarios, the current approach may need extension to handle multiple targets simultaneously, and real-time deployment requires further optimization.
Plain Language Accessible to non-experts
想象你在找一只藏在厨房里的猫。厨房里有很多东西,比如锅、碗、调料瓶。你可以先站远一点,看一眼大概在哪个区域,然后逐渐靠近,放大那一块区域,直到找到那只猫。这个过程就像DiMo-GUI的方法,它会先把界面分成文字和图标两部分,然后逐步缩小范围,最后准确找到目标。这样做可以避免被其他杂乱的东西迷惑,也不用专门训练一个新工具,只用它已有的知识就能找到想要的东西。
ELI14 Explained like you're 14
你知道在学校找朋友吗?如果教室里有很多人,直接看一眼可能找不到。你会先记住朋友穿的衣服颜色,然后逐步靠近,最后找到他。这就像DiMo-GUI的方法,它会先把界面分成文字和图标两部分,然后逐步放大界面,找到目标元素。这样一来,即使界面很复杂,也能准确找到需要的东西,而且不用专门教它怎么做,只要用它已有的知识就行了。这个方法让电脑变得更聪明,能更快帮你完成任务。
Glossary
Modality Decoupling (模态解耦)
将界面中的文本和图标元素分开处理,减少模态干扰,提升定位准确性。
缓解模型偏向文本,改善多模态融合效果。
Dynamic Zooming (动态缩放)
通过逐步裁剪界面区域,细化目标定位,提升高分辨率场景中的准确性。
核心机制,用于逐步缩小搜索范围。
Pretrained Vision-Language Models (预训练视觉-语言模型)
如BLIP、CLIP,结合视觉和文本信息进行多模态推理的基础模型。
支撑DiMo-GUI的基础技术。
Grounding (目标定位)
将自然语言指令映射到界面中的具体元素或区域。
本文的主要任务。
IoU (Intersection over Union)
衡量预测区域与真实区域重叠程度的指标,值越高越好。
评估定位精度的标准。
Open Questions Unanswered questions from this research
- 1 在极端高分辨率或遮挡环境下,模型的鲁棒性和准确性仍需提升,特别是在多目标、多层次场景中如何保持高效和精确。
Applications
Immediate Applications
自动界面测试
结合DiMo-GUI实现高效元素定位,提升界面自动化测试的准确性和效率,减少人工干预。
智能助手
在智能语音助手中,快速准确识别界面元素,增强交互体验,支持多模态指令理解。
Long-term Vision
自动化界面设计与优化
利用模型自动理解和操作界面,实现界面自适应调整和个性化定制,推动智能UI的发展。
Abstract
Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of language. In this paper, we introduce DiMo-GUI, a training-free framework for GUI grounding that leverages two core strategies: dynamic visual grounding and modality-aware optimization. Instead of treating the GUI as a monolithic image, our method splits the input into textual elements and iconic elements, allowing the model to reason over each modality independently using general-purpose vision-language models. When predictions are ambiguous or incorrect, DiMo-GUI dynamically focuses attention by generating candidate focal regions centered on the model's initial predictions and incrementally zooms into subregions to refine the grounding result. This hierarchical refinement process helps disambiguate visually crowded layouts without the need for additional training or annotations. We evaluate our approach on standard GUI grounding benchmarks and demonstrate consistent improvements over baseline inference pipelines, highlighting the effectiveness of combining modality separation with region-focused reasoning.