HalluClear: Diagnosing, Evaluating and Mitigating Hallucinations in GUI Agents
HalluClear employs classification, three-stage evaluation, and closed-loop reasoning to reduce GUI hallucinations effectively.
Chao Jin, Wenkui Yang, Hao Sun et al.
HalluClear employs classification, three-stage evaluation, and closed-loop reasoning to reduce GUI hallucinations effectively.
Chao Jin, Wenkui Yang, Hao Sun et al.
Proposed DeepInsightTheorem framework enhances informal theorem proving by identifying core techniques, significantly outperforming baselines.
Yunhe Li, Hao Shi, Bowen Deng et al.
Using the CompCQ framework, this study analyzes LLM-generated competency questions across domains, revealing generation characteristics.
Reham Alharbi, Valentina Tamma, Terry R. Payne et al.
The study shows language models exhibit strong spatial transfer in shortest path problems but fail in length scaling due to recursive instability.
Yao Tong, Jiayuan Ye, Anastasia Borovykh et al.
Diagnosing LLM judge reliability using transitivity analysis and conformal prediction sets, revealing 33%-67% documents with at least one 3-cycle.
Manan Gupta, Dhruv Kumar
Study reveals LLMs and VLMs struggle with viewpoint rotation understanding without vision, proposes VRUBench dataset, and improves performance via selective fine-tuning.
Zhen Yang, Ping Jian, Zhongbin Guo et al.
Introduces IRS framework, enhancing multimodal humor understanding with incongruity-resolution supervision, 72B model approaches expert level on NYCC.
Hatice Merve Vural, Doga Kukul, Ege Erdem Ozlu et al.
Policy-Guided Hybrid Simulation (PGHS) achieves 8.80% group simulation error on Meituan, improving over baselines by 45.8% and 40.9%.
Ziyang Chen, Renbing Chen, Daowei Li et al.
OpenMobile synthesizes environment-aware instructions and trajectories with policy switching, achieving 51.7% success on AndroidWorld.
Kanzhi Cheng, Zehao Li, Zheng Ma et al.
AIBuildAI employs hierarchical multi-agent architecture to automate AI model building from task description, achieving 63.1% medal rate, surpassing all baselines.
Ruiyi Zhang, Peijia Qin, Qi Cao et al.
RiskWebWorld offers 1,513 tasks showcasing challenges for GUI agents in e-commerce risk management, with top models achieving 49.1% success.
Renqi Chen, Zeyin Tao, Jianming Guo et al.
Proposes Fun-TSG, a function-driven multivariate time series generator with variable-level anomaly labels, enabling controllable, transparent synthetic data creation.
Pierre Lotte, André Péninou, Olivier Teste
Meerkat combines clustering with agentic search, finding roughly 4× more CyBench reward-hacking cases than prior audits.
Adam Stein, Davis Brown, Hamed Hassani et al.
Developed a scalable problem reduction library with 100+ problem types and 200+ rules, integrated via AI agents for automated contribution and verification.
Xi-Wei Pan, Shi-Wen An, Jin-Guo Liu
WebForge automates the end-to-end creation of realistic, reproducible, and scalable web environments using a four-stage pipeline, enabling multi-dimensional capability profiling.
Peng Yuan, Yuyang Yin, Yuxuan Cai et al.
CFMS integrates multimodal perception with hierarchical symbolic reasoning, boosting complex table understanding by 15% accuracy on benchmarks.
Qixian Huang, Hongqiang Lin, Tong Fu et al.
SkillClaw turns multi-user experience into shared skill evolution; on WildClawBench, Qwen3-Max gains up to 52% in Search & Retrieval.
Ziyu Ma, Shidong Yang, Yuxiang Ji et al.
UP-NRPA integrates user portraits with nested rollouts for dynamic strategy planning in goal-oriented dialogues, achieving a 56.41% improvement in negotiation SL.
Hui Wang, Fafa Zhang, Meng Liu et al.
Proposes a five-dimensional auditability framework, emphasizing that accountability depends on system auditability.
Yi Nian, Aojie Yuan, Haiyue Zhang et al.
Proposes a consequence-sensitive support compression method in belief arbitration, balancing information retention and resource costs for robust decision-making.
Mark Walsh