cs.AI 2607.24112

Scaling GUI Agents with Visual State Transitions

Proposes State Transition Pretraining (STP) using joint inverse and forward dynamics to enhance GUI agents, improving success rates by up to 6.2%.

Xiangyan Liu, Kaixin Li, Haonan Wang et al.

2026-07-27 33
cs.AI 2607.23700

Offline-Online Curriculum RL for Multimodal Reasoning

Proposes O²-CritiCuRL, combining offline analysis and online RL to identify critical reasoning steps, boosting multimodal reasoning accuracy and efficiency.

Wendi Deng, Hang Du, Guoshun Nan et al.

2026-07-26 31
cs.AI 2607.19592

Knowledge-Centric Self-Improvement

Knowledge-centric self-improvement uses a curated knowledge base to enhance task solving and transferability, reducing costs by 30%.

Xuefei Julie Wang, Lauren Hyoseo Yoon, Chengrui Qu et al.

2026-07-22 39