cs.RO 2509.21986

Developing Vision-Language-Action Model from Egocentric Videos

Using EgoScaler to automatically extract 6DoF object trajectories from unlabeled egocentric videos significantly improves VLA pre-training, achieving over 20% success rate gains.

Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura et al.

2025-09-26 48
cs.AI 2509.21651

Can AI Perceive Physical Danger and Intervene?

Proposed ASIMOV-2.0 benchmark integrates real injury reports and generative models to evaluate AI's physical risk perception and intervention.

Abhishek Jindal, Dmitry Kalashnikov, R. Alex Hofer et al.

2025-09-26 35
cs.LG 2509.21240

Tree Search for LLM Agent Reinforcement Learning

Introduces Tree-GRPO, a tree search-based policy optimization, boosting sample efficiency and process supervision in multi-turn LLM agent tasks.

Yuxiang Ji, Ziyu Ma, Yong Wang et al.

2025-09-25 46