cs.LG 2509.21240

Tree Search for LLM Agent Reinforcement Learning

Introduces Tree-GRPO, a tree search-based policy optimization, boosting sample efficiency and process supervision in multi-turn LLM agent tasks.

Yuxiang Ji, Ziyu Ma, Yong Wang et al.

2025-09-25 41
cs.LG 2509.20051

One Filters All: A Generalist Filter for State Estimation

Introduces LLM-Filter, leveraging pretrained large language models for state estimation, outperforming traditional and learning-based filters with strong generalization.

Shiqi Liu, Wenhan Cao, Chang Liu et al.

2025-09-24 51