cs.AI 2605.09163

FORTIS: Benchmarking Over-Privilege in Agent Skills

FORTIS benchmark quantifies over-privilege in large models' skill selection and execution, with failure rates exceeding 62.5%, revealing systemic permission control issues.

Shawn Li, Chenxiao Yu, Han Wang et al.

2026-05-10 7 citations 81
cs.LG 2605.08733

Generative Actor-Critic with Soft Bridge Policies

SoftGAC introduces path regularization with short Gaussian bridges, enabling efficient single-pass maximum entropy RL with improved performance.

Ke He, Le He, Shunpu Tang et al.

2026-05-09 51