cs.SE 2603.23448

Code Review Agent Benchmark

c-CRAB dataset evaluates code review agents' abilities; current agents solve only 40% of tasks.

Yuntong Zhang, Zhiyuan Pan, Imam Nur Bani Yusuf et al.

2026-03-25 4 citations 275
cs.IR 2603.23183

Reasoning over Semantic IDs Enhances Generative Recommendation

Proposes SIDReasoner, enhancing SID–language alignment and using outcome-driven reinforcement to improve reasoning in generative recommendation, boosting accuracy and interpretability.

Yingzhi He, Yan Sun, Junfei Tan et al.

2026-03-24 45
cs.RO 2603.22703

Learning Safe-Stoppability Monitors for Humanoid Robots

PRISM framework employs importance sampling and neural prediction to efficiently learn humanoid robot safe-stoppability boundaries, enabling proactive safety monitoring.

Yifan Sun, Yiyuan Pan, Shangtao Li et al.

2026-03-24 53