cs.CV 2605.14696

EponaV2: Driving World Model with Comprehensive Future Reasoning

EponaV2 employs future depth and semantic prediction with flow matching to enhance autonomous driving planning, outperforming perception-free models with +1.3PDMS and +5.5EPDMS improvements.

Jiawei Xu, Zhizhou Zhong, Zhijian Shu et al.

2026-05-14 39
cs.LG 2605.14477

Test-Time Learning with an Evolving Library

EvoLib constructs an evolving knowledge library via self-supervised abstraction extraction, boosting performance in mathematical reasoning and code generation without parameter updates.

Weijia Xu, Alessandro Sordoni, Chandan Singh et al.

2026-05-14 48
cs.SE 2605.15229

PBT-Bench: Benchmarking AI Agents on Property-Based Testing

PBT-Bench evaluates AI's property-based testing ability using 100 curated problems with 365 semantic bugs across 40 Python libraries, revealing significant model performance gaps.

Lucas Jing, Xinqi Wang, Liao Zhang et al.

2026-05-14 52