cs.LG 2505.23884

Test-Time Training Done Right

Large Chunk Test-Time Training (LaCT) enhances long-sequence modeling efficiency, handling up to 1M context length.

Tianyuan Zhang, Sai Bi, Yicong Hong et al.

2025-05-30 5
cs.SE 2505.23419

SWE-bench Goes Live!

SWE-bench-Live introduces an automated, real-time benchmark from GitHub issues, covering 93 repositories, enhancing dynamic evaluation of bug-fixing models.

Linghao Zhang, Shilin He, Chaoyun Zhang et al.

2025-05-29 33
cs.RO 2505.23189

TrackVLA: Embodied Visual Tracking in the Wild

TrackVLA introduces a unified VLA model trained on 1.7 million samples, outperforming SOTA in embodied visual tracking and recognition.

Shaoan Wang, Jiazhao Zhang, Minghan Li et al.

2025-05-29 47
cs.LG 2505.22988

Model-Preserving Adaptive Rounding

YAQA, a Hessian-structured adaptive rounding algorithm, reduces quantization error by ~30%, with theoretical end-to-end error bounds.

Albert Tseng, Zhaofeng Sun, Christopher De Sa

2025-05-29 42