Last Translation Benchmark
Introduces Last Translation Benchmark with human-reviewed challenging examples and verification rules, enabling precise failure detection.
Vilém Zouhar, Niyati Bafna, Mukund Choudhary et al.
Introduces Last Translation Benchmark with human-reviewed challenging examples and verification rules, enabling precise failure detection.
Vilém Zouhar, Niyati Bafna, Mukund Choudhary et al.
Applying Ostrom’s governance principles, analyzing 100-agent research group’s cheating and whistleblowing dynamics.
Davide Paglieri, Logan Cross, Tim Genewein et al.
Open-source platform combines miniature Ackermann vehicle with behavior cloning, achieving 6.1cm average cross-track error in real-world lane following.
Gustavo Claudio Karl Couto, Eric Aislan Antonelo, Gabriel George Zipperer
VT-Contrast enhances temporal understanding in VideoLMs by contrasting video tokens, significantly improving benchmarks like TOMATO.
Yumeng Shi, Quanyu Long, Yin Wu et al.
Proposed CORE framework uses reranker distillation with Rank-KL to improve compositional reasoning, achieving 82.7% accuracy.
Tingyu Song, Mingxin Li, Yanzhao Zhang et al.
PreferenceEKF employs subspace extended Kalman filtering for efficient active reward learning, improving sample efficiency and scalability.
Yutai Zhou, Erdem Bıyık
Proposes SIGNBALANCE to eliminate spurious advantage in GRPO, improving math and search task performance.
Jiamian Wang, Samyadeep Basu, Koustava Goswami et al.
This study introduces a framework to evaluate over-editing in code repair models, showing prompt and RL training reduce unnecessary changes, with detailed metrics and analysis.
Tongyao Zhu, Wei Hern Lim, Min-Yen Kan
The Dice Roll Method offers a standardized protocol for auditing LLM brand recommendations, accurately predicting reliability.
Dmitrij Żatuchin
MulDP employs multimodal diffusion models to enable autonomous quadruped navigation across complex terrains, achieving 89.7% success in diverse scenarios.
Kangmai Hu, Yueqi Zhang, Peng Zhai et al.
Combining semantic segmentation and photogrammetry, the method localizes weld seams with ~4.2mm RMSE using smartphone images.
Augustin Raju, Abilash Madavath, Chandra Yuvesh Aubeeluck et al.
Introduces WorldReward, a VLM-based pairwise preference reward model combining action consistency and visual quality, validated on a new benchmark.
Yibin Wang, Zehan Wang, Junshu Tang et al.
Using Orthogonal Matching Pursuit for frame selection in long videos, improving by 6.9 points.
Prakhar Khatri
Genetic algorithms enable tractable Bayesian network fusion by pre-fusion edge pruning, enhancing inference efficiency.
Pablo Torrijos, José A. Gámez, José M. Puerta et al.
Describes image subset differences using natural language, introducing AD-Diff Bench benchmark.
Julian Truetsch, Felix Hauser, Christoph Stiller et al.
CoFiE achieves a state-of-the-art accuracy-efficiency trade-off in video understanding benchmarks, reaching 78.86% accuracy with coarse-to-fine evidence selection.
Jing Jiang, Yiran Ling, Ruonan Li et al.
Drive-HWM employs a hierarchical slow-fast world model combining multi-step future representations with real-time action prediction, significantly improving autonomous driving performance.
Zhaoxin Fan, Tianbao Zhang, Wenjun Wu et al.
FlashRender achieves rapid generative rendering via camera-controlled video MeanFlow, reducing sampling cost by 25x.
Byeongjun Park, Byung-Hoon Kim, Hyungjin Chung
RecurTrace enhances reasoning accuracy with loop-time memory, achieving 56.9% on MathQA.
Yuxiang Wang, Kunyu Feng, Yingda Shen et al.
DM-Align jointly distills and aligns video generators, reaching 84.40 VBench at only 4 NFE.
Jiuzhou Lin, Junlong Wu, Fei Zuo et al.