HardTests: Synthesizing High-Quality Test Cases for LLM Coding
Proposes HARDTESTGEN, leveraging LLMs to synthesize high-quality test cases, improving verification precision by 11.3% and recall by 17.5%.
Zhongmou He, Yee Man Choi, Kexun Zhang et al.
Proposes HARDTESTGEN, leveraging LLMs to synthesize high-quality test cases, improving verification precision by 11.3% and recall by 17.5%.
Zhongmou He, Yee Man Choi, Kexun Zhang et al.
Introduces UCerF metric combining uncertainty to evaluate LLM fairness; SynthBias dataset enhances diversity testing.
Yinong Oliver Wang, Nivedha Sivakumar, Falaah Arif Khan et al.
InterMT aligns multi-turn multimodal preferences with human feedback, featuring 15.6k prompts and 52.6k dialogue instances.
Boyuan Chen, Donghai Hong, Jiaming Ji et al.
ScaleLong evaluates long video understanding across timescales, revealing a U-shaped performance curve.
David Ma, Huaqing Yuan, Xingjian Wang et al.
DeepTheorem enhances LLM reasoning for theorem proving using natural language and RL, featuring a 121K informal theorem dataset.
Ziyin Zhang, Jiahao Xu, Zhiwei He et al.
Proposed OWL framework with modular multi-agent architecture achieves 69.70% accuracy on GAIA, surpassing OpenAI Deep Research by 2.34%.
Mengkang Hu, Yuhang Zhou, Wendong Fan et al.
Large Chunk Test-Time Training (LaCT) enhances long-sequence modeling efficiency, handling up to 1M context length.
Tianyuan Zhang, Sai Bi, Yicong Hong et al.
Table-R1 enhances table reasoning via inference-time scaling using distillation and RLVR, surpassing GPT-4.1 with a 7B model.
Zheyuan Yang, Lyuhao Chen, Arman Cohan et al.
SWE-bench-Live introduces an automated, real-time benchmark from GitHub issues, covering 93 repositories, enhancing dynamic evaluation of bug-fixing models.
Linghao Zhang, Shilin He, Chaoyun Zhang et al.
InfiMed integrates reflective-patterned CoT and rule-based RLVR, boosting low-resource medical multimodal model performance, achieving 7 SOTA benchmarks.
Zeyu Liu, Zhitian Hou, Guanghao Zhu et al.
Proposed GenCAD-Self-Repairing enhances CAD generation feasibility from 10% to 97% using diffusion guidance and self-repair, fixing 65% of infeasible designs.
Chikaha Tsuji, Enrique Flores Medina, Harshit Gupta et al.
MathArena leverages newly released math competitions for real-time, uncontaminated evaluation of LLM reasoning and proof-writing, surpassing static benchmarks.
Mislav Balunović, Jasper Dekoninck, Ivo Petrov et al.
TrackVLA introduces a unified VLA model trained on 1.7 million samples, outperforming SOTA in embodied visual tracking and recognition.
Shaoan Wang, Jiazhao Zhang, Minghan Li et al.
LODGE introduces a hierarchical LOD framework for large-scale 3D Gaussian Splatting, reducing GPU memory by 50% and increasing rendering speed to 200 FPS.
Jonas Kulhanek, Marie-Julie Rakotosaona, Fabian Manhardt et al.
Infi-MMR employs a three-phase curriculum reinforcement learning framework to enhance multimodal small language models, achieving 43.68% on MathVerse.
Zeyu Liu, Yuhang Liu, Guanghao Zhu et al.
K2VAE combines KoopmanNet and KalmanNet for efficient long- and short-term probabilistic time series forecasting, outperforming state-of-the-art methods.
Xingjian Wu, Xiangfei Qiu, Hongfan Gao et al.
YAQA, a Hessian-structured adaptive rounding algorithm, reduces quantization error by ~30%, with theoretical end-to-end error bounds.
Albert Tseng, Zhaofeng Sun, Christopher De Sa
WorkForceAgent-R1 enhances LLM web agents' reasoning via R1-style RL, achieving 10.26-16.59% improvement.
Yuchen Zhuang, Di Jin, Jiaao Chen et al.
FAMA is the first large-scale open-science speech foundation model for English and Italian, achieving up to 8x speed improvement.
Sara Papi, Marco Gaido, Luisa Bentivogli et al.
VScan accelerates LVLM inference by 2.91× with only 4.6% performance loss, using a two-stage token pruning strategy.
Ce Zhang, Kaixin Ma, Tianqing Fang et al.