Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
ThinkARM framework reveals reasoning dynamics in language models, significantly improving mathematical problem-solving accuracy.
Ming Li, Chenrui Fan, Yize Cheng et al.
ThinkARM framework reveals reasoning dynamics in language models, significantly improving mathematical problem-solving accuracy.
Ming Li, Chenrui Fan, Yize Cheng et al.
Introduced OpenBench, a multi-sensor outdoor benchmark, revealing current models' poor generalization in real-world spatial reasoning, with performance gaps highlighted by specific metrics.
Mingrui Wu, Zhaozhi Wang, Fangjinhua Wang et al.
VA-π employs variational policy optimization to align autoregressive image models with pixel distribution, reducing FID from 14.36 to 7.65 with minimal data and time.
Xinyao Liao, Qiyuan He, Kai Xu et al.
Proposes an optimal-coupling observer framework with LMI-based design to detect and reject bounded sensor attacks, ensuring vehicle safety and ride comfort.
Farzam Tajdari, Georgios Papaioannou, Riender Happee
StoryMem employs Memory-to-Video (M2V) to convert pretrained single-shot diffusion models into multi-shot storytelling tools, significantly improving cross-shot consistency with 28.7% gain on ST-Bench.
Kaiwen Zhang, Liming Jiang, Angtian Wang et al.
This review discusses diffusion models in simulation-based inference, focusing on training, inference, and evaluation, highlighting guidance, score composition, and flow matching techniques.
Jonas Arruda, Niels Bracher, Ullrich Köthe et al.
Spectral shrinkage of Gaussian entropic OT enables direct algebraic computation, improving efficiency and understanding of infinite-dimensional degeneracy.
Ho Yun
EchoTrail-GUI enhances GUI agents by critic-guided self-exploration, significantly improving task success rates.
Runze Li, Yuwen Zhai, Bo Xu et al.
Signal-SGN++ achieves skeleton-based action recognition with topology-enhanced time-frequency spiking graph network, significantly reducing energy consumption.
Naichuan Zheng, Xiahai Lun, Weiyi Li et al.
Proposes GenSDR, leveraging generative models with conditional velocity fields for nonlinear SDR, ensuring full information recovery at both sample and population levels.
Shuntuo Xu, Zhou Yu, Jian Huang
Explores teaching and critiquing conceptualization and operationalization in NLP, emphasizing interdisciplinary reading and discussion.
Vagrant Gautam
SWE-EVO benchmark assesses AI coding agents' ability to perform long-term software evolution tasks; top models reach only 25%, exposing significant gaps.
Tue Le, Minh V. T. Thai, Dung Nguyen Manh et al.
RadarGen uses latent diffusion to generate realistic automotive radar point clouds from multi-view images, guided by BEV and pretrained models.
Tomer Borreda, Fangqiang Ding, Sanja Fidler et al.
YOTO framework combines differentiable subset selection and multi-task learning for efficient gene selection and prediction in single-cell transcriptomics.
Daphné Chopard, Jorge da Silva Gonçalves, Irene Cannistraci et al.
Seed-Prover 1.5 combines agentic RL and test-time scaling, solving 87.9% of PutnamBench.
Jiangjie Chen, Wenxiang Chen, Jiacheng Du et al.
Mitty leverages diffusion transformers for end-to-end human-to-robot video synthesis, achieving state-of-the-art results.
Yiren Song, Cheng Liu, Weijia Mao et al.
Diffusion Maps Kernel Ridge Regression (DM-KRR) enables long-term prediction of complex dynamical systems, outperforming state-of-the-art methods with higher accuracy and data efficiency.
Jiwoo Song, Daning Huang, John Harlim
DTDR leverages dynamic tool dependency modeling, boosting function calling success by 23%-104% through context-aware retrieval.
Bhrij Patel, Davide Belli, Amir Jalalirad et al.
Introduces LongShOTBench, a long-video multi-modal reasoning benchmark with rubric-based scoring, evaluating 105 models, with top score of 66.64%.
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Jinxing Zhou et al.
Posterior Behavioral Cloning (PostBC) models demonstrator behavior's posterior, improving coverage and RL finetuning efficiency.
Andrew Wagenmaker, Perry Dong, Raymond Tsao et al.