Process Reinforcement through Implicit Rewards
PRIME introduces implicit process rewards for online reward model updates, boosting multi-step reasoning by 15.1% on benchmarks.
Ganqu Cui, Lifan Yuan, Zefan Wang et al.
PRIME introduces implicit process rewards for online reward model updates, boosting multi-step reasoning by 15.1% on benchmarks.
Ganqu Cui, Lifan Yuan, Zefan Wang et al.
Proposes InfoBridge, an unbiased mutual information estimator using diffusion bridge matching, effective for high-dimensional complex data.
Sergei Kholkin, Ivan Butakov, Evgeny Burnaev et al.
STP employs self-play with iterative conjecturing and proving, achieving 28.5% proof rate, doubling prior best.
Kefan Dong, Tengyu Ma
This study analyzes the role of feed-forward layers in Transformer-based in-context nonlinear learning, proposing a GLU-based model for polynomial functions.
Haoyuan Sun, Ali Jadbabaie, Navid Azizan
HLE introduces a 2,500-question multi-modal academic benchmark, exposing current LLM limitations in complex reasoning and cross-disciplinary understanding.
Long Phan, Alice Gatti, Ziwen Han et al.
This survey reviews LLMs in EDA, focusing on architectures, model size effects, and customization techniques, highlighting their impact on chip design automation.
Jingyu Pan, Guanglei Zhou, Chen-Chia Chang et al.
TempoGPT uses vector quantization of temporal embeddings to enhance multimodal time series reasoning, outperforming existing models.
Haochuan Zhang, Chunhua Yang, Jie Han et al.
COMIX method explains ML model decisions by decomposing images into prototypes, achieving a 48.82% improvement in C-insertion score.
Sarath Sivaprasad, Dmitry Kangin, Plamen Angelov et al.
RoRA optimizes LoRA's scaling factor to α/√r, significantly improving fine-tuning accuracy on large and pruned models.
Jun Liu, Zhenglun Kong, Peiyan Dong et al.
Using Vaidya's plane cutting method, the paper proposes an optimal accuracy-communication-privacy trade-off algorithm for distributed DP convex optimization.
Sudeep Salgia, Nikola Pavlovic, Yuejie Chi et al.
Proposes InfAlign, optimizing inference win rate via reward transformation, achieving 3-8% improvements.
Ananth Balashankar, Ziteng Sun, Jonathan Berant et al.
Proposes SympFlow, a time-dependent symplectic neural network using parameterized Hamiltonian flows, improving energy conservation and long-term stability.
Priscilla Canizares, Davide Murari, Carola-Bibiane Schönlieb et al.
Proposed an input perturbation method 'forget vector' for machine unlearning without altering model weights.
Changchang Sun, Ren Wang, Yihua Zhang et al.
OREO method enhances LLM multi-step reasoning, outperforming on GSM8K and MATH datasets.
Huaijie Wang, Shibo Hao, Hanze Dong et al.
Graph mixing analysis reveals how local and global strategies affect membership inference attack vulnerability in decentralized learning.
Ousmane Touat, Jezekael Brunon, Yacine Belal et al.
Entropy-regularized process reward model (ER-PRM) with KL regularization improves mathematical reasoning by guiding multi-step inference.
Hanning Zhang, Pengcheng Wang, Shizhe Diao et al.
METIS jointly optimizes query scheduling and configuration adaptation, reducing RAG latency by 1.64-2.54× while maintaining high quality.
Siddhant Ray, Rui Pan, Zhuohan Gu et al.
HyperForget uses hypernetworks for machine unlearning, maintaining high accuracy on retain sets.
Jose Miguel Lara Rangel, Stefan Schoepf, Jack Foster et al.
CLOVER leverages cross-layer SVD-based orthogonal vectors for pruning and fine-tuning, achieving high compression with minimal performance loss.
Fanxu Meng, Pingzhi Tang, Fan jiang et al.
Post-hoc regression-based λ optimization reduces variance, improving PPI in few-label settings.
Benjamin Eyre, David Madras