Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
Pion optimizer preserves spectrum via orthogonal equivalence transformation, enhancing LLM training stability.
Kexuan Shi, Hanxuan Li, Zeju Qiu et al.
Pion optimizer preserves spectrum via orthogonal equivalence transformation, enhancing LLM training stability.
Kexuan Shi, Hanxuan Li, Zeju Qiu et al.
Proposes a sparse-to-dense reward principle combining GRPO and OPD to enhance language model post-training.
Yuanda Xu, Hejian Sang, Zhengze Zhou et al.
MEME evaluates multi-entity and evolving memory tasks, exposing dependency reasoning failures in current systems.
Seokwon Jung, Alexander Rubinstein, Arnas Uselis et al.
The paper introduces a parameter-free online K-Means router leveraging geometric coupling for effective expert assignment, reducing load imbalance with only a slight perplexity increase.
Sagi Ahrac, Noya Hochwald, Mor Geva
KV-Fold: A training-free protocol for long-context inference achieving 100% exact-match retrieval.
Alireza Nadali, Patrick Cooper, Ashutosh Trivedi et al.
Attractor Models enhance language modeling and reasoning via fixed-point solving, improving training efficiency by 46.6% and accuracy by 19.7%.
Jacob Fein-Ashley, Paria Rashidinejad
Multi-stream LLMs unlock language models with parallel streams of thoughts, inputs, and outputs, enhancing efficiency and security.
Guinan Su, Yanwu Yang, Xueyan Li et al.
Proposes Linearized Graph Sequence Models (LGSM), decoupling information propagation from nonlinear processing to enhance long-range dependency learning.
Joël Mathys, Basil Rohner, Saku Peltonen et al.
Proposes COMPOSE framework: training preserves object geometry, inference dynamically composes slots, achieving state-of-the-art generalization in continual few-shot learning.
Phu-Quy Nguyen-Lam, Phu-Hoa Pham, Dao Sy Duy Minh et al.
D-PACE uses dynamic position-aware weights to improve parallel speculative decoding, boosting speed and output length.
Tianyu Wu, Yu Yao, Zhenting Qi et al.
Proposes a block-based double decoder architecture combining full supervision training with inference efficiency, reducing memory and computation by over 66%.
Asher Labovich, Benjamin Bradley, Vanessa Alexander et al.
Proposes an adaptive correction scheduling method that balances constraint projection timing and sampling fidelity, improving trajectory consistency with 75% fewer corrections.
Noah Trupin, Yexiang Xue
V4FinBench excels in corporate bankruptcy prediction using TabPFN and Llama-3-8B models.
Marcin Kostrzewa, Sebastian Tomczak, Roman Furman et al.
Proves that the smallest weight norm of a neural network equals the Kolmogorov complexity of its output, explaining weight decay's effectiveness.
Tiberiu Musat
Proposed ERD framework addresses representation degradation in diffusion model training, significantly improving efficiency and generation quality.
Zhipeng Yao, Dazhou Li, Zitong Zhang et al.
CAQ-ZO method eliminates query-time residuals in low-precision evaluations, enhancing NF4 model performance.
Yao Shu, Zilin Zhu
BCJR-QAT uses BCJR algorithm for differentiable quantization training, reducing PPL by 0.084 on WikiText-2.
Venugopalan Iyengar
Proposes the DICOLA framework, using recursive decomposition to improve causal structure learning with latent variables, significantly boosting efficiency.
Zheng Li, Feng Xie, Shenglan Nie et al.
C2LT-3D introduces interface-centric generative states to improve assembly robustness in open-world 3D structures by separating geometry, ownership, and seam relations.
Xiang Chen, Alexander Binder
Introduces Structured Spectral Propagator (SSP) for long-horizon PDE forecasting, reducing errors by up to 48.9%.
Xiaoxiao Lu, Ye Yuan, Jiahao Shi