DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference
DepCap uses cross-step dependency and conflict-aware decoding, reaching up to 5.63× speedup with negligible quality loss.
Xiang Xia, Wuyang Zhang, Jiazheng Liu et al.
DepCap uses cross-step dependency and conflict-aware decoding, reaching up to 5.63× speedup with negligible quality loss.
Xiang Xia, Wuyang Zhang, Jiazheng Liu et al.
Muon optimizer outperforms AdamW in MLP-based tabular deep learning, recommended if training efficiency is acceptable.
Yury Gorishniy, Ivan Rubachev, Dmitrii Feoktistov et al.
Analyzes stability and generalization of looped transformers using a fixed-point framework, validated on chess, sudoku, and prefix-sum tasks.
Asher Labovich
Proposes DyMETER, integrating instance-aware parameter migration and dynamic thresholding to enhance online anomaly detection under concept drift.
Jiaqi Zhu, Shaofeng Cai, Jie Chen et al.
GUI-Perturbed employs domain randomization to reveal systematic weaknesses in GUI grounding models, especially in spatial reasoning and robustness, with accuracy drops of 27-56%.
Yangyue Wang, Harshvardhan Sikka, Yash Mathur et al.
π-Play combines self-play and privileged self-distillation, achieving 2-3x efficiency improvement in multi-agent evolution.
Yaocheng Zhang, Yuanheng Zhu, Wenyue Chong et al.
ASTER generates pseudo-anomalies in latent space, combining Transformer and pre-trained LLMs for unsupervised time-series anomaly detection, outperforming existing methods.
Romain Hermary, Samet Hicsonmez, Dan Pineau et al.
Proves that a single-layer linear Transformer is mathematically equivalent to Ordinary Least Squares (OLS) regression via spectral decomposition, enabling one-pass statistical inference.
Xiaojun Tan, Yuchen Zhao
Proposes Proxy Compression Hypothesis (PCH) to explain reward hacking, emphasizing goal compression, optimization amplification, and evaluator-policy co-adaptation.
Xiaohua Wang, Muzhao Tian, Yuqi Zeng et al.
Propose a lightweight framework using KL divergence to analyze quantization sensitivity in mixed-precision SSM-Transformer models.
Jason Kong, Nilesh Prasad Pandey, Flavio Ponzina et al.
This paper systematically analyzes the training dynamics of on-policy distillation (OPD) for large language models, revealing that success hinges on thinking-pattern alignment and new capabilities.
Yaxuan Li, Yuxin Zuo, Bingxiang He et al.
Parcae, a stable looped language model constrained by spectral norm limits, reduces perplexity by 6.3% and surpasses Transformer baselines at 1.3B parameters.
Hayden Prairie, Zachary Novack, Taylor Berg-Kirkpatrick et al.
K-STEMIT integrates geometric, temporal, and physical priors in a multi-branch GNN for ice layer thickness estimation, reducing RMSE by 21%.
Zesheng Liu, Maryam Rahnemoonfar
Proposed Neighborhood Transformer (NT) with switchable attention outperforms SOTA in node classification, especially on heterophilic graphs, reducing complexity significantly.
Yi Luo, Xu Sun, Guangchun Luo et al.
We propose a feedforward graph architecture using frozen large language models as nodes, achieving 87.3% accuracy on ARC-Challenge.
Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
MENO combines improved MeanFlow with neural operators for high-resolution dynamical system prediction, achieving 2× spectral accuracy and 14× faster inference.
Tianyue Yang, Xiao Xue
Proposed an expert-wise mixed precision quantization method based on router l2 norm changes, improving model accuracy and inference efficiency.
Mohammed Nowaz Rabbani Chowdhury, Kaoutar El Maghraoui, Hsinyu Tsai et al.
Proposes CMRM, a quantile-calibrated framework that enhances robustness under label noise without prior knowledge, improving accuracy by up to 3.39%.
Yuanjie Shi, Peihong Li, Zijian Zhang et al.
MO-RiskVAE enhances survival risk prediction in multiple myeloma by tuning latent regularization and structure, achieving higher C-index scores.
Zixuan Chen, Heng Zhang, YuPeng Qin et al.
ClawArena evaluates AI agents in evolving info environments, focusing on multi-source conflict reasoning, belief revision, and implicit personalization, across 12 scenarios with 337 rounds.
Haonian Ji, Kaiwen Xiong, Siwei Han et al.