xLSTM: Extended Long Short-Term Memory
xLSTM introduces exponential gating and novel memory structures, scaling to billions of parameters with superior performance over Transformers.
Maximilian Beck, Korbinian Pöppel, Markus Spanring et al.
xLSTM introduces exponential gating and novel memory structures, scaling to billions of parameters with superior performance over Transformers.
Maximilian Beck, Korbinian Pöppel, Markus Spanring et al.
ML combines deep neural networks and symbolic regression to uncover unknown physical laws, significantly advancing scientific discovery in complex systems.
Ricardo Vinuesa, Paola Cinnella, Jean Rabault et al.
VoxBind employs a 3D voxel score-based model for efficient, diverse molecule generation conditioned on protein pockets.
Pedro O. Pinheiro, Arian Jamasb, Omar Mahmood et al.
Proposes GOT-D, an OT-gradient-based data selection method, improving fine-tuning efficiency and performance by up to 13.9%.
Feiyang Kang, Hoang Anh Just, Yifan Sun et al.
Foundation models enhance time series analysis, addressing cross-task transfer issues.
Jiexia Ye, Yongzi Yu, Weiqi Zhang et al.
Proposes Large Tabular Model (LTM) as a foundational framework for structured data, emphasizing multi-task, cross-dataset, and generative capabilities.
Boris van Breugel, Mihaela van der Schaar
KAN, based on Kolmogorov-Arnold theorem, replaces fixed weights with learnable spline functions, outperforming MLP in accuracy and interpretability.
Ziming Liu, Yixuan Wang, Sachin Vaidya et al.
Proposes a multi-stage continual learning framework for LLMs, integrating vertical (task hierarchy) and horizontal (time/domain) dimensions to mitigate catastrophic forgetting.
Haizhou Shi, Zihao Xu, Hengyi Wang et al.
Proposes a supervised uncertainty estimation framework for LLMs using hidden layer activations and external features, outperforming unsupervised baselines.
Linyu Liu, Yu Pan, Xiaocheng Li et al.
Study on LLaMA3 quantization reveals performance drop at low bit-widths using methods like GPTQ and AWQ.
Wei Huang, Xingyu Zheng, Xudong Ma et al.
PROSE-PDE integrates multi-operator learning with symbolic encoding, enabling strong generalization and extrapolation for PDE solutions.
Jingmin Sun, Yuxuan Liu, Zecheng Zhang et al.
Proposes SOBER, a kernel quadrature-based batch Bayesian optimization framework with probabilistic lifting, supporting discrete, non-Euclidean spaces, and adaptive batch sizes.
Masaki Adachi, Satoshi Hayakawa, Martin Jørgensen et al.
Wasserstein Wormhole, a transformer-based autoencoder, embeds distributions into a latent space where Euclidean distances approximate Wasserstein distances, enabling linear-time OT computations for large datasets.
Doron Haviv, Russell Zhang Kunes, Thomas Dougherty et al.
Proposes DS-MoE: dense training with sparse inference, achieving high parameter and computational efficiency.
Bowen Pan, Yikang Shen, Haokun Liu et al.
SiD distills pretrained diffusion models into one-step generators with near-exponential FID drops and teacher-level quality.
Mingyuan Zhou, Huangjie Zheng, Zhendong Wang et al.
Mixture-of-Depths method dynamically allocates compute, boosting Transformer inference speed by 50%.
David Raposo, Sam Ritter, Blake Richards et al.
This survey introduces a four-role taxonomy of LLMs in RL—information processor, reward designer, decision-maker, generator—improving sample efficiency and generalization.
Yuji Cao, Huan Zhao, Yuheng Cheng et al.
PEFT fine-tunes large models by adjusting minimal parameters, reducing training costs by over 80%.
Zeyu Han, Chao Gao, Jinyang Liu et al.
Proposes a multi-model ensemble and multi-criteria decision approach to select optimal counterfactual explanations, improving trade-offs among quality metrics.
Ignacy Stępka, Mateusz Lango, Jerzy Stefanowski
Proposes a probabilistic forecasting framework using stochastic interpolants and Föllmer processes, leveraging SDEs for high-dimensional condition sampling.
Yifan Chen, Mark Goldstein, Mengjian Hua et al.