cs.LG 2405.04517

xLSTM: Extended Long Short-Term Memory

xLSTM introduces exponential gating and novel memory structures, scaling to billions of parameters with superior performance over Transformers.

Maximilian Beck, Korbinian Pöppel, Markus Spanring et al.

2024-05-08 30
cs.LG 2404.19756

KAN: Kolmogorov-Arnold Networks

KAN, based on Kolmogorov-Arnold theorem, replaces fixed weights with learnable spline functions, outperforming MLP in accuracy and interpretability.

Ziming Liu, Yixuan Wang, Sachin Vaidya et al.

2024-05-01 38
cs.LG 2404.09411

Wasserstein Wormhole: Scalable Optimal Transport Distance with Transformers

Wasserstein Wormhole, a transformer-based autoencoder, embeds distributions into a latent space where Euclidean distances approximate Wasserstein distances, enabling linear-time OT computations for large datasets.

Doron Haviv, Russell Zhang Kunes, Thomas Dougherty et al.

2024-04-15 19 citations 36