cs.LG 2305.17560

Scalable Transformer for PDE Surrogate Modeling

Proposed FactFormer uses axial factorized kernel integral in Transformer for scalable high-dimensional PDE surrogate modeling.

Zijie Li, Dule Shu, Amir Barati Farimani

2023-05-28 36
cs.LG 2305.17333

Fine-Tuning Language Models with Just Forward Passes

MeZO, a memory-efficient zeroth-order optimizer, enables fine-tuning of 30B models with 12× less memory, matching performance of backpropagation.

Sadhika Malladi, Tianyu Gao, Eshaan Nichani et al.

2023-05-27 461 citations 31
cs.LG 2305.17126

Large Language Models as Tool Makers

LATM framework uses GPT-4 to generate reusable Python tools, reducing inference costs significantly.

Tianle Cai, Xuezhi Wang, Tengyu Ma et al.

2023-05-27 60
cs.SD 2305.15719

Efficient Neural Music Generation

Proposes MeLoDy, an efficient diffusion-based neural music generator reducing inference steps by 95.7%/99.6%, maintaining high quality.

Max W. Y. Lam, Qiao Tian, Tang Li et al.

2023-05-25 94 citations 37
cs.LG 2305.15555

Deep Reinforcement Learning with Plasticity Injection

Plasticity injection enhances neural network adaptability in deep RL, boosting performance by 20% on Atari without increasing trainable parameters.

Evgenii Nikishin, Junhyuk Oh, Georg Ostrovski et al.

2023-05-25 48