cs.LG 2601.19897

Self-Distillation Enables Continual Learning

Introduces Self-Distillation Fine-Tuning (SDFT), enabling continual learning with reduced catastrophic forgetting and improved new-task accuracy.

Idan Shenfeld, Mehul Damani, Jonas Hübotter et al.

2026-01-28 46
cs.LG 2601.19895

Post-LayerNorm Is Back: Stable, ExpressivE, and Deep

Keel integrates Highway connections into Post-LayerNorm, enabling stable training of Transformer models over 1000 layers, outperforming Pre-LN in stability and expressivity.

Chen Chen, Lai Wei

2026-01-28 6 citations 32
cs.IR 2601.19501

Masked Diffusion Generative Recommendation

MDGR uses masked diffusion for SID generation, improving over SOTA by up to 10.78% and increasing online revenue by 1.20%.

Lingyu Mu, Hao Deng, Haibo Xing et al.

2026-01-27 35