cs.LG 2605.18807

Block-Based Double Decoders

Proposes a block-based double decoder architecture combining full supervision training with inference efficiency, reducing memory and computation by over 66%.

Asher Labovich, Benjamin Bradley, Vanessa Alexander et al.

2026-05-12 51
cs.LG 2605.10878

Neural Weight Norm = Kolmogorov Complexity

Proves that the smallest weight norm of a neural network equals the Kolmogorov complexity of its output, explaining weight decay's effectiveness.

Tiberiu Musat

2026-05-12 21