cs.CL 2601.01972

Hidden State Poisoning Attacks against Mamba-based Language Models

Introduces Hidden State Poisoning Attacks (HiSPA) against Mamba models, using short triggers to irreversibly overwrite hidden states and impair long-context retrieval, with experimental validation on ROBENCH-25.

Alexandre Le Mercier, Chris Develder, Thomas Demeester

2026-01-05 3 citations 37
cs.CL 2512.24880

mHC: Manifold-Constrained Hyper-Connections

mHC projects residual mappings onto the doubly stochastic matrix manifold, restoring identity properties and enhancing large-scale training stability.

Zhenda Xie, Yixuan Wei, Huanqi Cao et al.

2025-12-31 18
cs.CL 2512.16649

JustRL: Scaling a 1.5B LLM with a Simple RL Recipe

This study introduces JustRL, a simple fixed-hyperparameter RL training method for 1.5B models, surpassing complex approaches with more stability and less compute.

Bingxiang He, Zekai Qu, Zeyuan Liu et al.

2025-12-18 35 citations 38
cs.CL 2512.13564

Memory in the Age of AI Agents

Proposes a unified 'forms-functions-dynamics' framework for agent memory, covering token-level, parametric, and latent types with detailed functional and dynamic analysis.

Yuyang Hu, Shichun Liu, Yanwei Yue et al.

2025-12-16 31