e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
e3 trains models with chain-based exploration, enabling extrapolation to twice the training inference length, significantly improving reasoning performance beyond training limits.
Amrith Setlur, Matthew Y. R. Yang, Charlie Snell et al.