cs.LG 2510.02259

Transformers Discover Molecular Structure Without Graph Priors

This study demonstrates that a standard Transformer, trained directly on Cartesian coordinates without graph priors, can achieve energy and force prediction accuracy comparable to state-of-the-art equivariant GNNs on OMol25, with faster inference.

Tobias Kreiman, Yutong Bai, Fadi Atieh et al.

2025-10-03 49
cs.IR 2510.01149

ModernVBERT: Towards Smaller Visual Document Retrievers

ModernVBERT is a 250M-parameter vision-language encoder that outperforms larger models in document retrieval by optimizing attention masks, image resolution, and training strategies.

Paul Teiletche, Quentin Macé, Max Conti et al.

2025-10-02 37
cs.LG 2510.01132

A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

This study introduces a systematic multi-turn RL framework for large language models, emphasizing environment, reward, and policy pillars, validated across TextWorld, ALFWorld, and SWE-Gym with key improvements of up to 88%.

Ruiyi Wang, Prithviraj Ammanabrolu

2025-10-02 63