cs.CL 2505.11820

Chain-of-Model Learning for Language Model

CoLM enables progressive scaling and elastic inference; CoLM-Air reaches about 3× faster 1M-token prefilling.

Kaitao Song, Xiaohua Wang, Xu Tan et al.

2025-05-17 27
cs.CL 2505.13508

Time-R1: Towards Comprehensive Temporal Reasoning in LLMs

Time-R1 employs a three-stage RL fine-tuning framework to endow a 3B-parameter LLM with comprehensive temporal reasoning, outperforming models over 200 times larger in future event prediction and creative scenario generation.

Zijia Liu, Peixuan Han, Haofei Yu et al.

2025-05-16 27 citations 39
cs.CL 2505.09388

Qwen3 Technical Report

Qwen3 integrates thinking and non-thinking modes, with 235B parameters, enhancing multilingual and multi-task performance.

An Yang, Anfeng Li, Baosong Yang et al.

2025-05-14 34
cs.CL 2505.02387

RM-R1: Reward Modeling as Reasoning

RM-R1 formulates reward modeling as a reasoning task using chain-of-thought and Rubrics, outperforming larger models by up to 4.9%.

Xiusi Chen, Gaotang Li, Ziqi Wang et al.

2025-05-05 31
cs.CL 2505.00662

DeepCritic: Deliberate Critique with Large Language Models

DeepCritic introduces a two-stage framework leveraging Qwen2.5-72B-Instruct to enhance mathematical critique, outperforming GPT-4o and DeepSeek with significant accuracy gains.

Wenkai Yang, Jingwen Chen, Yankai Lin et al.

2025-05-02 59