cs.CL 2504.03553

Agentic Knowledgeable Self-awareness

Introduces KnowSelf for LLMs' situational self-awareness, using two-stage training to enhance planning with minimal external knowledge.

Shuofei Qiao, Zhisong Qiu, Baochang Ren et al.

2025-04-05 40
cs.CL 2504.02441

Cognitive Memory in Large Language Models

Proposes multi-layered memory combining text, KV cache, parameters, and hidden states to enhance long-term memory in LLMs.

Lianlei Shan, Shixian Luo, Zezhou Zhu et al.

2025-04-03 22
cs.CL 2504.00050

JudgeLRM: Large Reasoning Models as a Judge

JudgeLRM employs reinforcement learning with outcome-driven rewards to surpass SFT models, with 7B/8B models exceeding GPT-4 in F1 scores.

Nuo Chen, Zhiyuan Hu, Qingyun Zou et al.

2025-03-31 33
cs.CL 2503.21295

R-PRM: Reasoning-Driven Process Reward Modeling

Proposes R-PRM, leveraging stronger models for seed data, preference optimization, and multi-trajectory inference, boosting step evaluation F1 by 11.9 points.

Shuaijie She, Junxiao Liu, Yifeng Liu et al.

2025-03-27 48 citations 39
cs.CL 2503.19786

Gemma 3 Technical Report

Gemma 3 introduces multimodal, 128K tokens long context, and vision understanding, using architecture improvements and distillation, outperforming Gemma 2.

Gemma Team, Aishwarya Kamath, Johan Ferret et al.

2025-03-25 29
cs.CL 2503.17489

Judge Anything: MLLM as a Judge Across Any Modality

This paper introduces JudgeAnything, a benchmark leveraging multimodal large language models (MLLMs) as automated judges across 15 diverse tasks, revealing strong performance in understanding but limitations in generation.

Shu Pu, Yaochen Wang, Dongping Chen et al.

2025-03-22 33 citations 55