cs.CL 2502.10881

CiteCheck: Towards Accurate Citation Faithfulness Detection

CiteCheck introduces a large-scale Chinese citation faithfulness dataset using a two-stage manual annotation and LLM-based data augmentation, enabling efficient training of models with high robustness.

Ziyao Xu, Shaohang Wei, Zhuoheng Han et al.

2025-02-16 50
cs.CL 2502.03387

LIMO: Less is More for Reasoning

LIMO achieves complex reasoning with minimal data, scoring 63.3% on AIME24 and 95.6% on MATH500.

Yixin Ye, Zhen Huang, Yang Xiao et al.

2025-02-06 37
cs.CL 2501.19324

Reward-Guided Speculative Decoding for Efficient LLM Reasoning

Reward-Guided Speculative Decoding (RSD) combines lightweight draft models with reward signals to improve LLM inference efficiency, reducing up to 4.4× FLOPs while boosting accuracy.

Baohao Liao, Yuhui Xu, Hanze Dong et al.

2025-02-01 25
cs.CL 2501.16673

LLM-AutoDiff: Auto-Differentiate Any LLM Workflow

LLM-AutoDiff introduces a graph-based automatic prompt optimization framework using textual gradients, supporting multi-component and cyclic LLM workflows, significantly improving accuracy and efficiency.

Li Yin, Zhangyang Wang

2025-01-28 21 citations 55
cs.CL 2501.15915

Parametric Retrieval Augmented Generation

Proposes Parametric RAG, embedding external knowledge into model parameters for improved efficiency and performance.

Weihang Su, Yichen Tang, Qingyao Ai et al.

2025-01-27 28
cs.CL 2501.13912

Analysis of Indic Language Capabilities in LLMs

Analyzed 28 Indian language-supporting LLMs; found significant performance disparities, especially in low-resource languages.

Aatman Vaidya, Tarunima Prabhakar, Denny George et al.

2025-01-24 5 citations 54
cs.CL 2501.13956

Zep: A Temporal Knowledge Graph Architecture for Agent Memory

Zep employs Graphiti's time-aware knowledge graph, dynamically integrating conversational and business data, achieving 94.8% accuracy and 90% latency reduction, outperforming MemGPT.

Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais et al.

2025-01-21 39