GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CL 2512.11399

Minimal Clips, Maximum Salience: Long Video Summarization via Key Moment Extraction

Proposes lightweight clip selection combined with large language models to extract key moments, achieving near-reference summary quality with less than 6% video content.

Galann Pennec, Zhengyuan Liu, Nicholas Asher et al.

2025-12-12 46
cs.CL 2512.10791

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

The FACTS Leaderboard evaluates large language models' factuality using four sub-leaderboards, with an average score of 68.8.

Aileen Cheng, Alon Jacovi, Amir Globerson et al.

2025-12-12 4
cs.CL 2512.09742

Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs

Fine-tuning on narrow datasets causes broad, unpredictable behaviors and hidden backdoors, exposing security risks in LLMs.

Jan Betley, Jorio Cocola, Dylan Feng et al.

2025-12-10 35
cs.CL 2512.07075

Do Large Language Models Truly Understand Cross-cultural Differences?

SAGE benchmark evaluates LLMs' cross-cultural understanding via 9 dimensions and 4,530 tasks, exposing systematic deficiencies.

Shiwei Guo, Sihang Jiang, Qianxi He et al.

2025-12-08 34
cs.CL 2512.02556

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

DeepSeek-V3.2 employs sparse attention and large-scale reinforcement learning to match GPT-5's reasoning and agent capabilities, significantly narrowing the open-source gap.

DeepSeek-AI, Aixin Liu, Aoxue Mei et al.

2025-12-02 717 citations 56
cs.CL 2512.01725

Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks

Introducing reasoning overconfidence in LLMs; Long-CoT reduces it via iterative reflection, based on the cognitive rigidity hypothesis.

Jiannan Guan, Qiguang Chen, Libo Qin et al.

2025-12-01 51
cs.CL 2511.21686

Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework

Decentralized peer-to-peer framework Matrix boosts multi-agent synthetic data generation throughput by 2-15× via message-driven asynchronous scheduling.

Dong Wang, Yang Li, Ansong Ni et al.

2025-11-27 2 citations 31
cs.CL 2511.21437

A Systematic Study of In-the-Wild Model Merging for Large Language Models

Task Arithmetic is the only method that reliably improves LLM performance in 'in-the-wild' settings.

Oğuz Kağan Hitit, Leander Girrbach, Zeynep Akata

2025-11-26 4
cs.CL 2511.21066

Context-Aware Pragmatic Metacognitive Prompting for Sarcasm Detection

Proposed retrieval-aware PMP enhances sarcasm detection, achieving 9.87% macro-F1 improvement on Twitter Indonesia dataset.

Michael Iskandardinata, William Christian, Derwin Suhartono

2025-11-26 41
cs.CL 2511.20857

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

Evo-Memory evaluates self-evolving memory in LLMs during test-time, achieving an average performance gain of 0.65 across diverse tasks.

Tianxin Wei, Noveen Sachdeva, Benjamin Coleman et al.

2025-11-26 124 citations 42
cs.CL 2511.19399

DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Introduces RLER, combining search and evolving rubrics, training DR Tulu-8B to outperform open-source models with 15.6% gain on benchmarks.

Rulin Shao, Akari Asai, Shannon Zejiang Shen et al.

2025-11-25 52
cs.CL 2511.17432

SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation

SMILE integrates semantic and lexical precision, achieving efficient and accurate QA evaluation.

Shrikant Kendre, Austin Xu, Honglu Zhou et al.

2025-11-22 33
cs.CL 2511.16664

Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs

Nemotron Elastic achieves efficient many-in-one reasoning by embedding nested submodels, reducing training costs by 360x.

Ali Taghibakhshi, Sharath Turuvekere Sreenivas, Saurav Muralidharan et al.

2025-11-21 3
cs.CL 2511.15304

Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models

Poetic style as a universal single-turn jailbreak, achieving 62% attack success across 25 LLMs, with high cross-domain transferability.

Piercosma Bisconti, Matteo Prandi, Federico Pierucci et al.

2025-11-19 56
cs.CL 2511.14460

Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning

Agent-R1 introduces step-level trajectory abstraction and flexible context management, enabling multi-turn reinforcement learning with diverse optimization strategies, achieving state-of-the-art results.

Mingyue Cheng, Shuo Yu, Daoyu Wang et al.

2025-11-18 31 citations 29
cs.CL 2511.13043

Spark-Prover-X1: Formal Theorem Proving Through Diverse Data Training

Spark-Prover-X1 enhances formal theorem proving via diverse data training, solving 27 problems on PutnamBench.

Xinyuan Zhou, Yi Lei, Xiaoyu Zhou et al.

2025-11-17 1
cs.CL 2511.12116

LLMLagBench: Identifying Temporal Training Boundaries in Large Language Models

LLMLagBench uses PELT change-point detection to identify training data cutoffs in LLMs, revealing many models' knowledge ends earlier than declared.

Piotr Pęzik, Konrad Kaczyński, Maria Szymańska et al.

2025-11-15 47
cs.CL 2511.10643

Black-Box On-Policy Distillation of Large Language Models

Introduces GAD, an adversarial framework for black-box LLM distillation, with Qwen2.5-14B approaching GPT-5-Chat performance.

Tianzhu Ye, Li Dong, Zewen Chi et al.

2025-11-14 44 citations 60
cs.CL 2511.10070

ADI-20: Arabic Dialect Identification dataset and models

ADI-20 dataset expands ADI-17 for Arabic Dialect Identification using ECAPA-TDNN and Whisper models.

Haroun Elleuch, Salima Mdhaffar, Yannick Estève et al.

2025-11-13 6
cs.CL 2511.09865

In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback

InTRO employs token-level exploration and self-feedback via KL divergence, boosting reasoning accuracy by up to 20% and producing more concise rationales.

Mingye Zhu, Yi Liu, Zheren Fu et al.

2025-11-13 17
Prev 1 ... 26 27 28 29 30 31 32 ... 92 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home