Training Language Model to Critique for Better Refinement
RCO framework trains critic models via refinement signals, improving critique quality and response refinement by 15% on average.
Tianshu Yu, Chao Xiang, Mingchuan Yang et al.
RCO framework trains critic models via refinement signals, improving critique quality and response refinement by 15% on average.
Tianshu Yu, Chao Xiang, Mingchuan Yang et al.
FineWeb2 pipeline enables automatic filtering and deduplication supporting over 1000 languages, improving multilingual model performance.
Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec et al.
Cognitive models reveal value trade-offs in language models using RSA for polite speech analysis.
Sonia K. Murthy, Rosie Zhao, Jennifer Hu et al.
Proposes multi-perspective soft label models to improve inclusivity and performance in subjective NLP tasks.
Benedetta Muscato, Lucia Passaro, Gizem Gezici et al.
IndexTTS2 achieves precise duration control and emotional expressiveness in zero-shot autoregressive TTS, outperforming state-of-the-art models with specific metrics.
Siyi Zhou, Yiquan Zhou, Yi He et al.
This study systematically evaluates 17 Indic languages using BPE and Unigram LM, highlighting script normalization and clustering as key improvements.
Maharaj Brahma, N J Karthika, Rajat Verma et al.
The study reveals quality issues in multilingual speech datasets, emphasizing the need for sociolinguistic awareness and proactive language planning.
Mingfei Lau, Qian Chen, Yeming Fang et al.
MemBench: a multi-scenario, multi-level benchmark evaluating LLMs' memory via accuracy, recall, capacity, and efficiency.
Haoran Tan, Zeyu Zhang, Chen Ma et al.
Introduces autoregressive U-Net for multi-scale byte embedding, matching BPE performance, enhancing low-resource and multilingual NLP.
Mathurin Videau, Badr Youbi Idrissi, Alessandro Leite et al.
AgentSynth uses chain-based subtask generation to create over 6000 diverse, high-quality computer tasks, enabling scalable dataset expansion.
Jingxu Xie, Dylan Xu, Xuandong Zhao et al.
VL-GenRM enhances vision-language verification using vision experts and iterative training, significantly improving multimodal reasoning.
Jipeng Zhang, Kehao Miao, Renjie Pi et al.
MiniMax-M1 combines hybrid MoE architecture with Lightning Attention, enabling efficient long-context reasoning with 456B parameters and 1 million token support.
MiniMax, :, Aili Chen et al.
Konooz corpus covers 16 Arabic dialects and 10 domains, revealing a 38% performance drop in NER models.
Nagham Hamad, Mohammed Khalilia, Mustafa Jarrar
DeepResearch Bench offers 100 PhD-level tasks to evaluate deep research agents' report generation and information retrieval capabilities.
Mingxuan Du, Benfeng Xu, Chiwei Zhu et al.
ClaimSpect uses retrieval-augmented hierarchical analysis to decompose nuanced claims, identifying multi-angle evidence and perspectives.
Priyanka Kargupta, Runchu Tian, Jiawei Han
PersonaLens uses multi-task, multi-domain user profiles and two LLM agents to evaluate personalization in task-oriented dialogue systems.
Zheng Zhao, Clara Vania, Subhradeep Kayal et al.
Analyzed 40 LLM uncertainty quantification methods, highlighting evaluation on non-realistic benchmarks and advocating human-centered assessment.
Siddartha Devic, Tejas Srinivasan, Jesse Thomason et al.
Instruction-based text embedding framework using Mistral-7B, combining soft supervision and adaptive hard-negative mining for multi-task performance.
Jooyoung Choi, Hyun Kim, Hansol Jang et al.
Lingshu: a unified multimodal medical foundation model with extensive knowledge integration and state-of-the-art performance.
LASA Team, Weiwen Xu, Hou Pong Chan et al.
Proposes a five-dimensional framework distinguishing AI agents from LLM chatbots, based on evolutionary analysis of environment and capabilities.
Jiachen Zhu, Menghui Zhu, Renting Rui et al.