Micro Language Models Enable Instant Responses
Micro Language Models (μLMs) enable instant responses by generating the first 4-8 words on-device, with cloud models completing the response.
Wen Cheng, Tuochao Chen, Karim Helwani et al.
Micro Language Models (μLMs) enable instant responses by generating the first 4-8 words on-device, with cloud models completing the response.
Wen Cheng, Tuochao Chen, Karim Helwani et al.
Study shows large language models impact AI conference peer reviews, especially in linguistic complexity and evaluative focus.
Wenqing Wu, Chengzhi Zhang, Yi Zhao et al.
The study reveals dual alignment between language model layers and human sentence processing, with early layers suited for natural reading and later layers better modeling complex syntactic processing.
Tatsuki Kuribayashi, Alex Warstadt, Yohei Oseki et al.
GSQ achieves high-accuracy low-bit quantization using Gumbel-Softmax sampling, narrowing the accuracy gap with QTIP methods.
Alireza Dadgarnia, Soroush Tabesh, Mahdi Nikdan et al.
Transition-matrix regularization improves next dialogue act prediction in counseling conversations, boosting macro-F1 by 9-42%.
Eric Rudolph, Philipp Steigerwald, Jens Albrecht
This study introduces the Adversarial Humanities Benchmark (AHB), revealing that stylistic transformations increase attack success rates from 3.84% to up to 65%, exposing weaknesses in model safety robustness.
Marcello Galisai, Susanna Cifani, Francesco Giarrusso et al.
ArbGraph enhances long-form RAG reliability through conflict-aware evidence arbitration, reducing hallucinations.
Qingying Niu, Yuhao Wang, Ruiyang Ren et al.
Study shows current LLM agents lack environmental curiosity, discovering solutions in 79-81% of cases but utilizing them in only 37-50%.
Leon Engländer, Sophia Althammer, Ahmet Üstün et al.
Token-efficient, self-evolving LLM agent using context density maximization with four core mechanisms.
Jiaqing Liang, Jinyi Han, Weijia Li et al.
Proposes grounded multi-turn social simulation with SSBC to analyze LLM support strategy shifts based on estimated user distress.
Michelle Star, Andrew Aquilina, Yu-Ru Lin
A dual-aspect evaluation framework analyzes LLMs on Vietnamese legal text, revealing readability-accuracy trade-offs.
Van-Truong Le
BAGEL benchmark evaluates language models' performance on animal knowledge using closed-book questions on taxonomy, morphology, etc.
Jiacheng Shen, Masato Hagiwara, Milad Alizadeh et al.
Self-distillation reduces fine-tuning-induced hallucinations, lowering factual forgetting from 15% to 3%.
Guy Kaplan, Zorik Gekhman, Zhen Zhu et al.
SpecGuard enhances multi-step reasoning efficiency and accuracy using internal signals for step-level verification.
Kiran Purohit, Ramasuri Narayanam, Soumyabrata Pal
MADE benchmark enhances multi-label text classification accuracy with uncertainty quantification, especially in medical device adverse events.
Raunak Agarwal, Markus Wenzel, Simon Baur et al.
LLMs generate excessive content in translations; detection strategies improve translation quality.
Lisa Vasileva, Karin Sim
DiscoTrace structures answers as discourse acts paired with question interpretations, revealing human-AI strategy differences.
Neha Srikanth, Jordan Boyd-Graber, Rachel Rudinger
MARCA benchmarks multilingual web search with 52 questions, using two frameworks, revealing large model performance gaps.
Thales Sales Almeida, Giovana Kerche Bonás, Ramon Pires et al.
Using GPT-4.1 to predict overall experience scores from fan text with 67%±1 accuracy, demonstrating stable and reliable measurement.
Jason Potteiger, Andrew Hong, Ito Zapata
Doc-V* enhances multi-page document VQA by sequential evidence aggregation, outperforming RAG baseline by 47.9%.
Yuanlei Zheng, Pei Fu, Hang Li et al.