Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once
Proposes the Magic-Madness-Heaven-Sin framework, categorizing LLM output diversity by task goals, revealing trade-offs across contexts.
Harnoor Dhingra
Proposes the Magic-Madness-Heaven-Sin framework, categorizing LLM output diversity by task goals, revealing trade-offs across contexts.
Harnoor Dhingra
Using Qwen3:30b and others to annotate PersuasionForGood, guilt induction reduces donation rates by 23 percentage points.
Tatiana Petrova, Stanislav Sokol, Radu State
Exons-Detect uses hidden-state discrepancies to identify exonic tokens, boosting AI text detection robustness without training.
Xiaowei Zhu, Yubing Ren, Fang Fang et al.
Study finds RAG system improvements in retrieval do not guarantee better QA performance in AI policy analysis.
Saahil Mathur, Ryan David Rittner, Vedant Ajit Thakur et al.
MARCH framework significantly reduces LLM hallucination using multi-agent reinforced self-check, enhancing factual consistency in an 8B parameter model.
Zhuo Li, Yupeng Zhang, Pengyu Cheng et al.
Self-distillation can degrade LLMs' reasoning in math by suppressing uncertainty expression.
Jeonghye Kim, Xufang Luo, Minbeom Kim et al.
CAPITU uses literary texts to design 59 automatically verifiable instructions, evaluating LLMs' instruction-following in Portuguese with cultural and morphological constraints.
Giovana Kerche Bonás, Roseval Malaquias Junior, Marcos Piau et al.
TiCo method significantly enhances time control in dialogue models using Spoken Time Markers, reducing MAE to 4.54 seconds.
Kai-Wei Chang, Wei-Chih Chen, En-Pei Hu et al.
MemDLM embeds a simulated denoising process into training via bi-level optimization, enhancing DLM training efficiency and long-context understanding.
Zehua Pei, Hui-Ling Zhen, Weizhe Lin et al.
TableLong method enhances long-context reasoning with table data, achieving an average improvement of 8.24%.
Huaibing Xie, Guoliang Zhao, Yang Liu et al.
VARS introduces dual vectors and weak rewards for online user preference learning, improving multi-session interaction efficiency by 3.2% success over baselines.
Yuren Hao, Shuhaib Mehri, ChengXiang Zhai et al.
PUPPET framework predicts LLM-driven belief shifts with r=0.3-0.5, revealing systematic biases.
Jocelyn Shen, Amina Luvsanchultem, Jessica Kim et al.
Semantic Token Clustering (STC) method achieves efficient uncertainty quantification in large language models, significantly reducing computational overhead.
Qi Cao, Andrew Gambardella, Takeshi Kojima et al.
Study of SFT-DPO interaction in small models reveals full fine-tuning outperforms LoRA.
Yuming Feng, Christy Yang
Proposes Recoding-Decoding (RD) to enhance LLMs' sustained creativity and diversity, surpassing modal decoding limits.
Queenie Luo, Gary King, Michael Puett et al.
F2LLM-v2 offers efficient multilingual embeddings using a two-stage training and matryoshka learning, supporting over 200 languages.
Ziyin Zhang, Zihan Liao, Hang Yu et al.
Nemotron-Cascade 2 achieves top-tier reasoning with Cascade RL and multi-domain distillation in a 30B MoE model.
Zhuolin Yang, Zihan Liu, Yang Chen et al.
VEPO enhances translation quality and tokenization efficiency for low-resource languages using reinforcement learning with verifiable rewards.
Chonghan Liu, Yimin Du, Qi An et al.
MoRI employs reinforcement learning with entropy-aware information gain and semantic contrast to enhance scientific ideation depth and validity.
Chenyang Gu, Jiahao Cheng, Meicong Zhang et al.
RADIUS evaluates survey simulation alignment via ranking and distribution metrics with statistical significance, improving robustness over traditional methods.
Weronika Łajewska, Paul Missault, George Davidson et al.