Argument Collapse: LLMs Flatten Long-Form Public Debate
LLMs generate debate essays with only 3.4% unique main arguments, far below humans' 65.3%.
Yekyung Kim, Yapei Chang, Chau Minh Pham et al.
LLMs generate debate essays with only 3.4% unique main arguments, far below humans' 65.3%.
Yekyung Kim, Yapei Chang, Chau Minh Pham et al.
Introduced LongJudgeBench, a benchmark for evaluating LLMs as long-form judges, with an average output length over 9000 tokens, revealing significant instability.
Junjie Chen, Yuxi Dong, Haitao Li et al.
LongTraceRL enhances long-context reasoning by generating complex multi-hop questions via knowledge graph random walks, using search trajectories for layered distractors, and entity-level rubric rewards, achieving significant improvements.
Nianyi Lin, Jiajie Zhang, Lei Hou et al.
Using Cultural Consensus Theory (CCT) to analyze LLM alignment in single- and multi-cultural settings, revealing errors in diversity representation and over-homogenization.
Krishna Pothugunta, John P. Lalor
This paper introduces the 'Disagreeing Rationales' framework, systematically analyzing how diverse human annotations and explanations impact hate speech detection, emphasizing the benefits of soft labels and rationales.
Benedetta Muscato, Beiduo Chen, Gizem Gezici et al.
Proposed horizon-control strategies—Progressive OPD and Truncated OPD—significantly improve long-horizon on-policy distillation efficiency, achieving 3× faster training and comparable performance with only 10% rollout length.
Yaocheng Zhang, Jiajun Chai, Yuqian Fu et al.
BenHalluEval evaluates Bengali LLM hallucinations; BenHalluScore ranges from 7.72% to 55.42%.
Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham et al.
This study introduces a multi-turn multi-agent dialogue framework to evaluate VLMs in spatial reasoning, showing limited improvements mainly due to visual grounding challenges.
Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen
Introduced RHELM benchmark with multi-source heterogeneous data and dynamic user profiles to evaluate long-term memory in dialogue models.
Han Zhang, Zihao Tang, Xin Yu et al.
COFT employs counterfactual causal intervention with distribution-free marginal guarantees, reducing bias by 30-55% without retraining.
Arya Fayyazi, Mehdi Kamal, Massoud Pedram
LLMSurgeon formulates data mixture diagnosis as a label-shift inverse problem, achieving 94.46% accuracy on the LLMSurgeon benchmark.
Yaxin Luo, Jiacheng Cui, Xiaohan Zhao et al.
Proposes COMPOSE, a dual-graph framework combining citation and formal theorem graphs, generating plausible future theorems with 108K training pairs and 47K future papers tested.
David Busbib, Michael Werman
Proposes MedCase-Structured, a pipeline combining LLMs and terminology validation to generate HL7 FHIR R4 clinical datasets for diagnostic reasoning, with an 82.5% success rate.
Valentina Bui Muti, Eugénie Dulout, Ziquan Fu
Proposed dual-path block architecture enhances compute and capacity in LLMs at fixed FLOPs.
Markus Frey, Behzad Shomali, Joachim Koehler et al.
DirectorBench uses personalized multi-agent diagnosis and finds cross-shot transitions average only 0.256.
Jiamin Chen, Qianben Chen, Jiawen Zhang et al.
Proposed a causal intervention method for continuous variables, applied to verb bias, validating its role in in-context learning.
Zhenghao Herbert Zhou, R. Thomas McCoy, Robert Frank
Introduces PLANAHEAD framework, systematically evaluating natural language plan representations' impact on multimodal LLM web agent performance.
Alejandra Zambrano, Sara Vera Marjanovic, Imene Kerboua et al.
UKA introduces user-aware active knowledge acquisition using ToM uncertainty, boosting emotional support dialogue quality.
Mufan Xu, Kehai Chen, Jiahao Hu et al.
Proposes Draft-OPD, combining target supervision and error replay, achieving 5× speedup in large language model inference.
Haodi Lei, Yafu Li, Haoran Zhang et al.
STAMP framework trains explicit memory for mobile GUI agents using controllable virtual environments, enhancing Memory-World benchmark performance.
Junyang Wang, Haiyang Xu, Xi Zhang et al.