Depth-Attention: Cross-Layer Value Mixing for Language Models
Depth-Attention enhances language models by cross-layer value mixing, improving accuracy by up to 2.3 points.
Boyi Zeng, Yiqin Hao, Zitong Wang et al.
Depth-Attention enhances language models by cross-layer value mixing, improving accuracy by up to 2.3 points.
Boyi Zeng, Yiqin Hao, Zitong Wang et al.
LifeSide uses multi-session Memory-Emotion-Environment loops to evaluate lifelong digital companions' long-term memory, understanding, privacy, and emotional support, based on 2000 personas and 111K tasks.
Yuqian Wu, Zhijie Deng, Wei Chen et al.
RAMPART optimizes LLM agent memory with priority-aware runtime transformation, enhancing task success.
Nikodem Tomczak
VCIFBench evaluates complex video instruction following, covering 306 tasks with multi-constraint challenges, revealing significant gaps in current models.
Huangchen Xu, Yuan Wu, Yi Chang
Proposed a difficulty-aware SFT-then-RL framework to enhance small language model reasoning.
Chongyang He, Rui Zhang, Zixuan Wang et al.
SEA-Embedding employs contrastive learning and distribution matching, trained solely on public data, achieving SOTA in Southeast Asian multilingual embeddings.
Peerat Limkonchotiwat, Raymond Ng, Sarana Nutanong et al.
SubFit introduces non-contiguous submodule replacement in LLMs, achieving superior compression with 84.6% accuracy at 25% sparsity, using residual fitting without retraining.
Elia Cunegatti, Marcus Vukojevic, Erik Nielsen et al.
Proposed Script-Normalized WER (SN-WER), reducing script mismatch inflation by up to 12% across five Indic languages, enhancing multi-script ASR evaluation accuracy.
Priyaranjan Pattnayak
SimSD employs a plug-and-play masking strategy to enable token-level speculative decoding in diffusion LLMs, achieving up to 7.46× speedup while maintaining quality.
Junxia Cui, Haotian Ye, Runchu Tian et al.
PaSBench-Video is a 740-video benchmark for testing proactive safety warning capabilities of video MLLMs; no model exceeds 20.0% on strictest metric.
Yusong Zhao, Yuejin Xie, Youliang Yuan et al.
DFlare enhances draft model capacity with layer-wise fusion, achieving over 5x LLM inference speedup.
Jiebin Zhang, Zhenghan Yu, Song Liu et al.
DocFormFlow improves formatting accuracy by 72.53% while reducing token consumption.
Shihao Rao, Liang Li, Jiapeng Liu et al.
Proposes the 'Decan' metric, using single-pass log-probabilities to measure diversity in AI and human outputs, achieving 0.846 on McDiv benchmark.
Matthew Khoriaty, David Williams-King, Shi Feng
LLMs generate debate essays with only 3.4% unique main arguments, far below humans' 65.3%.
Yekyung Kim, Yapei Chang, Chau Minh Pham et al.
Introduced LongJudgeBench, a benchmark for evaluating LLMs as long-form judges, with an average output length over 9000 tokens, revealing significant instability.
Junjie Chen, Yuxi Dong, Haitao Li et al.
LongTraceRL enhances long-context reasoning by generating complex multi-hop questions via knowledge graph random walks, using search trajectories for layered distractors, and entity-level rubric rewards, achieving significant improvements.
Nianyi Lin, Jiajie Zhang, Lei Hou et al.
Using Cultural Consensus Theory (CCT) to analyze LLM alignment in single- and multi-cultural settings, revealing errors in diversity representation and over-homogenization.
Krishna Pothugunta, John P. Lalor
This paper introduces the 'Disagreeing Rationales' framework, systematically analyzing how diverse human annotations and explanations impact hate speech detection, emphasizing the benefits of soft labels and rationales.
Benedetta Muscato, Beiduo Chen, Gizem Gezici et al.
Proposed horizon-control strategies—Progressive OPD and Truncated OPD—significantly improve long-horizon on-policy distillation efficiency, achieving 3× faster training and comparable performance with only 10% rollout length.
Yaocheng Zhang, Jiajun Chai, Yuqian Fu et al.
BenHalluEval evaluates Bengali LLM hallucinations; BenHalluScore ranges from 7.72% to 55.42%.
Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham et al.