Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing
Efficient training-free multi-token prediction via embedding-space probing, improving LLaMA3 acceptance length by 12%.
Raghavv Goel, Mukul Gagrani, Mingu Lee et al.
Efficient training-free multi-token prediction via embedding-space probing, improving LLaMA3 acceptance length by 12%.
Raghavv Goel, Mukul Gagrani, Mingu Lee et al.
GSU dataset evaluates LLMs' spatial reasoning over grids, revealing limitations of visual modality in 3D understanding.
Risham Sidhu, Julia Hockenmaier
Mixture-of-Depths Attention (MoDA) improves downstream task performance by 2.11% on a 1.5B-parameter model with only a 3.7% increase in FLOPs.
Lianghui Zhu, Yuxin Fang, Bencheng Liao et al.
Correcting moral indifference in language models using Sparse Autoencoders, achieving a 75% win-rate on adversarial benchmarks.
Lingyu Li, Yan Teng, Yingchun Wang
Code-A1 enhances code and test generation through an adversarial co-evolution framework.
Aozhe Wang, Yuchen Yan, Nan Zhou et al.
AI learns scientific taste via Reinforcement Learning from Community Feedback, enhancing discovery efficiency.
Jingqi Tong, Mingzhe Li, Hangcheng Li et al.
PDPS method exposes long-tail safety failures in LLMs, reducing computational cost by 33%.
Suvadeep Hajra, Palash Nandi, Tanmoy Chakraborty
Develops an exactly solvable n-gram-based recursive framework analyzing drift and selection effects on public text ecosystems.
Søren Riis
AIM-SciQA employs large language models to automatically generate 13,672 multi-hop scientific QA pairs across documents, enhancing cross-document reasoning.
Seungmin Lee, Dongha Kim, Yuni Jeon et al.
NAIT framework selects efficient instruction tuning data via neuron activation patterns, enhancing LLM performance.
Xin Chen, Junchao Wu, Shu Yang et al.
ESG-Bench significantly reduces hallucinations in long-context ESG report analysis using task-specific Chain-of-Thought prompting strategies.
Siqi Sun, Ben Peng Wu, Mali Jin et al.
WALAR method enhances low-resource language translation using monolingual data, surpassing LLaMAX model.
Yifeng Liu, Siqi Ouyang, Yatish Hosmane Revanasiddappa et al.
Proposed a PCA sweep method to optimize dimension selection in SSD, enhancing interpretability and stability.
Hubert Plisiecki, Maria Leniarska, Jan Piotrowski et al.
Long-form RewardBench evaluates reward models for long-form generation, revealing current models' deficiencies in long-form reward modeling.
Hui Huang, Yancheng He, Wei Liu et al.
HMS-BERT uses hybrid multi-task self-training for multilingual, multi-label cyberbullying detection, achieving a macro F1-score of 0.9847.
Zixin Feng, Xinying Cui, Yifan Sun et al.
Using log-odds ratio analysis, this study reveals systematic biases in LLM-generated educational feedback conditioned on student attributes, highlighting overuse of praise and stereotype reinforcement.
Mei Tan, Lena Phalen, Dorottya Demszky
Idea-Catalyst framework boosts scientific creativity via interdisciplinary insights, improving novelty by 21% and insightfulness by 16%.
Priyanka Kargupta, Shuhaib Mehri, Dilek Hakkani-Tur et al.
CLASP model detects malicious tokens using XGBoost classifier, achieving 95.9% token-level F1 score.
Alexandre Le Mercier, Thomas Demeester, Chris Develder
IndexCache accelerates sparse attention by reusing cross-layer indices, reducing 75% of computations, achieving 1.82x speedup.
Yushi Bai, Qian Dong, Ting Jiang et al.
Introduced a Polish long-context encoder model handling up to 8192 tokens, significantly improving long-document task performance.
Sławomir Dadas, Rafał Poświata, Marek Kozłowski et al.