DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting
DeLS-Spec improves inference speed and acceptance length by combining long and short contexts.
Hong-Kai Zheng, Piji Li
DeLS-Spec improves inference speed and acceptance length by combining long and short contexts.
Hong-Kai Zheng, Piji Li
Pluralis v0.1 introduces a culture-first multimodal evaluation framework with 6,448 prompts, exposing VLM blind spots in cultural alignment.
Alicia Parrish, Rajat Shinde, Sanket Badhe et al.
This study systematically compares BPE and Unigram-LM tokenizers on fixed 165-token chemical bases, revealing near-disjoint vocabularies and emphasizing tokenizer choice as a key model design decision.
Hunter Heidenreich
Audex, built on Nemotron-Cascade-2-30B-A3B, unifies audio-text modeling with a single Transformer decoder, achieving state-of-the-art performance in audio understanding, speech recognition, translation, and generation, while maintaining text reasoning.
Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim et al.
Introduces MTEB-BR, a benchmark with 22 native Portuguese tasks, evaluating 93 models using rigorous statistical analysis to distinguish capability tiers.
Tardelli Ronan Coelho Stekel
dOPSD enhances reasoning in diffusion language models by self-distillation, excelling on Dream and LLaDA datasets.
Phuong Tuan Dat, Qi Li, Xinchao Wang
CrossHallu evaluates internal signals' cross-lingual and cross-domain generalization in LLMs, showing most models transfer hallucination signals effectively.
Aisha Alansari, Malak Alkhorasani, Hamzah Luqman
AdaptiveSpec boosts inference speed by 56% and recovers 93% accuracy by dynamically adjusting draft tree shapes and lossy verification.
Oszkár Urbán, Young D. Kwon, Stylianos I. Venieris et al.
DCCD combines document and token confidence to improve conflict resolution in multi-document QA, outperforming baselines on DRQA.
Raymond Li, Md Tawkat Islam Khondaker, Amirhossein Abaskohi et al.
ALEE framework uses AMR-based minimal pairs to evaluate 275+ languages' embedding models, revealing performance gaps related to resources and linguistic phenomena.
Andrianos Michail, Stylianos Psychias, Michelle Wastl et al.
GradeSQL framework enhances Text-to-SQL reliability using ORM, achieving a 4.33% gain on BIRD.
Mattia Tritto, Giuseppe Farano, Dario Di Palma et al.
Introduced STE and SPID frameworks to quantify semantic information flow and multi-source contributions in communication.
Leonardo S. Goodall, Andrea I. Luppi, Pedro A. M. Mediano
VISTA enables self-managed context in LLMs via a training-free, model-agnostic layer that visualizes and archives internal states, boosting long-horizon task performance.
Binyan Xu, Haitao Li, Kehuan Zhang
IHDec uses JSD-based role attribution and contrastive decoding for real-time hierarchy correction, improving multi-turn instruction adherence by 11.98pp.
Nicole Geumheon Liu, Haeun Jang, Yonghyun Jun et al.
EVLA fuses multimodal perception with vehicle physics via UCSE and ESRC, achieving energy-efficient driving decisions with +0.0871 score improvement.
Yuxin Liu, Zihan Chen, Haoyu Wang et al.
MultiHashFormer employs a multi-hash signature mechanism supporting causal language modeling, outperforming standard Transformers across 100M-3B parameters, with zero-parameter multilingual vocabulary expansion.
Huiyin Xue, Atsuki Yamaguchi, Nikolaos Aletras
VASAE introduces vocabulary-aligned anchoring to train SAE features with intrinsic token names, maintaining reconstruction quality and achieving over 90% feature-token alignment.
Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah et al.
Introduces Masked Language Flow Models combining masking and flow models to enhance multi-step reasoning.
Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang et al.
Introduces Ko-WideSearch, a Korean breadth-search benchmark with 228 tables, evaluating web agents' ability for exhaustive set enumeration and attribute filling.
Minbyul Jeong
PEEU framework enables small multimodal models to achieve 30.6% success in web navigation by autonomous exploration and hindsight experience, outperforming larger models.
Tianyi Men, Zhuoran Jin, Pengfei Cao et al.