The Foundations of Tokenization: Statistical and Computational Concerns
Proposes a unified stochastic map framework for tokenization, ensuring estimator consistency in language models.
Juan Luis Gastaldi, John Terilla, Luca Malagutti et al.
Proposes a unified stochastic map framework for tokenization, ensuring estimator consistency in language models.
Juan Luis Gastaldi, John Terilla, Luca Malagutti et al.
FLAMe, a large-scale general auto-evaluator trained on 5.3M human judgments, outperforms proprietary models in multiple benchmarks.
Tu Vu, Kalpesh Krishna, Salaheddin Alzubi et al.
This study evaluates LLM non-determinism, showing greedy decoding outperforms sampling; small models can match larger ones with best-of-N and reward models.
Yifan Song, Guoyin Wang, Sujian Li et al.
Proposed MTAD framework combines auxiliary models with joint multi-token decoding, boosting efficiency and quality.
Zongyue Qin, Ziniu Hu, Zifan He et al.
Introduced SPIQA dataset, combining multimodal large models for scientific figure understanding, with chain-of-thought reasoning, significantly advancing research comprehension.
Shraman Pramanick, Rama Chellappa, Subhashini Venugopalan
MAGNET employs script-specific boundary predictors with adaptive gradient optimization, reducing over-segmentation in non-Latin scripts and improving efficiency.
Orevaoghene Ahia, Sachin Kumar, Hila Gonen et al.
Proposed SEA framework integrates GPT-4 standardization, Mistral-7B evaluation, and self-correction, enhancing automated paper review quality.
Jianxiang Yu, Zichen Ding, Jiaqi Tan et al.
LLaMAX leverages continual multilingual pretraining and vocabulary preservation to support over 100 languages, achieving +10 spBLEU over open-source models, comparable to M2M-100-12B.
Yinquan Lu, Wenhao Zhu, Lei Li et al.
Using LTL to generate 2000 challenges, evaluating 12 LLMs, revealing three main issues in complex temporal reasoning tasks.
Weizhi Tang, Kwabena Nuamah, Vaishak Belle
Constructed MMSci dataset with 72 disciplines, enabling large models to understand complex scientific figures; fine-tuned Qwen2-VL-7B achieved 87.48% accuracy.
Zekun Li, Xianjun Yang, Kyuri Choi et al.
A comprehensive review of LLM evaluation challenges, emphasizing reproducibility, reliability, and robustness, with specific algorithm and dataset references.
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari et al.
ComplexBench benchmark with hierarchical taxonomy and rule-augmented evaluation reveals significant deficiencies of current LLMs in multi-constraint complex instructions.
Bosi Wen, Pei Ke, Xiaotao Gu et al.
Proposes active inheritance using non-differentiable metrics to steer synthetic data, improving attributes like lexical diversity and reducing toxicity in LLMs
Luísa Shimabucoro, Sebastian Ruder, Julia Kreutzer et al.
uDistil-Whisper achieves label-free knowledge distillation in low-data regimes, improving performance by 5-7 WER points.
Abdul Waheed, Karima Kadaoui, Bhiksha Raj et al.
FineSurE uses LLMs for fine-grained summarization evaluation, improving completeness and conciseness.
Hwanjun Song, Hang Su, Igor Shalyminov et al.
This study assesses the necessity of Bengali-specific LLMs, revealing challenges in tokenization and bias, with LLaMA-3 outperforming fine-tuned models in understanding tasks.
Tamzeed Mahfuz, Satak Kumar Dey, Ruwad Naswan et al.
Proposes a two-dimensional taxonomy for long-context task difficulty based on information diffusion and scope, highlighting under-explored high-diffusion, high-scope scenarios.
Omer Goldman, Alon Jacovi, Aviv Slobodkin et al.
Proposed STBench benchmark evaluates 13 LLMs across four spatio-temporal abilities, emphasizing knowledge comprehension and reasoning, with over 60,000 QA pairs.
Wenbin Li, Di Yao, Ruibo Zhao et al.
This paper systematically evaluates the current state of membership inference attacks (MIAs) on LLMs, highlighting dataset distribution shifts and proposing multiple mitigation strategies.
Matthieu Meeus, Igor Shilov, Shubham Jain et al.
Unified framework linking reward models, parameter updates, and prompts via six bidirectional transformations.
Deng Cai, Huayang Li, Tingchen Fu et al.