GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CL 2407.11606

The Foundations of Tokenization: Statistical and Computational Concerns

Proposes a unified stochastic map framework for tokenization, ensuring estimator consistency in language models.

Juan Luis Gastaldi, John Terilla, Luca Malagutti et al.

2024-07-16 65
cs.CL 2407.10817

Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

FLAMe, a large-scale general auto-evaluator trained on 5.3M human judgments, outperforms proprietary models in multiple benchmarks.

Tu Vu, Kalpesh Krishna, Salaheddin Alzubi et al.

2024-07-15 37
cs.CL 2407.10457

The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism

This study evaluates LLM non-determinism, showing greedy decoding outperforms sampling; small models can match larger ones with best-of-N and reward models.

Yifan Song, Guoyin Wang, Sujian Li et al.

2024-07-15 43
cs.CL 2407.09722

Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference

Proposed MTAD framework combines auxiliary models with joint multi-token decoding, boosting efficiency and quality.

Zongyue Qin, Ziniu Hu, Zifan He et al.

2024-07-13 39
cs.CL 2407.09413

SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers

Introduced SPIQA dataset, combining multimodal large models for scientific figure understanding, with chain-of-thought reasoning, significantly advancing research comprehension.

Shraman Pramanick, Rama Chellappa, Subhashini Venugopalan

2024-07-13 112 citations 37
cs.CL 2407.08818

MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based Tokenization

MAGNET employs script-specific boundary predictors with adaptive gradient optimization, reducing over-segmentation in non-Latin scripts and improving efficiency.

Orevaoghene Ahia, Sachin Kumar, Hila Gonen et al.

2024-07-12 34
cs.CL 2407.12857

Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis

Proposed SEA framework integrates GPT-4 standardization, Mistral-7B evaluation, and self-correction, enhancing automated paper review quality.

Jianxiang Yu, Zichen Ding, Jiaqi Tan et al.

2024-07-09 35
cs.CL 2407.05975

LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages

LLaMAX leverages continual multilingual pretraining and vocabulary preservation to support over 100 languages, achieving +10 spBLEU over open-source models, comparable to M2M-100-12B.

Yinquan Lu, Wenhao Zhu, Lei Li et al.

2024-07-08 27
cs.CL 2407.05434

LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models

Using LTL to generate 2000 challenges, evaluating 12 LLMs, revealing three main issues in complex temporal reasoning tasks.

Weizhi Tang, Kwabena Nuamah, Vaishak Belle

2024-07-08 74
cs.CL 2407.04903

MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding

Constructed MMSci dataset with 72 disciplines, enabling large models to understand complex scientific figures; fine-tuned Qwen2-VL-7B achieved 87.48% accuracy.

Zekun Li, Xianjun Yang, Kyuri Choi et al.

2024-07-06 41
cs.CL 2407.04069

A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations

A comprehensive review of LLM evaluation challenges, emphasizing reproducibility, reliability, and robustness, with specific algorithm and dataset references.

Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari et al.

2024-07-05 56
cs.CL 2407.03978

Benchmarking Complex Instruction-Following with Multiple Constraints Composition

ComplexBench benchmark with hierarchical taxonomy and rule-augmented evaluation reveals significant deficiencies of current LLMs in multi-constraint complex instructions.

Bosi Wen, Pei Ke, Xiaotao Gu et al.

2024-07-04 161 citations 29
cs.CL 2407.01490

LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives

Proposes active inheritance using non-differentiable metrics to steer synthetic data, improving attributes like lexical diversity and reducing toxicity in LLMs

Luísa Shimabucoro, Sebastian Ruder, Julia Kreutzer et al.

2024-07-02 37
cs.CL 2407.01257

uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes

uDistil-Whisper achieves label-free knowledge distillation in low-data regimes, improving performance by 5-7 WER points.

Abdul Waheed, Karima Kadaoui, Bhiksha Raj et al.

2024-07-01 6
cs.CL 2407.00908

FineSurE: Fine-grained Summarization Evaluation using LLMs

FineSurE uses LLMs for fine-grained summarization evaluation, improving completeness and conciseness.

Hwanjun Song, Hang Su, Igor Shalyminov et al.

2024-07-01 28
cs.CL 2407.00416

Too Late to Train, Too Early To Use? A Study on Necessity and Viability of Low-Resource Bengali LLMs

This study assesses the necessity of Bengali-specific LLMs, revealing challenges in tokenization and bias, with LLaMA-3 outperforming fine-tuned models in understanding tasks.

Tamzeed Mahfuz, Satak Kumar Dey, Ruwad Naswan et al.

2024-06-29 66
cs.CL 2407.00402

Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP

Proposes a two-dimensional taxonomy for long-context task difficulty based on information diffusion and scope, highlighting under-explored high-diffusion, high-scope scenarios.

Omer Goldman, Alon Jacovi, Aviv Slobodkin et al.

2024-06-29 28 citations 65
cs.CL 2406.19065

STBench: Assessing the Ability of Large Language Models in Spatio-Temporal Analysis

Proposed STBench benchmark evaluates 13 LLMs across four spatio-temporal abilities, emphasizing knowledge comprehension and reasoning, with over 60,000 QA pairs.

Wenbin Li, Di Yao, Ruibo Zhao et al.

2024-06-27 42
cs.CL 2406.17975

SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)

This paper systematically evaluates the current state of membership inference attacks (MIAs) on LLMs, highlighting dataset distribution shifts and proposing multiple mitigation strategies.

Matthieu Meeus, Igor Shilov, Shubham Jain et al.

2024-06-26 27
cs.CL 2406.16377

On the Transformations across Reward Model, Parameter Update, and In-Context Prompt

Unified framework linking reward models, parameter updates, and prompts via six bidirectional transformations.

Deng Cai, Huayang Li, Tingchen Fu et al.

2024-06-24 43
Prev 1 ... 48 49 50 51 52 53 54 ... 93 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home