cs.CL 2406.16253

LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing

This study evaluates LLMs (GPT-4, Gemini, Claude) in NLP peer review, revealing high defect rates and limited professional judgment, with scores averaging 7.45 vs. human 6.41.

Jiangshu Du, Yibo Wang, Wenting Zhao et al.

2024-06-24 31
cs.CL 2406.13560

Lexically Grounded Subword Segmentation

Introduces a lexically grounded subword segmentation method, improving morphological plausibility.

Jindřich Libovický, Jindřich Helcl

2024-06-19 3
cs.CL 2406.09393

Improving Autoregressive Training with Dynamic Oracles

Proposes dynamic oracle-guided autoregressive training, improving NER and summarization, with exact and approximate algorithms for specific metrics.

Jianing Yang, Harshine Visvanathan, Yilin Wang et al.

2024-06-14 35
cs.CL 2406.08446

OLMES: A Standard for Language Model Evaluations

OLMES standardizes large language model evaluation, covering prompt formatting, example selection, probability normalization, ensuring reproducibility and fairness.

Yuling Gu, Oyvind Tafjord, Bailey Kuehl et al.

2024-06-13 100 citations 31
cs.CL 2406.07524

Simple and Effective Masked Diffusion Language Models

This paper introduces Masked Diffusion Language Models (MDLM), achieving near state-of-the-art perplexity by combining efficient training, Rao-Blackwellized objectives, and semi-autoregressive sampling.

Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff et al.

2024-06-12 837 citations 26
cs.CL 2406.07496

TextGrad: Automatic "Differentiation" via Text

TextGrad employs natural language feedback as 'gradients' to optimize complex AI systems, improving performance across tasks like QA, coding, and drug design.

Mert Yuksekgonul, Federico Bianchi, Joseph Boen et al.

2024-06-12 45