Attention as a Guide for Simultaneous Speech Translation
Introduces EDAtt, an attention-based policy for SimulST, improving BLEU by up to 7 points and reducing latency by 1.4s.
Sara Papi, Matteo Negri, Marco Turchi
Introduces EDAtt, an attention-based policy for SimulST, improving BLEU by up to 7 points and reducing latency by 1.4s.
Sara Papi, Matteo Negri, Marco Turchi
Proposes LMQL, a scripting-based query language that constrains and optimizes large language model calls, reducing costs by up to 85%.
Luca Beurer-Kellner, Marc Fischer, Martin Vechev
Introduces Contrast-Consistent Search (CCS), an unsupervised method to discover latent knowledge in language models, improving question-answering accuracy by 4%.
Collin Burns, Haotian Ye, Dan Klein et al.
UniKGQA unifies retrieval and reasoning in a single architecture, achieving significant improvements in multi-hop KGQA accuracy, with over 75% Hits@1 on WebQSP.
Jinhao Jiang, Kun Zhou, Wayne Xin Zhao et al.
PromptInject framework reveals GPT-3's vulnerability to goal hijacking (58.6%) and prompt leaking (23.6%) via adversarial prompts.
Fábio Perez, Ian Ribeiro
Proposes TART, a task-aware retrieval system using multi-task instruction tuning, excelling in zero-shot and cross-domain scenarios.
Akari Asai, Timo Schick, Patrick Lewis et al.
Entity linking-based document count reveals that model accuracy correlates with relevant data, showing exponential growth needed for long-tail knowledge mastery.
Nikhil Kandpal, Haikang Deng, Adam Roberts et al.
CRINGE loss leverages contrastive negative generation with iterative self-labeling, significantly improving safety and coherence in language models, outperforming baselines.
Leonard Adolphs, Tianyu Gao, Jing Xu et al.
BLOOM is a 176B-parameter open-source multilingual language model based on Transformer, trained on ROOTS corpus, excelling in diverse NLP tasks.
BigScience Workshop, :, Teven Le Scao et al.
PASTA employs sentence-table cloze pre-training with six operation types, achieving 85.6% accuracy on TabFact's complex set, surpassing previous SOTA by 4.7%.
Zihui Gu, Ju Fan, Nan Tang et al.
Self-Correction method enhances sequence generation quality, achieving 99% accuracy in mathematical program synthesis.
Sean Welleck, Ximing Lu, Peter West et al.
SSD-LM is a diffusion-based language model excelling in text generation and modular control.
Xiaochuang Han, Sachin Kumar, Yulia Tsvetkov
Contrastive Decoding (CD) leverages likelihood differences between large and small models to improve open-ended text generation without extra training.
Xiang Lisa Li, Ari Holtzman, Daniel Fried et al.
Introduced SciFact-Open, leveraging model fusion and retrieval pooling on 500K abstracts, with F1 dropping over 15 points compared to closed-domain benchmarks.
David Wadden, Kyle Lo, Bailey Kuehl et al.
Proposed Code4Struct leverages code generation for few-shot event structure prediction, achieving 29.5% absolute F1 gain over SOTA.
Xingyao Wang, Sha Li, Heng Ji
Utilizing large language models for MCQA, MCP method narrows the gap with SOTA across 20 datasets.
Joshua Robinson, Christopher Michael Rytting, David Wingate
This paper introduces a syntactic surprisal measure based on CCG supertags within neural language models, revealing underestimation of human processing difficulty in garden path sentences.
Suhas Arehalli, Brian Dillon, Tal Linzen
ReasonFormer enhances performance across 11 datasets through modular reasoning.
Wanjun Zhong, Tingting Ma, Jiahai Wang et al.
Simple prompts enhance GPT-3's reliability in generalizability, social bias, calibration, and factuality.
Chenglei Si, Zhe Gan, Zhengyuan Yang et al.
MEMIT mass-edits 10,000 facts in GPT-J, reaching an 85.8 COUNTERFACT editing score at scale.
Kevin Meng, Arnab Sen Sharma, Alex Andonian et al.