cs.CL 2112.08608

QuALITY: Question Answering with Long Input Texts, Yes!

Introduces QuALITY, a long-text (≈5000 words) multiple-choice QA dataset, with models achieving only 55.4% accuracy versus 93.5% by humans.

Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi et al.

2021-12-16 36
cs.CL 2112.00861

A General Language Assistant as a Laboratory for Alignment

This paper compares prompting, imitation learning, and preference modeling across model scales, highlighting preference ranking's advantages and sample efficiency improvements via pretraining.

Amanda Askell, Yuntao Bai, Anna Chen et al.

2021-12-02 45
cs.CL 2110.08193

BBQ: A Hand-Built Bias Benchmark for Question Answering

Introduces BBQ, a handcrafted bias benchmark for question answering, revealing models' reliance on stereotypes especially under low-information conditions, with bias scores up to 100%.

Alicia Parrish, Angelica Chen, Nikita Nangia et al.

2021-10-16 42
cs.CL 2110.06341

Learning Compact Metrics for MT

Proposes RemBERT, a distilled multilingual evaluation metric reaching 92.6% of the large model's performance with only one-third parameters.

Amy Pu, Hyung Won Chung, Ankur P. Parikh et al.

2021-10-13 40
cs.CL 2110.03215

Towards Continual Knowledge Learning of Language Models

Proposes CKL framework with INVARIANTLAMA, UPDATEDLAMA, NEWLAMA datasets and FUAR metric; uses parameter expansion methods to improve knowledge retention and acquisition.

Joel Jang, Seonghyeon Ye, Sohee Yang et al.

2021-10-07 29
cs.CL 2109.06096

The Grammar-Learning Trajectories of Neural Language Models

This study tracks neural language models' (NLMs) learning trajectories, revealing highly consistent stages across architectures and data, driven by an underlying inductive bias.

Leshem Choshen, Guy Hacohen, Daphna Weinshall et al.

2021-09-14 31