cs.CL 1906.04341

What Does BERT Look At? An Analysis of BERT's Attention

This study analyzes BERT's attention heads, revealing their alignment with syntactic relations, and demonstrates dependency parsing with 77% accuracy using attention maps.

Kevin Clark, Urvashi Khandelwal, Omer Levy et al.

2019-06-11 54
cs.CL 1905.10650

Are Sixteen Heads Really Better than One?

This study reveals that many attention heads in multi-head attention are redundant; introduces gradient-based greedy pruning to improve efficiency.

Paul Michel, Omer Levy, Graham Neubig

2019-05-26 1481 citations 34
cs.CL 1905.07830

HellaSwag: Can a Machine Really Finish Your Sentence?

HellaSwag dataset leverages adversarial filtering to challenge state-of-the-art models, exposing their limitations in commonsense reasoning with only ~48% accuracy versus 95% humans.

Rowan Zellers, Ari Holtzman, Yonatan Bisk et al.

2019-05-20 24
cs.CL 1904.09751

The Curious Case of Neural Text Degeneration

Proposes Nucleus Sampling, a dynamic probability truncation method, to enhance diversity and coherence in neural text generation.

Ari Holtzman, Jan Buys, Li Du et al.

2019-04-22 4547 citations 36
cs.CL 1904.08920

Towards VQA Models That Can Read

Introduced TextVQA dataset and LoRRA model, achieving 27.63% accuracy on text-based VQA, surpassing SOTA methods.

Amanpreet Singh, Vivek Natarajan, Meet Shah et al.

2019-04-19 25