cs.CL 2004.09095

The State and Fate of Linguistic Diversity and Inclusion in the NLP World

This paper analyzes global linguistic resource distribution, proposes a six-category classification, highlights disparities affecting multilingual NLP, and urges community focus on low-resource languages.

Pratik Joshi, Sebastin Santy, Amar Budhiraja et al.

2020-04-20 1454 citations 50
cs.CL 2004.05150

Longformer: The Long-Document Transformer

Introduces Longformer with linear attention combining local windows and global tokens, enabling efficient long document modeling surpassing RoBERTa.

Iz Beltagy, Matthew E. Peters, Arman Cohan

2020-04-11 35
cs.CL 2004.04037

DynaBERT: Dynamic BERT with Adaptive Width and Depth

DynaBERT combines adaptive width and depth, trained via knowledge distillation, achieving performance comparable to BERT-base with improved efficiency.

Lu Hou, Zhiqi Huang, Lifeng Shang et al.

2020-04-08 44
cs.CL 2004.02644

Sparse Text Generation

Using entmax for training and sampling, reducing mismatch, improving diversity and coherence in text generation.

Pedro Henrique Martins, Zita Marinho, André F. T. Martins

2020-04-06 63
cs.CL 2003.07892

Calibration of Pre-trained Transformers

This study evaluates BERT and RoBERTa calibration, showing temperature scaling and label smoothing significantly improve in- and out-of-domain calibration.

Shrey Desai, Greg Durrett

2020-03-18 50
cs.CL 2002.12327

A Primer in BERTology: What we know about how BERT works

This paper systematically reviews BERT's internal mechanisms, analyzing its encoding of syntax, semantics, and world knowledge, and discusses model improvements and compression.

Anna Rogers, Olga Kovaleva, Anna Rumshisky

2020-02-28 50
cs.CL 2002.08909

REALM: Retrieval-Augmented Language Model Pre-Training

REALM integrates a trainable neural retriever with language models, boosting open-domain QA accuracy by 4-16%, with improved interpretability and modularity.

Kelvin Guu, Kenton Lee, Zora Tung et al.

2020-02-11 33