cs.CL 2006.14799

Evaluation of Text Generation: A Survey

This survey reviews three categories of NLG evaluation metrics: human, automatic, and machine-learned, analyzing their strengths, weaknesses, and future directions.

Asli Celikyilmaz, Elizabeth Clark, Jianfeng Gao

2020-06-26 60
cs.CL 2006.08328

ETHOS: an Online Hate Speech Detection Dataset

ETHOS dataset offers hate speech detection based on YouTube and Reddit comments, using binary and multi-label classification.

Ioannis Mollas, Zoe Chrysopoulou, Stamatis Karlos et al.

2020-06-11 1
cs.CL 2005.14165

Language Models are Few-Shot Learners

Scaling up to 175B parameters, GPT-3 achieves remarkable few-shot and zero-shot NLP performance, surpassing many fine-tuned models.

Tom B. Brown, Benjamin Mann, Nick Ryder et al.

2020-05-29 34
cs.CL 2005.07683

Movement Pruning: Adaptive Sparsity by Fine-Tuning

Movement pruning leverages first-order importance scores during fine-tuning, achieving high sparsity with minimal accuracy loss, outperforming magnitude pruning.

Victor Sanh, Thomas Wolf, Alexander M. Rush

2020-05-16 42
cs.CL 2005.04511

Finding Universal Grammatical Relations in Multilingual BERT

This paper uses structural probes to analyze multilingual BERT, revealing shared syntactic subspaces across languages, indicating learned universal grammatical relations.

Ethan A. Chi, John Hewitt, Christopher D. Manning

2020-05-10 38
cs.CL 2005.00187

Cross-Linguistic Syntactic Evaluation of Word Prediction Models

Introduced CLAMS, a multilingual syntactic evaluation suite, revealing that monolingual LSTM models excel in simple dependencies but struggle with complex structures; multilingual BERT performs well in English but poorly in other languages.

Aaron Mueller, Garrett Nicolai, Panayiota Petrou-Zeniou et al.

2020-05-01 54
cs.CL 2004.14373

ToTTo: A Controlled Table-To-Text Generation Dataset

Introduced ToTTo dataset with 120K samples, combining manual revision and highlighted cells for high-precision table-to-text generation.

Ankur P. Parikh, Xuezhi Wang, Sebastian Gehrmann et al.

2020-04-30 54