cs.CL 2009.10855

Controlling Style in Generated Dialogue

This work adapts three large-scale controllable dialogue architectures to manage about 200 styles, improving style consistency and diversity.

Eric Michael Smith, Diana Gonzalez-Rico, Emily Dinan et al.

2020-09-23 54
cs.CL 2007.01852

Language-agnostic BERT Sentence Embedding

This paper introduces LaBSE, a multilingual sentence embedding model combining MLM, TLM, dual encoders, and additive margin softmax, achieving 83.7% accuracy across 112 languages.

Fangxiaoyu Feng, Yinfei Yang, Daniel Cer et al.

2020-07-04 1365 citations 32
cs.CL 2006.15020

Pre-training via Paraphrasing

MARGE introduces a retrieval-based pretraining method for multilingual multi-document paraphrasing, achieving strong zero-shot performance on translation and summarization.

Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh et al.

2020-06-26 57
cs.CL 2006.14799

Evaluation of Text Generation: A Survey

This survey reviews three categories of NLG evaluation metrics: human, automatic, and machine-learned, analyzing their strengths, weaknesses, and future directions.

Asli Celikyilmaz, Elizabeth Clark, Jianfeng Gao

2020-06-26 63
cs.CL 2006.08328

ETHOS: an Online Hate Speech Detection Dataset

ETHOS dataset offers hate speech detection based on YouTube and Reddit comments, using binary and multi-label classification.

Ioannis Mollas, Zoe Chrysopoulou, Stamatis Karlos et al.

2020-06-11 7