cs.CL 1803.02324

Annotation Artifacts in Natural Language Inference Data

Using fastText model to analyze annotation artifacts in NLI datasets, finding 67% of SNLI and 53% of MultiNLI data can be classified by hypothesis alone.

Suchin Gururangan, Swabha Swayamdipta, Omer Levy et al.

2018-03-07 51
cs.CL 1803.02155

Self-Attention with Relative Position Representations

Introduces relative position representations into self-attention, improving machine translation BLEU scores by 1.3 and 0.3, replacing absolute encodings.

Peter Shaw, Jakob Uszkoreit, Ashish Vaswani

2018-03-06 55
cs.CL 1712.07040

The NarrativeQA Reading Comprehension Challenge

Introduces NarrativeQA dataset emphasizing deep story comprehension; models struggle with long, complex narratives.

Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom et al.

2017-12-20 56
cs.CL 1710.04087

Word Translation Without Parallel Data

Unsupervised cross-lingual word mapping via adversarial training and Procrustes refinement surpasses supervised methods, achieving 66.2% accuracy on English-Italian translation.

Alexis Conneau, Guillaume Lample, Marc'Aurelio Ranzato et al.

2017-10-11 61
cs.CL 1706.09254

The E2E Dataset: New Challenges For End-to-End Generation

Introduces the E2E dataset, ten times larger than previous, emphasizing lexical richness, syntactic diversity, and content selection challenges, advancing natural language generation.

Jekaterina Novikova, Ondřej Dušek, Verena Rieser

2017-06-28 47