The NarrativeQA Reading Comprehension Challenge
Introduces NarrativeQA dataset emphasizing deep story comprehension; models struggle with long, complex narratives.
Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom et al.
Introduces NarrativeQA dataset emphasizing deep story comprehension; models struggle with long, complex narratives.
Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom et al.
Leveraging temporal word embeddings to quantify a century of gender and ethnic stereotypes in the US, validated against Census data.
Nikhil Garg, Londa Schiebinger, Dan Jurafsky et al.
SQLNet uses dependency graphs to avoid order sensitivity, improving WikiSQL accuracy by 9-13%.
Xiaojun Xu, Chang Liu, Dawn Song
Constructed WIKIHOP and MEDHOP datasets for multi-hop cross-document QA; models reach 54.5% accuracy, human 85%.
Johannes Welbl, Pontus Stenetorp, Sebastian Riedel
Unsupervised cross-lingual word mapping via adversarial training and Procrustes refinement surpasses supervised methods, achieving 66.2% accuracy on English-Italian translation.
Alexis Conneau, Guillaume Lample, Marc'Aurelio Ranzato et al.
Arabic MGB-3 Challenge uses i-vector features for dialect identification, achieving 75% accuracy.
Ahmed Ali, Stephan Vogel, Steve Renals
Deep learning-based supervised speech separation uses time-frequency masks, achieving 15dB SNR improvement and high STOI/PESQ scores.
DeLiang Wang, Jitong Chen
Proposed sequence-based CNN (SCNN) for multiparty dialogue emotion detection; achieved 37.9% accuracy for fine-grained emotion classification.
Sayyed M. Zahiri, Jinho D. Choi
Edinburgh's WMT17 system uses Nematus with deep architectures, layer normalization, and BPE, achieving 2.2-5 BLEU gains.
Rico Sennrich, Alexandra Birch, Anna Currey et al.
Multi-model ensemble approach combining feature engineering and deep learning achieved an average Pearson correlation of 0.73 in multilingual STS tasks.
Daniel Cer, Mona Diab, Eneko Agirre et al.
Large-scale hyperparameter tuning reveals that well-regularized standard LSTM outperforms recent architectures on Penn and Wikitext-2 benchmarks.
Gábor Melis, Chris Dyer, Phil Blunsom
Introduces the E2E dataset, ten times larger than previous, emphasizing lexical richness, syntactic diversity, and content selection challenges, advancing natural language generation.
Jekaterina Novikova, Ondřej Dušek, Verena Rieser
Transformer uses solely attention mechanisms, achieving 28.4 BLEU on WMT 2014 English-German translation, with faster training and superior quality.
Ashish Vaswani, Noam Shazeer, Niki Parmar et al.
Proposes a multi-document extractive summarization framework combining sentence ranking, keyphrase extraction, and semantic-based ordering, outperforming SOTA on DUC 2004.
Mir Tafseer Nayeem, Yllias Chali
Cross-Alignment enables style transfer in non-parallel text, achieving 78.4% accuracy in sentiment modification.
Tianxiao Shen, Tao Lei, Regina Barzilay et al.
Introduces a end-to-end differentiable Key-Value Retrieval Network that outperforms baselines on multi-domain task-oriented dialogue tasks.
Mihail Eric, Christopher D. Manning
Proposed a deep model with intra-attention and reinforcement learning, achieving ROUGE-1 41.16 on CNN/Daily Mail.
Romain Paulus, Caiming Xiong, Richard Socher
TriviaQA is a large-scale distant supervision reading comprehension dataset with 650K QA triples, emphasizing complex reasoning and multi-source evidence.
Mandar Joshi, Eunsol Choi, Daniel S. Weld et al.
Proposes a fully convolutional sequence-to-sequence model with GLU and multi-step attention, outperforming LSTM-based models on WMT tasks with 10x faster training.
Jonas Gehring, Michael Auli, David Grangier et al.
InferSent, a supervised framework using SNLI data, achieves superior sentence representation transfer performance.
Alexis Conneau, Douwe Kiela, Holger Schwenk et al.