SentEval: An Evaluation Toolkit for Universal Sentence Representations
SentEval evaluates universal sentence representations via standardized tasks like classification and similarity.
Alexis Conneau, Douwe Kiela
SentEval evaluates universal sentence representations via standardized tasks like classification and similarity.
Alexis Conneau, Douwe Kiela
Constructed FEVER dataset with 185,445 claims; max accuracy 31.87%; highlights evidence retrieval as key bottleneck.
James Thorne, Andreas Vlachos, Christos Christodoulopoulos et al.
Using fastText model to analyze annotation artifacts in NLI datasets, finding 67% of SNLI and 53% of MultiNLI data can be classified by hypothesis alone.
Suchin Gururangan, Swabha Swayamdipta, Omer Levy et al.
Introduces relative position representations into self-attention, improving machine translation BLEU scores by 1.3 and 0.3, replacing absolute encodings.
Peter Shaw, Jakob Uszkoreit, Ashish Vaswani
Proposes an attention-based CNN model (CAML) for ICD code prediction with 0.54 micro-F1 and interpretable text snippets.
James Mullenbach, Sarah Wiegreffe, Jon Duke et al.
Texygen integrates multiple models and metrics, enabling comprehensive evaluation of text generation quality.
Yaoming Zhu, Sidi Lu, Lei Zheng et al.
Generating Wikipedia articles by summarizing long sequences using a decoder-only architecture and ROUGE scores.
Peter J. Liu, Mohammad Saleh, Etienne Pot et al.
Proposes a method to map natural language to math expressions, improving arithmetic problem-solving accuracy.
Subhro Roy, Dan Roth
Introduces NarrativeQA dataset emphasizing deep story comprehension; models struggle with long, complex narratives.
Tomáš Kočiský, Jonathan Schwarz, Phil Blunsom et al.
Leveraging temporal word embeddings to quantify a century of gender and ethnic stereotypes in the US, validated against Census data.
Nikhil Garg, Londa Schiebinger, Dan Jurafsky et al.
SQLNet uses dependency graphs to avoid order sensitivity, improving WikiSQL accuracy by 9-13%.
Xiaojun Xu, Chang Liu, Dawn Song
Constructed WIKIHOP and MEDHOP datasets for multi-hop cross-document QA; models reach 54.5% accuracy, human 85%.
Johannes Welbl, Pontus Stenetorp, Sebastian Riedel
Unsupervised cross-lingual word mapping via adversarial training and Procrustes refinement surpasses supervised methods, achieving 66.2% accuracy on English-Italian translation.
Alexis Conneau, Guillaume Lample, Marc'Aurelio Ranzato et al.
Arabic MGB-3 Challenge uses i-vector features for dialect identification, achieving 75% accuracy.
Ahmed Ali, Stephan Vogel, Steve Renals
Deep learning-based supervised speech separation uses time-frequency masks, achieving 15dB SNR improvement and high STOI/PESQ scores.
DeLiang Wang, Jitong Chen
Proposed sequence-based CNN (SCNN) for multiparty dialogue emotion detection; achieved 37.9% accuracy for fine-grained emotion classification.
Sayyed M. Zahiri, Jinho D. Choi
Edinburgh's WMT17 system uses Nematus with deep architectures, layer normalization, and BPE, achieving 2.2-5 BLEU gains.
Rico Sennrich, Alexandra Birch, Anna Currey et al.
Multi-model ensemble approach combining feature engineering and deep learning achieved an average Pearson correlation of 0.73 in multilingual STS tasks.
Daniel Cer, Mona Diab, Eneko Agirre et al.
Large-scale hyperparameter tuning reveals that well-regularized standard LSTM outperforms recent architectures on Penn and Wikitext-2 benchmarks.
Gábor Melis, Chris Dyer, Phil Blunsom
Introduces the E2E dataset, ten times larger than previous, emphasizing lexical richness, syntactic diversity, and content selection challenges, advancing natural language generation.
Jekaterina Novikova, Ondřej Dušek, Verena Rieser