Hierarchical Neural Story Generation
Hierarchical neural story generation with fusion and multi-scale self-attention significantly improves coherence and relevance, outperforming baselines.
Angela Fan, Mike Lewis, Yann Dauphin
Hierarchical neural story generation with fusion and multi-scale self-attention significantly improves coherence and relevance, outperforming baselines.
Angela Fan, Mike Lewis, Yann Dauphin
Using Luong’s seq2seq model, automatically translating informal LaTeX math texts into formal Mizar language achieved 65.73% accuracy on test data.
Qingxiang Wang, Cezary Kaliszyk, Josef Urban
Proposed a hypothesis-only baseline for NLI, significantly outperforming majority class baselines.
Adam Poliak, Jason Naradowsky, Aparajita Haldar et al.
Introduces Winogender schemas to evaluate gender bias in coreference systems, revealing systematic bias correlated with real-world occupational gender statistics.
Rachel Rudinger, Jason Naradowsky, Brian Leonard et al.
This paper advocates using WMT standard for BLEU reporting, introducing SACREBLEU to ensure parameter consistency and comparability.
Matt Post
GLUE benchmark with 9 tasks, multi-task training improves generalization; baseline scores still far from human performance.
Alex Wang, Amanpreet Singh, Julian Michael et al.
Pre-trained word embeddings improve low-resource NMT by up to 20 BLEU points.
Ye Qi, Devendra Singh Sachan, Matthieu Felix et al.
Proposed a discourse-aware model for long document summarization, significantly outperforming existing models.
Arman Cohan, Franck Dernoncourt, Doo Soon Kim et al.
Introduces Speech Commands dataset with 105,829 samples for keyword spotting; achieves up to 88.2% accuracy with CNN models.
Pete Warden
Using neural network-based word embeddings to analyze cultural dimensions via geometric relationships in high-dimensional space.
Austin C. Kozlowski, Matt Taddy, James A. Evans
Proposes a web-based complex question answering framework with question decomposition, improving precision@1 from 20.8% to 27.5%.
Alon Talmor, Jonathan Berant
Constructed GYAFC dataset with 110K sentence pairs; applied PBMT and NMT models for formality transfer; evaluated with automatic and human metrics.
Sudha Rao, Joel Tetreault
SentEval evaluates universal sentence representations via standardized tasks like classification and similarity.
Alexis Conneau, Douwe Kiela
Constructed FEVER dataset with 185,445 claims; max accuracy 31.87%; highlights evidence retrieval as key bottleneck.
James Thorne, Andreas Vlachos, Christos Christodoulopoulos et al.
Using fastText model to analyze annotation artifacts in NLI datasets, finding 67% of SNLI and 53% of MultiNLI data can be classified by hypothesis alone.
Suchin Gururangan, Swabha Swayamdipta, Omer Levy et al.
Introduces relative position representations into self-attention, improving machine translation BLEU scores by 1.3 and 0.3, replacing absolute encodings.
Peter Shaw, Jakob Uszkoreit, Ashish Vaswani
Proposes an attention-based CNN model (CAML) for ICD code prediction with 0.54 micro-F1 and interpretable text snippets.
James Mullenbach, Sarah Wiegreffe, Jon Duke et al.
Texygen integrates multiple models and metrics, enabling comprehensive evaluation of text generation quality.
Yaoming Zhu, Sidi Lu, Lei Zheng et al.
Generating Wikipedia articles by summarizing long sequences using a decoder-only architecture and ROUGE scores.
Peter J. Liu, Mohammad Saleh, Etienne Pot et al.
Proposes a method to map natural language to math expressions, improving arithmetic problem-solving accuracy.
Subhro Roy, Dan Roth