cs.CL 2104.14337

Dynabench: Rethinking Benchmarking in NLP

Dynabench introduces a dynamic adversarial data collection platform, significantly reducing model error rates across four NLP tasks by leveraging multi-round human-model interactions.

Douwe Kiela, Max Bartolo, Yixin Nie et al.

2021-04-08 38
cs.CL 2104.02112

Efficient Attentions for Long Document Summarization

HEPOS employs head-wise positional strides, reducing complexity and enabling processing of ten times longer documents for improved ROUGE scores.

Luyang Huang, Shuyang Cao, Nikolaus Parulian et al.

2021-04-06 43
cs.CL 2104.00369

FeTaQA: Free-form Table Question Answering

FeTaQA introduces a 10K Wikipedia-based dataset for complex, free-form table question answering, emphasizing multi-fact reasoning.

Linyong Nan, Chiachun Hsieh, Ziming Mao et al.

2021-04-01 42
cs.CL 2103.07191

Are NLP Models really able to Solve Simple Math Word Problems?

This study reveals that current NLP models for math word problems rely heavily on shallow heuristics, with the SVAMP challenge set exposing their fragility, dropping accuracy significantly.

Arkil Patel, Satwik Bhattamishra, Navin Goyal

2021-03-12 1341 citations 36
cs.CL 2103.06332

Hurdles to Progress in Long-form Question Answering

Identifies fundamental challenges in LFQA evaluation and datasets; uses sparse Transformer and contrastive retrieval, achieving SOTA but with answers not truly grounded.

Kalpesh Krishna, Aurko Roy, Mohit Iyyer

2021-03-11 46
cs.CL 2103.02143

Random Feature Attention

Proposes RFA, a linear-time softmax approximation using random features, boosting efficiency in long sequence tasks.

Hao Peng, Nikolaos Pappas, Dani Yogatama et al.

2021-03-03 34
cs.CL 2102.13249

Chess as a Testbed for Language Model State Tracking

Proposes training GPT-2 on UCI chess notation with RAP to enhance state tracking, achieving over 97% accuracy in move and piece position prediction.

Shubham Toshniwal, Sam Wiseman, Karen Livescu et al.

2021-02-26 55