AQuaMuSe: Automatically Generating Datasets for Query-Based Multi-Document Summarization
AQuaMuSe method automatically generates 5,519 query-based multi-document summarization datasets.
Sayali Kulkarni, Sheide Chammas, Wan Zhu et al.
AQuaMuSe method automatically generates 5,519 query-based multi-document summarization datasets.
Sayali Kulkarni, Sheide Chammas, Wan Zhu et al.
TweetEval unifies benchmarks to improve tweet classification accuracy.
Francesco Barbieri, Jose Camacho-Collados, Leonardo Neves et al.
Proposes a regularized Mahalanobis-based differential privacy mechanism for text, balancing privacy and utility with elliptical noise.
Zekun Xu, Abhinav Aggarwal, Oluwaseyi Feyisetan et al.
Proposes XOR QA, a cross-lingual open-retrieval question answering framework using a 40K-question dataset across 7 languages, achieving up to 18.7 F1 score.
Akari Asai, Jungo Kasai, Jonathan H. Clark et al.
Proposes SapBERT, a self-alignment pretraining method leveraging UMLS for biomedical entity embeddings, achieving SOTA in entity linking.
Fangyu Liu, Ehsan Shareghi, Zaiqiao Meng et al.
Proposed SAU-Solver using UET to solve diverse math word problems, achieving 44.83% accuracy.
Jinghui Qin, Lihui Lin, Xiaodan Liang et al.
Constructed Tatoeba benchmark with 555 languages, 2961 pairs, using OPUS and crowd-sourced data; trained Transformer models achieving high scores in low-resource scenarios.
Jörg Tiedemann
ReviewRobot uses knowledge graph comparison to automatically generate paper review scores and comments, achieving 71.4% accuracy.
Qingyun Wang, Qi Zeng, Lifu Huang et al.
Reformulating unsupervised style transfer as paraphrase generation using pretrained GPT-2, achieving 2-3x improvements over SOTA on automatic metrics.
Kalpesh Krishna, John Wieting, Mohit Iyyer
Introduced TG-ReDial dataset, combining SASRec and BERT to enhance conversational recommender systems.
Kun Zhou, Yuanhang Zhou, Wayne Xin Zhao et al.
ALFWorld integrates text and visual environments, enabling abstract policy transfer to embodied tasks with 26% success in unseen scenes.
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté et al.
Span-Fact model corrects factual errors in summaries via span selection, significantly improving consistency.
Yue Dong, Shuohang Wang, Zhe Gan et al.
Imitating implicit scenarios enhances dialogue diversity and relevance, outperforming SOTA by 10% on key metrics.
Shaoxiong Feng, Xuancheng Ren, Hongshen Chen et al.
Using 68 probing tasks, the study finds BERT loses linguistic information after fine-tuning, while more readable representations improve NLI accuracy.
Alessio Miaschi, Dominique Brunato, Felice Dell'Orletta et al.
kNN-MT enhances translation via nearest neighbor search, improving German-English BLEU by 1.5.
Urvashi Khandelwal, Angela Fan, Dan Jurafsky et al.
QAEval uses question-answer pairs generated from references and pre-trained QA models to evaluate summary content, outperforming ROUGE and BERTScore on benchmarks.
Daniel Deutsch, Tania Bedrax-Weiss, Dan Roth
CrowS-Pairs dataset quantifies social biases in MLMs via 1508 sentence pairs across nine bias types, revealing widespread bias.
Nikita Nangia, Clara Vania, Rasika Bhalerao et al.
Introduces MedQA dataset, combining retrieval and comprehension models, with test accuracy of only 36.7%-70.1%, highlighting challenges in medical open-domain QA.
Di Jin, Eileen Pan, Nassim Oufattole et al.
SPARTA employs sparse Transformer matching with inverted index for efficient open-domain QA, outperforming dense vector methods.
Tiancheng Zhao, Xiaopeng Lu, Kyusong Lee
Proposed ClariQ task with deep learning models for generating and ranking clarifying questions in open-domain dialogue.
Mohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin et al.