ASQA: Factoid Questions Meet Long-Form Answers
ASQA dataset addresses ambiguity in long-form QA with a new evaluation metric.
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra et al.
ASQA dataset addresses ambiguity in long-form QA with a new evaluation metric.
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra et al.
Debate-style explanations trained with long contexts did not significantly improve human answer accuracy; snippets outperformed debates.
Alicia Parrish, Harsh Trivedi, Ethan Perez et al.
Pathways系统支持下的540B参数PaLM模型,显著提升少样学习能力,超越多项自然语言任务的SOTA,展现出大规模模型的潜力。
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin et al.
SpecDec leverages speculative execution to accelerate seq2seq generation by 5x with quality comparable to beam search.
Heming Xia, Tao Ge, Peiyi Wang et al.
This study introduces compute-optimal training for large language models, showing model size and data should scale together; trained 70B Chinchilla surpasses larger models.
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch et al.
MedMCQA is a large-scale, multi-subject medical QA dataset with over 194k questions, designed to advance deep reasoning in medical AI.
Ankit Pal, Logesh Kumar Umapathi, Malaikannan Sankarasubbu
BARCOR: A BART-based unified framework for conversational recommendation systems achieving state-of-the-art performance in the movie domain.
Ting-Chun Wang, Shang-Yu Su, Yun-Nung Chen
SAL uses spectral decomposition to remove protected attribute info from neural representations, effective for linear and nonlinear cases.
Shun Shao, Yftah Ziser, Shay B. Cohen
ScienceWorld uses interactive text environments to train small agents (150k params) that outperform large static models (11B params) in elementary science reasoning.
Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté et al.
SummaReranker improves abstractive summarization by 5.44% ROUGE-1 on CNN/DM using a multi-task mixture-of-experts re-ranking framework.
Mathieu Ravaut, Shafiq Joty, Nancy F. Chen
Chart-to-Text provides a large-scale benchmark for chart summarization with 44,096 charts.
Shankar Kantharaj, Rixie Tiffany Ko Leong, Xiang Lin et al.
Fine-tuning GPT-3 with human preferences and PPO reinforcement learning improves instruction-following, truthfulness, and safety; parameters reduced from 175B to 1.3B.
Long Ouyang, Jeff Wu, Xu Jiang et al.
DeepNorm stabilizes extremely deep Transformers, enabling training up to 1000 layers with significant performance gains in multilingual translation.
Hongyu Wang, Shuming Ma, Li Dong et al.
Proposed a trainable subgraph retriever (SR) that enhances multi-hop KBQA by decoupling retrieval and reasoning, achieving state-of-the-art results.
Jing Zhang, Xiaokang Zhang, Jifan Yu et al.
This study reveals that in in-context learning, models rely more on demonstration format and distribution than on correct labels.
Sewon Min, Xinxi Lyu, Ari Holtzman et al.
Contrastive training (SimCTG) and contrastive search improve diversity and coherence in neural text generation, outperforming SOTA methods.
Yixuan Su, Tian Lan, Yan Wang et al.
Proposes causal intervention and ROME method to locate and edit factual associations in GPT, achieving high success rates and better control over knowledge editing.
Kevin Meng, David Bau, Alex Andonian et al.
Survey of hallucination in NLG, covering metrics, mitigation, and task-specific research.
Ziwei Ji, Nayeon Lee, Rita Frieske et al.
Proposes automated red teaming using language models to generate and detect harmful outputs in 280B parameter chatbots, employing multi-strategy approaches.
Ethan Perez, Saffron Huang, Francis Song et al.
Locally typical sampling enforces per-word information content constraints, reducing repetition and improving text quality in language generation.
Clara Meister, Tiago Pimentel, Gian Wiher et al.