Leveraging Large Language Models for Multiple Choice Question Answering
Utilizing large language models for MCQA, MCP method narrows the gap with SOTA across 20 datasets.
Joshua Robinson, Christopher Michael Rytting, David Wingate
Utilizing large language models for MCQA, MCP method narrows the gap with SOTA across 20 datasets.
Joshua Robinson, Christopher Michael Rytting, David Wingate
This paper introduces a syntactic surprisal measure based on CCG supertags within neural language models, revealing underestimation of human processing difficulty in garden path sentences.
Suhas Arehalli, Brian Dillon, Tal Linzen
ReasonFormer enhances performance across 11 datasets through modular reasoning.
Wanjun Zhong, Tingting Ma, Jiahai Wang et al.
Simple prompts enhance GPT-3's reliability in generalizability, social bias, calibration, and factuality.
Chenglei Si, Zhe Gan, Zhengyuan Yang et al.
MEMIT mass-edits 10,000 facts in GPT-J, reaching an 85.8 COUNTERFACT editing score at scale.
Kevin Meng, Arnab Sen Sharma, Alex Andonian et al.
This study investigates GPT-3's few-shot table reasoning ability, leveraging Chain of Thoughts prompts to achieve near state-of-the-art performance.
Wenhu Chen
ERNIE-Layout enhances pre-training with layout knowledge, achieving significant improvements in document understanding tasks.
Qiming Peng, Yinxu Pan, Wenjin Wang et al.
This study compares transformers' generalization from stored weights versus in-context information, revealing larger models exhibit more rule-based in-context reasoning.
Stephanie C. Y. Chan, Ishita Dasgupta, Junkyung Kim et al.
Proposes RA-VQA, an end-to-end framework combining differentiable DPR and answer generation, achieving 54.48 VQA score on OK-VQA.
Weizhe Lin, Bill Byrne
Auto-CoT automatically generates diverse reasoning demonstrations, leveraging clustering and prompting to outperform manual methods in reasoning tasks.
Zhuosheng Zhang, Aston Zhang, Mu Li et al.
The study narrows the compositionality gap in language models using the self-ask method, improving accuracy on GPT-3.
Ofir Press, Muru Zhang, Sewon Min et al.
Pix2Struct pretrains by parsing web screenshots into simplified HTML, achieving state-of-the-art results in six tasks.
Kenton Lee, Mandar Joshi, Iulia Turc et al.
This study analyzes added toxicity in multilingual MT using HOLISTICBIAS, ALTI+ attribution, revealing low-resource languages and demographic axes prone to toxicity.
Marta R. Costa-jussà, Eric Smith, Christophe Ropers et al.
Binder combines GPT-3 Codex in a training-free neural-symbolic framework, achieving state-of-the-art results in complex question answering with minimal examples.
Zhoujun Cheng, Tianbao Xie, Peng Shi et al.
Decomposed Prompting employs modular task decomposition, improving GPT-3 few-shot performance on complex reasoning tasks.
Tushar Khot, Harsh Trivedi, Matthew Finlayson et al.
Introduces RL4LMs library, GRUE benchmark, and NLPO algorithm, improving language model preference alignment.
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley et al.
ProPose PrOntoQA, a logic-based synthetic dataset, to systematically analyze LLM reasoning, revealing strengths in single-step deduction but weaknesses in proof planning.
Abulhair Saparov, He He
DecAF jointly decodes logical forms and direct answers using text retrieval, achieving SOTA on WebQSP, FreebaseQA, GrailQA with 79-82% Hits@1.
Donghan Yu, Sheng Zhang, Patrick Ng et al.
Proposes 'Context Distillation' to internalize reasoning, instructions, and examples, boosting model performance by 9-30% on key benchmarks.
Charlie Snell, Dan Klein, Ruiqi Zhong
LiMBeR method linearly maps image features to text prompts, enhancing visual question answering performance.
Jack Merullo, Louis Castricato, Carsten Eickhoff et al.