cs.CL 2109.01652

Finetuned Language Models Are Zero-Shot Learners

This paper introduces instruction tuning on a 137B parameter model, significantly improving zero-shot performance across 60 NLP tasks, outperforming GPT-3 on many benchmarks.

Jason Wei, Maarten Bosma, Vincent Y. Zhao et al.

2021-09-04 5316 citations 49
cs.CL 2106.09685

LoRA: Low-Rank Adaptation of Large Language Models

LoRA introduces low-rank matrices to freeze pre-trained weights, reducing trainable parameters by 10,000x, with performance comparable or better than full fine-tuning.

Edward J. Hu, Yelong Shen, Phillip Wallis et al.

2021-06-18 22282 citations 39
cs.CL 2106.01229

Lower Perplexity is Not Always Human-Like

This study evaluates the relationship between perplexity and human-like reading behavior across Japanese and English, revealing language-specific differences.

Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito et al.

2021-06-02 36