cs.CL 2204.02311

PaLM: Scaling Language Modeling with Pathways

Pathways系统支持下的540B参数PaLM模型,显著提升少样学习能力,超越多项自然语言任务的SOTA,展现出大规模模型的潜力。

Aakanksha Chowdhery, Sharan Narang, Jacob Devlin et al.

2022-04-06 8306 citations 48
cs.CL 2203.15556

Training Compute-Optimal Large Language Models

This study introduces compute-optimal training for large language models, showing model size and data should scale together; trained 70B Chinchilla surpasses larger models.

Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch et al.

2022-03-29 23
cs.CL 2203.07540

ScienceWorld: Is your Agent Smarter than a 5th Grader?

ScienceWorld uses interactive text environments to train small agents (150k params) that outperform large static models (11B params) in elementary science reasoning.

Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté et al.

2022-03-15 66
cs.CL 2203.00555

DeepNet: Scaling Transformers to 1,000 Layers

DeepNorm stabilizes extremely deep Transformers, enabling training up to 1000 layers with significant performance gains in multilingual translation.

Hongyu Wang, Shuming Ma, Li Dong et al.

2022-03-01 20
cs.CL 2202.06417

A Contrastive Framework for Neural Text Generation

Contrastive training (SimCTG) and contrastive search improve diversity and coherence in neural text generation, outperforming SOTA methods.

Yixuan Su, Tian Lan, Yan Wang et al.

2022-02-14 53
cs.CL 2202.05262

Locating and Editing Factual Associations in GPT

Proposes causal intervention and ROME method to locate and edit factual associations in GPT, achieving high success rates and better control over knowledge editing.

Kevin Meng, David Bau, Alex Andonian et al.

2022-02-11 30
cs.CL 2202.03286

Red Teaming Language Models with Language Models

Proposes automated red teaming using language models to generate and detect harmful outputs in 280B parameter chatbots, employing multi-strategy approaches.

Ethan Perez, Saffron Huang, Francis Song et al.

2022-02-07 22
cs.CL 2202.00666

Locally Typical Sampling

Locally typical sampling enforces per-word information content constraints, reducing repetition and improving text quality in language generation.

Clara Meister, Tiago Pimentel, Gian Wiher et al.

2022-02-02 47