cs.LG 2402.19464

Curiosity-driven Red-teaming for Large Language Models

Curiosity-driven red teaming (CRT) leverages exploration rewards to enhance test coverage, successfully eliciting toxic responses from LLaMA2, with a 19.6% toxicity rate compared to 10.2% by baseline methods.

Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang et al.

2024-03-01 96 citations 32
cs.LG 2402.10260

A StrongREJECT for Empty Jailbreaks

StrongREJECT benchmark uses high-quality forbidden prompts and an automated evaluator to accurately measure jailbreak effectiveness, reducing overestimation issues.

Alexandra Souly, Qingyuan Lu, Dillon Bowen et al.

2024-02-16 36
cs.LG 2402.10240

A Dynamical View of the Question of Why

Proposes a causal reasoning framework based on reinforcement learning, utilizing two key lemmas to quantify causality in multivariate stochastic processes.

Mehdi Fatemi, Sindhu Gowda

2024-02-15 42
cs.LG 2402.08667

Target Score Matching

Introduces Target Score Matching (TSM), leveraging known target scores to improve low-noise score estimates, enhancing diffusion models’ accuracy.

Valentin De Bortoli, Michael Hutchinson, Peter Wirnsberger et al.

2024-02-14 35
cs.LG 2402.07871

Scaling Laws for Fine-Grained Mixture of Experts

Introduces granular hyperparameter G and scaling laws for MoE, optimizing training to outperform dense Transformers under various compute budgets.

Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski et al.

2024-02-13 32