cs.CL 2401.10020

Self-Rewarding Language Models

Self-Rewarding language model using iterative DPO training, surpassing Claude 2, Gemini Pro, and GPT-4 0613 on AlpacaEval 2.0 with Llama 2 70B.

Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho et al.

2024-01-18 693 citations 27
cs.CL 2401.06954

Bridging the Preference Gap between Retrievers and LLMs

Proposes BGM, a bridge model trained via supervised and reinforcement learning, to optimize the connection between retrievers and LLMs, boosting QA and personalized generation by 10-15%.

Zixuan Ke, Weize Kong, Cheng Li et al.

2024-01-13 90 citations 35
cs.CL 2401.03462

Long Context Compression with Activation Beacon

Activation Beacon introduces progressive activation compression, enabling efficient long-text processing with up to 8x compression, doubling inference speed.

Peitian Zhang, Zheng Liu, Shitao Xiao et al.

2024-01-07 57