cs.CL 2305.14788

Adapting Language Models to Compress Contexts

Proposes AutoCompressors, an unsupervised method compressing long texts into summary vectors, improving perplexity and in-context learning efficiency.

Alexis Chevalier, Alexander Wettig, Anirudh Ajith et al.

2023-05-24 33
cs.LG 2305.14314

QLoRA: Efficient Finetuning of Quantized LLMs

QLoRA combines 4-bit quantization, LoRA, and paging to finetune 65B models on a single GPU, achieving near ChatGPT performance.

Tim Dettmers, Artidoro Pagnoni, Ari Holtzman et al.

2023-05-24 57