DataComp-LM: In search of the next generation of training sets for language models
Introduces DataComp-LM, a benchmark with 240T tokens, using model-based filtering to improve training data quality, achieving 64% MMLU 5-shot accuracy with a 7B model.
Jeffrey Li, Alex Fang, Georgios Smyrnis et al.