LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
Introduces LFQA-HP-1M dataset with nine scoring metrics, achieving model performance comparable to human preferences in long-form QA evaluation.
Rafid Ishrak Jahan, Fahmid Shahriar Iqbal, Sagnik Ray Choudhury