UQ: Assessing Language Models on Unsolved Questions
Proposes UQ, a dynamic benchmark assessing LLMs on 500 unsolved questions sourced from Stack Exchange, integrating validator-assisted screening and community verification.
Fan Nie, Ken Ziyu Liu, Zihao Wang et al.