cs.CL 2508.17580

UQ: Assessing Language Models on Unsolved Questions

Proposes UQ, a dynamic benchmark assessing LLMs on 500 unsolved questions sourced from Stack Exchange, integrating validator-assisted screening and community verification.

Fan Nie, Ken Ziyu Liu, Zihao Wang et al.

2025-08-25 24