cs.AI 2607.28367

How Benchmarks Mis-Score Computer-Use Agents

The study reveals benchmark scoring errors for computer-use agents, proposing a four-stage reliability framework with 15.3% incorrect failure verdicts.

Zihan Dong, Zhiyuan Ma, Zekun Wang et al.

2026-07-30 1
cs.AI 2607.27824

STEREODISCO: Discovering Stereotypicality in LLMs

Proposed STEREODISCO framework applies semantic differential method to discover and quantify stereotypical axes in LLM internal representations, revealing new bias dimensions beyond prior social psychology research.

Farane Jalali Farahani, Corina Dima, Mojtaba Nayyeri et al.

2026-07-30 37
cs.AI 2607.24647

Efficiency Matters in Autonomous Research

Proposes AUC of Pareto frontier for efficiency evaluation; introduces fluid search, outperforming fixed strategies in 12 tasks.

Haiqian Yang, Yuan Cao

2026-07-28 44