cs.AI 2607.28367

How Benchmarks Mis-Score Computer-Use Agents

The study reveals benchmark scoring errors for computer-use agents, proposing a four-stage reliability framework with 15.3% incorrect failure verdicts.

Zihan Dong, Zhiyuan Ma, Zekun Wang et al.

2026-07-30 14
cs.CV 2607.28211

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

Large-scale evaluation of 194 VLMs shows model size correlates with overall accuracy (ρ=0.68), but not with robustness to multi-attribute bias; data quality is more critical.

Ioannis Sarridis, Ioannis Kompatsiaris, Symeon Papadopoulos

2026-07-30 48
cs.LG 2607.28022

Flux-OPD: On-Policy Distillation with Evolving Contexts

Flux-OPD uses evolving contexts and reverse KL decomposition to stabilize open-domain distillation, outperforming existing methods with +2-3 score improvements.

Yuran Wang, Zekun Wang, Bohan Zeng et al.

2026-07-30 44
cs.AI 2607.27824

STEREODISCO: Discovering Stereotypicality in LLMs

Proposed STEREODISCO framework applies semantic differential method to discover and quantify stereotypical axes in LLM internal representations, revealing new bias dimensions beyond prior social psychology research.

Farane Jalali Farahani, Corina Dima, Mojtaba Nayyeri et al.

2026-07-30 41