cs.AI 2504.18530

Scaling Laws For Scalable Oversight

Proposes a framework modeling oversight as a capability-mismatch game using Elo scores, validated on multiple oversight scenarios, deriving scaling laws for success probabilities.

Joshua Engels, David D. Baek, Subhash Kantamneni et al.

2025-04-26 37
cs.AI 2504.05299

SmolVLM: Redefining small and efficient multimodal models

SmolVLM employs architectural and tokenization innovations to create resource-efficient multimodal models with under 1GB GPU memory, outperforming much larger models.

Andrés Marafioti, Orr Zohar, Miquel Farré et al.

2025-04-08 48
cs.AI 2503.23037

Agentic Large Language Models, a survey

Agentic LLMs enhance decision-making via reasoning, acting, and interacting, significantly improving medical diagnosis and logistics analysis.

Aske Plaat, Max van Duijn, Niki van Stein et al.

2025-03-29 0
cs.AI 2503.16416

Survey on Evaluation of LLM-based Agents

This survey systematically analyzes evaluation methods for LLM-based agents, covering core capabilities, application benchmarks, evaluation frameworks, and key dimensions.

Asaf Yehudai, Lilach Eden, Alan Li et al.

2025-03-21 214 citations 37
cs.AI 2502.18864

Accelerating scientific discovery with Co-Scientist

Multi-agent Gemini-based system accelerates hypothesis generation, improving Elo scores from 3.0 to 4.2 over iterative cycles, validated in drug repurposing and target discovery.

Juraj Gottweis, Wei-Hung Weng, Alexander Daryin et al.

2025-02-26 470 citations 38