cs.CL 2608.06370

The Bitter Lesson of Tool Calling

This study compares programmatic tool calling (PTC) with native JSON calls across 14 models on BFCL v4, showing a 10.6% accuracy boost for PTC, demonstrating robustness and scalability.

Ishan Patel, Sahil Sen, Elias Lumer et al.

2026-08-07 102
cs.AI 2608.06301

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

HarnessOpt-Bench benchmark evaluates 5 frontier LLMs in costly, stochastic harness optimization, revealing significant model differences and room for improvement.

Varun Ursekar, Apaar Shanker, Yash Maurya et al.

2026-08-07 141
math.DS 2608.06155

Verifiable Regularity Criterion for Conditional Expectation Operators and Conditional Mean Embeddings with Applications to Nonparametric Regression, Bayesian Inverse Problems, and Koopman Operators

Proposes a verifiable regularity criterion based on the Sobolev regularity of the Radon–Nikodym density, ensuring boundedness and Hilbert–Schmidt properties of conditional expectation operators (CEOs) and embeddings, with applications in nonparametric regression, Bayesian inverse problems, and Koopman operators.

Maximiliano Hertel, Ilja Klebanov, Manuel Schaller et al.

2026-08-06 89