An Empirical Study of Harness Design for Coding Agents
Study improves coding agents' long-term performance via lightweight harness design; context management proves most effective.
Run-Ze Fan, Zihao Zhang, Simin Ma et al.
Study improves coding agents' long-term performance via lightweight harness design; context management proves most effective.
Run-Ze Fan, Zihao Zhang, Simin Ma et al.
FINSKILLOPS improves SEC filing QA correctness through self-evolving multi-agent system, raising it from 3.70 to 4.55.
Yanzhang Ma, Zhenghan Tai, Hanwei Wu et al.
Introduces Mahalanobis-Ensemble Decoding to optimize LLM decoding via ensemble pruning, enhancing semantic diversity and generation quality.
Dunyao Xue, Chengshuo Du, Zhengbo Wang et al.
AALT selects demonstrations to maximize start-goal connectivity, improving task success rate.
Maxwell J. Jacobson, Ahmed H Qureshi, Yexiang Xue
The study introduces META-AGENT for task-agnostic preprocessing in unknown environments, achieving highest Avg@3 reward on five benchmarks.
Vinay Samuel, Varun Ursekar, Vijay S. Kalmath et al.
JarvisGUI evaluates cross-device GUI agents with dynamic task composition, revealing capability gaps in real-world workflows.
Zixiang Chen, Yuheng Lu, Zihao Cheng et al.
ConvMem reformulates long-context reasoning as hierarchical convolution, outperforming training-free baselines on RULER-HotpotQA.
Hongming Zhang, Zhaozhen Gu, Fengshuo Bai et al.
Fortunate Recall uses a 10+1 behavioral ontology and lifecycle policies to enhance LLM memory management, achieving 76.9% on LifecycleBench.
Ansuman Mullick, Eray Tüzün
LLM digital twins reduce human measurement via statistical substitutability, but behavioral fidelity is insufficient.
Steven Wang, Kyle Hunt, Shaojie Tang et al.
Introduced IIns-GAN for generating labeled wireless signals, enhancing model training.
Yuxiao Li, Keke Hu, Santiago Mazuelas et al.
EDGE framework synthesizes tool-calling data using dynamic graphs, enhancing performance.
Dain Kim, Eungi Cho, Kyumin Kim et al.
Evaluates LLM explanations' necessity and sufficiency using behavioral evidence, finding limited correlation with model behavior.
Urja Pawar, Rajitha Ramanayake, Nabeel Kemal et al.
Study finds widespread verbatim retrieval in LLMs on molecular regression benchmarks, affecting prediction accuracy.
Matthias Busch, Marius Tacke, Sviatlana V. Lamaka et al.
CUA-Universe uses App-Forge, Task-Weave, and Path-Steer to create hybrid GUI+CLI environments, enhancing efficiency and success rates.
Haoting Shi, Wenhao Wang, Weicheng Fang et al.
The study examines memory portability during model upgrades, finding fixed-schema knowledge graphs remain stable.
Ankit Goyal, Jaideep Ray
Toolkit for measuring contextual individuation in Transformer models using bridge forms.
José Luciano Verçosa Marques, Frederico Jorge Heitmann, Daniel Omar Perez et al.
Large language models for HVAC operations in building energy systems; none ready for industry deployment.
Alexander Neubauer, Tianzhen Hong, Han Li et al.
Achieved efficient ETO evaluation via accumulation and blending matrix reformulations, with up to 256.72x speedup.
Yanchen Li, Xiaoming Xue, Kay Chen Tan
APT-RAG framework excels in evidence-intensive QA, achieving a 40.69% F1 score improvement.
Songeun Lee, Kyungjin Min, Injae Na et al.
Proposed a cost-aware single-generation architecture achieving 91.7% accuracy on the DevRev benchmark.
Yoga Sri Varshan Varadharajan, Ajay Yadav, Ritesh Goru et al.