From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
LLM-as-a-judge uses large language models for scoring and selection, enhancing evaluation accuracy.
Dawei Li, Bohan Jiang, Liangjie Huang et al.
LLM-as-a-judge uses large language models for scoring and selection, enhancing evaluation accuracy.
Dawei Li, Bohan Jiang, Liangjie Huang et al.
Proposes CATP-LLM framework combining TPL and CAORL for cost-aware tool planning, improving plan quality by up to 93.9%.
Duo Wu, Jinghe Wang, Yuan Meng et al.
FEPS combines active inference with interpretable graph models, avoiding deep neural networks, to enable flexible, transparent decision-making in partially observable environments.
Joséphine Pazem, Marius Krumm, Alexander Q. Vining et al.
This study uses self-report grounded LLM agents, combining interviews and surveys, to predict individual responses across multiple outcomes with 83-86% accuracy without task-specific training.
Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst et al.
FrontierMath benchmark tests advanced mathematical reasoning; AI solves less than 2%, exposing a large gap with human experts.
Elliot Glazer, Ege Erdil, Tamay Besiroglu et al.
SWE-Search combines MCTS with self-improvement, achieving 23% performance gains in software tasks.
Antonis Antoniades, Albert Örwall, Kexun Zhang et al.
Reflection-Bench, inspired by cognitive psychology, evaluates LLMs' epistemic agency across seven dimensions, revealing significant limitations in meta-reflection with scores below 50 in many models.
Lingyu Li, Yixu Wang, Haiquan Zhao et al.
Proposes a sparse, hierarchical Bayesian network world model supporting open-ended learning with interpretability and scalability.
Lancelot Da Costa
Proposes a standardized framework for tool integration in LLMs, combining fine-tuning and in-context learning to enhance complex task performance.
Zhuocheng Shen
Windows Agent Arena evaluates multi-modal OS agents; Navi achieves 19.5% success in Windows domain.
Rogerio Bonatti, Dan Zhao, Francesco Bonacci et al.
Categorizes concept models into four types—Abstractionism, Similarity, Functional, Invariance—highlighting their mathematical structures.
Jun Otsuka
Proposes Meta Agent Search, an automated code-based agent discovery method outperforming manual designs with strong cross-domain transfer.
Shengran Hu, Cong Lu, Jeff Clune
VAB benchmark evaluates large multimodal models as visual agents across diverse scenarios; behavior cloning significantly improves open model performance.
Xiao Liu, Tianjie Zhang, Yu Gu et al.
Integrating Active Inference with topological mapping enables rapid environment structure learning in navigation agents, outperforming CSCG with fewer steps and no prior environment knowledge.
Daria de Tinguy, Tim Verbelen, Bart Dhoedt
Proposes inference scaling laws showing smaller models with advanced strategies outperform larger ones under limited compute, achieving Pareto optimality.
Yangzhen Wu, Zhiqing Sun, Shanda Li et al.
Llama 3 employs a 405B-parameter dense Transformer supporting 128K tokens, multi-modal, and multi-language capabilities, rivaling GPT-4.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri et al.
MoMa employs modality-aware mixture-of-experts, achieving 3.7× FLOPs savings on 1.4B parameters with minimal performance loss.
Xi Victoria Lin, Akshat Shrivastava, Liang Luo et al.
Proposes CMR, combining neural rule memory with symbolic evaluation for interpretable, verifiable AI, achieving high accuracy and logical rule discovery.
David Debot, Pietro Barbiero, Francesco Giannini et al.
IGOR combines LLM subtask planning with goal-conditioned asynchronous PPO, outperforming reported IGLU and Crafter baselines.
Zoya Volovikova, Alexey Skrynnik, Petr Kuderov et al.
Proposes Trace framework converting complex workflows into OPTO problems, leveraging execution traces and LLMs for automatic parameter optimization, outperforming specialized optimizers.
Ching-An Cheng, Allen Nie, Adith Swaminathan