The Importance of Directional Feedback for LLM-based Optimizers
OptAgent uses directional feedback for text optimization; across four functions and 10-step trials, it is more stable and efficient.
Allen Nie, Ching-An Cheng, Andrey Kolobov et al.
OptAgent uses directional feedback for text optimization; across four functions and 10-step trials, it is more stable and efficient.
Allen Nie, Ching-An Cheng, Andrey Kolobov et al.
MuDreamer predicts rewards, value, and actions in latent space without pixel reconstruction, boosting robustness and training speed.
Maxime Burchi, Radu Timofte
ChatScene employs LLMs and knowledge retrieval to generate safety-critical traffic scenarios, enhancing autonomous vehicle testing.
Jiawei Zhang, Chejian Xu, Bo Li
Proposes nondeterministic structural causal models (NSCM) using multi-valued functions to enhance counterfactual semantics.
Sander Beckers
STAR benchmark evaluates real-world situated reasoning via hypergraph abstraction and programmatic QA, revealing current models' limitations in dynamic videos.
Bo Wu, Shoubin Yu, Zhenfang Chen et al.
This study surveys emerging AI agent architectures for reasoning, planning, and tool calling, revealing their capabilities and limitations.
Tula Masterman, Sandi Besen, Mason Sawtell et al.
Proposes FAC framework with context-specific independence and JACI algorithm for automatic causal discovery in complex environments.
Caleb Chuck, Sankaran Vaidyanathan, Stephen Giguere et al.
This study evaluates VLMs' ability to infer private attributes from inconspicuous images, achieving up to 77.6% accuracy, highlighting privacy risks.
Batuhan Tömekçe, Mark Vero, Robin Staab et al.
I-Design employs multi-agent LLMs to generate personalized 3D interior scenes, integrating scene graph layout and asset retrieval for high-quality results.
Ata Çelen, Guo Han, Konrad Schindler et al.
Eurus models, finetuned with UltraInteract preference trees, achieve state-of-the-art open-source reasoning performance, surpassing GPT-3.5 Turbo in complex tasks.
Lifan Yuan, Ganqu Cui, Hanbin Wang et al.
AgentStudio offers a unified multi-modal environment with tools and benchmarks to evaluate general virtual agents' core abilities, including GUI grounding, video learning, and success detection.
Longtao Zheng, Zhiyuan Huang, Zhenghai Xue et al.
Proposes LoraRetriever, an input-aware LoRA retrieval and composition framework, achieving significant performance gains over baselines in multi-task scenarios, with detailed experimental validation.
Ziyu Zhao, Leilei Gan, Guoyin Wang et al.
Proposes an equitable R&D framework based on Scrum, integrating tools like model cards and diversity metrics to mitigate bias in AI and robotics projects.
Andrew Hundt, Julia Schuller, Severin Kacianka
Proposes debate framework to improve weak model evaluation of strong models, achieving 76% accuracy with non-expert judges.
Akbir Khan, John Hughes, Dan Valentine et al.
BetterV fine-tunes LLMs with discriminative guidance for controlled Verilog generation, outperforming GPT-4 on VerilogEval.
Zehua Pei, Hui-Ling Zhen, Mingxuan Yuan et al.
Marabou 2.0 employs optimized architecture and DeepSoI algorithm to support multiple activation functions, boosting neural network formal verification efficiency.
Haoze Wu, Omri Isac, Aleksandar Zeljić et al.
Proposes MMToM-QA benchmark integrating Bayesian inverse planning and language models to enhance machine Theory of Mind in multimodal scenarios.
Chuanyang Jin, Yutong Wu, Jing Cao et al.
Extends causal models with impossible worlds to formalize mathematical explanations, addressing the issue that all mathematical facts are true in all models.
Joseph Y. Halpern
Proposes NPHardEval, a dynamic benchmark based on complexity classes, with 900 algorithmic questions to evaluate LLM reasoning, updated monthly.
Lizhou Fan, Wenyue Hua, Lingyao Li et al.
Proposes Meta-CPO, integrating differentiable convex programming for safe, rapid adaptation in non-stationary environments.
Minjae Cho, Chuangchuang Sun