UserBench: An Interactive Gym Environment for User-Centric Agents
UserBench evaluates user interaction capabilities, finding models align with user intent only 20% of the time.
Cheng Qian, Zuxin Liu, Akshara Prabhakar et al.
UserBench evaluates user interaction capabilities, finding models align with user intent only 20% of the time.
Cheng Qian, Zuxin Liu, Akshara Prabhakar et al.
Agentic Web leverages AI agents for automated web interactions, enhancing user experience.
Yingxuan Yang, Mulei Ma, Yuxuan Huang et al.
Integrating RLHF and Constitutional AI, achieving 20% safety improvement over baseline, with high preference alignment and robustness.
Haoran Lu, Luyang Fang, Ruidong Zhang et al.
Using GPT-4-based generative agents in large-scale social simulations, evaluating their Theory of Mind and biases, with implications for AI-driven social modeling.
Patrick Taillandier, Jean Daniel Zucker, Arnaud Grignard et al.
WGRAMMAR precompiles static structures and uses FSMs, achieving 250× speedup in constrained decoding.
Ran Wang, Xiaoxuan Liu, Hao Ren et al.
RCT study shows 2025 AI tools unexpectedly increased developer task time by 19%, contradicting expectations.
Joel Becker, Nate Rush, Elizabeth Barnes et al.
GTA1 employs test-time scaling and RL-based coordinate prediction, achieving state-of-the-art GUI interaction accuracy with 83.1% on ScreenSpot-Pro.
Yan Yang, Dongxu Li, Yutong Dai et al.
Introduced J1-ENVS and J1-EVAL for dynamic legal environment assessment; models scored below 60%, highlighting procedural gaps.
Zheng Jia, Shengbin Yue, Wei Chen et al.
EvoAgentX integrates TextGrad, AFlow, and MIPRO for automated multi-agent workflow evolution, achieving up to 20% performance gains.
Yingxu Wang, Siwei Liu, Jinyuan Fang et al.
Proposes multimodal perception and world modeling framework for autonomous embodied AI, achieving 85% action prediction accuracy.
Pascale Fung, Yoram Bachrach, Asli Celikyilmaz et al.
Introduces Mind2Web 2 benchmark with Agent-as-a-Judge, evaluating 130 long-horizon tasks achieving 50-70% of human performance.
Boyu Gou, Zanming Huang, Yuting Ning et al.
jina-embeddings-v4 is a 3.8B parameter unified multimodal embedding model supporting single/multi-vector representations for cross-modal retrieval.
Michael Günther, Saba Sturua, Mohammad Kalim Akram et al.
Deep Research agents enhance complex task handling via dynamic reasoning and multi-hop retrieval.
Yuxuan Huang, Yihang Chen, Haozheng Zhang et al.
HeurAgenix leverages large language models (LLMs) for automatic heuristic evolution and adaptive selection, outperforming existing hyper-heuristics on classic benchmarks.
Xianliang Yang, Ling Zhang, Haolong Qian et al.
AlphaEvolve combines evolutionary algorithms with large models to discover efficient algorithms, achieving breakthroughs like a 48-multiplication matrix algorithm after 56 years.
Alexander Novikov, Ngân Vũ, Marvin Eisenberger et al.
ContextBench evaluates context modification methods for targeted latent activation, enhanced with model-assisted EPO achieving better activation-fluency balance.
Robert Graham, Edward Stevinson, Leo Richter et al.
Proposes MM-R5, a reinforcement learning-based multimodal reranker with explicit reasoning chains, achieving over 4% recall@1 improvement on MMDocIR.
Mingjun Xu, Jinhan Dong, Jue Hou et al.
DiMo-GUI combines modality decoupling and dynamic zooming to improve GUI grounding accuracy without extra training, boosting performance over baseline models.
Hang Wu, Hongkai Chen, Yujun Cai et al.
GUI-Critic-R1 model with S-GRPO optimizes pre-operation error diagnosis in GUI automation, achieving 85% critic accuracy and improving success rate from 22.4% to 27.6%.
Yuyang Wanyan, Xi Zhang, Haiyang Xu et al.
This paper establishes theoretical bounds on predicting agent behavior from observed actions, based on causal models, highlighting fundamental limits in inferring beliefs in unseen environments.
Alexis Bellot, Jonathan Richens, Tom Everitt