Detecting Safety Violations Across Many Agent Traces
Meerkat combines clustering with agentic search, finding roughly 4× more CyBench reward-hacking cases than prior audits.
Adam Stein, Davis Brown, Hamed Hassani et al.
Meerkat combines clustering with agentic search, finding roughly 4× more CyBench reward-hacking cases than prior audits.
Adam Stein, Davis Brown, Hamed Hassani et al.
Developed a scalable problem reduction library with 100+ problem types and 200+ rules, integrated via AI agents for automated contribution and verification.
Xi-Wei Pan, Shi-Wen An, Jin-Guo Liu
WebForge automates the end-to-end creation of realistic, reproducible, and scalable web environments using a four-stage pipeline, enabling multi-dimensional capability profiling.
Peng Yuan, Yuyang Yin, Yuxuan Cai et al.
CFMS integrates multimodal perception with hierarchical symbolic reasoning, boosting complex table understanding by 15% accuracy on benchmarks.
Qixian Huang, Hongqiang Lin, Tong Fu et al.
SkillClaw turns multi-user experience into shared skill evolution; on WildClawBench, Qwen3-Max gains up to 52% in Search & Retrieval.
Ziyu Ma, Shidong Yang, Yuxiang Ji et al.
UP-NRPA integrates user portraits with nested rollouts for dynamic strategy planning in goal-oriented dialogues, achieving a 56.41% improvement in negotiation SL.
Hui Wang, Fafa Zhang, Meng Liu et al.
Proposes a five-dimensional auditability framework, emphasizing that accountability depends on system auditability.
Yi Nian, Aojie Yuan, Haiyue Zhang et al.
Proposes a consequence-sensitive support compression method in belief arbitration, balancing information retention and resource costs for robust decision-making.
Mark Walsh
InfoSeeker uses Host–Manager–Worker parallelism, reaching 8.38% WideSearch success and 3–5× faster execution.
Ka Yiu Lee, Yuxuan Huang, Zhiyuan He et al.
HippoCamp benchmarks multimodal file management agents, revealing limitations in user environments with top accuracy only 48.3%.
Zhe Yang, Shulin Tian, Kairui Hu et al.
C-TRAIL integrates LLMs with trust mechanisms via a closed-loop Recall-Plan-Update framework for autonomous driving trajectory planning.
Zhihong Cui, Haoran Tang, Tianyi Li et al.
GISTBench uses IG and IS to test whether LLM user profiles are supported by behavior; survey alignment reaches ρ=0.67.
Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan et al.
Proposed a Markovian framework for auditing agentic AI reliability and oversight cost, improving state-action blind mass by 12.53%.
Biplab Pal, Santanu Bhattacharya
Transformer-based DRL policy achieves 12.89-15.12% gap on large OSSP instances, generalizing from small training data.
Faezeh Ardali, Mwembezi A. Nyelele, Gerald M. Knapp
Proposes LLM Olympiad evaluation to address transparency and trust issues in model evaluation.
Jan Christian Blaise Cruz, Alham Fikri Aji
LongCat-Flash-Prover enhances Lean4 formal reasoning via tool-integrated RL, achieving 97.1% pass rate on MiniF2F-Test.
Jianing Wang, Jianfei Zhang, Qi Guo et al.
OS-Themis framework improves GUI agent performance by 10.3% on AndroidWorld using a multi-agent critic mechanism.
Zehao Li, Zhenyu Wu, Yibo Zhao et al.
Box Maze framework reduces LLM reasoning error rate to below 1% through memory grounding, structured inference, and boundary enforcement.
Zou Qiang
This study employs multi-modal prompt engineering and multi-agent generate-critique-revise frameworks to analyze and mitigate dialect-induced biases in LLM outputs, demonstrating significant bias reduction.
Martina Ullasci, Marco Rondina, Riccardo Coppola et al.
Proposes a reference-free simulation framework by training independent user and recommender simulators for more realistic dialogues.
Jerome Ramos, Feng Xia, Xi Wang et al.