InfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information Seeking
InfoSeeker uses Host–Manager–Worker parallelism, reaching 8.38% WideSearch success and 3–5× faster execution.
Ka Yiu Lee, Yuxuan Huang, Zhiyuan He et al.
InfoSeeker uses Host–Manager–Worker parallelism, reaching 8.38% WideSearch success and 3–5× faster execution.
Ka Yiu Lee, Yuxuan Huang, Zhiyuan He et al.
HippoCamp benchmarks multimodal file management agents, revealing limitations in user environments with top accuracy only 48.3%.
Zhe Yang, Shulin Tian, Kairui Hu et al.
C-TRAIL integrates LLMs with trust mechanisms via a closed-loop Recall-Plan-Update framework for autonomous driving trajectory planning.
Zhihong Cui, Haoran Tang, Tianyi Li et al.
GISTBench uses IG and IS to test whether LLM user profiles are supported by behavior; survey alignment reaches ρ=0.67.
Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan et al.
Proposed a Markovian framework for auditing agentic AI reliability and oversight cost, improving state-action blind mass by 12.53%.
Biplab Pal, Santanu Bhattacharya
Transformer-based DRL policy achieves 12.89-15.12% gap on large OSSP instances, generalizing from small training data.
Faezeh Ardali, Mwembezi A. Nyelele, Gerald M. Knapp
Proposes LLM Olympiad evaluation to address transparency and trust issues in model evaluation.
Jan Christian Blaise Cruz, Alham Fikri Aji
LongCat-Flash-Prover enhances Lean4 formal reasoning via tool-integrated RL, achieving 97.1% pass rate on MiniF2F-Test.
Jianing Wang, Jianfei Zhang, Qi Guo et al.
OS-Themis framework improves GUI agent performance by 10.3% on AndroidWorld using a multi-agent critic mechanism.
Zehao Li, Zhenyu Wu, Yibo Zhao et al.
Box Maze framework reduces LLM reasoning error rate to below 1% through memory grounding, structured inference, and boundary enforcement.
Zou Qiang
This study employs multi-modal prompt engineering and multi-agent generate-critique-revise frameworks to analyze and mitigate dialect-induced biases in LLM outputs, demonstrating significant bias reduction.
Martina Ullasci, Marco Rondina, Riccardo Coppola et al.
SCALe method enhances reasoning and answer accuracy in vision-language models through dynamic weight adjustment.
Shaked Perek, Ben Wiesel, Avihu Dekel et al.
Proposes a reference-free simulation framework by training independent user and recommender simulators for more realistic dialogues.
Jerome Ramos, Feng Xia, Xi Wang et al.
Understanding DNNs through differential equations to enhance performance and applications.
Hongjue Zhao, Yizhuo Chen, Yuchen Wang et al.
This review discusses how synthetic data, virtual environments, and domain adaptation improve autonomous driving perception and planning, emphasizing digital twins and vision-language models.
A. Humnabadkar, A. Sikdar, B. Cave et al.
Adaptive Domain Models leverage Bayesian distillation and warm rotation for efficient training in geometric and neuromorphic AI.
Houston Haynes
LEAFE framework internalizes recovery agency from reflective experience, enhancing Pass@k performance in long-horizon tasks.
Rui Ge, Yichao Fu, Yuyang Qian et al.
The study finds that counterfactual explanation metrics do not align with user perception, necessitating more human-centered evaluation methods.
Felix Liedeker, Basil Ell, Philipp Cimiano et al.
OpenSeeker democratizes frontier search agents by fully open-sourcing training data, utilizing controllable QA synthesis and denoised trajectory synthesis.
Yuwen Du, Rui Ye, Shuo Tang et al.
Proposes a cognitive architecture viewing the psyche as an operating system for constructing AGI.
Anton Kolonin, Vladimir Krykov