WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
WebOperator employs action-aware tree search with safe backtracking, achieving 54.6% success on WebArena.
Mahir Labib Dihan, Tanzima Hashem, Mohammed Eunus Ali et al.
WebOperator employs action-aware tree search with safe backtracking, achieving 54.6% success on WebArena.
Mahir Labib Dihan, Tanzima Hashem, Mohammed Eunus Ali et al.
Proposes AdaptiveDetector, combining YOLOv11 and VLM with adaptive thresholds and GRPO, boosting zero-shot polyp recall by 14-22% under challenging conditions.
Shengkai Xu, Hsiang Lun Kao, Tianxiang Xu et al.
WeDetect achieves fast open-vocabulary object detection via retrieval, achieving SOTA across 15 benchmarks with high inference efficiency.
Shenghao Fu, Yukun Su, Fengyun Rao et al.
Proposes lightweight clip selection combined with large language models to extract key moments, achieving near-reference summary quality with less than 6% video content.
Galann Pennec, Zhengyuan Liu, Nicholas Asher et al.
UFVideo unifies multi-scale video understanding, integrating global, pixel, and temporal info, outperforming GPT-4 with 7.3% improvement across benchmarks.
Hewen Pan, Cong Wei, Dashuang Liang et al.
Proposes Group Diffusion, leveraging cross-sample attention to improve image generation, achieving up to 32.2% FID reduction.
Sicheng Mo, Thao Nguyen, Richard Zhang et al.
Proposed SEPL method enhances noisy domain generalization via feature probing and prediction ensemble.
Wang Lu, Jindong Wang
The FACTS Leaderboard evaluates large language models' factuality using four sub-leaderboards, with an average score of 68.8.
Aileen Cheng, Alon Jacovi, Amir Globerson et al.
ReMe employs multi-faceted distillation, scenario-aware indexing, and utility-based refinement, enabling small models to outperform larger ones without memory, with 8.83% gain on BFCL-V3.
Zouying Cao, Jiaji Deng, Li Yu et al.
Proposes KHRONOS, a kernel-based neural surrogate for multi-fidelity aerodynamic prediction, reducing parameters by 94% and accelerating training/inference.
Apurba Sarker, Reza T. Batley, Darshan Sarojini et al.
SEMDICE is an off-policy state entropy maximization method using stationary distribution correction, enabling learning from arbitrary datasets with theoretical guarantees.
Jongmin Lee, Meiqi Sun, Pieter Abbeel
UniUGP framework enhances autonomous driving by integrating understanding, generation, and planning, achieving superior decision accuracy in long-tail scenarios.
Hao Lu, Ziyang Liu, Guangfeng Jiang et al.
Integrates conformal prediction into multi-armed bandits, ensuring finite-sample coverage and reward efficiency under weak arm separability.
Simone Cuonzo, Nina Deliu
Fine-tuning on narrow datasets causes broad, unpredictable behaviors and hidden backdoors, exposing security risks in LLMs.
Jan Betley, Jorio Cocola, Dylan Feng et al.
d-TreeRPO enhances policy optimization reliability for diffusion language models, achieving 86.2% improvement on Sudoku.
Leyi Pan, Shuchang Tao, Yunpeng Zhai et al.
A YOLO–TensorRT–Isaac ROS QuadPlane landing stack achieved 32.7 ms per frame, enabling over 30 FPS edge inference.
Ashik E Rasul, Humaira Tasnim, Ji Yu Kim et al.
METRO algorithm minimizes activated experts instead of token counts, reducing latency and boosting throughput in memory-bound MoE inference.
Yanpeng Yu, Haiyue Ma, Krish Agarwal et al.
Tri-Bench tests VLM spatial reasoning under camera tilt and object interference, with ~69% average accuracy.
Amit Bendkhale
InfiniteVL synergizes linear and sparse attention for efficient unlimited-input vision-language models, achieving 1.7x decoding speedup.
Hongyuan Tao, Bencheng Liao, Shaoyu Chen et al.
AEGIS integrates control barrier functions into VLA models, achieving over 50% improvement in obstacle avoidance and nearly 10% higher task success.
Songqiao Hu, Zeyi Liu, Shuang Liu et al.