Memory in the Loop: In-Process Retrieval as Extended Working Memory for Language Agents
Proposes in-loop memory for language agents, reducing latency to 100μs and improving task efficiency and accuracy.
Yusuf Khan, Carlo Lipizzi
Proposes in-loop memory for language agents, reducing latency to 100μs and improving task efficiency and accuracy.
Yusuf Khan, Carlo Lipizzi
This study introduces M3Bench, a benchmark for evaluating model editing in medical VLMs, measuring reliability, locality, and generalization across 16,276 clinical questions.
Guli Zhu, Chenwei Wu, Liyue Shen
MetaSkill-Evolve introduces recursive self-improvement via two-timescale meta-skill evolution, boosting task accuracy by +23.54 points on OfficeQA.
Zefeng Wang, Minxi Yan, Jinhe Bi et al.
This paper introduces GPUSimBench, a benchmark system to evaluate GPU-based robotic simulators' scalability, physical fidelity, and non-determinism, revealing key limitations.
Huzhenyu Zhang, Shenghai Yuan, Wenrui Yan et al.
EvoAgentBench evaluates agent self-evolution via ability transfer across four domains, highlighting limitations of automatic methods.
Xingze Gao, Chuanrui Hu, Hongda Chen et al.
Audex, built on Nemotron-Cascade-2-30B-A3B, unifies audio-text modeling with a single Transformer decoder, achieving state-of-the-art performance in audio understanding, speech recognition, translation, and generation, while maintaining text reasoning.
Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim et al.
UNIVERSE employs a single mask-modulated Diffusion Transformer for joint video and trajectory prediction, achieving 4.3× inference speedup and state-of-the-art generalization.
Mengmeng Liu, Diankun Zhang, Jiuming Liu et al.
Proposes multi-scale convolution and capsule routing for video-text localization, achieving 42.9% [email protected] on ActivityNet.
Gengtian Shi, Jinze Yu, Chenhao Wu et al.
RADIANCE uses CLIP feedback for real-time adaptive denoising, boosting rare concept synthesis without retraining.
Zi-Xiang Ni, Bo-Lun Huang, Teng-Fang Hsiao et al.
HunyuanOCR-1.5 enhances OCR performance using DFlash acceleration and Agentic Data Flow, achieving 6.37x inference speedup.
Gengluo Li, Xingyu Wan, Shangpin Peng et al.
Unsupervised detection of underground tunnels using depth-restricted reconstruction scoring, achieving AUC of 0.994.
Muhammad Junaid, Shoab A. Khan, Nisar Ahmed
PRISM generates personalized robotic datasets from a single image and instruction, constructing semantically aligned digital scenes with instance diversity.
Dogyu Ko, Haneul Kim, Chanyoung Yeo et al.
Proposes MECo-WAM, integrating 4D geometric priors during training, boosting manipulation success to 98.2% without increasing inference cost.
Jianjun Zhang, Jian Zhu, Taiyi Su et al.
GeoMoLa learns motion latents by predicting point cloud evolution, achieving SOTA performance with single-view RGB-D input.
Yunchao Zhang, Yijia Weng, Ruizhe Liu et al.
FocusGS uses targeted structure completion to reduce Gaussian count by 74%, improving sparse-view 3D reconstruction efficiency and quality.
Guoqing Wang, Pin Tang, Xiangxuan Ren et al.
PixelPilot decouples 2D planning from 3D lifting, enabling scalable vision-language-driven autonomous driving with state-of-the-art accuracy.
Pin Tang, Guoqing Wang, Xiangxuan Ren et al.
HIEVI-RAG framework improves long document understanding accuracy by 8.05% through hierarchical evidence-driven reasoning.
Junyu Xiong, Yonghui Wang, Rongjian Gu et al.
G2VD integrates counterfactual intervention and causal disentanglement, achieving over 90% accuracy in cross-domain AI video forgery detection.
Meng Du, Hongchang Chen, Ran Li et al.
Using Classifier Discrimination Score (CDS) to address class overlap in single-cell perturbation data, significantly improving identification accuracy.
Youssef Marrakchi, Davide D'Ascenzo, Sebastiano Cultrera di Montesano
Introduces MTEB-BR, a benchmark with 22 native Portuguese tasks, evaluating 93 models using rigorous statistical analysis to distinguish capability tiers.
Tardelli Ronan Coelho Stekel