SE-GA: Memory-Augmented Self-Evolution for GUI Agents
SE-GA integrates hierarchical memory and self-evolution to boost GUI task success to 89.0%, enabling long-term planning and continuous learning.
Shilong Jin, Lanjun Wang, Zhuosheng Zhang
SE-GA integrates hierarchical memory and self-evolution to boost GUI task success to 89.0%, enabling long-term planning and continuous learning.
Shilong Jin, Lanjun Wang, Zhuosheng Zhang
QuadLink generates high-quality quad-dominant meshes via point-relation learning, improving geometric fidelity and topological quality.
Yiheng Zhang, Zhe Zhu, Tingrui Shen et al.
EgoKit offers unified low-cost egocentric data collection across six heterogeneous devices.
Liuchuan Yu, Erdem Murat, Beichen Wang et al.
EVA01 integrates 3D mesh understanding and generation using Mixture-of-Transformers.
Zongyuan Yang, Mingjing Yi, Wanli Ma et al.
This study compares MRL-trained and non-MRL models under various truncation levels, showing non-MRL models outperform MRL ones below 80% truncation.
Sotaro Takeshita, Yurina Takeshita, Simone Paolo Ponzetto et al.
MetaEns uses unsupervised meta-learning to select outlier detection ensembles, improving precision and reducing model count across 39 datasets.
Hong-Phuc Phan, Tuan-Anh Vu, Tung Kieu et al.
ARIA framework decomposes musical aspects for training data attribution, enhancing diagnostic reliability.
Changheon Han, Ashkan Panahi, Kıvanç Tatar
ShopGym combines ShopArena and ShopGuru to create realistic, scalable e-commerce environments for benchmarking AI agents, ensuring structural fidelity and reproducibility.
Chinmay Savadikar, Mingyu Zhao, Yuanzheng Zhu et al.
MO-CAPO introduces a multi-objective, cost-aware prompt optimization algorithm, improving Pareto front diversity and efficiency.
Jan Büssing, Moritz Schlager, Timo Heiß et al.
PAGER enhances geometric GUI control precision with topology-aware agents, achieving 4.1x task success improvement.
Jingxuan Wei, Xi Bai, Shan Liu et al.
Detect unsupervised domain shifts using interpretable subspace attribution, identifying subtle differences in dataset probability distributions.
Sebastian Springer, Alessandro Laio
RoadmapBench evaluates long-horizon software development; top model Claude-Opus-4.7 resolves only 39.1% of tasks.
Xinbo Xu, Ruihan Yang, Haiyang Shen et al.
VLMs fail in visual path following, especially under local similarity interference.
Hyesoo Hong, Minsoo Kim, Wonje Jeung et al.
VAGS adaptively adjusts guidance scale based on velocity direction similarity, improving image fidelity and semantic control without extra training.
Yan Luo, Ahmadou Aidara, Jingyi Lu et al.
PSD enhances diffusion LLM inference efficiency via Parallel Speculative Decoding, achieving up to 5.5× tokens per forward pass.
Shengyin Sun, Yiming Li, Renxi Liu et al.
Derives transformer-like inference architecture via optimal control, covering nonlinear discrete and linear Gaussian models.
Aditya Kudre, Heng-Sheng Chang, Prashant G. Mehta
NavRL++ introduces a system-level framework with perturbation-aware fine-tuning and Transformer-based temporal reasoning, achieving zero-shot sim-to-real transfer in robot navigation.
Zhefan Xu, Hanyu Jin, Kenji Shimada
Proposes Ghosted Layers, a training-free method deriving a closed-form linear operator to recover pruned layer activations, improving accuracy and perplexity.
Vincent-Daniel Yun, Junhyuk Jo, Sai Praneeth Karimireddy et al.
LifeSentence leverages pretrained language models with structured event data to predict and analyze human life trajectories, outperforming baselines with 35.3% joint accuracy.
Samuel Liu, Muchen Xi, William Yeoh et al.
DriveCtrl narrows the sim-to-real gap in driving videos using depth-conditioned generation.
Haonan Zhao, Yiting Wang, Jingkun Chen et al.