Agents in Software Engineering: Survey, Landscape, and Vision
A survey of 115 studies proposes a Perception–Memory–Action framework for LLM-based software-engineering agents.
Yanlin Wang, Wanjun Zhong, Yanxian Huang et al.
A survey of 115 studies proposes a Perception–Memory–Action framework for LLM-based software-engineering agents.
Yanlin Wang, Wanjun Zhong, Yanxian Huang et al.
VLTP leverages MLLM-guided dynamic token pruning to reduce ViT computation by 25%, maintaining performance in task-oriented segmentation.
Hanning Chen, Yang Ni, Wenjun Huang et al.
Windows Agent Arena evaluates multi-modal OS agents; Navi achieves 19.5% success in Windows domain.
Rogerio Bonatti, Dan Zhao, Francesco Bonacci et al.
BLens uses contrastive learning to generate binary function names, achieving an F1 score of 0.79.
Tristan Benoit, Yunru Wang, Moritz Dannehl et al.
This survey reviews methods for handling missing modalities in deep multimodal learning, enhancing model robustness.
Renjie Wu, Hu Wang, Hsiang-Ting Chen et al.
Proposes Agent Workflow Memory (AWM), which induces and stores reusable workflows, boosting web navigation success rates by 24.6% and 51.1% on Mind2Web and WebArena respectively.
Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried et al.
Proposes multi-category negative sampling strategies, enhancing detection of hard negatives in recommendation models.
Haokai Ma, Ruobing Xie, Lei Meng et al.
Proposed RFedAGS, a Riemannian federated learning algorithm based on gradient stream averaging, effectively handles partial participation and data heterogeneity.
Zhenwei Huang, Wen Huang, Pratik Jawanpuria et al.
PingPong benchmark evaluates role-playing language models using user emulation and multi-model evaluation, validating over 40 models.
Ilya Gusev
E2LLM combines chunk-based soft prompts with pre-trained encoders and decoders, overcoming the 'impossible triangle' of performance, efficiency, and compatibility for long-context tasks.
Zihan Liao, Jun Wang, Hang Yu et al.
Proposes Semantic Norm Behavior Analysis using ontology for traceable behavior specifications in automated driving.
Nayel Fabian Salem, Marcus Nolte, Veronica Haber et al.
DOLCE framework parameterizes problems with λ and k to identify retrieval and holistic understanding tasks in long contexts.
Zi Yang
OmniLens combines EAOD-generated LensLib with unsupervised domain adaptation for blind lens aberration correction, achieving +1.81dB PSNR improvement.
Qi Jiang, Yao Gao, Shaohua Gao et al.
EndoOmni employs teacher-student self-learning with confidence-guided robust loss to achieve zero-shot cross-dataset endoscopy depth estimation, reducing absolute relative error by 33%.
Qingyao Tian, Zhen Chen, Huai Liao et al.
This paper introduces GenCAD, integrating Transformer, contrastive learning, and diffusion models to generate editable CAD programs from images, outperforming SOTA methods.
Md Ferdous Alam, Faez Ahmed
Vibration-based omni-directional sliding control enables ultra-thin card-shaped robots with tactile feedback, integrating sensing and wireless communication.
Aditya Retnanto, Emilie Faracci, Anup Sathya et al.
Diffusion models enhance recommender systems' generative capabilities and stability, significantly improving user preference prediction.
Jianghao Lin, Jiaqi Liu, Jiachen Zhu et al.
POINTS improves vision-language models with perplexity-based data filtering and model weight fusion, achieving SOTA performance.
Yuan Liu, Zhongyin Zhao, Ziyuan Zhuang et al.
Proposes URE for unbiased Recall@K estimation in recommendation evaluation, correcting bias from randomly exposed data.
Chengbing Wang, Wentao Shi, Jizhi Zhang et al.
VILA-U employs a unified autoregressive framework integrating visual and language understanding and generation, achieving near state-of-the-art results.
Yecheng Wu, Zhuoyang Zhang, Junyu Chen et al.