Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
First efficiency-focused VLA survey: four-way taxonomy, from 55B to 0.24B models.
Weifan Guan, Qinghao Hu, Aosheng Li et al.
First efficiency-focused VLA survey: four-way taxonomy, from 55B to 0.24B models.
Weifan Guan, Qinghao Hu, Aosheng Li et al.
Comparative analysis of GPT-4o and Gemini 2.5 Flash in usability inspection shows high AI performance but still benefits from human-AI collaboration.
Luis F. G. Campos, Leonardo C. Marques, Walter T. Nakamura
Proposes a system-level, multi-modal, multi-agent framework for interpretable and trustworthy time series reasoning, emphasizing structured multi-step inference.
Kanghui Ning, Zijie Pan, Yushan Jiang et al.
SAKE introduces the first benchmark for editing auditory attribute knowledge in LALMs, revealing reliability and generalization challenges in current methods.
Chih-Kai Yang, Yen-Ting Piao, Tzu-Wen Hsu et al.
Study finds many-shot prompting degrades functional correctness in code translation; optimal examples are 5-25.
Amirkia Rafiei Oskooei, Kaan Baturalp Cosdan, Husamettin Isiktas et al.
SkipV1Former introduces skip connections from first-layer Value heads, reducing KV cache by 25% while improving perplexity.
Zhoutong Wu, Yuan Zhang, Yiming Dong et al.
Proposes LAC layout for generative recommendation, balancing signal maximization, causal fidelity, and efficiency, reducing FLOPs by 40%.
Xiaokai Wei, Jiajun Wu, Daiyao Yi et al.
Using LLMs to generate optimization algorithms, achieving an average 72.4% performance improvement.
Floris-Jan Willemsen, Niki van Stein, Ben van Werkhoven
Reinforcement learning enhances adaptive search strategies, improving accuracy by 15% on HotpotQA with multi-step decision models.
Minhua Lin, Zongyu Wu, Zhichao Xu et al.
A knapsack-inspired online framework dynamically tests and optimizes agent component selection, boosting success rates and reducing costs.
Michelle Yuan, Khushbu Pahwa, Shuaichen Chang et al.
Proposed NP-ENGINE framework integrates instance generation, rule verification, and heuristics, boosting LLMs' NP-hard optimization reasoning.
Xiaozhe Li, Xinyu Fang, Shengyuan Ding et al.
Introduces Feedback-Based Conformal Prediction (Fb-CP) for trajectory optimization, leveraging realized trajectories to adapt risk regions online, ensuring safety and improving performance.
Han Wang, Chao Ning
Proposes a relation-centric procedural language and a symbol-based error correction method for open-universe scene layout generation, outperforming declarative approaches.
Maxim Gumin, Do Heon Han, Seung Jean Yoo et al.
VISTA employs multi-agent self-improvement, achieving up to 60% win rate in video quality enhancement.
Do Xuan Long, Xingchen Wan, Hootan Nakhost et al.
Chronos-2 is a pretrained universal forecasting model employing group attention, supporting zero-shot predictions for univariate, multivariate, and covariate-informed tasks, achieving SOTA results.
Abdul Fatir Ansari, Oleksandr Shchur, Jaris Küken et al.
Proposes Critic-LLM-RS combining collaborative filtering with pre-trained LLMs for improved recommendation without fine-tuning.
Zhisheng Yang, Xiaofei Xu, Ke Deng et al.
XModBench evaluates Gemini 2.5 Pro's cross-modal consistency, revealing less than 60% accuracy in spatial and temporal reasoning.
Xingrui Wang, Jiang Liu, Chao Huang et al.
DLER employs batch normalization and dynamic sampling to optimize RL, reducing output length by over 70% while surpassing baseline accuracy.
Shih-Yang Liu, Xin Dong, Ximing Lu et al.
SPA framework uses self-play finetuning to internalize world models, boosting RL performance in OOD environments by 2x on Sokoban.
Shiqi Chen, Tongyao Zhu, Zian Wang et al.
Proposes NEO, a native vision-language model trained on 390M image-text pairs, achieving pixel-word alignment and surpassing modular models in various benchmarks.
Haiwen Diao, Mingxuan Li, Silei Wu et al.