A Survey on Efficient Vision-Language-Action Models
Proposes resource-efficient VLA design combining model pruning, sparse attention, and training strategies, reducing inference latency to under 50ms.
Zhaoshu Yu, Bo Wang, Pengpeng Zeng et al.
Proposes resource-efficient VLA design combining model pruning, sparse attention, and training strategies, reducing inference latency to under 50ms.
Zhaoshu Yu, Bo Wang, Pengpeng Zeng et al.
Video-Thinker employs reinforcement learning to enable autonomous grounding and captioning for video reasoning, outperforming baselines with 80.69% accuracy on Video-Holmes.
Shijian Wang, Jiarui Jin, Xingjian Wang et al.
Proposes max@k gradient estimation to optimize Best-of-N sampling, improving code generation performance.
Farid Bagirov, Mikhail Arkhipov, Ksenia Sycheva et al.
Proposes a deep active inference framework combining diffusion policy and multi-timescale world model, improving robotic exploration and navigation.
Riko Yokozawa, Kentaro Fujii, Yuta Nomura et al.
Introduces TIR-Judge, an RL framework integrating tools to improve LLM judging, surpassing multiple benchmarks with only 8B parameters.
Ran Xu, Jingjing Chen, Jiayu Ye et al.
EchoMind introduces a multi-level benchmark to evaluate SLMs in speech understanding, vocal cue perception, reasoning, and empathetic dialogue generation.
Li Zhou, Lutong Yu, You Lyu et al.
E2Rank unifies retrieval and listwise reranking using continued training on a single embedding model, achieving state-of-the-art results with high efficiency.
Qi Liu, Yanzhao Zhang, Mingxin Li et al.
Layer pruning impacts LLM reasoning, especially long-chain reasoning.
Keyu Wang, Tian Lyu, Guinan Su et al.
Proposes Huxley-Gödel Machine using CMP estimation to achieve human-level coding agent self-improvement.
Wenyi Wang, Piotr Piękos, Li Nanbo et al.
FlexIO employs prompt vectors and array-agnostic channel communication to enable flexible multi-microphone and multi-speaker speech separation, achieving SDR up to 9.7dB.
Yoshiki Masuyama, Kohei Saijo, Francesco Paissan et al.
SSoT prompts LLMs with random strings and operations, achieving near-ideal distribution fidelity and enhanced diversity.
Kou Misaki, Takuya Akiba
SIREN, MFN-Gabor, and MHE compressed thoracic-aorta fields up to 230×, with ~1 mmHg pressure and sub-1.6 mm anatomical errors.
Jubilee Lee, Daniele E. Schiavazzi
ColMAD reframes multi-agent debate as a cooperative game, boosting error detection accuracy by 10% through truthful, informative messaging.
Yongqiang Chen, Gang Niu, James Cheng et al.
HoloCine employs Window Cross-Attention and Sparse Inter-Shot Self-Attention to generate coherent multi-shot long videos with high narrative consistency.
Yihao Meng, Hao Ouyang, Yue Yu et al.
Open-o3-Video integrates explicit spatio-temporal evidence, achieving 14.4% mAM improvement on V-STAR, with high-quality dataset STGR and two-stage training.
Jiahao Meng, Xiangtai Li, Haochen Wang et al.
UI-Ins enhances GUI grounding with multi-perspective reasoning, achieving 87.3% on UI-I2E-Bench.
Liangyu Chen, Hanzhang Zhou, Chenglin Cai et al.
Proposes GEMs with diverse architectures and high-quality corpora, achieving up to 3.6% accuracy improvement in Greek NLP tasks.
Alexandra Apostolopoulou, Konstantinos Kanaris, Athanasios Koursaris et al.
Proposes WorldTest protocol to evaluate AI models' generality via environment-level queries; humans outperform AI in experiments.
Archana Warrier, Dat Nguyen, Michelangelo Naim et al.
SEMPO is a lightweight foundation model for time series forecasting, combining energy-aware spectral decomposition and prompt-based Transformer.
Hui He, Kun Yi, Yuanchi Ma et al.
Cosine-warming Temperature Sampling raises Libero UniVLA success from 0.76 to 0.85 under severe task imbalance.
Basavasagar Patil, Sydney Belt, Jayjun Lee et al.