$π_0$: A Vision-Language-Action Flow Model for General Robot Control
Pre-trained vision-language flow model enables zero-shot multi-robot control with high-frequency continuous actions.
Kevin Black, Noah Brown, Danny Driess et al.
Pre-trained vision-language flow model enables zero-shot multi-robot control with high-frequency continuous actions.
Kevin Black, Noah Brown, Danny Driess et al.
Sparsh uses self-supervised learning on 460k+ tactile images, improving TacBench performance by 95.1%.
Carolina Higuera, Akash Sharma, Chaithanya Krishna Bodduluri et al.
IC-LoRA activates DiT in-context generation with 20–100 image sets, producing high-fidelity multi-image outputs without architectural changes.
Lianghua Huang, Wei Wang, Zhi-Fan Wu et al.
SuctionPrompt integrates VLMs and 3D detection for zero-shot robotic picking, achieving 65% success in diverse environments.
Tomohiro Motoda, Takahide Kitamura, Ryo Hanai et al.
Proposes MM-Det leveraging LMM-generated multi-modal forgery representations, achieving 92.0% AUC on DVF for diffusion video detection.
Xiufeng Song, Xiao Guo, Jiache Zhang et al.
Survey of cultural awareness data collection and benchmarking methods in multimodal LLMs, emphasizing cross-cultural datasets and evaluation metrics.
Siddhesh Pawar, Junyeong Park, Jiho Jin et al.
LCoDeepNEAT combines Lamarckian genetic algorithms and gradient transfer to optimize CNN architectures and last-layer weights, boosting accuracy and efficiency.
Zaniar Sharifi, Khabat Soltanian, Ali Amiri
Senna integrates LVLM with end-to-end models, reducing planning error by 27.12% and collision rate by 33.33%.
Bo Jiang, Shaoyu Chen, Bencheng Liao et al.
Proposes an adaptive environment-shaping framework using secondary RL (SAC) to enable drone racing in unseen tracks with 100% success rate.
Hongze Wang, Jiaxu Xing, Nico Messikommer et al.
DynaMath evaluates the robustness of vision-language models in mathematical reasoning through dynamic question generation, revealing instability in handling variants.
Chengke Zou, Xingang Guo, Rui Yang et al.
Proposes a sighted guide-based VR framework with Shared Movement and Flying, improving navigation and social interaction for BLV users.
Jazmin Collins, Crescentia Jung, Yeonju Jang et al.
Safety cases for frontier AI propose structured argumentation with evidence support.
Marie Davidsen Buhl, Gaurav Sett, Leonie Koessler et al.
Fed-POE combines local and federated models with dynamic selection for adaptive prediction and fine-tuning, achieving significant performance gains.
Pouya M. Ghari, Yanning Shen
Graph spectral bandit approach with graph Fourier coefficients enhances real-time voltage monitoring under bandwidth constraints.
Samuel Talkington, Rahul Gupta, Richard Asiamah et al.
LARP employs holistic queries and AR prior to achieve state-of-the-art FVD 57 in video generation.
Hanyu Wang, Saksham Suri, Yixuan Ren et al.
Energy-based diffusion language model (EDLM) leverages residual energy models to improve text generation, outperforming existing diffusion models and approaching autoregressive perplexity.
Minkai Xu, Tomas Geffner, Karsten Kreis et al.
Constructed M2RC-EVAL, a multilingual repository-level code completion benchmark with AST-based fine-grained annotations for 18 languages.
Jiaheng Liu, Ken Deng, Congnan Liu et al.
CycleResearcher employs reinforcement learning to automate research from literature review to peer review, reducing prediction MAE by 26.89%.
Yixuan Weng, Minjun Zhu, Guangsheng Bao et al.
ProtoViT combines Vision Transformers for interpretable image classification, outperforming existing prototype models.
Chiyu Ma, Jon Donnelly, Wenjun Liu et al.
Proposed Relaxed Recursive Transformers with layer-wise LoRA for effective parameter sharing, achieving near-full model performance.
Sangmin Bae, Adam Fisch, Hrayr Harutyunyan et al.