dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
dVLA, a diffusion-based multimodal model with chain-of-thought, achieves 96.4% success in robot tasks.
Junjie Wen, Minjie Zhu, Jiaming Liu et al.
dVLA, a diffusion-based multimodal model with chain-of-thought, achieves 96.4% success in robot tasks.
Junjie Wen, Minjie Zhu, Jiaming Liu et al.
Introduces QFrBLiMP, a human-annotated Quebec-French minimal pairs benchmark, evaluating LLMs' grammatical competence with detailed analysis.
David Beauchemin, Pier-Luc Veilleux, Johanna-Pascale Roy et al.
Introduces Hybrid Reward Normalization (PPR) achieving state-of-the-art performance across benchmarks.
Peiran Xu, Zhuohao Li, Xiaoying Xing et al.
Forensic-Chat framework enhances fake image detection by seeing before reasoning, improving generalization and explainability.
Kaiqing Lin, Zhiyuan Yan, Ruoxin Chen et al.
The study introduces a 'Perception to Cognition' framework to analyze bottlenecks in MLLMs for vision-language interaction.
Chenyue Zhou, Mingxuan Wang, Yanbiao Ma et al.
Proposes SARM: a video-based stage-aware reward model that improves long-horizon robot manipulation, achieving 83% success in T-shirt folding.
Qianzhong Chen, Justin Yu, Mac Schwager et al.
YOLO26 achieves end-to-end NMS-free detection, removing DFL, with ProgLoss, STAL, and MuSGD, boosting speed and accuracy for edge deployment.
Ranjan Sapkota, Rahul Harsha Cheppally, Ajay Sharda et al.
Cogito, ergo ludo (CEL) learns games through reasoning and planning, mastering diverse grid-world tasks.
Sai Wang, Yu Wu, Zhongwen Xu
Proposes GCM, leveraging multi-model history to improve calibration and generalization, surpassing self-prediction.
Hanqi Xiao, Vaidehi Patil, Hyunji Lee et al.
ROVER evaluates fixed uniform policy Q-values with softmax sampling, outperforming complex RL methods in LLM reasoning benchmarks.
Haoran He, Yuxiao Ye, Qingpeng Cai et al.
Proposed a Spectral-Grassmann Wasserstein metric, significantly improving efficiency in comparing operator representations of dynamical systems.
Thibaut Germain, Rémi Flamary, Vladimir R. Kostic et al.
DelRec employs surrogate gradient with differentiable interpolation to train recurrent synaptic delays, achieving state-of-the-art accuracy on temporal datasets.
Alexandre Queant, Ulysse Rançon, Benoit R Cottereau et al.
T-POP algorithm achieves real-time personalization using online preference feedback, significantly enhancing LLM performance.
Zikun Qu, Min Zhang, Mingze Kong et al.
InfLLM-V2 introduces a dense-sparse switchable attention mechanism, achieving 4× speedup with 98%+ performance on long-sequence tasks.
Weilin Zhao, Zihan Zhou, Zhou Su et al.
SPLICE benchmark reveals significant gaps in VLM visual reasoning compared to human performance.
Mohamad Ballout, Okajevo Wilfred, Seyedalireza Yaghoubi et al.
Introduced NeMo task for long video understanding; automated data pipeline; built NeMoBench with 31,378 QA pairs from 13,486 videos.
Zi-Yuan Hu, Shuo Liang, Duo Zheng et al.
Samiksha method evaluates LLMs in India's healthcare using community feedback, enhancing multilingual model performance.
Hamna Hamna, Gayatri Bhat, Sourabrata Mukherjee et al.
Latent Visual Reasoning (LVR) enables end-to-end reasoning in visual embedding space, boosting perception-intensive visual question answering by 5% (71.67% vs 66.67%).
Bangzheng Li, Ximeng Sun, Jiang Liu et al.
Introduced Grounding IDs to improve multimodal binding via external cues, enhancing vision-language model performance.
Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari et al.
Introduces EvolveCast framework to evaluate LLMs' forecast updates; finds models are overly conservative and inconsistent in belief revision.
Zhangdie Yuan, Zifeng Ding, Andreas Vlachos