Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing
FastVim halves Mamba scan depth through alternating spatial pooling, achieving up to 72.5% faster inference on 2048² images.
Saarthak Kapse, Robin Betz, Srinivasan Sivanandan
FastVim halves Mamba scan depth through alternating spatial pooling, achieving up to 72.5% faster inference on 2048² images.
Saarthak Kapse, Robin Betz, Srinivasan Sivanandan
EgoMe dataset with 7902 pairs, multimodal data, enhances robot imitation via cross-view alignment.
Heqian Qiu, Zhaofeng Shi, Lanxiao Wang et al.
OmniPhysGS employs learnable constitutive Gaussians for multi-material 3D dynamic scene synthesis, achieving 3-16% better visual and text alignment metrics.
Yuchen Lin, Chenguo Lin, Jianjin Xu et al.
This work introduces human-preference-guided reinforcement learning for video generation, leveraging VideoReward to improve flow-based models, achieving significant quality gains.
Jie Liu, Gongye Liu, Jiajun Liang et al.
VideoLLaMA3 employs a vision-centric multi-stage training framework, significantly improving image and video understanding performance.
Boqiang Zhang, Kehan Li, Zesen Cheng et al.
Using DiverseAR dataset, VLMs like GPT achieve 93% TPR in AR scene recognition.
Lin Duan, Yanming Xiu, Maria Gorlatova
InternLM-XComposer2.5-Reward enhances LVLMs' generation quality with a multi-modal reward model, achieving 70% accuracy.
Yuhang Zang, Xiaoyi Dong, Pan Zhang et al.
VipDiff uses optical flow and training-free diffusion models for video inpainting, enhancing spatio-temporal coherence.
Chaohao Xie, Kai Han, Kwan-Yee K. Wong
Hunyuan3D 2.0 employs flow-based diffusion transformers for high-res textured 3D asset creation, integrating ShapeVAE and multi-view texture synthesis.
Zibo Zhao, Zeqiang Lai, Qingxiang Lin et al.
Proposes One-D-Piece, a variable-length image tokenizer with Tail Token Drop, achieving high-quality, flexible compression surpassing JPEG/WebP.
Keita Miwa, Kento Sasaki, Hidehisa Arai et al.
WMamba combines wavelet analysis with dynamic contour convolution, achieving SOTA face forgery detection with high robustness and efficiency.
Siran Peng, Tianshuo Zhang, Li Gao et al.
VRS-HQ employs hierarchical multimodal tokens with dynamic temporal aggregation, achieving 63.3% J&F on ReVOS, surpassing VISA.
Sitong Gong, Yunzhi Zhuge, Lu Zhang et al.
Using conformal prediction to ensure at least 90% coverage in deep learning-based weed detection for precision spraying.
Paul Melki, Lionel Bombrun, Boubacar Diallo et al.
VideoRAG leverages LVLM for dynamic video retrieval and multimodal response generation, outperforming baselines with significant improvements in ROUGE-L and BLEU scores.
Soyeong Jeong, Kangsan Kim, Jinheon Baek et al.
OVO-Bench evaluates Video-LLMs' temporal awareness, highlighting gaps with human understanding.
Yifei Li, Junbo Niu, Ziyang Miao et al.
Proposes a Token-level shuffling and mixing framework for unbiased deepfake detection, significantly improving cross-dataset generalization with state-of-the-art AUC scores.
Xinghe Fu, Zhiyuan Yan, Taiping Yao et al.
Sa2VA unifies SAM-2 and MLLM for dense, multi-task image/video understanding, achieving over 15% improvement in complex scene segmentation.
Haobo Yuan, Xiangtai Li, Tao Zhang et al.
STAR method enhances video super-resolution using text-to-video models, improving spatio-temporal consistency.
Rui Xie, Yinhong Liu, Penghao Zhou et al.
SceneVTG++ combines TLCG and CLTD to generate realistic, controllable multilingual scene text, surpassing SOTA in fidelity and utility.
Jiawei Liu, Yuanzhi Zhu, Feiyu Gao et al.
AVTrustBench assesses AVLLM reliability; CAVPref improves performance by 30.19%.
Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta et al.