Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
DyLaR enhances video QA accuracy to 58.2% with under 20 tokens per query using dynamic latent reasoning.
Haotian Xia, Zilin Xiao, Junbo Zou et al.
DyLaR enhances video QA accuracy to 58.2% with under 20 tokens per query using dynamic latent reasoning.
Haotian Xia, Zilin Xiao, Junbo Zou et al.
AgenticVAU employs multi-agent explore-verify reasoning, surpassing zero-shot and RL baselines in video anomaly understanding with significant accuracy gains.
Yuxiang Duan, Huining Li, Ao Li et al.
SEER exposes query-specific evidence for frozen VLMs, improving frozen GQA-Train900 accuracy by 3.94 points over Full.
Feixiang Liu, Likun Wang, Qiang Qiu et al.
Proposes SGFormer with Triple-Structure-Attention, significantly reducing attention divergence and boosting local feature matching accuracy by 2% on MegaDepth-1500.
Runyu Zhu
EditFlow3D automates local 3D asset editing with trajectory preservation.
Rui Nie, Chuang Wang, Haitao Zhou et al.
Unified fusion framework based on Rank-enhanced linear attention with FiLM conditioning achieves 34-40dB PSNR across multiple optical slit configurations.
Yiming Gong, Kai Wang
V-FIND localizes sparse neurons to improve video forgery detection, achieving state-of-the-art results with frozen backbones and linear classifiers.
Shichao Kan, Chengpeng Hong, Jingtong Dou et al.
Identifies in-context collapse in vision-language models, localizes it to the fusion interface, and introduces CircA adapter for cross-task robustness, boosting 16-shot accuracy from 0.39 to 0.91.
Mohammad Rostami
ChunkVAE enables efficient 3D modeling via local chunk compression and stitching, scaling from 512³ to 1536³ resolution.
Kaiyi Zhang, Zhihao Liang, Haolin Liu et al.
AdaThinkV uses adaptive reinforcement learning to optimize token usage, achieving 40.79% accuracy with 22.7% fewer tokens in video reasoning.
Jingqi Tian, Haoji Zhang, Lin Chen et al.
SpatioLM enhances vision-language models' spatial reasoning using a plug-and-play module, achieving 71.6 on VSI-Bench without extra 3D inputs.
Jing Wu, Jianhua Wu, Jiayi Guan et al.
PartMat employs a single global latent to achieve efficient, material-aware 3D part decomposition, surpassing existing methods in accuracy and scalability.
Guangming Fu, Jin Song, Yiyun Fei et al.
LiveLight enables real-time video relighting with interactive 3D lighting control, significantly enhancing user experience.
Yue Ma, Jiangming Wang, Yucheng Wang et al.
G-Skin leverages 2D generative priors to learn skeleton binding for 3D Gaussian representations, addressing data scarcity with high-fidelity animation.
Yuxin Yao, Kendong Liu, Shiqi Zhou et al.
MeanFlow achieves unified RAW restoration under extreme low-light and motion blur, improving PSNR by up to 7.42 dB.
Zepu Wang, Jingze Liang, Weijie Xiao et al.
Proposes LIA-MTR, a linear O(N) cross-modal bridge with multi-timescale retention, enabling infinite-context vision-language processing.
Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam
Proposed D²-4DGS fuses dual-depth priors for sparse-camera 4D Gaussian scene synthesis, improving PSNR by 1.33dB.
Jijian Zhao
STAR-VLM uses nuScenes radar supervision to train VLMs, reaching 0.94 motion accuracy and 1.20 m/s radial-velocity MAE.
Pou-Chun Kung, Aryaman Rao, Utkrisht Sahai et al.
Pre-trained detection transformers (e.g., DETR) encode depth and 3D position info in object embeddings without explicit 3D supervision, as shown by probing experiments.
Robin Kim, Colin Samplawski, Benjamin M. Marlin
VaRS-Doc enhances visual document retrieval by diversifying document representations via latent self-probing, achieving state-of-the-art performance.
Haocheng Wang, Tongkun Guan, Wei Shen et al.