GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2608.04124

Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering

DyLaR enhances video QA accuracy to 58.2% with under 20 tokens per query using dynamic latent reasoning.

Haotian Xia, Zilin Xiao, Junbo Zou et al.

2026-08-05 28
cs.CV 2608.03779

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding

AgenticVAU employs multi-agent explore-verify reasoning, surpassing zero-shot and RL baselines in video anomaly understanding with significant accuracy gains.

Yuxiang Duan, Huining Li, Ao Li et al.

2026-08-04 39
cs.CV 2608.03631

SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification

SEER exposes query-specific evidence for frozen VLMs, improving frozen GQA-Train900 accuracy by 3.94 points over Full.

Feixiang Liu, Likun Wang, Qiang Qiu et al.

2026-08-04 16
cs.CV 2608.03423

SGFormer: Structure-Guided Transformer for Robust Local Feature Matching

Proposes SGFormer with Triple-Structure-Attention, significantly reducing attention divergence and boosting local feature matching accuracy by 2% on MegaDepth-1500.

Runyu Zhu

2026-08-04 49
cs.CV 2608.03179

EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation

EditFlow3D automates local 3D asset editing with trajectory preservation.

Rui Nie, Chuang Wang, Haitao Zhou et al.

2026-08-04 35
cs.CV 2608.03107

A Unified Resolution-Conditioned Framework for Orthogonal Line-Scanning Image Fusion

Unified fusion framework based on Rank-enhanced linear attention with FiLM conditioning achieves 34-40dB PSNR across multiple optical slit configurations.

Yiming Gong, Kai Wang

2026-08-04 37
cs.CV 2608.03008

V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors

V-FIND localizes sparse neurons to improve video forgery detection, achieving state-of-the-art results with frozen backbones and linear classifiers.

Shichao Kan, Chengpeng Hong, Jingtong Dou et al.

2026-08-04 63
cs.CV 2608.02830

In-Context Collapse in Vision-Language Models and How to Mitigate it?

Identifies in-context collapse in vision-language models, localizes it to the fusion interface, and introduces CircA adapter for cross-task robustness, boosting 16-shot accuracy from 0.39 to 0.91.

Mohammad Rostami

2026-08-04 40
cs.CV 2608.02016

Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling

ChunkVAE enables efficient 3D modeling via local chunk compression and stitching, scaling from 512³ to 1536³ resolution.

Kaiyi Zhang, Zhihao Liang, Haolin Liu et al.

2026-08-03 32
cs.CV 2608.01980

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

AdaThinkV uses adaptive reinforcement learning to optimize token usage, achieving 40.79% accuracy with 22.7% fewer tokens in video reasoning.

Jingqi Tian, Haoji Zhang, Lin Chen et al.

2026-08-03 38
cs.CV 2608.01899

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

SpatioLM enhances vision-language models' spatial reasoning using a plug-and-play module, achieving 71.6 on VSI-Bench without extra 3D inputs.

Jing Wu, Jianhua Wu, Jiayi Guan et al.

2026-08-03 62
cs.CV 2608.01825

PartMat: Material-Aware 3D Part Decomposition with a Single Global Latent

PartMat employs a single global latent to achieve efficient, material-aware 3D part decomposition, surpassing existing methods in accuracy and scalability.

Guangming Fu, Jin Song, Yiyun Fei et al.

2026-08-03 18
cs.CV 2608.01771

LiveLight: Real-time Streaming Video Relighting with Interactive Control

LiveLight enables real-time video relighting with interactive 3D lighting control, significantly enhancing user experience.

Yue Ma, Jiangming Wang, Yucheng Wang et al.

2026-08-03 4
cs.CV 2608.01726

G-Skin: Learning to Bind 3D Gaussians with Generative Visual Priors

G-Skin leverages 2D generative priors to learn skeleton binding for 3D Gaussian representations, addressing data scarcity with high-fidelity animation.

Yuxin Yao, Kendong Liu, Shiqi Zhou et al.

2026-08-03 42
cs.CV 2608.01720

When Extreme Darkness Meets Motion Blur: MeanFlow for Unified RAW Restoration

MeanFlow achieves unified RAW restoration under extreme low-light and motion blur, improving PSNR by up to 7.42 dB.

Zepu Wang, Jingze Liang, Weijie Xiao et al.

2026-08-03 24
cs.CV 2608.01614

Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge

Proposes LIA-MTR, a linear O(N) cross-modal bridge with multi-timescale retention, enabling infinite-context vision-language processing.

Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam

2026-08-03 40
cs.CV 2608.01588

D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting

Proposed D²-4DGS fuses dual-depth priors for sparse-camera 4D Gaussian scene synthesis, improving PSNR by 1.33dB.

Jijian Zhao

2026-08-03 37
cs.CV 2608.01535

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision

STAR-VLM uses nuScenes radar supervision to train VLMs, reaching 0.94 motion accuracy and 1.20 m/s radial-velocity MAE.

Pou-Chun Kung, Aryaman Rao, Utkrisht Sahai et al.

2026-08-03 20
cs.CV 2608.01495

Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers

Pre-trained detection transformers (e.g., DETR) encode depth and 3D position info in object embeddings without explicit 3D supervision, as shown by probing experiments.

Robin Kim, Colin Samplawski, Benjamin M. Marlin

2026-08-03 50
cs.CV 2608.01211

VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval

VaRS-Doc enhances visual document retrieval by diversifying document representations via latent self-probing, achieving state-of-the-art performance.

Haocheng Wang, Tongkun Guan, Wei Shen et al.

2026-08-02 31
Prev 1 ... 8 9 10 11 12 13 14 ... 134 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home