GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2608.28455

ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT

ARC-CT employs anatomy-guided contrastive learning, combining regional features to achieve 0.86 macro AUC on chest CT abnormalities.

Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol et al.

2026-08-28 124
cs.CV 2608.28404

How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion Models

Scaling laws for video diffusion models in driving data show validation loss follows power laws; a 9B-parameter model predicts loss of ~0.0753 with 3.6% error.

Victor Besnier, Anh-Quan Cao, Elias Ramzi et al.

2026-08-28 23
cs.CV 2608.27881

StreamEMS: Streaming Video Understanding with Self-Evolving Memory Scheme for Vision-Language Models

StreamEMS enhances video streaming understanding in vision-language models with a self-evolving memory scheme, excelling on OVO-Bench.

Yuxin Liu, Peiqin Zhuang, Yali Wang

2026-08-28 7
cs.CV 2608.28706

Variable-Granularity Tokenization for High-Resolution Object Detection

VGTok optimizes high-resolution object detection using regional separability, achieving 48.38 AP on VisDrone dataset.

Khayrul Islam

2026-08-28 22
cs.CV 2608.27407

Reconstructing Humans and Objects in Interaction using Large Reconstruction Models

This paper introduces MILO, leveraging Large Reconstruction Models (LRMs) to reconstruct detailed 3D human-object interactions from a single image, outperforming state-of-the-art methods.

Agniv Chatterjee, Georgios Pavlakos

2026-08-28 67
cs.CV 2608.27395

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

LeVJEPA uses collapse-free invariance loss and SIGReg regularization, reducing pretraining compute by up to 20x while maintaining or surpassing state-of-the-art performance.

Lukas Kuhn, Lucas Maes, Giuseppe Serra et al.

2026-08-28 64
cs.CV 2608.27282

TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object Detection

Proposed TADP: task-aware deformable prediction for single-stage 3D detection, achieving 80.91% mAP on KITTI.

Su Wang, Yaochen Li, Min Yang et al.

2026-08-27 75
cs.CV 2608.27226

DINOcular: Self-Supervised Visuospatial Representations

DINOcular integrates depth priors with self-supervised vision transformers, boosting 3D spatial understanding by 15% on Probe3D benchmarks.

Farkhat Almukhamedov, Sami Azirar, Hermann Blum

2026-08-27 13
cs.CV 2608.27206

PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference

PACE achieves 93.8% performance retention while reducing visual tokens by 90%, enabling a 3.1x speedup via pixel compression and dual-attention extraction.

Junjie Liu, Shengyuan Ye, Xu Chen

2026-08-27 27
cs.CV 2608.26495

Video-FLAIR: Not Whether to Reason, But How

Video-FLAIR uses reinforcement learning to select reasoning modes, improving accuracy and reducing computation.

Yogesh Kulkarni, Pooyan Fazli

2026-08-27 3
cs.CV 2608.26105

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

VBVR-Pro introduces 300 procedurally generated tasks, rule-based reward scorers, and multi-modal evaluation, advancing scalable and reliable visual reasoning.

Junxiang Xu, Ruisi Wang, Fanyi Pu et al.

2026-08-27 98
cs.CV 2608.26095

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

VDA framework with VC-OT and VMA effectively mitigates visual dependence drift in unsupervised continual multimodal learning, achieving 62.5% accuracy across six tasks.

Kaichen Li, Zhilin Zhu, Jianhao Huang et al.

2026-08-27 178
cs.CV 2608.26067

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

StreamPI employs instruction-anchored streaming multimodal temporal modeling, leveraging pre-trained LLM length extrapolation for parameter-free multi-frame inference, outperforming pi0.5 in robot tasks.

Zhe Liu, Jinghua Hou, Yuxiang Lu et al.

2026-08-27 97
cs.CV 2608.25479

4DStreamCtrl: Interactive Video Generation with Online 4D Control

4DStreamCtrl unifies 3D point-track representation for real-time video generation and control, enhancing motion precision.

Shiqian Li, Chenguo Lin, Zhiguang Liu et al.

2026-08-26 34
cs.CV 2608.25386

Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

MTAR framework achieves efficient autoregressive image generation on ImageNet with multi-token prediction, reducing FID by 0.95.

Guo Niu, Xiongfei Yao, Teng Wang et al.

2026-08-26 38
cs.CV 2608.25332

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

ProViP enhances VLM inference efficiency via head-aware pruning, retaining 95.9% performance with 1.62x speedup.

Chaofang Ma, Lin Jiang, Carol Jingyi Li et al.

2026-08-26 29
cs.CV 2608.25308

V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

V-Link restores visual cues in Action DiT, improving GR00T N1.6 by 31.2% on LIBERO-Plus and 18.8% on RoboTwin 2.0.

Yehao Lu, Jiarui Yang, Yuning Su et al.

2026-08-26 18
cs.CV 2608.24877

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

Unified framework models smart glasses as closed-loop systems with eight hardware capability axes, connecting perception, state, and action.

Jiangning Zhang, Haojun Chen, Yong Liu

2026-08-26 81
cs.CV 2608.24845

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

LAION-BVD is a 10M-hour open video dataset with 1.3B URLs, supporting large-scale multimodal pretraining with content-aware scene detection.

Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti et al.

2026-08-26 111
cs.CV 2608.24783

MoE-based Feature Adapter for Prompt-free Binary Coronary Artery Segmentation in X-ray Angiography

Proposed MoE-based feature adapter for prompt-free binary coronary artery segmentation, significantly improving cross-dataset generalization.

Lin Xi, Yingliang Ma

2026-08-26 107
Prev 1 2 3 4 5 6 7 8 ... 134 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home