GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2607.04884

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

HunyuanOCR-1.5 enhances OCR performance using DFlash acceleration and Agentic Data Flow, achieving 6.37x inference speedup.

Gengluo Li, Xingyu Wan, Shangpin Peng et al.

2026-07-06 22
cs.CV 2607.04882

Unsupervised Detection of Underground Tunnels in Ground-Penetrating Radar Using Depth-Restricted Reconstruction Scoring

Unsupervised detection of underground tunnels using depth-restricted reconstruction scoring, achieving AUC of 0.994.

Muhammad Junaid, Shoab A. Khan, Nisar Ahmed

2026-07-06 1
cs.CV 2607.04661

Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving

FocusGS uses targeted structure completion to reduce Gaussian count by 74%, improving sparse-view 3D reconstruction efficiency and quality.

Guoqing Wang, Pin Tang, Xiangxuan Ren et al.

2026-07-06 39
cs.CV 2607.04637

PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving

PixelPilot decouples 2D planning from 3D lifting, enabling scalable vision-language-driven autonomous driving with state-of-the-art accuracy.

Pin Tang, Guoqing Wang, Xiangxuan Ren et al.

2026-07-06 47
cs.CV 2607.04607

G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement

G2VD integrates counterfactual intervention and causal disentanglement, achieving over 90% accuracy in cross-domain AI video forgery detection.

Meng Du, Hongchang Chen, Ran Li et al.

2026-07-06 45
cs.CV 2607.04461

Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

Flash-BoN combines timestep truncation, layer skipping, and activation proxies to optimize inference, achieving +8% AUC under fixed budgets.

Ruchit Rawal, Reza Shirkavand, Sayak Paul et al.

2026-07-06 29
cs.CV 2607.04423

Transferability Between Understanding and Generation in Unified Multimodal Models

This study investigates cross-task transferability in unified multimodal models, showing fully shared transformer architectures (like Lumina-DiMOO) achieve up to 9% accuracy gains in counting tasks.

Jiwon Kang, Heeji Yoon, Jaewoo Jung et al.

2026-07-06 38
cs.CV 2607.04330

Framework and Multi-modal Dataset for Roadwork Zone Detection and Geo-localization

Proposed RZDG dataset and a tracker-based geo-localization pipeline significantly improve roadwork zone detection and global positioning accuracy.

Zhiran Yan, Yutong Xin, S Shyam Shenoi et al.

2026-07-05 31
cs.CV 2607.04256

AdaptiveSplat:Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction

AdaptiveSplat uses texture-aware pruning and adaptive Gaussian prediction to improve sparse 3D scene reconstruction without fine-tuning.

Badrinath Singhal, Srihari K G, Sreehari Iyer et al.

2026-07-05 41
cs.CV 2607.04101

The Multipath Blind Spot: $K$-Agnostic Robust Calibration for Sparse-Anchor Metric Depth from Frozen Foundations

MRAC method significantly improves depth estimation robustness without adding parameters.

Sohag Roy, Rajesh Misra, Swami Shastravidyananda et al.

2026-07-05 30
cs.CV 2607.03817

Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection

GLLS introduces a training-free dual-stream framework combining global logical verification and active local search, achieving 84% accuracy on MMAD-QA.

Runzhi Deng, Yundi Hu, Yiming Zhong et al.

2026-07-04 34
cs.CV 2607.03765

Sparse-View Surface Reconstruction using Gaussian Splatting through High-Confidence Depth Propagation with Normal Priors

Proposes a Gaussian Splatting-based scene optimization with normal-guided depth propagation, outperforming SOTA in sparse-view 3D surface reconstruction.

Liang Han, Bangcai Wei, Junsheng Zhou et al.

2026-07-04 67
cs.CV 2607.03256

A Decomposable Probe for Few-Step Diffusion Models: Prompt, Latent, and Score Selectivity across Backbone Families and Distillation Paradigms

Decomposable Probe separates prompt, latent, and score responses, detecting rectified flow through latent selectivity across 23 models.

Patrick Mu Haojie

2026-07-03 15
cs.CV 2607.03184

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

BVS combines Bayesian optimization and multimodal large language models for fine-grained perception in ultra-high-resolution images, significantly improving accuracy and efficiency.

Geng Li, Yuxin Peng

2026-07-03 32
cs.CV 2607.03069

SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection

SafeGuard employs multi-agent perception and reasoning, achieving 18.7% accuracy improvement in social-risk video detection on SafeVid dataset.

Wenlin Wu, Sheng Zhou, Peipei Song et al.

2026-07-03 47
cs.CV 2607.02998

CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training

CONFLUX combines 3D latent flow generation with RL control, reaching tri-planar FID 32.3 versus MAISI’s 74.6.

Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge et al.

2026-07-03 17
cs.CV 2607.02988

Cross-device Collaborative Test-time Adaptation with Zeroth-order Optimization and Model Merging

Proposed cross-device collaborative test-time adaptation using zeroth-order optimization and model merging, reducing error rate by 18.3% on CIFAR10-C.

Yu Mitsuzumi, Akisato Kimura, Yasuhiro Fujiwara et al.

2026-07-03 31
cs.CV 2607.02968

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

EMA uses DiT Massive Activations: DG improves details and MREP improves dense features; MA disruption leaves BLIP/CLIP win rates at 0.462/0.512.

Chaofan Gan, Zicheng Zhao, Yuanpeng Tu et al.

2026-07-03 18
cs.CV 2607.02921

R3D: Quantitative 3D Spatial Reasoning for Egocentric Wearables

R3D constructs 3D scenes from egocentric RGB-D video and uses tool calls to achieve 73.5% accuracy in quantitative 3D reasoning tasks.

Maxwell Horton, Wei Lu, Quan Tran et al.

2026-07-03 30
cs.CV 2607.02798

Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

UniCaMo constructs a 3D-grounded noise space for joint control of object and camera motion, improving video coherence and quality.

Long Vu, Tan Ngo, Animesh Karnewar et al.

2026-07-03 33
Prev 1 ... 14 15 16 17 18 19 20 ... 135 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home