GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2606.22197

Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation

Multi4D employs multi-level competitive allocation for high-fidelity dynamic Gaussian splatting, balancing motion consistency and detail preservation.

Rui Wang, Quentin Lohmeyer, Siyu Tang et al.

2026-06-21 36
cs.CV 2606.21661

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

UnityShots employs dual-slot memory and boundary-aware gating in LTX-2.3 to generate coherent multi-shot audio-video sequences, outperforming baselines on cross-shot consistency.

Jiehui Huang, Yuechen Zhang, Bin Xia et al.

2026-06-20 18
cs.CV 2606.21623

A DVDrive Approach for doScenes Instructed Driving Challenge

DVDrive introduces divided-view perception with Transformer, improving instruction-conditioned trajectory prediction, reducing ADE from 0.372 to 0.358.

Zijian Fu, Xiangyang Chu, Mengshi Qi et al.

2026-06-20 37
cs.CV 2606.21590

Radial Basis Function Networks as Projection Heads in Self-Supervised Learning

Replace MLP with Radial Basis Function Networks to enhance representation quality in self-supervised learning.

Andreas Schliebitz, Heiko Tapken, Martin Atzmueller

2026-06-20 2
cs.CV 2606.21373

FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction

FLM-Occ enhances indoor occupancy prediction efficiency via feed-forward likelihood maximization, achieving 3.7x speedup with only 32 superquadrics.

Guangcheng Chen, Lihuang Fang, Huaqi Tao et al.

2026-06-19 31
cs.CV 2606.20563

JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising

JanusMesh is a fast, zero-shot framework for generating dual-semantic 3D illusions using cross-space denoising, completing in 3-5 minutes with high geometric and semantic fidelity.

Siang-Ling Zhang, Huai-Hsun Cheng, Tsung-Ju Yang et al.

2026-06-19 206
cs.CV 2606.20559

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

UNIEGO employs proxy-mediated hierarchical distillation from nine heterogeneous teachers to unify egocentric video representations, achieving state-of-the-art results.

Wenhao Chi, Arkaprava Sinha, Dominick Reilly et al.

2026-06-19 172
cs.CV 2606.20543

SSD: Spatially Speculative Decoding Accelerates Autoregressive Image Generation

Proposes Spatially Speculative Decoding (SSD), leveraging 2D spatial prediction to accelerate autoregressive image generation by up to 13.3×.

Shilong Xiang, Zirui Zhang, Lijun Yu et al.

2026-06-19 206
cs.CV 2606.20542

CalTennis: Large Multi-View Tennis Video Dataset and Benchmark of Monocular-to-3D Pose Estimation

CalTennis is a large multi-view tennis video dataset with over 11 million frames, used to evaluate monocular-to-3D pose estimation, highlighting challenges in depth and foot contact accuracy.

Ilona Demler, Xinran Xie, Blake Werner et al.

2026-06-19 1 citations 231
cs.CV 2606.20521

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

Using egocentric human videos with a Mixture-of-Transformers model outperforms robot trajectories in embodied pretraining, especially on out-of-distribution tasks.

Juncheng Ma, Jianxin Bi, Yufan Deng et al.

2026-06-19 45
cs.CV 2606.20083

Holo-World: Unified Camera, Object and Weather Control for Video World Model

Holo-World employs a unified control framework with residual subspaces to generate videos from a single image, enabling precise camera, object, and weather manipulation, trained on HoloStateData.

Xiangchen Yin, Wenzhang Sun, Jiahui Yuan et al.

2026-06-18 41
cs.CV 2606.20077

The Hidden Evolution of Disguised Visual Context inside the VLM

The study reveals the hidden evolution of visual tokens in LLMs, comparing in-context and layer-wise injection methods.

Wish Suharitdamrong, Tony Alex, Xiatian Zhu et al.

2026-06-18 1
cs.CV 2606.19927

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs

CARE optimizes reasoning length in video-MLLMs via competence-aware reward shaping, enhancing accuracy and efficiency.

Chengwen Liu, Hao Peng, Jisheng Dang et al.

2026-06-18 27
cs.CV 2606.19776

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding

Occ-VLM achieves state-of-the-art multi-view occupancy prediction using a single 2D encoder for indoor scene understanding.

Jianing Li, Zhou Fang, Yijiang Liu et al.

2026-06-18 32
cs.CV 2606.18846

From Bounding Boxes to Visual Reasoning: An On-Policy Data Annotation Tool for Vision-Language Models

ScreenAnnotator tool enhances annotation efficiency for vision-language models using unified annotation atoms and Bayesian verifier.

Like Zhang, Runliang Niu, Shiqi Wang et al.

2026-06-17 3
cs.CV 2606.18702

UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation

UniTemp enables bidirectional video generation via joint distillation, addressing boundary flickering caused by causal 3D VAE, with strong performance on short and long videos.

Lin Zhang, Sicheng Mo, Zefan Cai et al.

2026-06-17 31
cs.CV 2606.18591

Bridging Creative Intent and Visual Quality: Creator-Driven Recurrent Video Generation with Agentic Feedback Loops

CHIEF is a creator-driven video loop that extends student-made films from 1 to 10 minutes.

Denis Savytski, Aiden Lei, Heding Liu et al.

2026-06-17 42
cs.CV 2606.18441

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs

Introduces CF-GRPO framework to enhance video reasoning performance with Consensus Frame Reward, significantly improving multiple benchmarks.

Chengwen Liu, Zhe Huang, Jisheng Dang et al.

2026-06-17 24
cs.CV 2606.18249

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

UniAR introduces a unified autoregressive framework with a single discrete visual tokenizer, achieving state-of-the-art results in image generation and understanding.

Wujian Peng, Lingchen Meng, Yuxuan Cai et al.

2026-06-17 197
cs.CV 2606.18242

EventDrive: Event Cameras for Vision-Language Driving Intelligence

EventDrive integrates event cameras with vision-language models, significantly improving perception, understanding, prediction, and planning in autonomous driving.

Dongyue Lu, Rong Li, Ao Liang et al.

2026-06-17 195
Prev 1 ... 19 20 21 22 23 24 25 ... 135 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home