GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2507.01467

Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think

REG method entangles high-level class tokens with low-level latents, boosting diffusion training 63× faster with improved quality.

Ge Wu, Shen Zhang, Ruijing Shi et al.

2025-07-02 26
cs.CV 2507.00992

UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis

UniGlyph uses pixel-level text masks and adaptive guidance to improve high-fidelity visual text synthesis in Chinese and English, surpassing prior methods by 15%.

Yuanrui Wang, Cong Han, Yafei Li et al.

2025-07-02 26
cs.CV 2507.00676

A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation

Proposes a Transformer-based framework for whole-body grasping motion generation, excelling on the GRAB dataset.

Edward Effendy, Kuan-Wei Tseng, Rei Kawakami

2025-07-01 14
cs.CV 2507.00603

World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

World4Drive employs an intention-aware latent world model, reducing L2 error by 18.1%, collision rate by 46.7%, without perception annotations, achieving state-of-the-art results.

Yupeng Zheng, Pengxuan Yang, Zebin Xing et al.

2025-07-01 40
cs.CV 2507.02978

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models

Proposed Inf-Bench framework evaluates spatial deformation reasoning in VLMs, revealing poor 3D task performance.

Jiahuan Zhang, Shunwen Bai, Tianheng Wang et al.

2025-07-01 7
cs.CV 2507.00372

Efficient Depth- and Spatially-Varying Image Simulation for Defocus Deblur

Proposed an efficient depth- and spatially-varying image simulation method, enhancing 12MP real-world deblur performance.

Xinge Yang, Chuong Nguyen, Wenbin Wang et al.

2025-07-01 31
cs.CV 2507.00339

Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video

MOVi-MC-AC supplies 5.8M instances across six-camera videos for amodal segmentation, content completion, and cross-view re-identification.

Alexander Moore, Amar Saini, Kylie Cancilla et al.

2025-07-01 22
cs.CV 2506.23468

NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments

NavMorph employs a self-evolving world model with continuous latent space and scene memory, boosting VLN-CE performance and online adaptation.

Xuan Yao, Junyu Gao, Changsheng Xu

2025-06-30 28
cs.CV 2506.23361

OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions

OmniVCus employs a diffusion Transformer with multimodal control, enabling multi-subject video customization without labels, using LE and TAE mechanisms.

Yuanhao Cai, He Zhang, Xi Chen et al.

2025-06-30 38
cs.CV 2506.23236

VolumetricSMPL: A Neural Volumetric Body Model for Efficient Interactions, Contacts, and Collisions

VolumetricSMPL achieves 10x faster inference using Neural Blend Weights.

Marko Mihajlovic, Siwei Zhang, Gen Li et al.

2025-06-29 4
cs.CV 2506.21891

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025

DIVE employs iterative reasoning with semantic decomposition, achieving 81.44% accuracy on CVRR-ES, winning the challenge.

Umihiro Kamoto, Tatsuya Ishibashi, Noriyuki Kugo

2025-06-27 41
cs.CV 2506.21249

Temporal Rate Reduction Clustering for Human Motion Segmentation

TR²C combines MCR² and temporal continuity, reaching 97.96% ACC and 98.96% NMI across five HMS benchmarks.

Xianghan Meng, Zhengyu Tong, Zhiyuan Huang et al.

2025-06-26 18
cs.CV 2506.20967

DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing

DFVEdit uses Conditional Delta Flow Vectors for zero-shot video editing, achieving 20× speed-up and 85% memory reduction, maintaining state-of-the-art quality.

Lingling Cai, Kang Zhao, Hangjie Yuan et al.

2025-06-26 32
cs.CV 2507.00049

AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training

AdaDeDup combines density clustering and model feedback for adaptive data pruning, reducing 20% data with minimal performance loss.

Feiyang Kang, Nadine Chang, Maying Shen et al.

2025-06-25 39
cs.CV 2506.18903

VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory

Proposes Surfel-Indexed View Memory (VMem) for long-term scene consistency, reducing computational cost while maintaining scene coherence.

Runjia Li, Philip Torr, Andrea Vedaldi et al.

2025-06-24 35
cs.CV 2506.16802

Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation

This study introduces wavelet-based frequency augmentation to improve AI-generated video detection, boosting cross-model accuracy by over 12%.

Riccardo Corvi, Davide Cozzolino, Ekta Prashnani et al.

2025-06-20 21 citations 44
cs.CV 2506.16058

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation

Proposes OpenBench for evaluating open-vocabulary segmentation, with OVSNet achieving SOTA, outperforming existing methods by ~10%.

Yong Liu, SongLi Wu, Sule Bai et al.

2025-06-19 41
cs.CV 2506.15871

Visual symbolic mechanisms: Emergent symbol processing in vision language models

This paper uncovers emergent symbolic mechanisms in VLMs using spatial indexing (position IDs) to solve the binding problem, validated via causal mediation.

Rim Assouel, Declan Campbell, Yoshua Bengio et al.

2025-06-19 35
cs.CV 2506.15564

Show-o2: Improved Native Unified Multimodal Models

Show-o2 model enhances multimodal understanding and generation via autoregressive modeling and flow matching.

Jinheng Xie, Zhenheng Yang, Mike Zheng Shou

2025-06-18 5
cs.CV 2506.15524

NTIRE 2025 Image Shadow Removal Challenge Report

Proposed a multi-scale Transformer with frequency fusion, achieving PSNR 25.90 on WSRD+ dataset.

Florin-Alexandru Vasluianu, Tim Seizinger, Zhuyun Zhou et al.

2025-06-18 41
Prev 1 ... 59 60 61 62 63 64 65 ... 138 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home