GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2603.13215

Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models

STEVO-Bench evaluates video world models' ability to evolve state during observation interruptions, revealing limitations.

Ziqi Ma, Mengzhan Liufu, Georgia Gkioxari

2026-03-14 5 citations 436
cs.CV 2603.13082

InterEdit: Navigating Text-Guided Multi-Human 3D Motion Editing

InterEdit uses Semantic-Aware Plan Token Alignment and Interaction-Aware Frequency Token Alignment for multi-human 3D motion editing.

Yebin Yang, Di Wen, Lei Qi et al.

2026-03-13 293
cs.CV 2603.12829

coDrawAgents: A Multi-Agent Dialogue Framework for Compositional Image Generation

coDrawAgents framework improves compositional text-to-image generation with 94% overall accuracy on GenEval via multi-agent collaboration.

Chunhan Li, Qifeng Wu, Jia-Hui Pan et al.

2026-03-13 31
cs.CV 2603.12811

OARS: Process-Aware Online Alignment for Generative Real-World Image Super-Resolution

OARS framework uses COMPASS reward for real-time image super-resolution, enhancing perceptual quality and fidelity.

Shijie Zhao, Xuanyu Zhang, Bin Chen et al.

2026-03-13 1
cs.CV 2603.12655

VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model

VGGT-World predicts scene geometry evolution via autoregressive modeling of frozen GFM features, achieving 21% better depth accuracy with 0.43B params, 3.6-5× faster.

Xiangyu Sun, Shijie Wang, Fengyi Zhang et al.

2026-03-13 28
cs.CV 2603.12354

Alternating Gradient Flow Utility: A Unified Metric for Structural Pruning and Dynamic Routing in Deep Networks

Introduces Alternating Gradient Flow (AGF) to prevent structural collapse under 75% compression on ImageNet-1K.

Tianhao Qian, Zhuoxuan Li, Jinde Cao et al.

2026-03-13 173
cs.CV 2603.12310

VQQA: An Agentic Approach for Video Evaluation and Quality Improvement

VQQA employs multi-agent QA with semantic gradients to improve video quality by +11.57%.

Yiwen Song, Tomas Pfister, Yale Song

2026-03-13 46
cs.CV 2603.12267

EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation

EVATok achieves efficient visual autoregressive generation with adaptive video tokenization, saving 24.4% tokens on average.

Tianwei Xiong, Jun Hao Liew, Zilong Huang et al.

2026-03-13 2 citations 234
cs.CV 2603.12266

MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning

MM-CondChain uses VPIR for visually grounded deep compositional reasoning, with top model achieving only 53.33 Path F1.

Haozhan Shen, Shilin Yan, Hongwei Xue et al.

2026-03-13 226
cs.CV 2603.12265

OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams

OmniStream achieves perception, reconstruction, and action in visual streams using causal spatiotemporal attention and 3D-RoPE, excelling across 29 datasets.

Yibin Yan, Jilan Xu, Shangzhe Di et al.

2026-03-13 7 citations 429
cs.CV 2603.12257

DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning

DreamVideo-Omni achieves multi-subject video customization with latent identity reinforcement learning, enhancing identity fidelity and motion control precision.

Yujie Wei, Xinyu Liu, Shiwei Zhang et al.

2026-03-13 193
cs.CV 2603.12254

Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing

AutoGaze autoregressively selects multi-scale video patches, reducing redundancy and enhancing efficiency, enabling 1K-frame 4K video processing.

Baifeng Shi, Stephanie Fu, Long Lian et al.

2026-03-13 6 citations 326
cs.CV 2603.12252

EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models

EndoCoT activates MLLMs' reasoning potential, achieving 92.1% accuracy, 8.3% higher than the baseline.

Xuanlang Dai, Yujie Zhou, Long Xing et al.

2026-03-13 1 citations 240
cs.CV 2603.12240

BiGain: Unified Token Compression for Joint Generation and Classification

BiGain enhances diffusion models by frequency separation, improving classification accuracy by 7.15% and FID by 0.34.

Jiacheng Liu, Shengkun Tang, Jiacheng Cui et al.

2026-03-13 234
cs.CV 2603.12215

RDNet: Region Proportion-Aware Dynamic Adaptive Salient Object Detection Network in Optical Remote Sensing Images

RDNet enhances salient object detection in optical remote sensing images using dynamic adaptive modules.

Bin Wan, Runmin Cong, Xiaofei Zhou et al.

2026-03-13 4 citations 180
cs.CV 2603.12146

FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance

FlashMotion introduces a three-stage training framework combining diffusion and adversarial objectives, achieving 47× faster controllable video generation with high quality.

Quanhao Li, Zhen Xing, Rui Wang et al.

2026-03-13 33
cs.CV 2603.12144

O3N: Omnidirectional Open-Vocabulary Occupancy Prediction

O3N framework achieves state-of-the-art performance on QuadOcc and Human360Occ benchmarks using polar-spiral topology for 360° spatial representation.

Mengfei Duan, Hao Shi, Fei Teng et al.

2026-03-13 272
cs.CV 2603.12138

HATS: Hardness-Aware Trajectory Synthesis for GUI Agents

HATS employs hardness-aware trajectory synthesis, enhancing GUI agents' generalization in ambiguous interactions.

Rui Shao, Ruize Gao, Bin Xie et al.

2026-03-13 32
cs.CV 2603.12083

Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysis

Introduced UniCAC benchmark to evaluate 24 algorithms under various optical aberrations.

Xiaolong Qian, Qi Jiang, Yao Gao et al.

2026-03-12 238
cs.CV 2603.12078

Node-RF: Learning Generalized Continuous Space-Time Scene Dynamics with Neural ODE-based NeRFs

Node-RF integrates Neural ODE with NeRF for continuous-time scene dynamics, achieving superior long-range extrapolation and generalization.

Hiran Sarkar, Liming Kuang, Yordanka Velikova et al.

2026-03-12 38
Prev 1 ... 38 39 40 41 42 43 44 ... 138 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home