GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2503.18783

Frequency Dynamic Convolution for Dense Image Prediction

Proposes Frequency Dynamic Convolution (FDConv) that learns diverse frequency responses with minimal parameter increase, boosting dense image prediction tasks.

Linwei Chen, Lin Gu, Liang Li et al.

2025-03-24 40
cs.CV 2503.17973

PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos

PhysTwin combines spring-mass physics with generative shape models, reconstructing deformable objects from sparse videos with high fidelity.

Hanxiao Jiang, Hao-Yu Hsu, Kaifeng Zhang et al.

2025-03-23 40
cs.CV 2503.17574

Is there anything left? Measuring semantic residuals of objects removed from 3D Gaussian Splatting

Proposes a quantitative evaluation to measure semantic residuals in 3D Gaussian Splatting, validated for privacy-preserving mapping scenarios.

Simona Kocour, Assia Benbihi, Aikaterini Adam et al.

2025-03-22 32
cs.CV 2503.17109

Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval

PrediCIR uses a world model to predict missing target features, boosting zero-shot image retrieval by up to 4.45%.

Yuanmin Tang, Jing Yu, Keke Gai et al.

2025-03-21 43
cs.CV 2503.16930

Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks

VLU-Net uses a vision-language model for unified multi-degradation image restoration, improving 3.74 dB on the SOTS dehazing dataset.

Haijin Zeng, Xiangming Wang, Yongyong Chen et al.

2025-03-21 10
cs.CV 2503.16795

DCEdit: Dual-Level Controlled Image Editing via Precisely Localized Semantics

DCEdit employs dual-level control and precise semantic localization, achieving superior image editing accuracy with no extra training, outperforming existing diffusion-based methods.

Yihan Hu, Jianing Peng, Yiheng Lin et al.

2025-03-21 26
cs.CV 2503.16421

MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance

MagicMotion employs dense-to-sparse trajectory guidance within a diffusion transformer framework, outperforming SOTA with a new dataset and benchmark, achieving precise multi-object control.

Quanhao Li, Zhen Xing, Rui Wang et al.

2025-03-21 62 citations 37
cs.CV 2503.15898

Reconstructing In-the-Wild Open-Vocabulary Human-Object Interactions

Proposes Gaussian-HOI optimizer for 3D human-object interaction reconstruction; introduces Open3DHOI dataset with 133 object categories and 120 interactions.

Boran Wen, Dingbang Huang, Zichen Zhang et al.

2025-03-20 42
cs.CV 2503.15451

MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space

MotionStreamer integrates continuous causal latent space with autoregressive diffusion, enabling real-time streaming human motion synthesis.

Lixing Xiao, Shunlin Lu, Huaijin Pi et al.

2025-03-20 26
cs.CV 2503.15369

EfficientLLaVA:Generalizable Auto-Pruning for Large Vision-language Models

EfficientLLaVA employs structural risk minimization for auto-pruning large vision-language models, achieving 83.05% accuracy with only 64 samples and 1.8× speedup.

Yinan Liang, Ziwei Wang, Xiuwei Xu et al.

2025-03-20 37
cs.CV 2503.15293

Test-Time Backdoor Detection for Object Detection Models

Proposes TRACE, a semantic-aware transformation method, achieving 30% AUROC improvement for test-time backdoor detection in object detection.

Hangtao Zhang, Yichen Wang, Shihui Yan et al.

2025-03-19 62
cs.CV 2503.14853

Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake Detection

Proposes a knowledge-guided LVLM framework achieving 99.53% AUC for deepfake detection, enhancing generalization and explainability.

Peipeng Yu, Jianwei Fei, Hui Gao et al.

2025-03-19 43
cs.CV 2503.14324

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

DualToken employs dual visual vocabularies to unify understanding and generation, achieving 82.0% zero-shot accuracy on ImageNet with 0.25 rFID.

Wei Song, Yuran Wang, Zijia Song et al.

2025-03-18 42 citations 36
cs.CV 2503.14151

Concat-ID: Towards Universal Identity-Preserving Video Synthesis

Concat-ID leverages VAE and 3D self-attention for universal identity-preserving video synthesis across multiple scenarios.

Yong Zhong, Zhuoyi Yang, Jiayan Teng et al.

2025-03-18 41
cs.CV 2503.14106

Reliable uncertainty quantification for 2D/3D anatomical landmark localization using multi-output conformal prediction

Introduces multi-output conformal prediction (M-R2CCP, M-R2C2R) for reliable uncertainty quantification in 2D/3D anatomical landmark localization.

Jef Jonkers, Frank Coopman, Luc Duchateau et al.

2025-03-18 35
cs.CV 2503.13966

FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks

FlexVLN combines an LLM planner with supervised execution, achieving strong zero-shot transfer across REVERIE, SOON, and CVDN-target.

Siqi Zhang, Yanyuan Qiao, Qunbo Wang et al.

2025-03-18 20
cs.CV 2503.13429

Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning

CAVE combines 3D neural object volumes with sparse concept learning, achieving robust out-of-distribution performance and faithful interpretability, with 77.4% OOD accuracy.

Nhi Pham, Artur Jesslen, Bernt Schiele et al.

2025-03-18 41
cs.CV 2503.13086

Gaussian On-the-Fly Splatting: A Progressive Framework for Robust Near Real-Time 3DGS Optimization

Proposes On-the-Fly GS, a progressive framework enabling near real-time 3DGS optimization, reducing per-image optimization to seconds.

Yiwei Xu, Yifei Yu, Wentian Gan et al.

2025-03-17 26
cs.CV 2503.12271

Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection

Reflect-DiT employs in-context reflection to improve text-to-image diffusion, achieving +0.19 score improvement with only 20 samples on GenEval.

Shufan Li, Konstantinos Kallidromitis, Akash Gokul et al.

2025-03-16 45
cs.CV 2503.11806

Human-in-the-Loop Local Corrections of 3D Scene Layouts via Infilling

Multi-task SceneScript model with infilling enhances 3D scene layout local correction, achieving 98.6% F1 in local tasks.

Christopher Xie, Armen Avetisyan, Henry Howard-Jenkins et al.

2025-03-15 39
Prev 1 ... 67 68 69 70 71 72 73 ... 138 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home