GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CV 2404.00172

Universal Bovine Identification via Depth Data and Deep Metric Learning

Deep metric learning with depth data achieves 99% accuracy in cattle identification, eliminating the need for breed-specific markings.

Asheesh Sharma, Lucy Randewich, William Andrew et al.

2024-03-30 44
cs.CV 2403.20126

ECLIPSE: Efficient Continual Learning in Panoptic Segmentation with Visual Prompt Tuning

ECLIPSE leverages Visual Prompt Tuning, freezing backbone parameters and fine-tuning prompts, achieving state-of-the-art continual panoptic segmentation with only 1.3% trainable parameters.

Beomyoung Kim, Joonsang Yu, Sung Ju Hwang

2024-03-29 44
cs.CV 2403.19651

MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

MagicLens uses self-supervised learning on 36.7M web triplets to support open-ended image retrieval, outperforming SOTA.

Kai Zhang, Yi Luan, Hexiang Hu et al.

2024-03-29 44
cs.CV 2403.19213

Learning Multiple Representations with Inconsistency-Guided Detail Regularization for Mask-Guided Matting

Introduced inconsistency-guided detail regularization, enhancing mask-guided matting accuracy, surpassing SOTA methods.

Weihao Jiang, Zhaozhi Xie, Yuxiang Lu et al.

2024-03-28 33
cs.CV 2403.18818

ObjectDrop: Bootstrapping Counterfactuals for Photorealistic Object Removal and Insertion

ObjectDrop uses counterfactual datasets and diffusion model fine-tuning to achieve photorealistic object removal and insertion, outperforming prior methods.

Daniel Winter, Matan Cohen, Shlomi Fruchter et al.

2024-03-28 36
cs.CV 2403.18814

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Mini-Gemini employs dual visual encoders and patch info mining to enhance high-res visual understanding, surpassing private models in zero-shot benchmarks.

Yanwei Li, Yuechen Zhang, Chengyao Wang et al.

2024-03-28 34
cs.CV 2403.18036

Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance

Proposed a two-stage diffusion-based framework utilizing scene affordance for language-guided human motion synthesis, outperforming baselines on benchmark datasets.

Zan Wang, Yixin Chen, Baoxiong Jia et al.

2024-03-27 48
cs.CV 2403.16370

GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-aware Panoramic Semantic Segmentation

Proposed GoodSAM framework combines DAR and MKA modules to transfer SAM's zero-shot instance segmentation for panoramic semantic segmentation, achieving +3.75% mIoU improvement.

Weiming Zhang, Yexin Liu, Xu Zheng et al.

2024-03-25 48
cs.CV 2403.15388

LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models

PruMerge adaptively prunes and merges visual tokens, achieving 14× compression while maintaining performance.

Yuzhang Shang, Mu Cai, Bingxin Xu et al.

2024-03-23 35
cs.CV 2403.14624

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Introduced MathVerse, a multi-version benchmark with chain-of-thought analysis, revealing most models rely heavily on text rather than visual understanding in math problems.

Renrui Zhang, Dongzhi Jiang, Yichi Zhang et al.

2024-03-22 46
cs.CV 2403.14198

Unleashing Unlabeled Data: A Paradigm for Cross-View Geo-Localization

Unsupervised cross-view geo-localization using projection and re-ranking achieves over 40% pseudo-label accuracy.

Guopeng Li, Ming Qian, Gui-Song Xia

2024-03-21 34
cs.CV 2403.12895

mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

mPLUG-DocOwl 1.5 improves document understanding via Unified Structure Learning, achieving >10-point SOTA gains on 5/10 benchmarks.

Anwen Hu, Haiyang Xu, Jiabo Ye et al.

2024-03-20 31
cs.CV 2403.13438

SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

SpatialPIN enhances VLM's 3D reasoning via multi-model prompting, improving spatial VQA and robotics tasks without training.

Chenyang Ma, Kai Lu, Ta-Ying Cheng et al.

2024-03-19 42
cs.CV 2403.11481

VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

VideoAgent integrates structured memory with foundation models, boosting long video understanding by 6.6% on NExT-QA.

Yue Fan, Xiaojian Ma, Rujie Wu et al.

2024-03-18 46
cs.CV 2403.11373

Reconstruct before Query: Continual Missing Modality Learning with Decomposed Prompt Collaboration

RebQ framework improves average precision from 20.00 to 50.92 and reduces forgetting from 75.95 to 8.56 in missing modality learning.

Shu Zhao, Xiaohan Zou, Tan Yu et al.

2024-03-18 6
cs.CV 2403.11157

Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion Model

DiffUIR employs a selective hourglass mapping strategy with strong condition guidance and shared distribution terms, achieving state-of-the-art performance in multi-task image restoration with only 0.89M parameters.

Dian Zheng, Xiao-Ming Wu, Shuzhou Yang et al.

2024-03-17 122 citations 75
cs.CV 2403.10517

VideoAgent: Long-form Video Understanding with Large Language Model as Agent

VideoAgent uses a large language model as an agent for multi-round retrieval, achieving 54.1% zero-shot accuracy with only 8.4 frames on average.

Xiaohan Wang, Yuhui Zhang, Orr Zohar et al.

2024-03-16 40
cs.CV 2403.09611

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Proposes MM1, a multimodal pretraining framework combining contrastive and reconstructive losses, achieving SOTA with 30B parameters.

Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier et al.

2024-03-15 41
cs.CV 2403.09412

OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments

OpenGraph introduces a hierarchical 3D graph framework using vision-language models for large-scale outdoor scene understanding, achieving 73.4% IoU without fine-tuning.

Yinan Deng, Jiahui Wang, Jingyu Zhao et al.

2024-03-14 55
cs.CV 2403.08629

Scaling Up Dynamic Human-Scene Interaction Modeling

Proposes TRUMANS dataset and diffusion-based autoregressive model for arbitrary-length human-scene interaction motion synthesis.

Nan Jiang, Zhiyuan Zhang, Hongjie Li et al.

2024-03-13 34
Prev 1 ... 83 84 85 86 87 88 89 ... 138 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home