LISA: Reasoning Segmentation via Large Language Model
LISA integrates large language models for reasoning segmentation, achieving over 56% gIoU on ReasonSeg with minimal fine-tuning.
Xin Lai, Zhuotao Tian, Yukang Chen et al.
LISA integrates large language models for reasoning segmentation, achieving over 56% gIoU on ReasonSeg with minimal fine-tuning.
Xin Lai, Zhuotao Tian, Yukang Chen et al.
Proposes ADTrans framework using semantic prototypes to address bias in PSG, improving R@100 by 3-4% and balancing long-tail relations.
Li Li, Wei Ji, Yiming Wu et al.
PointOdyssey constructs a large-scale synthetic dataset using real motion capture and structure-from-motion trajectories, improving long-term point tracking accuracy.
Yang Zheng, Adam W. Harley, Bokui Shen et al.
CLIP-KD employs feature alignment and contrastive learning to improve small models, maximizing similarity with teacher features, achieving significant performance gains.
Chuanguang Yang, Zhulin An, Libo Huang et al.
Combines OpenPose and CNN to detect suspicious exam behaviors by analyzing arm angles and duration, achieving 82% accuracy in identifying object exchanges.
Reuben Moyo, Stanley Ndebvu, Michael Zimba et al.
DNA-Rendering creates a large-scale, high-fidelity human dataset with 1500+ subjects, 67.5M frames, boosting neural human rendering performance.
Wei Cheng, Ruixiang Chen, Wanqi Yin et al.
InternVid leverages large-scale video-text data and multi-scale captioning to train ViCLIP, achieving state-of-the-art zero-shot action recognition.
Yi Wang, Yinan He, Yizhuo Li et al.
Generate coherent storytelling videos using retrieval-augmented video generation.
Yingqing He, Menghan Xia, Haoxin Chen et al.
NaViT uses sequence packing to handle images of any resolution, enhancing training efficiency and performance.
Mostafa Dehghani, Basil Mustafa, Josip Djolonga et al.
MMBench introduces a CircularEval-based benchmark with 3000+ questions across 20 abilities, supporting bilingual evaluation of vision-language models.
Yuan Liu, Haodong Duan, Yuanhan Zhang et al.
Proposes a differentiable rendering framework for scene decomposition using textured superquadrics, achieving accurate, interpretable 3D representations from multi-view images.
Tom Monnier, Jake Austin, Angjoo Kanazawa et al.
Objaverse-XL, with over 10 million 3D models, significantly advances large-scale 3D vision tasks.
Matt Deitke, Ruoshi Liu, Matthew Wallingford et al.
AnimateDiff employs transfer learning and MotionLoRA, enabling personalized animation synthesis without model fine-tuning, significantly improving animation quality and diversity.
Yuwei Guo, Ceyuan Yang, Anyi Rao et al.
SVIT introduces a 4.2M high-quality visual instruction dataset, significantly outperforming state-of-the-art multimodal models.
Bo Zhao, Boya Wu, Muyang He et al.
Hierarchical neural code trees enable controllable CAD generation with high complexity, outperforming SOTA in diversity and realism metrics.
Xiang Xu, Pradeep Kumar Jayaraman, Joseph G. Lambourne et al.
Proposes Michelangelo framework utilizing shape-image-text aligned latent space with diffusion models, achieving superior quality and diversity in cross-modal 3D shape generation.
Zibo Zhao, Wen Liu, Xin Chen et al.
Introduces MESS benchmark for zero-shot semantic segmentation across 22 diverse datasets, leveraging CLIP-based models with detailed performance analysis.
Benedikt Blumenstiel, Johannes Jakubik, Hilde Kühne et al.
Shikra model enables multimodal LLMs to perform referential dialogue, improving localization tasks.
Keqin Chen, Zhao Zhang, Weili Zeng et al.
FunQA enhances VLM understanding of counter-intuitive videos via multi-turn dialogues, featuring 312K QA pairs.
Binzhu Xie, Sicheng Zhang, Zitang Zhou et al.
Reduce hallucination in multi-modal models using LRV-Instruction dataset and GAVIE method.
Fuxiao Liu, Kevin Lin, Linjie Li et al.