cs.CV 2305.06355

VideoChat: Chat-Centric Video Understanding

VideoChat integrates video foundation models with LLMs via learnable interfaces, enabling advanced spatiotemporal reasoning and causal inference.

KunChang Li, Yinan He, Yi Wang et al.

2023-05-11 44
cs.CV 2305.05665

ImageBind: One Embedding Space To Bind Them All

ImageBind learns a unified embedding space for six modalities, enabling zero-shot cross-modal retrieval, composition, detection, and generation, surpassing specialized models.

Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu et al.

2023-05-10 1776 citations 33
cs.CL 2305.05280

VCSUM: A Versatile Chinese Meeting Summarization Dataset

VCSUM creates a large-scale Chinese meeting dataset with topic segmentation, multi-granularity summaries, and salient sentence annotations, enabling multi-task learning.

Han Wu, Mingjie Zhan, Haochen Tan et al.

2023-05-09 52
cs.CV 2305.04789

AvatarReX: Real-time Expressive Full-body Avatars

AvatarReX employs NeRF with structured local implicit fields and geometry-appearance disentanglement for real-time, expressive full-body avatar synthesis, achieving high fidelity and speed.

Zerong Zheng, Xiaochen Zhao, Hongwen Zhang et al.

2023-05-08 37