cs.CV 2408.00714

SAM 2: Segment Anything in Images and Videos

SAM 2 employs streaming memory within a transformer framework for promptable image and video segmentation, achieving 3x fewer interactions and 6x faster speeds than prior models.

Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu et al.

2024-08-02 49
cs.CV 2408.00712

MotionFix: Text-Driven 3D Human Motion Editing

Proposes MotionFix dataset and TMED diffusion model for text-driven 3D human motion editing, achieving significant improvements in editing accuracy.

Nikos Athanasiou, Alpár Cseke, Markos Diomataris et al.

2024-08-02 47
cs.CV 2408.00203

OmniParser for Pure Vision Based GUI Agent

OmniParser significantly enhances GPT-4V performance on ScreenSpot benchmark using pure vision-based UI parsing.

Yadong Lu, Jianwei Yang, Yelong Shen et al.

2024-08-01 17
cs.AI 2407.21783

The Llama 3 Herd of Models

Llama 3 employs a 405B-parameter dense Transformer supporting 128K tokens, multi-modal, and multi-language capabilities, rivaling GPT-4.

Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri et al.

2024-08-01 63
cs.LG 2407.21243

Informed Correctors for Discrete Diffusion Models

Informed corrector using model guidance and hollow Transformer architecture improves discrete diffusion sampling efficiency and quality.

Yixiu Zhao, Jiaxin Shi, Feng Chen et al.

2024-07-31 35
cs.LG 2407.17437

Nerva: a Truly Sparse Implementation of Neural Networks

Nerva leverages sparse matrix operations via Intel MKL to accelerate neural network training, reducing time by 4× at 99% sparsity with comparable accuracy.

Wieger Wesselink, Bram Grooten, Qiao Xiao et al.

2024-07-25 43