cs.CV 2312.08914

CogAgent: A Visual Language Model for GUI Agents

CogAgent is an 18-billion-parameter visual language model that excels in GUI understanding and navigation, achieving state-of-the-art results on multiple VQA benchmarks.

Wenyi Hong, Weihan Wang, Qingsong Lv et al.

2023-12-14 44
cs.CV 2312.05915

Diffusion for Natural Image Matting

DiffMatte employs diffusion models for multi-step iterative natural image matting, reducing SAD by 8%.

Yihan Hu, Yiheng Lin, Wei Wang et al.

2023-12-10 31
cs.CV 2312.05251

Reconstructing Hands in 3D with Transformers

HaMeR employs a Transformer-based architecture with large-scale data, achieving state-of-the-art 3D hand mesh reconstruction from monocular images, outperforming previous methods.

Georgios Pavlakos, Dandan Shan, Ilija Radosavovic et al.

2023-12-09 40