cs.CV 2312.08914

CogAgent: A Visual Language Model for GUI Agents

CogAgent is an 18-billion-parameter visual language model that excels in GUI understanding and navigation, achieving state-of-the-art results on multiple VQA benchmarks.

Wenyi Hong, Weihan Wang, Qingsong Lv et al.

2023-12-14 873 citations 48
cs.CL 2312.06550

LLM360: Towards Fully Transparent Open-Source LLMs

Proposes LLM360 framework, fully open-sourcing 7B models, training code, data, checkpoints, and analysis for transparency.

Zhengzhong Liu, Aurick Qiao, Willie Neiswanger et al.

2023-12-12 122 citations 48
cs.CV 2312.06720

Audio-Visual LLM for Video Understanding

Audio-Visual LLM achieves 53.7% accuracy on MSRVTT-QA using modality-augmented training for video understanding.

Fangxun Shu, Lei Zhang, Hao Jiang et al.

2023-12-11 1