InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
PLoRA adapts only image tokens, enabling InternLM2-7B to compose and understand free-form interleaved text-image content.
Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.