cs.CV 2207.11871

Towards Complex Document Understanding By Discrete Reasoning

Proposes MHST, a multi-modal transformer model integrating text, layout, and visual features, achieving significant improvements on the TAT-DQA dataset for complex document VQA.

Fengbin Zhu, Wenqiang Lei, Fuli Feng et al.

2022-07-25 47
cs.CV 2207.03552

An Embedding-Dynamic Approach to Self-supervised Learning

MSBReg combines delayed-parameter embeddings, multiview centroid, singular value regularization, and Brownian motion to improve self-supervised learning, achieving 81.56% top-1 accuracy on ImageNet.

Suhong Moon, Domas Buracas, Seunghyun Park et al.

2022-07-08 30
cs.CV 2207.02504

Dual Decision Improves Open-Set Panoptic Segmentation

Dual decision framework combining known class discriminator and class-agnostic object head improves open-set panoptic segmentation PQ by over 30%.

Hai-Ming Xu, Hao Chen, Lingqiao Liu et al.

2022-07-06 9 citations 31
cs.CV 2206.08916

Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

UNIFIED-IO employs a unified sequence-to-sequence transformer architecture, converting multi-modal inputs and outputs into discrete tokens, trained on over 90 datasets, achieving state-of-the-art multi-task performance.

Jiasen Lu, Christopher Clark, Rowan Zellers et al.

2022-06-18 544 citations 26