cs.CV 2207.11871

Towards Complex Document Understanding By Discrete Reasoning

Proposes MHST, a multi-modal transformer model integrating text, layout, and visual features, achieving significant improvements on the TAT-DQA dataset for complex document VQA.

Fengbin Zhu, Wenqiang Lei, Fuli Feng et al.

2022-07-25 54
cs.CL 2207.10397

CodeT: Code Generation with Generated Tests

CodeT leverages pre-trained models to generate test cases, using dual execution agreement to significantly improve code solution selection accuracy.

Bei Chen, Fengji Zhang, Anh Nguyen et al.

2022-07-21 607 citations 44
cs.LG 2207.09090

Actor-Critic based Improper Reinforcement Learning

Proposes Actor-Critic-based improper RL algorithms for combining multiple controllers to optimize unknown MDPs, with proven convergence rates.

Mohammadi Zaki, Avinash Mohan, Aditya Gopalan et al.

2022-07-19 36
math.FA 2207.08266

Highly symmetric lines

Construct highly symmetric line systems using twisted spherical functions of finite groups, improving kissing number lower bounds for d=10,11,14.

Mikhail Ganzhinov

2022-07-18 6
cs.CL 2207.07061

Confident Adaptive Language Modeling

Introduced Confident Adaptive Language Modeling (CALM) to reduce computation while maintaining performance, achieving up to 3x speedup.

Tal Schuster, Adam Fisch, Jai Gupta et al.

2022-07-15 15
cs.CL 2207.06881

Recurrent Memory Transformer

RMT enhances long-sequence processing with memory, surpassing Transformer-XL.

Aydar Bulatov, Yuri Kuratov, Mikhail S. Burtsev

2022-07-14 0
cs.IR 2207.05969

Bootstrap Latent Representations for Multi-modal Recommendation

BM3 introduces a self-supervised multi-modal recommendation framework using dropout-based contrastive views, achieving 2-9x faster training and outperforming state-of-the-art on large datasets.

Xin Zhou, Hongyu Zhou, Yong Liu et al.

2022-07-13 401 citations 37