Bayesian Generative Active Deep Learning
Proposes a Bayesian Generative Active Deep Learning method, enhancing classification efficiency on datasets like MNIST.
Toan Tran, Thanh-Toan Do, Ian Reid et al.
Proposes a Bayesian Generative Active Deep Learning method, enhancing classification efficiency on datasets like MNIST.
Toan Tran, Thanh-Toan Do, Ian Reid et al.
Sparse Transformer reduces attention complexity to O(n√n), enabling modeling of sequences tens of thousands long with hundreds of layers.
Rewon Child, Scott Gray, Alec Radford et al.
This paper introduces a free-form video inpainting model using 3D Gated Convolution and Temporal PatchGAN, significantly improving spatial detail and temporal consistency.
Ya-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee et al.
Proposes Nucleus Sampling, a dynamic probability truncation method, to enhance diversity and coherence in neural text generation.
Ari Holtzman, Jan Buys, Li Du et al.
PullNet uses iterative subgraph building with Graph CNNs, achieving state-of-the-art multi-hop QA performance, especially with incomplete KBs.
Haitian Sun, Tania Bedrax-Weiss, William W. Cohen
Social Ways uses Info-GAN to preserve multimodal pedestrian futures, reaching 0.39/0.64 ADE/FDE on ETH.
Javad Amirian, Jean-Bernard Hayet, Julien Pettre
DiffNet models recursive social influence diffusion, improving recommendation accuracy by over 13% on real datasets.
Le Wu, Peijie Sun, Yanjie Fu et al.
Behavior cloning achieves state-of-the-art in complex driving but faces generalization and data bias issues.
Felipe Codevilla, Eder Santana, Antonio M. López et al.
Introduced TextVQA dataset and LoRRA model, achieving 27.63% accuracy on text-based VQA, surpassing SOTA methods.
Amanpreet Singh, Vivek Natarajan, Meet Shah et al.
SpecAugment applies time warping and frequency masking to spectrograms, achieving 6.8% WER on LibriSpeech test-other without language models.
Daniel S. Park, William Chan, Yu Zhang et al.
KPConv introduces flexible, deformable point cloud convolution, achieving 92.9% accuracy on ModelNet40.
Hugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud et al.
Proposes MinkowskiNet, a 4D sparse convolutional neural network leveraging generalized sparse convolutions, achieving 67.9% mIoU on ScanNet, outperforming prior methods by 19%.
Christopher Choy, JunYoung Gwak, Silvio Savarese
Proposed a deep depth estimation method using event streams, achieving 85% accuracy improvement and microsecond-level real-time processing.
Guillermo Gallego, Tobi Delbruck, Garrick Orchard et al.
Proposes a recurrent neural network for event-to-video reconstruction, improving image quality by over 20%, enabling application of standard vision algorithms.
Henri Rebecq, René Ranftl, Vladlen Koltun et al.
SAGPool enhances graph classification using self-attention, achieving a 5% improvement on the D&D dataset.
Junhyun Lee, Inyeop Lee, Jaewoo Kang
ContactDB analyzes and predicts grasp contact via thermal imaging, featuring 3750 3D meshes and 375K frames of RGB-D+thermal images.
Samarth Brahmbhatt, Cusuh Ham, Charles C. Kemp et al.
BERT4Rec employs bidirectional Transformer with Cloze training, outperforming SOTA in sequential recommendation.
Fei Sun, Jun Liu, Jian Wu et al.
GA-Net introduces semi-global and local guided aggregation layers, achieving state-of-the-art accuracy with significantly reduced computational cost.
Feihu Zhang, Victor Prisacariu, Ruigang Yang et al.
A face de-occlusion method using 3D Morphable Model and GAN, enhancing face reconstruction accuracy.
Xiaowei Yuan, In Kyu Park
Proposed Graph Wavelet Neural Network (GWNN) using wavelet transform for efficient graph convolution, achieving state-of-the-art accuracy on Cora and other datasets.
Bingbing Xu, Huawei Shen, Qi Cao et al.