VIGAN: Missing View Imputation with Generative Adversarial Networks
VIGAN combines CycleGAN and multi-modal autoencoder for missing view imputation in multi-view data.
Chao Shang, Aaron Palmer, Jiangwen Sun et al.
VIGAN combines CycleGAN and multi-modal autoencoder for missing view imputation in multi-view data.
Chao Shang, Aaron Palmer, Jiangwen Sun et al.
Proposes a regularization method based on uniform distribution to maximize feature spread, combined with triplet loss, significantly improving local descriptor performance.
Xu Zhang, Felix X. Yu, Sanjiv Kumar et al.
Proposes an end-to-end learned multi-view stereo system leveraging differentiable geometric projections, achieving high-quality 3D reconstructions from few views, tested on ShapeNet.
Abhishek Kar, Christian Häne, Jitendra Malik
Proposed LARS algorithm enables AlexNet and ResNet-50 to train with large batches without accuracy loss.
Yang You, Igor Gitman, Boris Ginsburg
Unsupervised sequence sorting via CNN enables rich visual features, improving action recognition and detection benchmarks.
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh et al.
A Siamese relative-pose CNN achieves 0.21 m and 9.30° mean median error on 7 Scenes while generalizing across scenes.
Zakaria Laskar, Iaroslav Melekhov, Surya Kalia et al.
The method converts node embeddings into multi-channel histograms for vanilla 2D CNNs, reaching 48.13% on REDDIT-12K.
Antoine Jean-Pierre Tixier, Giannis Nikolentzos, Polykarpos Meladianos et al.
Proposes a cascaded refinement network (CRN) for high-res photo synthesis from semantic layouts, achieving 2MP resolution with superior realism over GANs.
Qifeng Chen, Vladlen Koltun
Up-Down attention combines Faster R-CNN region features with task-conditioned attention, reaching CIDEr 117.9 and winning the 2017 VQA Challenge.
Peter Anderson, Xiaodong He, Chris Buehler et al.
GridNet, a multi-scale residual grid architecture, achieves 69.45% IoU on Cityscapes, outperforming many baselines.
Damien Fourure, Rémi Emonet, Elisa Fromont et al.
Proposes deep CNN-based point detector MagicPoint and homography estimator MagicWarp for robust, real-time SLAM.
Daniel DeTone, Tomasz Malisiewicz, Andrew Rabinovich
Proposes NASNet architecture via small-scale search on CIFAR-10, transferred to ImageNet, achieving 82.7% top-1 accuracy with 28% fewer FLOPS.
Barret Zoph, Vijay Vasudevan, Jonathon Shlens et al.
Deep Layer Aggregation (DLA) uses iterative and hierarchical fusion to improve recognition with fewer parameters, outperforming traditional skip connections.
Fisher Yu, Dequan Wang, Evan Shelhamer et al.
MoCoGAN decomposes content and motion in a latent space, enabling controllable, unsupervised video generation with improved content consistency and dynamic diversity.
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang et al.
Proposes a 2D convolution-based point cloud generation framework with differentiable pseudo-rendering for dense 3D reconstruction, outperforming volumetric methods.
Chen-Hsuan Lin, Chen Kong, Simon Lucey
Comparing deep neural networks (DNNs) and humans in object recognition under degraded signals, humans show greater robustness.
Robert Geirhos, David H. J. Janssen, Heiko H. Schütt et al.
DeepLabv3 integrates multi-scale atrous convolution and global features, achieving 85.7% mIOU on Pascal VOC 2012 without DenseCRF.
Liang-Chieh Chen, George Papandreou, Florian Schroff et al.
Large-batch synchronized SGD with linear scaling and warmup trains ResNet-50 on ImageNet in 1 hour on 256 GPUs, maintaining accuracy.
Priya Goyal, Piotr Dollár, Ross Girshick et al.
PointNet++ employs a hierarchical architecture with multi-scale neighborhood features, significantly improving point cloud classification and segmentation accuracy.
Charles R. Qi, Li Yi, Hao Su et al.
Layer-wise fine-tuning of pre-trained CNNs outperforms training from scratch in medical image analysis, with performance gains of 5-12%.
Nima Tajbakhsh, Jae Y. Shin, Suryakanth R. Gurudu et al.