Deep Visual Foresight for Planning Robot Motion
Combining deep video prediction with MPC enables robots to manipulate unseen objects without calibration or physical models.
Chelsea Finn, Sergey Levine
Combining deep video prediction with MPC enables robots to manipulate unseen objects without calibration or physical models.
Chelsea Finn, Sergey Levine
Proposes a VAE framework integrating DGDN decoder and CNN encoder for joint image, label, and caption modeling, achieving high efficiency and accuracy.
Yunchen Pu, Zhe Gan, Ricardo Henao et al.
Uses 3D CNN trained on 440,000+ models for fast shape completion, boosting robotic grasping success to 93.33%.
Jacob Varley, Chad DeChant, Adam Richardson et al.
Deep LSTM with attention and subword units achieves state-of-the-art BLEU scores (38.95/24.17) on WMT'14 en-fr/en-de, reducing errors by 60%.
Yonghui Wu, Mike Schuster, Zhifeng Chen et al.
Demonstrates nonstoquastic Hamiltonians outperform stoquastic in quantum annealing on long-range Ising spin glasses, especially for hard instances, via spectral analysis and success probability metrics.
L. Hormozi, E. W. Brown, G. Carleo et al.
Proposes PDE-FIND, a sparse regression algorithm that automatically identifies key PDE terms from spatiotemporal data.
Samuel H. Rudy, Steven L. Brunton, Joshua L. Proctor et al.
SeqGAN combines GAN with policy gradient reinforcement learning to generate high-quality discrete sequences, overcoming gradient issues.
Lantao Yu, Weinan Zhang, Jun Wang et al.
Style imitation and chord invention in polyphonic music using maximum entropy principle, evaluated on Bach chorales.
Gaëtan Hadjeres, Jason Sakellariou, François Pachet
Proposed SRGAN employs GAN with perceptual loss for 4× super-resolution, significantly enhancing perceptual quality.
Christian Ledig, Lucas Theis, Ferenc Huszar et al.
Proposes a novel unsupervised monocular depth estimation method with left-right consistency, outperforming supervised methods on KITTI.
Clément Godard, Oisin Mac Aodha, Gabriel J. Brostow
Uses GAN-based latent space modeling with constrained optimization for realistic, user-controlled image editing, achieving near real-time performance.
Jun-Yan Zhu, Philipp Krähenbühl, Eli Shechtman et al.
Decentralized mirror descent algorithm for dynamic environment tracking, with theoretical regret bounds based on network spectral gap and target deviation.
Shahin Shahrampour, Ali Jadbabaie
Deep GRU-based multi-task learning significantly improves text recommendation accuracy, especially in cold-start scenarios, outperforming traditional models by up to 34%.
Trapit Bansal, David Belanger, Andrew McCallum
Attention-based encoder-decoder and bidirectional RNN models achieve state-of-the-art results on ATIS for joint intent detection and slot filling.
Bing Liu, Ian Lane
Proposes an autonomous robot system using RGBD-based implicit mapping and active learning for dense clutter sorting, achieving over 85% success and 80% classification accuracy.
Janne V. Kujala, Tuomas J. Lukka, Harri Holopainen
This study reveals the existence of bad local maxima in Gaussian Mixture Models and its impact on EM algorithm convergence.
Chi Jin, Yuchen Zhang, Sivaraman Balakrishnan et al.
Deep learning framework combining RNN and CNN decodes EEG signals to classify 40 ImageNet categories with 83% accuracy, enabling brain-driven image recognition.
Concetto Spampinato, Simone Palazzo, Isaak Kavasidis et al.
Proposes filter pruning based on `1-norm` importance, reducing FLOPs by up to 34% for VGG-16 and 38% for ResNet-110 on CIFAR-10, with minimal accuracy loss.
Hao Li, Asim Kadav, Igor Durdanovic et al.
PGA and SGLD poison MovieLens factorization recommenders; at β=0.6, SGLD reaches detection-test p-values above 0.7.
Bo Li, Yining Wang, Aarti Singh et al.
Using GloVe embeddings and WEAT/WEFAT, the study quantifies cultural biases embedded in language, replicating psychological IAT results with high effect sizes.
Aylin Caliskan, Joanna J. Bryson, Arvind Narayanan