Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation
pSp framework uses StyleGAN for image-to-image translation, directly embedding into W+ space.
Elad Richardson, Yuval Alaluf, Or Patashnik et al.
pSp framework uses StyleGAN for image-to-image translation, directly embedding into W+ space.
Elad Richardson, Yuval Alaluf, Or Patashnik et al.
Using a chopstick teleoperation interface, the study analyzes human strategies, achieving the highest success rate in three out of five objects.
Liyiming Ke, Ajinkya Kamat, Jingqiang Wang et al.
Proposed low-rank tensor bandit algorithms with finite regret bounds, outperforming existing methods in high-dimensional online decision tasks.
Jie Zhou, Botao Hao, Zheng Wen et al.
LEMMA dataset benchmarks multi-agent, multi-task activities, enabling goal-directed behavior and temporal reasoning research.
Baoxiong Jia, Yixin Chen, Siyuan Huang et al.
Lie algebra-based conditional VAE enables diverse, realistic 3D human motion generation with improved physical consistency.
Chuan Guo, Xinxin Zuo, Sen Wang et al.
Proposes G-MC and G-TLS models; proves their inapproximability; develops minimally tuned algorithms ADAPT and GNC; demonstrates robustness up to 90% outliers in robotics tasks.
Pasquale Antonante, Vasileios Tzoumas, Heng Yang et al.
Face2Face achieves real-time monocular face reenactment using non-rigid model-based bundling and dense photometric tracking, reaching 28Hz with high realism.
Justus Thies, Michael Zollhöfer, Marc Stamminger et al.
BigBird introduces sparse attention with global, local, and random tokens, reducing complexity from quadratic to linear, enabling long sequence modeling.
Manzil Zaheer, Guru Guruganesh, Avinava Dubey et al.
HITNet employs multi-resolution hierarchical refinement without 3D convolutions, achieving real-time stereo matching with high accuracy.
Vladimir Tankovich, Christian Häne, Yinda Zhang et al.
Proposed a multi-modal behavioral cloning controller for peg-in-hole tasks, achieving 100% success rate.
Yifang Liu, Diego Romeres, Devesh K. Jha et al.
Extremely low-precision neural networks for speech mask estimation enable resource-efficient multi-channel speech enhancement.
Lukas Pfeifenberger, Matthias Zöhrer, Günther Schindler et al.
PackIt environment evaluates geometric planning using evolutionary algorithm, challenging task dataset validates effectiveness.
Ankit Goyal, Jia Deng
REC employs CNN-LSTM and graph models to localize objects, achieving 75% accuracy on RefCOCO+ dataset.
Yanyuan Qiao, Chaorui Deng, Qi Wu
F3-Net leverages frequency-aware clues via DCT-based decomposition and local statistics, outperforming SOTA in low-quality face forgery detection with 90.43% accuracy.
Yuyang Qian, Guojun Yin, Lu Sheng et al.
PyTorch3D introduces modular, differentiable rendering and operators, achieving up to 10× speedup in 3D deep learning tasks on ShapeNet.
Nikhila Ravi, Jeremy Reizenstein, David Novotny et al.
Proposes eSL-Net, a sparse learning framework, achieving 7-12dB PSNR improvement for high-quality event camera image reconstruction.
Bishan Wang, Jingwei He, Lei Yu et al.
Proposes InfoXLM, an information-theoretic framework utilizing mutual information maximization and contrastive learning to significantly improve cross-lingual transfer.
Zewen Chi, Li Dong, Furu Wei et al.
Proposes Self-Predictive Representations (SPR) to improve pixel-based deep RL sample efficiency, surpassing human scores within 100k steps.
Max Schwarzer, Ankesh Anand, Rishab Goel et al.
Proposes LayerPrune framework, using layer-level pruning based on importance metrics to surpass filter pruning in latency reduction while maintaining comparable accuracy.
Sara Elkerdawy, Mostafa Elhoushi, Abhineet Singh et al.
SUNRISE introduces a unified ensemble framework with weighted Bellman backups and UCB exploration, boosting deep RL stability and efficiency.
Kimin Lee, Michael Laskin, Aravind Srinivas et al.