GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
GODIVA combines VQ-VAE and 3D sparse attention to generate open-domain videos; RM reaches 93.48 on MSR-VTT.
Chenfei Wu, Lun Huang, Qianxi Zhang et al.
GODIVA combines VQ-VAE and 3D sparse attention to generate open-domain videos; RM reaches 93.48 on MSR-VTT.
Chenfei Wu, Lun Huang, Qianxi Zhang et al.
Proposes a layered neural representation (ST-NeRF) for editable, high-quality free-viewpoint video of large dynamic scenes using sparse 16 cameras.
Jiakai Zhang, Xinhang Liu, Xinyi Ye et al.
Introduced P3M-10k dataset and P3M-Net for privacy-preserving portrait matting, achieving state-of-the-art results with face-blurred images.
Jizhizi Li, Sihan Ma, Jing Zhang et al.
ViLD employs vision-language knowledge distillation to enable open-vocabulary object detection, achieving 16.1 mask AP_r on LVIS.
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo et al.
Proposes 8 sanity checks for reward functions, revealing widespread flaws in autonomous driving RL reward design.
W. Bradley Knox, Alessandro Allievi, Holger Banzhaf et al.
Proposes a method using dense correlation volumes to estimate extreme rotation in RGB image pairs, successfully handling non-overlapping images.
Ruojin Cai, Bharath Hariharan, Noah Snavely et al.
Introduces FRANK benchmark with fine-grained factuality error classification, evaluating summarization models and metrics.
Artidoro Pagnoni, Vidhisha Balachandran, Yulia Tsvetkov
MDETR performs text-conditioned end-to-end detection, pre-trained on 1.3M aligned image-text pairs for open-vocabulary grounding.
Aishwarya Kamath, Mannat Singh, Yann LeCun et al.
Introduces InfographicVQA dataset and evaluates Transformer-based models, revealing significant performance gaps in understanding complex infographics.
Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito et al.
Proposes importance sampling to accelerate PINNs training, improving convergence by ~30% and reducing training time by 20%.
Mohammad Amin Nabian, Rini Jasmine Gladstone, Hadi Meidani
DeepImpact leverages semantic impact scores with BERT and DocT5Query to improve first-stage retrieval by 17%, enabling faster and more accurate search.
Antonio Mallia, Omar Khattab, Nicola Tonellotto et al.
Proposes a deep video matting framework with spatio-temporal feature aggregation, outperforming traditional and deep image methods.
Yanan Sun, Guanzhi Wang, Qiao Gu et al.
Reset-Free Reinforcement Learning via Multi-Task Learning achieves complex dexterous manipulation without human intervention.
Abhishek Gupta, Justin Yu, Tony Z. Zhao et al.
Constructed H2O dataset and proposed a graph convolutional network for joint 3D hand-object interaction recognition from RGB images.
Taein Kwon, Bugra Tekin, Jan Stuhmer et al.
Proposes a transferable deep learning framework combining GFNet and Mosaic predictor for unseen PDE domains, achieving 3 orders speedup.
Hengjie Wang, Robert Planas, Aparna Chandramowlishwaran et al.
Proposes dataset inference leveraging model memorization to verify ownership with over 99% confidence, using statistical tests and distance estimation techniques.
Pratyush Maini, Mohammad Yaghini, Nicolas Papernot
FIERY employs end-to-end deep learning to predict multi-modal future trajectories in bird’s-eye view from monocular cameras, outperforming baselines with IoU of 57.8%.
Anthony Hu, Zak Murez, Nikhil Mohan et al.
Using GPS traces and microscopic models, the study reveals heavy-tailed emission distributions and identifies 'gross polluters' in European cities, guiding targeted policies.
Matteo Böhm, Mirco Nanni, Luca Pappalardo
This study quantifies the energy use and carbon footprint of large NLP models like GPT-3, proposing strategies for greener training.
David Patterson, Joseph Gonzalez, Quoc Le et al.
Introduces Waymo Motion Dataset with 570 hours of multi-agent interaction data, enabling joint prediction models for autonomous driving.
Scott Ettinger, Shuyang Cheng, Benjamin Caine et al.