LEMMA: A Multi-view Dataset for Learning Multi-agent Multi-task Activities
LEMMA dataset benchmarks multi-agent, multi-task activities, enabling goal-directed behavior and temporal reasoning research.
Baoxiong Jia, Yixin Chen, Siyuan Huang et al.
LEMMA dataset benchmarks multi-agent, multi-task activities, enabling goal-directed behavior and temporal reasoning research.
Baoxiong Jia, Yixin Chen, Siyuan Huang et al.
Lie algebra-based conditional VAE enables diverse, realistic 3D human motion generation with improved physical consistency.
Chuan Guo, Xinxin Zuo, Sen Wang et al.
Proposes G-MC and G-TLS models; proves their inapproximability; develops minimally tuned algorithms ADAPT and GNC; demonstrates robustness up to 90% outliers in robotics tasks.
Pasquale Antonante, Vasileios Tzoumas, Heng Yang et al.
Face2Face achieves real-time monocular face reenactment using non-rigid model-based bundling and dense photometric tracking, reaching 28Hz with high realism.
Justus Thies, Michael Zollhöfer, Marc Stamminger et al.
HITNet employs multi-resolution hierarchical refinement without 3D convolutions, achieving real-time stereo matching with high accuracy.
Vladimir Tankovich, Christian Häne, Yinda Zhang et al.
REC employs CNN-LSTM and graph models to localize objects, achieving 75% accuracy on RefCOCO+ dataset.
Yanyuan Qiao, Chaorui Deng, Qi Wu
F3-Net leverages frequency-aware clues via DCT-based decomposition and local statistics, outperforming SOTA in low-quality face forgery detection with 90.43% accuracy.
Yuyang Qian, Guojun Yin, Lu Sheng et al.
PyTorch3D introduces modular, differentiable rendering and operators, achieving up to 10× speedup in 3D deep learning tasks on ShapeNet.
Nikhila Ravi, Jeremy Reizenstein, David Novotny et al.
Proposes eSL-Net, a sparse learning framework, achieving 7-12dB PSNR improvement for high-quality event camera image reconstruction.
Bishan Wang, Jingwei He, Lei Yu et al.
Proposes LayerPrune framework, using layer-level pruning based on importance metrics to surpass filter pruning in latency reduction while maintaining comparable accuracy.
Sara Elkerdawy, Mostafa Elhoushi, Abhineet Singh et al.
Proposed a three-stage framework using scene context for long-term human motion prediction, significantly improving accuracy.
Zhe Cao, Hang Gao, Karttikeya Mangalam et al.
Swapping Autoencoder enables efficient image editing by encoding structure and texture, improving generation quality and speed.
Taesung Park, Jun-Yan Zhu, Oliver Wang et al.
Proposed Goal-Oriented Semantic Exploration system combines explicit semantic maps and reinforcement learning, achieving 54.4% success in unseen environments, outperforming baselines.
Devendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta et al.
Using fSIM-NET to predict 3D object functionality, complemented by iGEN-NET for interaction scene generation.
Ruizhen Hu, Zihao Yan, Jingwen Zhang et al.
Proposed a calibration-free, real-time 3D object tracking system combining neural detection and planar tracking on mobile devices, achieving 26FPS.
Adel Ahmadyan, Tingbo Hou, Jianing Wei et al.
Pix2Vox++ employs multi-scale context-aware fusion for accurate multi-view 3D reconstruction, outperforming SOTA with IoU of 85.2%.
Haozhe Xie, Hongxun Yao, Shengping Zhang et al.
CenterPoint uses point-based detection with heatmaps and regression, achieving 65.5 NDS on nuScenes with a single model.
Tianwei Yin, Xingyi Zhou, Philipp Krähenbühl
BlazePose is a lightweight CNN for real-time human pose estimation on mobile devices, outputting 33 keypoints at over 30fps.
Valentin Bazarevsky, Ivan Grishchenko, Karthik Raveendran et al.
SwAV introduces a clustering-based contrastive method for unsupervised visual representation, achieving 75.3% top-1 accuracy on ImageNet without memory banks.
Mathilde Caron, Ishan Misra, Julien Mairal et al.
Proposed Global-Local Attention Transformer (GLAT) learns structured visual commonsense, improving scene graph robustness without external knowledge.
Alireza Zareian, Zhecan Wang, Haoxuan You et al.