Ditto: Building Digital Twins of Articulated Objects from Interaction
Ditto reconstructs articulated digital twins from before/after point clouds, reaching 0.72 whole-object Chamfer Distance on Shape2Motion.
Zhenyu Jiang, Cheng-Chun Hsu, Yuke Zhu
Ditto reconstructs articulated digital twins from before/after point clouds, reaching 0.72 whole-object Chamfer Distance on Shape2Motion.
Zhenyu Jiang, Cheng-Chun Hsu, Yuke Zhu
Wukong: 100M Chinese cross-modal dataset with contrastive and token interaction techniques, boosting zero-shot classification accuracy to 73.03%.
Jiaxi Gu, Xiaojun Meng, Guansong Lu et al.
MaskGIT accelerates image generation by 64x using bidirectional transformers, improving quality.
Huiwen Chang, Han Zhang, Lu Jiang et al.
ShapeFormer uses sparse representation with Transformer to generate high-quality 3D shape completions, outperforming existing methods.
Xingguang Yan, Liqiang Lin, Niloy J. Mitra et al.
Proposes a multi-scale attention-based semantic-guided VPR algorithm, achieving 92.2% Recall@1 on Oxford RobotCar, outperforming SOTA.
Valerio Paolicelli, Antonio Tavera, Carlo Masone et al.
Proposes multiresolution hash encoding, enabling seconds-level training and tens-of-milliseconds rendering for neural graphics primitives with high quality.
Thomas Müller, Alex Evans, Christoph Schied et al.
Proposes physics-constrained adversarial attack boosting trajectory prediction error by 150%, highlighting safety risks in autonomous driving.
Qingzhao Zhang, Shengtuo Hu, Jiachen Sun et al.
LSeg achieves 52.3% mIoU in zero-shot segmentation on PASCAL-5i, outperforming existing methods.
Boyi Li, Kilian Q. Weinberger, Serge Belongie et al.
Proposes ITSA, an information-theoretic regularization, to automatically restrict shortcut features, improving domain generalization of stereo networks.
WeiQin Chuah, Ruwan Tennakoon, Reza Hoseinnezhad et al.
ReferFormer uses Transformer with language as queries, achieving 55.6 J&F on Ref-Youtube-VOS, surpassing previous SOTA by 8.4 points, enabling end-to-end video object segmentation and tracking.
Jiannan Wu, Yi Jiang, Peize Sun et al.
SurfGen employs differentiable spherical projection and adversarial training to directly optimize surface geometry, producing diverse high-fidelity 3D shapes.
Andrew Luo, Tianqin Li, Wen-Hao Zhang et al.
StyleGAN-V extends StyleGAN2 for continuous video generation, achieving high quality at 1024×1024 resolution with 30% better FVD scores and arbitrary length/frame rate.
Ivan Skorokhodov, Sergey Tulyakov, Mohamed Elhoseiny
SLIP combines self-supervised contrastive learning with CLIP-style language-image pretraining, boosting ImageNet linear accuracy by 8.1%.
Norman Mu, Alexander Kirillov, David Wagner et al.
OpenSeg uses image-level captions to achieve open-vocabulary segmentation, improving mIoU by 19.9 points.
Golnaz Ghiasi, Xiuye Gu, Yin Cui et al.
GOAL uses GNet and MNet to generate full-body 4D motions for grasping unseen objects, outperforming baselines with high realism and generalization.
Omid Taheri, Vasileios Choutas, Michael J. Black et al.
Proposes an unsupervised multi-view learning framework combining explicit ellipsoid and implicit neural representations to discover 3D joints across unseen categories, achieving high-precision re-posing.
Atsuhiro Noguchi, Umar Iqbal, Jonathan Tremblay et al.
Mega-NeRF achieves scalable NeRFs for large scenes with 3x faster training and 12% PSNR improvement.
Haithem Turki, Deva Ramanan, Mahadev Satyanarayanan
GAMMA generates realistic motions for diverse 3D bodies using body surface markers, enhancing motion control and scene interaction.
Yan Zhang, Siyu Tang
MaskFeat introduces feature prediction via masked regions, achieving 86.7% top-1 accuracy on Kinetics-400 without supervision, outperforming previous methods.
Chen Wei, Haoqi Fan, Saining Xie et al.
Proposes an efficient geometry-aware 3D GAN with hybrid explicit-implicit architecture, enabling real-time high-res multi-view consistent image synthesis.
Eric R. Chan, Connor Z. Lin, Matthew A. Chan et al.