Language-driven Semantic Segmentation
LSeg achieves 52.3% mIoU in zero-shot segmentation on PASCAL-5i, outperforming existing methods.
Boyi Li, Kilian Q. Weinberger, Serge Belongie et al.
LSeg achieves 52.3% mIoU in zero-shot segmentation on PASCAL-5i, outperforming existing methods.
Boyi Li, Kilian Q. Weinberger, Serge Belongie et al.
Proposes ITSA, an information-theoretic regularization, to automatically restrict shortcut features, improving domain generalization of stereo networks.
WeiQin Chuah, Ruwan Tennakoon, Reza Hoseinnezhad et al.
Studied 'grokking' phenomenon in small algorithm datasets; regularization (like weight decay) accelerates sudden generalization after overfitting.
Alethea Power, Yuri Burda, Harri Edwards et al.
Derived asymptotic differential equations describe how shallow nonlinear autoencoders sequentially learn principal components in high-dimensional data.
Maria Refinetti, Sebastian Goldt
Proposes ECOLog, an efficient and statistically optimal Logistic Bandit algorithm, achieving O(d^2) per-round complexity with regret matching the lower bound.
Louis Faury, Marc Abeille, Kwang-Sung Jun et al.
Proposes FedLRGD leveraging data smoothness for lower federated oracle complexity than FedAve.
Ali Jadbabaie, Anuran Makur, Devavrat Shah
Proposed C2-CRS employs coarse-to-fine contrastive learning to fuse multi-type external data for improved conversational recommendation.
Yuanhang Zhou, Kun Zhou, Wayne Xin Zhao et al.
Proposes translation invariant Sinkhorn and 1-D Frank-Wolfe algorithms for faster unbalanced optimal transport, demonstrating over 2x speedup in experiments.
Thibault Séjourné, François-Xavier Vialard, Gabriel Peyré
ReferFormer uses Transformer with language as queries, achieving 55.6 J&F on Ref-Youtube-VOS, surpassing previous SOTA by 8.4 points, enabling end-to-end video object segmentation and tracking.
Jiannan Wu, Yi Jiang, Peize Sun et al.
SurfGen employs differentiable spherical projection and adversarial training to directly optimize surface geometry, producing diverse high-fidelity 3D shapes.
Andrew Luo, Tianqin Li, Wen-Hao Zhang et al.
StyleGAN-V extends StyleGAN2 for continuous video generation, achieving high quality at 1024×1024 resolution with 30% better FVD scores and arbitrary length/frame rate.
Ivan Skorokhodov, Sergey Tulyakov, Mohamed Elhoseiny
Proposes AIS framework with two-stage human annotation to evaluate if NLG outputs are source-supported, applicable across multiple tasks.
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm et al.
SLIP combines self-supervised contrastive learning with CLIP-style language-image pretraining, boosting ImageNet linear accuracy by 8.1%.
Norman Mu, Alexander Kirillov, David Wagner et al.
OpenSeg uses image-level captions to achieve open-vocabulary segmentation, improving mIoU by 19.9 points.
Golnaz Ghiasi, Xiuye Gu, Yin Cui et al.
GOAL uses GNet and MNet to generate full-body 4D motions for grasping unseen objects, outperforming baselines with high realism and generalization.
Omid Taheri, Vasileios Choutas, Michael J. Black et al.
Proposes an unsupervised multi-view learning framework combining explicit ellipsoid and implicit neural representations to discover 3D joints across unseen categories, achieving high-precision re-posing.
Atsuhiro Noguchi, Umar Iqbal, Jonathan Tremblay et al.
Mega-NeRF achieves scalable NeRFs for large scenes with 3x faster training and 12% PSNR improvement.
Haithem Turki, Deva Ramanan, Mahadev Satyanarayanan
Proving theorems using Incremental Learning and Hindsight Experience Replay surpasses traditional provers on the TPTP dataset.
Eser Aygün, Laurent Orseau, Ankit Anand et al.
GAMMA generates realistic motions for diverse 3D bodies using body surface markers, enhancing motion control and scene interaction.
Yan Zhang, Siyu Tang
MaskFeat introduces feature prediction via masked regions, achieving 86.7% top-1 accuracy on Kinetics-400 without supervision, outperforming previous methods.
Chen Wei, Haoqi Fan, Saining Xie et al.