Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP
Proposes mask-adapted CLIP with mask prompt tuning, achieving 29.6% mIoU on ADE20K-150, outperforming state-of-the-art by 8.5%.
Feng Liang, Bichen Wu, Xiaoliang Dai et al.
Proposes mask-adapted CLIP with mask prompt tuning, achieving 29.6% mIoU on ADE20K-150, outperforming state-of-the-art by 8.5%.
Feng Liang, Bichen Wu, Xiaoliang Dai et al.
TEG-Track combines tactile and visual data to enhance 6D pose tracking of unseen objects, reducing rotation error by 30.9%.
Yun Liu, Xiaomeng Xu, Weihang Chen et al.
Proposes RA-VQA, an end-to-end framework combining differentiable DPR and answer generation, achieving 54.48 VQA score on OK-VQA.
Weizhe Lin, Bill Byrne
Auto-CoT automatically generates diverse reasoning demonstrations, leveraging clustering and prompting to outperform manual methods in reasoning tasks.
Zhuosheng Zhang, Aston Zhang, Mu Li et al.
The study narrows the compositionality gap in language models using the self-ask method, improving accuracy on GPT-3.
Ofir Press, Muru Zhang, Sewon Min et al.
Pix2Struct pretrains by parsing web screenshots into simplified HTML, achieving state-of-the-art results in six tasks.
Kenton Lee, Mandar Joshi, Iulia Turc et al.
Proposes content-based deep generative model retrieval using multi-modal feature contrastive learning and probability optimization.
Daohan Lu, Sheng-Yu Wang, Nupur Kumari et al.
This study analyzes added toxicity in multilingual MT using HOLISTICBIAS, ALTI+ attribution, revealing low-resource languages and demographic axes prone to toxicity.
Marta R. Costa-jussà, Eric Smith, Christophe Ropers et al.
metaPNS employs Bayesian meta-learning with set-conditioned generative models to achieve personalized cardiac surrogates with few samples, improving accuracy and efficiency.
Xiajun Jiang, Zhiyuan Li, Ryan Missel et al.
Binder combines GPT-3 Codex in a training-free neural-symbolic framework, achieving state-of-the-art results in complex question answering with minimal examples.
Zhoujun Cheng, Tianbao Xie, Peng Shi et al.
Proposed a unified hard-constraint framework to solve geometrically complex PDEs, significantly improving accuracy.
Songming Liu, Zhongkai Hao, Chengyang Ying et al.
DexGraspNet, a large-scale dataset of 1.32 million diverse, physics-validated dexterous grasps generated via a deep-accelerated optimization method for ShadowHand.
Ruicheng Wang, Jialiang Zhang, Jiayi Chen et al.
Decomposed Prompting employs modular task decomposition, improving GPT-3 few-shot performance on complex reasoning tasks.
Tushar Khot, Harsh Trivedi, Matthew Finlayson et al.
Phenaki generates variable-length videos from text using causal attention and a bidirectional masked transformer.
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans et al.
Bayesian Prompt Learning regularizes prompt distributions through variational inference and improves unseen-prompt generalization across 15 benchmarks.
Mohammad Mahdi Derakhshani, Enrique Sanchez, Adrian Bulat et al.
Proposes FAST model combining latent factors and sparse nonparametric regression via deep ReLU networks, achieving minimax optimal rates in high dimensions.
Jianqing Fan, Yihong Gu
HULC++ combines visual affordance models with hierarchical language-conditioned policies, achieving 65.2% success on CALVIN with minimal annotations.
Oier Mees, Jessica Borja-Diaz, Wolfram Burgard
DiffDock employs a diffusion generative model for molecular docking, achieving a 38% top-1 success rate (RMSD<2Å), surpassing traditional and deep learning methods.
Gabriele Corso, Hannes Stärk, Bowen Jing et al.
VICRegL combines global and local feature learning, boosting detection and segmentation while maintaining classification accuracy.
Adrien Bardes, Jean Ponce, Yann LeCun
Introduces RL4LMs library, GRUE benchmark, and NLPO algorithm, improving language model preference alignment.
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley et al.