In-Context Learning Unlocked for Diffusion Models
Prompt Diffusion enables in-context learning in diffusion models, trained on six vision-language tasks for strong generalization.
Zhendong Wang, Yifan Jiang, Yadong Lu et al.
Prompt Diffusion enables in-context learning in diffusion models, trained on six vision-language tasks for strong generalization.
Zhendong Wang, Yifan Jiang, Yadong Lu et al.
Proposes CoMOT with meta-OT and GumMS for fast, generalizable fair re-ranking, outperforming BvND in speed and storage.
Andrés Hoyos-Idrobo
mPLUG-Owl enhances large language models' multimodal capabilities through modular learning, significantly improving instruction and visual understanding.
Qinghao Ye, Haiyang Xu, Guohai Xu et al.
Introduces DataComp benchmark, utilizing 128B image-text pairs with filtering strategies, achieving 79.2% ImageNet zero-shot accuracy with ViT-L/14, outperforming prior datasets.
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang et al.
Citizen Science recruited 98 volunteers to re-annotate 1481 instances in NLP, achieving high-quality results comparable to expert labels.
Jan-Christoph Klie, Ji-Ung Lee, Kevin Stowe et al.
Proposes Evol-Instruct for automatic multi-level instruction generation, significantly boosting LLaMA's performance.
Can Xu, Qingfeng Sun, Kai Zheng et al.
TGNN combines message passing and graph kernel modules with consistency loss for semi-supervised graph classification, outperforming baselines.
Wei Ju, Xiao Luo, Meng Qu et al.
Predicting memorization behavior in large language models to reduce sensitive data retention.
Stella Biderman, USVSN Sai Prashanth, Lintang Sutawika et al.
Using natural logic intermediate features, the study applies Amnesic and Mnestic interventions to analyze causal roles in high-dimensional model representations, revealing limitations and improvements.
Julia Rozanova, Marco Valentino, Lucas Cordeiro et al.
Neurosymbolic models combine symbolic programs and neural networks for 2D/3D shape and material generation.
Daniel Ritchie, Paul Guerrero, R. Kenny Jones et al.
Proposes citation recall and precision metrics to evaluate verifiability; finds only 51.5% sentences fully supported in top search engines.
Nelson F. Liu, Tianyi Zhang, Percy Liang
Behavior Expectation Bounds (BEB) framework reveals fundamental limits of LLM alignment against adversarial prompts.
Yotam Wolf, Noam Wies, Oshri Avnery et al.
AMT achieves efficient video frame interpolation using bidirectional pixel correlations and multi-field transforms, enhancing PSNR.
Zhen Li, Zuo-Liang Zhu, Ling-Hao Han et al.
NeuralField-LDM employs hierarchical latent diffusion models to generate complex, realistic 3D scenes, outperforming state-of-the-art methods with significant improvements in FID and FVD scores.
Seung Wook Kim, Bradley Brown, Kangxue Yin et al.
Proposes behavior retrieval using learned similarity metrics to select relevant offline behaviors, boosting robot imitation learning by over 20%.
Maximilian Du, Suraj Nair, Dorsa Sadigh et al.
Estimate joint probability distribution using low-rank tensor decomposition and Radon transforms, significantly reducing sample complexity.
Pranava Singhal, Waqar Mirza, Ajit Rajwade et al.
VRB leverages human videos to learn visual affordances, enabling multi-task robot manipulation via multi-modal contact and trajectory prediction.
Shikhar Bahl, Russell Mendonca, Lili Chen et al.
LLaVA: leveraging GPT-4 generated multimodal instruction data, achieves 85.1% relative score, with 92.53% accuracy on Science QA, pioneering multimodal instruction tuning.
Haotian Liu, Chunyuan Li, Qingyang Wu et al.
Proposes a multi-view transformer-based method with epipolar sampling for high-quality novel view synthesis from a single wide-baseline stereo pair, outperforming prior sparse observation approaches.
Yilun Du, Cameron Smith, Ayush Tewari et al.
VALOR model achieves tri-modal learning through MGA and MGC tasks for vision, audio, and language.
Jing Liu, Sihan Chen, Xingjian He et al.