TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs
Proposes TaskMatrix.AI, integrating foundation models with millions of APIs for multi-task, multi-modal automation.
Yaobo Liang, Chenfei Wu, Ting Song et al.
Proposes TaskMatrix.AI, integrating foundation models with millions of APIs for multi-task, multi-modal automation.
Yaobo Liang, Chenfei Wu, Ting Song et al.
ARMBench is a large-scale, object-centric dataset for warehouse robotic manipulation, with 235K+ activities and 190K+ objects, supporting segmentation, recognition, and defect detection.
Chaitanya Mitash, Fan Wang, Shiyang Lu et al.
Proposes Diffusion Classifier, leveraging diffusion model density estimates for zero-shot image classification, outperforming traditional discriminative methods.
Alexander C. Li, Mihir Prabhudesai, Shivam Duggal et al.
LLaMA-Adapter uses zero-initialized attention with only 1.2M parameters to achieve instruction tuning comparable to full fine-tuning.
Renrui Zhang, Jiaming Han, Chris Liu et al.
Introducing a global error correction algorithm in machine-learned PDE solvers to preserve discrete invariants, enhancing stability and accuracy at large time steps.
Nick McGreivy, Ammar Hakim
CARTO achieves category- and joint-agnostic 3D reconstruction from a single stereo image, improving mAP 3D IOU50 by 20.4%.
Nick Heppert, Muhammad Zubair Irshad, Sergey Zakharov et al.
Proposes Anti-DreamBooth using adversarial noise to disrupt personalized diffusion model training, reducing fake image generation by over 85%.
Thanh Van Le, Hao Phung, Thuan Hoang Nguyen et al.
EVA-CLIP combines pre-trained EVA, LAMB optimizer, and masking techniques to achieve 82.0% zero-shot ImageNet accuracy with 5.0B parameters, trained on 9B samples.
Quan Sun, Yuxin Fang, Ledell Wu et al.
WinCLIP achieves 91.8% AUROC in zero-shot anomaly classification on MVTec-AD.
Jongheon Jeong, Yang Zou, Taewan Kim et al.
Proposes OKDPH, a parameter hybridization framework, to flatten loss minima and enhance online knowledge distillation generalization.
Tianli Zhang, Mengqi Xue, Jiangtao Zhang et al.
YOSO introduces a real-time panoptic segmentation framework using dynamic convolution for unified instance and semantic masks, achieving 46.4 PQ and 45.6 FPS on COCO.
Jie Hu, Linyan Huang, Tianhe Ren et al.
Chat-Rec enhances recommendation systems' interactivity and explainability by converting user data into prompts.
Yunfan Gao, Tao Sheng, Youlin Xiang et al.
DBARF self-supervises camera pose optimization using implicit cost functions, enhancing GeNeRF generalization.
Yu Chen, Gim Hee Lee
Make-It-3D leverages 2D diffusion priors with NeRF in a two-stage pipeline to produce high-fidelity 3D models from a single image, outperforming prior methods.
Junshu Tang, Tengfei Wang, Bo Zhang et al.
DreamBooth3D combines DreamBooth and DreamFusion to generate personalized 3D models from 3-6 images, achieving high-quality, editable assets.
Amit Raj, Srinivas Kaza, Ben Poole et al.
Zero-shot text-to-video synthesis leveraging Stable Diffusion with motion encoding and cross-frame attention, achieving high temporal consistency.
Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan et al.
Proposes domain alignment and spatio-temporal regularization for knowledge transfer, boosting SNN generalization on event data.
Xiang He, Dongcheng Zhao, Yang Li et al.
GPT-4 demonstrates near-human multi-domain intelligence, excelling in mathematics, coding, vision, medicine, and law, indicating progress toward AGI.
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan et al.
This study compares asexual/sexual brain reproduction within Darwinian/Lamarckian frameworks, finding asexual+Lamarckian yields best robot performance.
Jie Luo, Carlo Longhi, Agoston E. Eiben
TextKG method enhances video captioning with knowledge graphs, achieving an 18.7% CIDEr score improvement on the YouCookII dataset.
Xin Gu, Guang Chen, Yufei Wang et al.