API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
API-Bank benchmark evaluates tool-augmented LLMs with 73 APIs, boosting tool utilization by training and testing models like Lynx, GPT-3.5, GPT-4.
Minghao Li, Yingxiu Zhao, Bowen Yu et al.
API-Bank benchmark evaluates tool-augmented LLMs with 73 APIs, boosting tool utilization by training and testing models like Lynx, GPT-3.5, GPT-4.
Minghao Li, Yingxiu Zhao, Bowen Yu et al.
This study evaluates pre-trained LLMs in task-oriented dialogue, showing limited belief state tracking but effective dialogue guidance, improved by true belief states and few-shot examples.
Vojtěch Hudeček, Ondřej Dušek
Proposed certified black-box defense using Robust UNet denoiser with 35% improvement on CIFAR-10.
Astha Verma, A V Subramanyam, Siddhesh Bangar et al.
AGIEval benchmark assesses foundation models on human-standard exams; GPT-4 surpasses average human scores with 95% in SAT Math and 92.5% in Chinese college entrance English.
Wanjun Zhong, Ruixiang Cui, Yiduo Guo et al.
InterGen employs diffusion models with cooperative transformers and a large-scale multimodal dataset to generate diverse two-person interaction motions guided by text.
Han Liang, Wenqian Zhang, Wenxuan Li et al.
Evaluates ChatGPT's zero-shot performance across 37 languages and 7 tasks, revealing limitations in multilingual capabilities.
Viet Dac Lai, Nghia Trung Ngo, Amir Pouran Ben Veyseh et al.
L3MVN leverages large language models for zero-shot visual navigation, achieving success rates over 76% on Gibson and 50% on HM3D datasets.
Bangguo Yu, Hamidreza Kasaei, Ming Cao
Single-gate MoE achieves comparable efficiency and accuracy to complex models, outperforming non-mixture baselines.
Amelie Royer, Ilia Karmanov, Andrii Skliar et al.
Autonomous AI agent using GPT-4 for complex chemical synthesis planning and execution, achieving >85% success in drug synthesis tasks.
Daniil A. Boiko, Robert MacKnight, Gabe Gomes
This survey systematically reviews deep graph representation learning architectures, paradigms, and applications, highlighting GNN innovations and future challenges.
Wei Ju, Zheng Fang, Yiyang Gu et al.
StillFast introduces an end-to-end model for short-term object interaction anticipation, achieving 13.29% Top-5 mAP on EGO4D v2, surpassing SOTA.
Francesco Ragusa, Giovanni Maria Farinella, Antonino Furnari
Proposes GeneRec, integrating generative models and user instructions for personalized content creation, boosting recommendation diversity.
Wenjie Wang, Xinyu Lin, Fuli Feng et al.
Proposes AMS-DRL for drone navigation in multi-pursuer pursuit-evasion game, achieving over 85% success rate in simulations.
Jiaping Xiao, Mir Feroskhan
A novel architecture combining large language models with dynamic memory, reflection, and planning enables believable human-like agent behavior with long-term coherence.
Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai et al.
Using dynamical mean field theory, this paper quantifies finite-width neural network kernel and prediction fluctuations, revealing how feature learning dynamically reduces variance.
Blake Bordelon, Cengiz Pehlevan
Combining Chinchilla scaling laws and μP, trained models from 111M to 13B parameters achieving state-of-the-art compute efficiency.
Nolan Dey, Gurpreet Gosal, Zhiming et al.
Proposes a three-step prompting method using GPT-3 for zero-shot next-item recommendation, achieving HR@10 of 0.1187 on MovieLens 100K.
Lei Wang, Ee-Peng Lim
A novel algorithm based on convex body chasing (CBC) achieves online stabilization of unknown linear time-varying systems, ensuring BIBO stability.
Jing Yu, Varun Gupta, Adam Wierman
SAM model supports prompt-based image segmentation with over 1 billion masks, achieving state-of-the-art zero-shot performance.
Alexander Kirillov, Eric Mintun, Nikhila Ravi et al.
GINA-3D generates tri-plane implicit assets from Waymo sensor data, achieving FID 59.5 on WOD-Vehicle.
Bokui Shen, Xinchen Yan, Charles R. Qi et al.