A Survey of Large Language Models
Survey of large language models revealing their powerful capabilities in NLP tasks.
Wayne Xin Zhao, Kun Zhou, Junyi Li et al.
Survey of large language models revealing their powerful capabilities in NLP tasks.
Wayne Xin Zhao, Kun Zhou, Junyi Li et al.
KALE uses post-training KL alignment and asymmetric pruning to accelerate dense retrieval models, outperforming DistilBERT with 3x faster inference.
Daniel Campos, Alessandro Magnani, ChengXiang Zhai
HuggingGPT uses ChatGPT as a controller to coordinate models from Hugging Face, solving complex multimodal AI tasks autonomously.
Yongliang Shen, Kaitao Song, Xu Tan et al.
GPT-4 demonstrates near-human multi-domain intelligence, excelling in mathematics, coding, vision, medicine, and law, indicating progress toward AGI.
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan et al.
Enhance LLMs' contextual faithfulness using opinion-based prompts and counterfactual demonstrations, significantly reducing memorization ratio.
Wenxuan Zhou, Sheng Zhang, Hoifung Poon et al.
GPT-4 surpasses USMLE passing scores by over 20 points without domain-specific tuning, demonstrating strong reasoning and calibration.
Harsha Nori, Nicholas King, Scott Mayer McKinney et al.
Using TextFlint across 21 datasets, the study finds RLHF improves human-like dialogue but not overall NLU accuracy or robustness.
Junjie Ye, Xuanting Chen, Nuo Xu et al.
Introduces MATH 401 dataset, evaluates GPT-4, ChatGPT on arithmetic tasks, showing GPT-4's superior performance with 83.54% accuracy.
Zheng Yuan, Hongyi Yuan, Chuanqi Tan et al.
ART enables automatic multi-step reasoning and tool use in LLMs, improving unseen task performance by over 22%.
Bhargavi Paranjape, Scott Lundberg, Sameer Singh et al.
SelfCheckGPT detects factual hallucinations via multi-sample response consistency without external data.
Potsawee Manakul, Adian Liusie, Mark J. F. Gales
GPT-4 is a multimodal Transformer trained with predictive and RLHF methods, achieving human-level performance and scalable predictability.
OpenAI, Josh Achiam, Steven Adler et al.
Proposes Hidden Markov Transformer (HMT), modeling translation start times as hidden states, achieving state-of-the-art SiMT performance.
Shaolei Zhang, Yang Feng
Almanac framework enhances clinical language models with retrieval capabilities, improving factuality by 18%.
Cyril Zakka, Akash Chaurasia, Rohan Shad et al.
GEMBA leverages GPT-3.5+ for state-of-the-art translation quality evaluation, outperforming existing metrics on WMT22 data.
Tom Kocmi, Christian Federmann
AugGPT leverages ChatGPT for data augmentation, significantly boosting few-shot text classification accuracy by generating diverse, semantically consistent samples.
Haixing Dai, Zhengliang Liu, Wenxiong Liao et al.
This study systematically evaluates GPT models' performance in multilingual machine translation, revealing strong results in high-resource languages but limitations in low-resource scenarios.
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf et al.
STREET benchmark evaluates multi-task reasoning and explanation; GPT-3 and T5 models lag behind human performance.
Danilo Ribeiro, Shen Wang, Xiaofei Ma et al.
Evaluated GPT-3 (text-davinci-003) on the SARA dataset for legal reasoning, using various prompting methods, achieving better results but with notable errors and knowledge gaps.
Andrew Blair-Stanek, Nils Holzenberger, Benjamin Van Durme
Toolformer enables self-supervised learning for language models to call external APIs, boosting zero-shot performance with only 6.7B parameters.
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì et al.
DSIR framework uses hashed n-gram importance resampling to efficiently select 100M documents, boosting downstream performance.
Sang Michael Xie, Shibani Santurkar, Tengyu Ma et al.