Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions
Developed JAMA Clinical Challenge and Medbullets datasets to evaluate 7 LLMs on complex medical QA.
Hanjie Chen, Zhouxiang Fang, Yash Singla et al.
Developed JAMA Clinical Challenge and Medbullets datasets to evaluate 7 LLMs on complex medical QA.
Hanjie Chen, Zhouxiang Fang, Yash Singla et al.
BitNet b1.58 uses ternary quantization to achieve 1.58-bit LLMs, matching FP16 performance with lower costs.
Shuming Ma, Hongyu Wang, Lingxiao Ma et al.
This study introduces a new benchmark for continual pretraining of large language models, analyzing effects of model size and domain order on knowledge retention and transfer.
Çağatay Yıldız, Nishaanth Kanna Ravichandran, Nitin Sharma et al.
This survey reviews data selection methods for language models, emphasizing distribution matching and diversity, unified under a probabilistic framework across training stages.
Alon Albalak, Yanai Elazar, Sang Michael Xie et al.
Proposes temporal alignment of LLaMa2 via finetuning and prompting, achieving up to 62% performance boost on 2022 data.
Bowen Zhao, Zander Brumbaugh, Yizhong Wang et al.
SelectIT leverages LLM's intrinsic uncertainty through multi-granularity self-reflection to select high-quality instruction tuning data, boosting model performance.
Liangxin Liu, Xuebo Liu, Derek F. Wong et al.
This study systematically reviews LLMs like GPT and BERT in health, urban planning, climate, and disaster management, highlighting their transformative potential.
Pravneet Kaur, Gautam Siddharth Kashyap, Ankit Kumar et al.
Proposed CDD detects data contamination via output distribution peakedness, improving detection accuracy by 21.8%-30.2%.
Yihong Dong, Xue Jiang, Huanyu Liu et al.
Proposes tinyBenchmarks using 100 examples with IRT and clustering to estimate LLM performance on benchmarks within 2% error.
Felipe Maia Polo, Lucas Weber, Leshem Choshen et al.
CriticBench benchmarks 17 LLMs' critique and correction abilities across five reasoning domains, revealing linear relationships and size-dependent knowledge consistency.
Zicheng Lin, Zhibin Gou, Tian Liang et al.
LLMBind framework integrates multimodal tasks with dual-pathway mechanism, achieving superior performance and expandability.
Bin Zhu, Munan Ning, Peng Jin et al.
OlympiadBench, with 8,476 high-level bilingual scientific problems, evaluates GPT-4V's reasoning, scoring 17.97%, highlighting current AI limitations in complex science tasks.
Chaoqun He, Renjie Luo, Yuzhuo Bai et al.
This study introduces a GPT-4-based multi-stage data annotation framework, covering generation, assessment, and application, significantly improving efficiency.
Zhen Tan, Dawei Li, Song Wang et al.
Introduced MIRAGE benchmark and MedRAG toolkit, boosting GPT-3.5 accuracy to 71.57%, approaching GPT-4 performance.
Guangzhi Xiong, Qiao Jin, Zhiyong Lu et al.
Proposes restart-incremental Transformer analysis, revealing internal state updates during local ambiguity resolution.
Brielen Madureira, Patrick Kahardipraja, David Schlangen
Introduces FLenQA dataset to analyze input length effects on LLM reasoning, showing performance drops at much shorter lengths than maximum capacity.
Mosh Levy, Alon Jacoby, Yoav Goldberg
Erasing shortcut neurons significantly reduces multi-hop knowledge editing failures and improves reasoning reliability.
Tianjie Ju, Yijin Chen, Xinwei Yuan et al.
ALLaVA uses GPT-4V to generate 1.3M high-quality vision-language samples, enabling lightweight models to match large-scale model performance.
Guiming Hardy Chen, Shunian Chen, Ruifei Zhang et al.
BitDistiller enhances sub-4-bit LLM performance via self-distillation, significantly reducing data and training resource needs.
Dayou Du, Yijia Zhang, Shijie Cao et al.
Proposes a data engineering strategy using 80K sequence continual pretraining to extend language models' context to 128K, achieving near GPT-4 performance.
Yao Fu, Rameswar Panda, Xinyao Niu et al.