DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
DeepSeek-V2 is a 236B parameter MoE model using MLA and DeepSeekMoE, boosting inference speed and training efficiency.
DeepSeek-AI, Aixin Liu, Bei Feng et al.
DeepSeek-V2 is a 236B parameter MoE model using MLA and DeepSeekMoE, boosting inference speed and training efficiency.
DeepSeek-AI, Aixin Liu, Bei Feng et al.
This study evaluates GPT-3.5 Turbo and Gemini 1.5 Pro on Bangla NLI, outperforming traditional Transformer models in few-shot settings with accuracy up to 92%.
Fatema Tuj Johora Faria, Mukaffi Bin Moin, Asif Iftekher Fahim et al.
This review summarizes large language models' applications, methodologies, and ethical challenges in finance, healthcare, and law.
Zhiyu Zoey Chen, Jing Ma, Xinlu Zhang et al.
Analyzes five major generative language models' biases in educational stories, identifying types of representational harms and their psychological impacts.
Faye-Marie Vassel, Evan Shieh, Cassidy R. Sugimoto et al.
Prometheus 2 merges models trained on direct assessment and pairwise ranking, achieving top correlation with human and GPT-4 judgments, across multiple benchmarks.
Seungone Kim, Juyoung Suk, Shayne Longpre et al.
WildChat collected 1 million real user-chatbot interactions, revealing diverse use cases and toxicity issues, enabling model fine-tuning and safety analysis.
Wenting Zhao, Xiang Ren, Jack Hessel et al.
Proposes multi-token prediction to improve sample efficiency and inference speed; 13B model achieves 12% better problem-solving on code tasks.
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière et al.
Proposes Scaffold-BPE, a dynamic token removal method that alleviates frequency imbalance, improving large language model training and performance.
Haoran Lian, Yizhe Xiong, Jianwei Niu et al.
IndicGenBench evaluates multilingual generation across 29 Indic languages, using cross-lingual summarization, translation, and QA tasks.
Harman Singh, Nitish Gupta, Shikhar Bharadwaj et al.
LayerSkip combines layer dropout and early exit with self-speculative decoding, achieving up to 2.16× speedup in large language model inference.
Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich et al.
GraphRAG leverages hierarchical knowledge graphs and community detection to enable large-scale global understanding, outperforming traditional RAG by 25% on complex datasets.
Darren Edge, Ha Trinh, Newman Cheng et al.
SnapKV compresses KV cache by analyzing attention patterns from observation windows, enabling 3.6x faster decoding and handling 380K tokens without fine-tuning.
Yuhong Li, Yingbing Huang, Bowen Yang et al.
phi-3-mini is a 3.8B parameter model trained on 3.3T tokens, rivaling GPT-3.5, deployable on smartphones.
Marah Abdin, Jyoti Aneja, Hany Awadalla et al.
Proposed a multisense consistency framework based on Fregean sense to evaluate GPT-3.5’s semantic stability across five languages, revealing significant inconsistencies.
Xenia Ohmer, Elia Bruni, Dieuwke Hupkes
Language Ranker quantifies multilingual performance by comparing internal Transformer representations with an English baseline, revealing resource-based disparities.
Zihao Li, Yucheng Shi, Zirui Liu et al.
This study analyzes automatic metrics for meeting summarization, revealing their limited sensitivity to meeting-specific errors and model architecture differences.
Frederic Kirstein, Jan Philip Wahle, Terry Ruas et al.
MiniCheck achieves GPT-4 level fact-checking using synthetic data, reducing costs by 400x.
Liyan Tang, Philippe Laban, Greg Durrett
SPAG enhances LLM reasoning by 5% through self-playing adversarial language games.
Pengyu Cheng, Tianhao Hu, Han Xu et al.
Large language models like GPT-4 exhibit self-recognition, with a linear link to their self-preference bias, validated through fine-tuning experiments.
Arjun Panickssery, Samuel R. Bowman, Shi Feng
LLoCO combines offline context compression with LoRA fine-tuning, expanding a 4k model to 128k tokens, significantly improving long document QA performance.
Sijun Tan, Xiuyu Li, Shishir Patil et al.