Phi-4 Technical Report
phi-4 is a 14B parameter model leveraging synthetic data, surpassing GPT-4 in STEM QA, with innovative training strategies.
Marah Abdin, Jyoti Aneja, Harkirat Behl et al.
phi-4 is a 14B parameter model leveraging synthetic data, surpassing GPT-4 in STEM QA, with innovative training strategies.
Marah Abdin, Jyoti Aneja, Harkirat Behl et al.
Coconut leverages continuous latent space reasoning to outperform CoT in logical tasks like ProsQA, achieving higher accuracy and efficiency.
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su et al.
Proposes Gated DeltaNet, combining gating and delta rules, significantly enhancing long-sequence task performance.
Songlin Yang, Jan Kautz, Ali Hatamizadeh
This paper systematically reviews LLMs-as-judges, covering functionality, methodologies, applications, meta-evaluation, and limitations, highlighting innovations and challenges.
Haitao Li, Qian Dong, Junjie Chen et al.
Aya Expanse combines multi-model data arbitration, preference tuning, and merging, achieving top performance in 8B/32B multilingual models, surpassing state-of-the-art open models.
John Dang, Shivalika Singh, Daniel D'souza et al.
Relign framework uses reliability alignment to reduce tool hallucinations, boosting task success and efficiency.
Hongshen Xu, Zichen Zhu, Lei Pan et al.
Survey on social simulation driven by large language model-based agents, categorized into individual, scenario, and society simulations.
Xinyi Mou, Xuanwen Ding, Qi He et al.
Proposes Pragmatic Metacognitive Prompting (PMP), integrating pragmatics and reflection, boosting LLM sarcasm detection; GPT-4o achieves SOTA on MUStARD and SemEval2018 with +20% accuracy.
Joshua Lee, Wyatt Fong, Alexander Le et al.
MiniKV achieves 86% KV cache compression with 98.5% accuracy via 2-bit layer-discriminative KV cache.
Akshat Sharma, Hangliang Ding, Jianping Li et al.
Proposed a multimodal VQA framework integrating external knowledge and reasoning, achieving 78.5% accuracy on VQA 2.0.
Jiayi Kuang, Jingyou Xie, Haohao Luo et al.
Defines LLM-as-a-Judge with focus on reliability, bias mitigation, and multi-scenario adaptation, proposing a standardized evaluation framework.
Jiawei Gu, Xuhui Jiang, Zhichao Shi et al.
XGrammar accelerates CFG-based structured decoding with adaptive caching and persistent stacks, reaching up to 100x lower latency.
Yixin Dong, Charlie F. Ruan, Yaxing Cai et al.
Introduces MorphScore to evaluate morphological alignment across 22 languages; finds dataset size and encoding efficiency outweigh morphological complexity in influencing model performance.
Catherine Arnett, Benjamin K. Bergen
Hymba employs a hybrid-head architecture combining transformer attention and state space models, achieving superior efficiency and performance with fewer parameters, surpassing comparable small models.
Xin Dong, Yonggan Fu, Shizhe Diao et al.
Analyzed 100,000 user logs to reveal domain preferences and translation directions in Tetun low-resource MT, emphasizing community needs.
Raphael Merx, Adérito José Guterres Correia, Hanna Suominen et al.
Proposes RandSymKL, combining symmetric KL and cross-entropy, to mitigate extrinsic gender bias in Bangla classification tasks, achieving 90.66% accuracy and bias reduction.
Sajib Kumar Saha Joy, Arman Hassan Mahy, Meherin Sultana et al.
Using large language models as user agents to evaluate task-oriented dialogue systems, enhancing diversity and task completion rates.
Taaha Kazi, Ruiliang Lyu, Sizhe Zhou et al.
Spider 2.0 framework evaluates language models on enterprise text-to-SQL workflows, solving only 21.3% of tasks.
Fangyu Lei, Jixuan Chen, Yuxiao Ye et al.
Contextualized evaluations synthesize context to assess language model responses to underspecified queries, significantly altering evaluation conclusions.
Chaitanya Malaviya, Joseph Chee Chang, Dan Roth et al.
Introduces continual memorization framework; REMIX data mixing significantly reduces language model forgetting.
Howard Chen, Jiayi Geng, Adithya Bhaskar et al.