Qwen Technical Report
Qwen series models excel in multiple tasks, notably Qwen-Chat with RLHF technology.
Jinze Bai, Shuai Bai, Yunfei Chu et al.
Qwen series models excel in multiple tasks, notably Qwen-Chat with RLHF technology.
Jinze Bai, Shuai Bai, Yunfei Chu et al.
Proposes a benchmark-based model routing framework using binary classifiers to improve large language model selection, achieving significant performance gains across 29 datasets.
Tal Shnitzer, Anthony Ou, Mírian Silva et al.
Ragas uses multi-dimensional, prompt-based, model-driven metrics for reference-free RAG evaluation.
Shahul Es, Jithin James, Luis Espinosa-Anke et al.
This study evaluates GPT-3.5, LLaMA-2-13b-chat, and Claude-2 on seven Bengali NLP tasks in zero-shot settings, revealing significant performance gaps compared to fine-tuned models.
Mohsinul Kabir, Mohammed Saidul Islam, Md Tahmid Rahman Laskar et al.
MINT evaluates LLMs' multi-turn tool use and feedback leveraging, with performance gains of 2-17% across 20 models, revealing training impacts and evaluation gaps.
Xingyao Wang, Zihan Wang, Jiateng Liu et al.
Adding 3% safety examples to LLaMA significantly improves safety without reducing capability.
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio et al.
C-Pack introduces C-MTEB, C-MTP, and BGE, significantly advancing Chinese text embedding performance.
Shitao Xiao, Zheng Liu, Peitian Zhang et al.
Analyzing training trajectories reveals that Syntactic Attention Structure (SAS) causes phase transitions, crucial for grammar acquisition in MLMs.
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho et al.
SafetyBench evaluates LLM safety; GPT-4 excels.
Zhexin Zhang, Leqi Lei, Lindong Wu et al.
MAmmoTH models improve math reasoning by 16%-32% using hybrid instruction tuning.
Xiang Yue, Xingwei Qu, Ge Zhang et al.
SignRound optimizes LLM quantization using SignSGD, achieving 6.91%-33.22% accuracy improvement at 2 bits.
Wenhua Cheng, Weiwei Zhang, Haihao Shen et al.
Controlled experiment shows InstructGPT significantly reduces content diversity, increasing similarity by ~8% and decreasing lexical diversity by ~15%.
Vishakh Padmakumar, He He
MADLAD-400 is a manually audited multilingual monolingual dataset covering 419 languages, enabling high-quality pretraining for multilingual models.
Sneha Kudugunta, Isaac Caswell, Biao Zhang et al.
This paper analyzes how detoxification methods (fine-tuning and RLHF) influence language models' prompt dependence using attribution entropy metrics.
Daniel Scalena, Gabriele Sarti, Malvina Nissim et al.
RLAIF uses AI feedback instead of human feedback, achieving performance comparable to RLHF.
Harrison Lee, Samrat Phatale, Hassan Mansoor et al.
LongBench is the first bilingual, multi-task benchmark for long context understanding, with 21 datasets averaging 6,711 words, advancing long-sequence model evaluation.
Yushi Bai, Xin Lv, Jiajie Zhang et al.
Introduces SWIE and OVERMISS to improve translation faithfulness, achieving significant BLEU and trustworthiness gains.
Yijie Chen, Yijin Liu, Fandong Meng et al.
Introduces the Instruction-Following Difficulty (IFD) metric for self-guided data selection, boosting LLM instruction tuning with only 10% data, outperforming full-data models.
Ming Li, Yong Zhang, Zhitao Li et al.
This study analyzes LLMs' sensitivity to option order in MCQs, revealing up to 75% performance variation, and proposes calibration strategies to improve robustness.
Pouya Pezeshkpour, Estevam Hruschka
Graph of Thoughts (GoT) models reasoning as arbitrary graphs, boosting sorting quality by 62% and reducing costs by 31%, surpassing ToT.
Maciej Besta, Nils Blach, Ales Kubicek et al.