Flames: Benchmarking Value Alignment of LLMs in Chinese
Introduces FLAMES benchmark, evaluating 17 Chinese LLMs' value alignment; all models perform poorly, highlighting safety gaps.
Kexin Huang, Xiangyang Liu, Qianyu Guo et al.
Introduces FLAMES benchmark, evaluating 17 Chinese LLMs' value alignment; all models perform poorly, highlighting safety gaps.
Kexin Huang, Xiangyang Liu, Qianyu Guo et al.
Introduces LRP2 modules to significantly improve factual knowledge retrieval accuracy in multilingual models.
Shaoyang Xu, Junzhuo Li, Deyi Xiong
Proposes DARE, a method to merge homologous language models without retraining, achieving performance surpassing individual models on the Open LLM Leaderboard.
Le Yu, Bowen Yu, Haiyang Yu et al.
This study reveals vulnerabilities of BERTScore, BLEURT, and COMET under adversarial attacks, proposing methods to enhance their robustness.
Yichen Huang, Timothy Baldwin
ChipNeMo combines DAPT, instruction fine-tuning, and RAG to enhance chip design tasks, outperforming GPT-4 in key benchmarks.
Mingjie Liu, Teodor-Dumitru Ene, Robert Kirby et al.
FollowBench introduces a multi-level, fine-grained benchmark for LLM instruction following, covering five constraint types, revealing models' limitations at higher difficulty levels with detailed metrics.
Yuxin Jiang, Yufei Wang, Xingshan Zeng et al.
Convex function-based training sharpens output distribution, improving BLEU/ROUGE scores by over 9 points in text generation tasks.
Chenze Shao, Zhengrui Ma, Min Zhang et al.
Proposes MIN-K% PROB, a reference-free method for detecting pretraining data, improving 7.4% AUC on WIKIMIA.
Weijia Shi, Anirudh Ajith, Mengzhou Xia et al.
DP-Prompt combines pretrained LLMs and zero-shot prompting with exponential mechanism to achieve effective local differential privacy in text generation.
Saiteja Utpala, Sara Hooker, Pin Yu Chen
MuSR employs neurosymbolic algorithms to generate complex multi-step reasoning narratives, challenging GPT-4 and similar models.
Zayne Sprague, Xi Ye, Kaj Bostrom et al.
Introduces 'information value', a measure based on neural language models quantifying utterance predictability via multi-dimensional distances, outperforming token surprisal.
Mario Giulianelli, Sarenne Wallbridge, Raquel Fernández
AgentTuning enhances LLMs' generalized agent abilities; AgentLM-70B matches GPT-3.5-turbo on unseen tasks.
Aohan Zeng, Mingdao Liu, Rui Lu et al.
This study uses interpretability methods to analyze gender bias in instruction-tuned models, proposing a few-shot bias mitigation approach that significantly improves translation fairness.
Giuseppe Attanasio, Flor Miriam Plaza-del-Arco, Debora Nozza et al.
SPEED accelerates Transformer decoding via speculative execution, significantly reducing latency.
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh et al.
Self-RAG employs self-reflection to dynamically retrieve, generate, and critique, significantly improving factual accuracy and controllability.
Akari Asai, Zeqiu Wu, Yizhong Wang et al.
NeMo Guardrails toolkit enables controllable and safe LLM applications with programmable rails.
Traian Rebedea, Razvan Dinu, Makesh Sreedhar et al.
CLIN employs causal abstraction-based persistent memory and reflection to enable parameter-free continual learning, outperforming SOTA with 23-point gains in ScienceWorld.
Bodhisattwa Prasad Majumder, Bhavana Dalvi Mishra, Peter Jansen et al.
Proposes a metric to quantify verbosity bias in LLMs, finds GPT-4 prefers longer answers with bias value 0.328.
Keita Saito, Akifumi Wachi, Koki Wataoka et al.
Large Language Model Unlearning uses gradient ascent to reduce undesirable behaviors efficiently.
Yuanshun Yao, Xiaojun Xu, Yang Liu
This study quantifies gender bias in LLM-generated recommendation letters using content and style metrics, revealing significant biases in ChatGPT and Alpaca.
Yixin Wan, George Pu, Jiao Sun et al.