LongNet: Scaling Transformers to 1,000,000,000 Tokens
LongNet introduces dilated attention, enabling linear complexity for over 1 billion tokens, supporting scalable long-sequence modeling.
Jiayu Ding, Shuming Ma, Li Dong et al.
LongNet introduces dilated attention, enabling linear complexity for over 1 billion tokens, supporting scalable long-sequence modeling.
Jiayu Ding, Shuming Ma, Li Dong et al.
This paper evaluates two belief probing methods in LLMs, finds poor generalization, and discusses philosophical and empirical issues surrounding model beliefs.
B. A. Levinstein, Daniel A. Herrmann
This paper introduces a quantitative framework using the GlobalOpinionQA dataset to evaluate how well language models represent diverse global opinions, revealing biases and stereotypes.
Esin Durmus, Karina Nguyen, Thomas I. Liao et al.
Position Interpolation (PI) extends RoPE-based LLaMA models to 32768 tokens with minimal fine-tuning, greatly improving long-context tasks.
Shouyuan Chen, Sherman Wong, Liangjian Chen et al.
InterCode: Standardizing interactive coding with execution feedback to enhance code generation.
John Yang, Akshara Prabhakar, Karthik Narasimhan et al.
KOSMOS-2 is a multimodal large language model with grounding and referring capabilities, achieving state-of-the-art performance on spatial understanding and vision-language tasks.
Zhiliang Peng, Wenhui Wang, Li Dong et al.
ToolQA employs an automated three-phase process to evaluate LLMs' external tool usage across 8 domains, showing significant performance gains over internal knowledge-only models.
Yuchen Zhuang, Yue Yu, Kuan Wang et al.
Proposes a black-box confidence estimation framework combining prompting, sampling, and aggregation, improving calibration and failure prediction metrics.
Miao Xiong, Zhiyuan Hu, Xinyang Lu et al.
LMFlow is a lightweight toolkit for efficient fine-tuning and inference of large foundation models.
Shizhe Diao, Rui Pan, Hanze Dong et al.
Wanda uses input activation and weight magnitude product for training-free pruning, outperforming magnitude pruning significantly.
Mingjie Sun, Zhuang Liu, Anna Bair et al.
Block-State Transformer combines SSM and Block Transformer, improving language modeling performance and speed.
Mahan Fathi, Jonathan Pilault, Orhan Firat et al.
MiniLLM uses reverse-KL on-policy distillation; LLaMA-7B reaches 73.1 GPT-4 and 23.2 Rouge-L on SelfInst.
Yuxian Gu, Li Dong, Furu Wei et al.
Proposes three frameworks—KG-enhanced LLMs, LLM-augmented KGs, and synergized models—to improve knowledge reasoning.
Shirui Pan, Linhao Luo, Yufei Wang et al.
WebGLM integrates GLM and web search, enhancing QA efficiency; 10B model outperforms 13B WebGPT.
Xiao Liu, Hanyu Lai, Hao Yu et al.
Mind2Web is a large-scale dataset with 137 real websites and 2000+ tasks, enabling web task automation via combined small and large language models.
Xiang Deng, Yu Gu, Boyuan Zheng et al.
PandaLM is an automatic evaluation benchmark for LLM instruction tuning, achieving 93.75% GPT-3.5 evaluation capability with multi-dimensional subjective metrics.
Yidong Wang, Zhuohao Yu, Zhengran Zeng et al.
CMExam is the first Chinese medical exam dataset with comprehensive annotations; GPT-4 achieved 61.6% accuracy on it.
Junling Liu, Peilin Zhou, Yining Hua et al.
Whisper outperforms XLS-R in diverse Arabic conditions but struggles with unseen dialects.
Bashar Talafha, Abdul Waheed, Muhammad Abdul-Mageed
Video-LLaMA integrates Video Q-former and Audio Q-former for audio-visual understanding.
Hang Zhang, Xin Li, Lidong Bing
Proposes Fine-Grained RLHF, using dense, multi-category rewards to improve language model training, reducing toxicity and factual errors.
Zeqiu Wu, Yushi Hu, Weijia Shi et al.