Exploring Length Generalization in Large Language Models
Combining pretrained large models' in-context learning with scratchpad prompting significantly enhances length generalization.
Cem Anil, Yuhuai Wu, Anders Andreassen et al.
Combining pretrained large models' in-context learning with scratchpad prompting significantly enhances length generalization.
Cem Anil, Yuhuai Wu, Anders Andreassen et al.
WebShop employs RL and imitation learning on a dataset of 1.18 million products, achieving 29% success in complex web tasks, surpassing rule-based methods (9.6%).
Shunyu Yao, Howard Chen, John Yang et al.
PlanBench evaluates LLMs' planning and reasoning abilities using IPC-based multi-task benchmarks, revealing significant gaps in key capabilities.
Karthik Valmeekam, Matthew Marquez, Alberto Olmo et al.
Proposes UniCRS, a knowledge-enhanced prompt learning framework, unifying recommendation and dialogue tasks, significantly improving semantic consistency.
Xiaolei Wang, Kun Zhou, Ji-Rong Wen et al.
Fine-tuning large language models via behavioral cloning to generate natural critiques, improving human evaluation by 50%.
William Saunders, Catherine Yeh, Jeff Wu et al.
GPT-3 model expresses uncertainty in natural language, tested with CalibratedMath suite.
Stephanie Lin, Jacob Hilton, Owain Evans
Diffusion-LM significantly enhances controllable text generation using continuous diffusion models, outperforming existing methods.
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani et al.
Quark algorithm uses quantized reward conditioning to generate text with reduced undesirable traits, improving quality.
Ximing Lu, Sean Welleck, Jack Hessel et al.
NaturalProver leverages background references with constrained decoding to generate and suggest mathematical proofs, achieving over 40% correctness and usefulness in human evaluations.
Sean Welleck, Jiacheng Liu, Ximing Lu et al.
Neural demographic perturbation enhances NLP fairness by reducing bias sensitivity, validated on GLUE and bias datasets.
Rebecca Qian, Candace Ross, Jude Fernandes et al.
Proposes Zero-shot-CoT, using simple prompts to boost large LLMs' zero-shot reasoning, improving MultiArith accuracy from 17.7% to 78.7%.
Takeshi Kojima, Shixiang Shane Gu, Machel Reid et al.
This paper introduces ALTI+, a layer-wise token contribution method for Transformer-based encoder-decoder models, quantifying source and target influences on translation outputs.
Javier Ferrando, Gerard I. Gállego, Belen Alastruey et al.
StreamingQA uses time-stamped news data to evaluate model adaptation, showing fine-tuning and retrieval updates improve performance in dynamic knowledge environments.
Adam Liška, Tomáš Kočiský, Elena Gribovskaya et al.
BBTv2 optimizes large models' prompts using a gradient-free algorithm, reducing parameters with performance comparable to full model tuning.
Tianxiang Sun, Zhengfu He, Hong Qian et al.
Proposes Multi2WOZ dataset and conversational pretraining framework TOD-XLMR, significantly improving cross-lingual transfer for task-oriented dialogue.
Chia-Chien Hung, Anne Lauscher, Ivan Vulić et al.
Proposes ClaimDecomp dataset and T5-based models for generating literal and implied subquestions, improving explainability in complex claim fact-checking.
Jifan Chen, Aniruddh Sriram, Eunsol Choi et al.
UniMorph 4.0 introduces hierarchical features and multi-source data, enhancing multi-language morphological annotation and derivational analysis.
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa et al.
Task-guided scientific knowledge retrieval enhances scientific discovery efficiency.
Tom Hope, Doug Downey, Oren Etzioni et al.
OPT suite (125M-175B parameters) achieves GPT-3-level performance with 1/7th carbon footprint.
Susan Zhang, Stephen Roller, Naman Goyal et al.
Introduced IMCS-21 dataset to support five tasks in automatic medical consultation systems.
Wei Chen, Zhiwei Li, Hongyi Fang et al.