排序: 最新 热门 引用
cs.CL 1706.03762

Attention Is All You Need

提出Transformer模型,完全基于注意力机制,显著提升机器翻译性能,训练时间缩短。

Ashish Vaswani, Noam Shazeer, Niki Parmar 等

2017-06-13 189620 引用 51
cs.CL 1705.03122

Convolutional Sequence to Sequence Learning

提出完全卷积的序列到序列模型,采用门控线性单元,在WMT'14英德、英法任务中超越LSTM,速度提升十倍。

Jonas Gehring, Michael Auli, David Grangier 等

2017-05-09 56