Dynamic layer aggregation for neural machine translation with routing-by-agreement

Zi Yi Dou; Zhaopeng Tu; Xing Wang; Longyue Wang; Shuming Shi; Tong Zhang

Conference ProceedingsOPEN ACCESS

Dynamic layer aggregation for neural machine translation with routing-by-agreement

33rd AAAI Conference on Artificial Intelligence, AAAI 2019, 31st Innovative Applications of Artificial Intelligence Conference, IAAI 2019 and the 9th AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019 (2019) 86-93

DOI: 10.1609/aaai.v33i01.330186

32Citations

60Readers

Abstract

With the promising progress of deep neural networks, layer aggregation has been used to fuse information across layers in various fields, such as computer vision and machine translation. However, most of the previous methods combine layers in a static fashion in that their aggregation strategy is independent of specific hidden states. Inspired by recent progress on capsule networks, in this paper we propose to use routing-by-agreement strategies to aggregate layers dynamically. Specifically, the algorithm learns the probability of a part (individual layer representations) assigned to a whole (aggregated representations) in an iterative way and combines parts accordingly. We implement our algorithm on top of the state-of-the-art neural machine translation model TRANSFORMER and conduct experiments on the widely-used WMT14 English-German and WMT17 Chinese-English translation datasets. Experimental results across language pairs show that the proposed approach consistently outperforms the strong baseline model and a representative static aggregation model.

Cite

CITATION STYLE

APA

Dou, Z. Y., Tu, Z., Wang, X., Wang, L., Shi, S., & Zhang, T. (2019). Dynamic layer aggregation for neural machine translation with routing-by-agreement. In 33rd AAAI Conference on Artificial Intelligence, AAAI 2019, 31st Innovative Applications of Artificial Intelligence Conference, IAAI 2019 and the 9th AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019 (pp. 86–93). AAAI Press. https://doi.org/10.1609/aaai.v33i01.330186

Dynamic layer aggregation for neural machine translation with routing-by-agreement

Abstract

Cite

Register to see more suggestions