Abstract
In recent years, the application of machine translation has become more and more widely. Currently, the neural multimodal translation models have made attractive progress, which combines images into deep learning networks, such as Transformer and RNN. When considering images in translation models, they directly apply gate structure or image attention to introduce image feature to enhance the translation effect. We argue that it may mismatch the text and image features since they are in different semantic space. In this paper, we propose a coordinated representation learning enhanced multimodal machine translation approach with multimodal attention. Our approach accepts the text data and its relevant image data as the input. The image features are fed into the decoder side of the basic Transformer model. Moreover, the Coordinated Representation Learning is utilized to map the different text and image modal features into their semantic representations. The mapped representations are linearly related in a shared semantic space. Finally, the sum of the image and text representations, called Coordinated Visual-Semantic Representation (CVSR), will be sent to a Multimodal Attention Layer (MAL) in our Transformer based translation approach. Experimental results show that our approach achieves the state-of-art performance on the public Multi30k dataset.
Author supplied keywords
Cite
CITATION STYLE
Han, Y., Li, L., & Zhang, J. (2020). A coordinated representation learning enhanced multimodal machine translation approach with multi-attention. In ICMR 2020 - Proceedings of the 2020 International Conference on Multimedia Retrieval (pp. 571–577). Association for Computing Machinery. https://doi.org/10.1145/3372278.3390717
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.