Abstract
Self-attention mechanisms are widely used in current encoder-decoder frameworks of image-captioning tasks. These mechanisms use the transformer module as the basic unit to compute the similarity between query and key vectors, and normalize them to obtain the attention distribution matrix, weighting and reconstructing the rich semantic query vectors. However, a single attention distribution matrix cannot express more complex intrinsic relations between query and key vectors. Some very important intrinsic relations can thus become lost. We therefore propose a hybrid attention distribution (HAD) that allows multiple distributions to be reconstructed to express deeper internal relations and avoids a single shallow attention distribution. In addition, to reduce the number of model parameters, we factorize the word-embedding matrix. This effectively increases the training efficiency and prevents the metrics from decreasing. We extensively evaluate our proposals using the MS-COCO image-captioning dataset. Our results outperform existing state-of-the-art methods on some metrics.
Author supplied keywords
Cite
CITATION STYLE
Wang, J., & Feng, J. (2020). Hybrid Attention Distribution and Factorized Embedding Matrix in Image Captioning. IEEE Access, 8, 154453–154460. https://doi.org/10.1109/ACCESS.2020.3018546
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.