Simple and Effective Multimodal Learning Based on Pre-Trained Transformer Models

N/ACitations
Citations of this article
44Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Transformer-based models have garnered attention because of their success in natural language processing, and in several other fields, such as image and automatic speech recognition. In addition to them being trained on unimodal information, many transformer-based models have been proposed for multimodal information. In multimodal learning, a common problem encountered is the insufficiency of multimodal training data. In this study, to address this problem, a simple and effective method is proposed by using 1) unimodal pre-trained transformer models as encoders for each modal input and 2) a set of transformer layers to fuse their output representations. Further, the proposed method is evaluated by conducting several experiments on two common benchmarks: CMU multimodal opinion sentiment intensity dataset and multimodal internet movie database. The proposed model exhibits state-of-the-art performances on both benchmarks and is robust against the reduction in the amount of training data.

Cite

CITATION STYLE

APA

Miyazawa, K., Kyuragi, Y., & Nagai, T. (2022). Simple and Effective Multimodal Learning Based on Pre-Trained Transformer Models. IEEE Access, 10, 29821–29833. https://doi.org/10.1109/ACCESS.2022.3159346

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free