Transformer-based Automatic Music Mood Classification Using Multi-modal Framework

N/ACitations
Citations of this article
21Readers
Mendeley users who have this article in their library.

Abstract

According to studies, music affects our moods, and we are also inclined to choose a theme based on our current moods. Audio-based techniques can achieve promising results, but lyrics also give relevant information about the moods of a song which may not be present in the audio part. So a multi-modal with both textual features and acoustic features can provide enhanced accuracy. Sequential networks such as long short-term memory networks (LSTM) and gated recurrent unit networks (GRU) are widely used in the most state-of-the-art natural language processing (NLP) models. A transformer model uses self-attention to compute representations of its inputs and outputs, unlike recurrent unit networks (RNNs) that use sequences and transformers that can parallelize over input positions during training. In this work, we proposed a multi-modal music mood classification system based on transformers and compared the system’s performance using a bi-directional GRU (Bi-GRU)-based system with and without attention. The performance is also analyzed for other state-of-the-art ap-proaches. The proposed transformer-based model ac-quired higher accuracy than the Bi-GRU-based multi-modal system with single-layer attention by providing a maximum accuracy of 77.94%.

Cite

CITATION STYLE

APA

Suresh Kumar, S. A., & Rajan, R. (2023). Transformer-based Automatic Music Mood Classification Using Multi-modal Framework. Journal of Computer Science and Technology(Argentina), 23(1), 18–34. https://doi.org/10.24215/16666038.23.e02

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free