Abstract
Deep learning has been applied to optical music sheet recognition (OMR). However, OMR processing from various sheet-music images still lacks precision to be widely applicable. We propose a measure-based multimodal deep-learning-driven assembly (MMdA) method enabling end-to-end OMR processing from various images including inclined photo images. Using this method, measures are extracted using a deep-learning model, aligned, and resized to be used for inference of given musical-symbol components by using multiple deep-learning models in sequence or in parallel. The use of each standardized measure enables efficient training of the deep-learning models and accurate adjustment of five staff lines in each measure, which enables locally inclined sheet-music images to be precisely positioned. Thus, a score can be reproduced from the inclined image with the proposed MMdA method while current OMR applications cannot. Multiple musical-symbol-component deep-learning feature-category models with a small number of feature types can represent a diverse set of notes and other musical symbols including chords. The proposed MMdA method provides a solution to end-to-end OMR processing and enhances the utility of OMR of mobile phone-based sheet-music photo images.
Author supplied keywords
Cite
CITATION STYLE
Shishido, T., Fati, F., Tokushige, D., Ono, Y., & Kumazawa, I. (2023). Production of MusicXML from Locally Inclined Sheet-music Photo Image by Using Measure-based Multi-modal Deep-learning-driven Assembly Method. Transactions of the Japanese Society for Artificial Intelligence, 38(3). https://doi.org/10.1527/tjsai.38-3_A-MA3
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.