Deep Learning for Heart Sound Abnormality of Infants: Proof-of-Concept Study of 1D and 2D Representations †

1Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.

Abstract

Highlights: What are the main findings? This study proposes a novel multimodal deep learning framework that uniquely integrates both 1D time-series and 2D image representations of infant heart sounds, achieving a state-of-the-art accuracy of 98.91% for abnormality detection. The study systematically demonstrates that this fusion model not only improves classification performance over using either data representation alone but also significantly reduces the computational training cost compared to purely image-based models. What are the implications of the main findings? This work’s primary implication is the potential for a more accessible and cost-effective method for early screening of congenital heart defects in infants, as the model’s high accuracy is achieved using audio from stethoscopes, which are more readily available than specialized equipment like ECGs. Academically, the research provides a significant insight for the field by demonstrating that a multimodal deep learning approach can surpass the performance of single-modality models while also being more computationally efficient, offering a promising direction for the development of future diagnostic AI. Introduction: Advanced identification and intervention for Congenital Heart Defects (CHDs) in pediatric populations are crucial, as approximately 1% of neonates worldwide present with these conditions. Traditional methods of diagnosing CHDs often rely on stethoscope auscultation, which heavily depends on the clinician’s expertise and may lead to the oversight of subtle acoustic indicators. Objectives: This study introduces an innovative deep-learning framework designed for the early diagnosis of congenital heart disease. It utilizes time-series data obtained from cardiac auditory signals captured through stethoscopes. Methods: The audio signals were processed into time–frequency representations using Mel-Frequency Cepstral Coefficients (MFCCs). The architecture of the model combines Convolutional Neural Networks (CNNs) for effective feature extraction with Long Short-Term Memory (LSTM) networks to accurately model temporal dependencies. Impressively, the model achieved an accuracy of 98.91% in the early detection of CHDs. Results: While traditional diagnostic tools like Electrocardiograms (ECG) and Phonocardiograms (PCG) remain indispensable for confirming diagnoses, many AI studies have primarily targeted ECG and PCG datasets. This approach emphasizes the potential of cardiac acoustics for the early diagnosis of CHDs, which could lead to improved clinical outcomes for infants. Notably, the dataset used in this research is publicly available, enabling wider application and model training within the research community.

Cite

CITATION STYLE

APA

Wazed, E., Lee, J., & Jeong, H. (2025). Deep Learning for Heart Sound Abnormality of Infants: Proof-of-Concept Study of 1D and 2D Representations †. Children, 12(9). https://doi.org/10.3390/children12091221

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free