Abstract
Highlights: What are the main findings? This study proposes a novel multimodal deep learning framework that uniquely integrates both 1D time-series and 2D image representations of infant heart sounds, achieving a state-of-the-art accuracy of 98.91% for abnormality detection. The study systematically demonstrates that this fusion model not only improves classification performance over using either data representation alone but also significantly reduces the computational training cost compared to purely image-based models. What are the implications of the main findings? This work’s primary implication is the potential for a more accessible and cost-effective method for early screening of congenital heart defects in infants, as the model’s high accuracy is achieved using audio from stethoscopes, which are more readily available than specialized equipment like ECGs. Academically, the research provides a significant insight for the field by demonstrating that a multimodal deep learning approach can surpass the performance of single-modality models while also being more computationally efficient, offering a promising direction for the development of future diagnostic AI. Introduction: Advanced identification and intervention for Congenital Heart Defects (CHDs) in pediatric populations are crucial, as approximately 1% of neonates worldwide present with these conditions. Traditional methods of diagnosing CHDs often rely on stethoscope auscultation, which heavily depends on the clinician’s expertise and may lead to the oversight of subtle acoustic indicators. Objectives: This study introduces an innovative deep-learning framework designed for the early diagnosis of congenital heart disease. It utilizes time-series data obtained from cardiac auditory signals captured through stethoscopes. Methods: The audio signals were processed into time–frequency representations using Mel-Frequency Cepstral Coefficients (MFCCs). The architecture of the model combines Convolutional Neural Networks (CNNs) for effective feature extraction with Long Short-Term Memory (LSTM) networks to accurately model temporal dependencies. Impressively, the model achieved an accuracy of 98.91% in the early detection of CHDs. Results: While traditional diagnostic tools like Electrocardiograms (ECG) and Phonocardiograms (PCG) remain indispensable for confirming diagnoses, many AI studies have primarily targeted ECG and PCG datasets. This approach emphasizes the potential of cardiac acoustics for the early diagnosis of CHDs, which could lead to improved clinical outcomes for infants. Notably, the dataset used in this research is publicly available, enabling wider application and model training within the research community.
Author supplied keywords
Cite
CITATION STYLE
Wazed, E., Lee, J., & Jeong, H. (2025). Deep Learning for Heart Sound Abnormality of Infants: Proof-of-Concept Study of 1D and 2D Representations †. Children, 12(9). https://doi.org/10.3390/children12091221
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.