Vocal Fold Disorders Classification and Optimization of a Custom Video Laryngoscopy Dataset Through Structural Similarity Index and a Deep Learning-Based Approach

4Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

Background/Objectives: Video laryngoscopy is one of the primary methods used by otolaryngologists for detecting and classifying laryngeal lesions. However, the diagnostic process of these images largely relies on clinicians’ visual inspection, which can lead to overlooked small structural changes, delayed diagnosis, and interpretation errors. Methods: AI-based approaches are becoming increasingly critical for accelerating early-stage diagnosis and improving reliability. This study proposes a hybrid Convolutional Neural Network (CNN) architecture that eliminates repetitive and clinically insignificant frames from videos, utilizing only meaningful key frames. Video data from healthy individuals, patients with vocal fold nodules, and those with vocal fold polyps were summarized using three different threshold values with the Structural Similarity Index Measure (SSIM). Results: The resulting key frames were then classified using a hybrid CNN. Experimental findings demonstrate that selecting an appropriate threshold can significantly reduce the model’s memory usage and processing load while maintaining accuracy. In particular, a threshold value of 0.90 provided richer information content thanks to the selection of a wider variety of frames, resulting in the highest success rate. Fine-tuning the last 20 layers of the MobileNetV2 and Xception backbones, combined with the fusion of extracted features, yielded an overall classification accuracy of 98%. Conclusions: The proposed approach provides a mechanism that eliminates unnecessary data and prioritizes only critical information in video-based diagnostic processes, thus helping physicians accelerate diagnostic decisions and reduce memory requirements.

Cite

CITATION STYLE

APA

Emre, E., Cetintas, D., Yildirim, M., & Emre, S. (2025). Vocal Fold Disorders Classification and Optimization of a Custom Video Laryngoscopy Dataset Through Structural Similarity Index and a Deep Learning-Based Approach. Journal of Clinical Medicine, 14(19). https://doi.org/10.3390/jcm14196899

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free