Abstract
Glaucoma is a serious eye-debilitating disease affecting around 76 million people worldwide. In this work, we apply deep learning models, CNN and Vision Transformer (ViT), to understand their performance in discriminating the normal and glaucomatous fundus images. We further study the performance of Swin Transformer, a hierarchical Vision Transformer, relative to CNN and ViT. In the proposed approach, fundus pre-processing is carried out before training the pre-trained ViT and Swin variants, and deep CNN models with transfer learning under tuned hyperparameter settings. Additionally, the explainable AI (XAI) technique -Grad-CAM- is used to understand the model that performs better in terms of visualizing the areas of damage due to glaucoma in the fundus image and thus has better explainability. The experimental results obtained appear to indicate that ViT with a pretrained ViT-B32 model performs better both in terms of the quality metrics such as accuracy, sensitivity, specificity and area under the receiver operating characteristic curve (AUC), as well as the explainability of results. This (ViT-B32) model exhibits a remarkable accuracy of 96.19%, along with a sensitivity of 99.01%, an F1 score of 0.96, and an AUC of 0.995, indicating that it could be a better model for improved accuracy and sensitivity in the detection of glaucoma.
Author supplied keywords
Cite
CITATION STYLE
Jisy, N. K., Radhika, S., Senthil, S., & Srinivas, M. B. (2025). A Comparative Study of Performance of Various Deep Learning Models and Their Explainability in Detection of Glaucoma. IEEE Access, 13, 192891–192905. https://doi.org/10.1109/ACCESS.2025.3629624
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.