Abstract
Retinal disease diagnosis is vital in ophthalmology, particularly for the early detection of various conditions to prevent vision impairment. Traditional models often focus on single-disease classification, overlooking the complexity of multiple co-existing retinal diseases and their interrelated features. In this work, we propose a novel Multi-Resolution Hypergraph Vision Transformer (MR-HGViT) framework for multi-label retinal disease classification and lesion characterization. Our approach constructs multi-resolution hypergraphs to capture both global anatomical structures and fine-grained lesion details in retinal images. Dynamic Hypergraph Convolutional Networks (DHGCNs) are employed to propagate features across different resolutions while dynamically adjusting hyperedge weights to emphasize critical retinal regions. Additionally, Vision Transformers are used to capture long-range dependencies efficiently. The model's interpretability is further enhanced through attention maps and Grad-CAM visualizations, providing insights into the key regions influencing the predictions. MR-HGViT is evaluated on three benchmark datasets - IDRiD, REFUGE, and MuReD - achieving state-of-the-art accuracies of 94.37%, 94.12%, and 93.78%, respectively. These results demonstrate the effectiveness of MR-HGViT in multi-label classification and its potential for clinical interpretability, making it a valuable tool for retinal disease diagnosis and lesion characterization.
Author supplied keywords
Cite
CITATION STYLE
Jothi Prakash, V., Arul Antran Vijay, S., & Sundaram, G. (2026). A Multi-Resolution Hypergraph Transformer for Explainable Retinal Disease Prediction. IEEE Journal of Biomedical and Health Informatics, 30(6), 4608–4621. https://doi.org/10.1109/JBHI.2025.3586292
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.