Hybrid ResNet-ViT Transfer Learning Approach for Brain Stroke Classification on Computed Tomography Images

10Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

—This study investigates the utilization of a hybrid Convolutional Neural Network (CNN) and Vision Transformer (ViT) model, employing transfer learning methods, to enhance brain stroke detection and classification of CT images. The objective is to integrate ResNet-101’s local feature extraction capabilities with ViT’s global context comprehension to develop a resilient model for precise detection and categorization of brain strokes using non-contrast-enhanced brain computed tomography (NCCT) data. Data from two hospitals in Sri Lanka comprising 11,300 images were retrospectively collected. ViT and ResNet-101 architectures were modified for multi-class classification of brain normal, ischemic, and hemorrhagic conditions, and further differentiated ischemic acute, subacute, and chronic conditions in two-step ways including two models of the proposed architecture. We developed two models, incorporating the ResNet-101 component with MC dropout layer and a fully connected layer by removing the final two layers and the ViT component modifying multi-layer perceptron, incorporating three classes by adding a fully connected layer. The proposed model 01 training, and testing accuracy were 99.69%, 99.16% whereas model 02 achieved 97.31%, and 95.33% respectively. The hybrid models offer a robust method for impartial stroke diagnosis, potentially enabling tailored treatment approaches based on stroke type and severity. Further validation of the proposed approach on larger and more diverse datasets in clinical settings is required. Impact Statement—Prompt diagnosis of brain strokes can significantly improve outcomes and reduce morbidity and mortality. Moreover, accurate diagnosis is crucial for tailoring specific treatment, as it allows for targeted approaches based on the type and severity of the stroke. Deep learning models can analyze complex imaging data with high precision, often surpassing traditional methods by identifying patterns and abnormalities in medical images that may not be visible to the naked eye. By leveraging the strengths of multiple architectures, hybrid deep learning models enhance both diagnostic accuracy and robustness. The ResNet-101 architecture excels at extracting local features from medical images, and capturing details, whereas the ViT 16 base model is proficient at capturing global context and relationships between different parts of the medical images. Thus, by combining both local and global features, the model has the potential to achieve higher accuracy and robustness in brain stroke classification tasks.

Cite

CITATION STYLE

APA

Kulathilake, C. D., Udupihille, J., & Senoo, A. (2025). Hybrid ResNet-ViT Transfer Learning Approach for Brain Stroke Classification on Computed Tomography Images. IEEE Transactions on Artificial Intelligence. https://doi.org/10.1109/TAI.2025.3596534

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free