Abstract
—This study investigates the utilization of a hybrid Convolutional Neural Network (CNN) and Vision Transformer (ViT) model, employing transfer learning methods, to enhance brain stroke detection and classification of CT images. The objective is to integrate ResNet-101’s local feature extraction capabilities with ViT’s global context comprehension to develop a resilient model for precise detection and categorization of brain strokes using non-contrast-enhanced brain computed tomography (NCCT) data. Data from two hospitals in Sri Lanka comprising 11,300 images were retrospectively collected. ViT and ResNet-101 architectures were modified for multi-class classification of brain normal, ischemic, and hemorrhagic conditions, and further differentiated ischemic acute, subacute, and chronic conditions in two-step ways including two models of the proposed architecture. We developed two models, incorporating the ResNet-101 component with MC dropout layer and a fully connected layer by removing the final two layers and the ViT component modifying multi-layer perceptron, incorporating three classes by adding a fully connected layer. The proposed model 01 training, and testing accuracy were 99.69%, 99.16% whereas model 02 achieved 97.31%, and 95.33% respectively. The hybrid models offer a robust method for impartial stroke diagnosis, potentially enabling tailored treatment approaches based on stroke type and severity. Further validation of the proposed approach on larger and more diverse datasets in clinical settings is required. Impact Statement—Prompt diagnosis of brain strokes can significantly improve outcomes and reduce morbidity and mortality. Moreover, accurate diagnosis is crucial for tailoring specific treatment, as it allows for targeted approaches based on the type and severity of the stroke. Deep learning models can analyze complex imaging data with high precision, often surpassing traditional methods by identifying patterns and abnormalities in medical images that may not be visible to the naked eye. By leveraging the strengths of multiple architectures, hybrid deep learning models enhance both diagnostic accuracy and robustness. The ResNet-101 architecture excels at extracting local features from medical images, and capturing details, whereas the ViT 16 base model is proficient at capturing global context and relationships between different parts of the medical images. Thus, by combining both local and global features, the model has the potential to achieve higher accuracy and robustness in brain stroke classification tasks.
Author supplied keywords
Cite
CITATION STYLE
Kulathilake, C. D., Udupihille, J., & Senoo, A. (2025). Hybrid ResNet-ViT Transfer Learning Approach for Brain Stroke Classification on Computed Tomography Images. IEEE Transactions on Artificial Intelligence. https://doi.org/10.1109/TAI.2025.3596534
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.