Abstract
This study introduces the Medical Vision Attention Generation (MedVAG) model, a novel framework designed to facilitate the automated generation of medical reports. MedVAG integrates Vision Transformer (ViT)-based visual feature extraction and GPT-2 language modeling, enhanced by graph-based feature fusion and multiple attention mechanisms (co-attention, cross-attention, memory-guided attention), to ensure semantic coherence and diagnostic accuracy. Evaluated on IU X-Ray and COV-CTR datasets, the model achieved state-of-the-art performance across natural language generation metrics (BLEU, METEOR, ROUGE, CIDEr) and clinical effectiveness measures. Ablation studies highlighted the critical role of attention mechanisms and feature fusion in aligning visual and textual features. MedVAG demonstrates strong potential as an assistive technology, aiming to support radiologists by reducing workload and enhancing diagnostic accuracy.
Author supplied keywords
Cite
CITATION STYLE
Varol Arısoy, M., Arısoy, A., & Uysal, İ. (2025). A vision attention driven Language framework for medical report generation. Scientific Reports, 15(1). https://doi.org/10.1038/s41598-025-95666-8
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.