Leveraging pretrained vision-language model for enhanced breast cancer diagnosis with multi-view mammography

N/ACitations
Citations of this article
9Readers
Mendeley users who have this article in their library.

Abstract

Background: Although fusion of information from multiple views of mammograms plays an important role to increase accuracy of breast cancer detection, developing multi-view mammograms-based computer-aided diagnosis (CAD) schemes still faces big challenges and no such CAD schemes have been used in clinical practice. Purpose: To overcome these challenges, we investigate a new approach based on the concept of contrastive language-image pre-training (CLIP), which has sparked interest across various medical imaging tasks. The aim is to solve the challenges in: (1) effectively adapting the single-view CLIP for multi-view feature fusion and (2) efficiently fine-tuning this parameter-dense model with limited samples and computational resources. Methods: We introduce a unique Mammo-CLIP, the first multi-modal framework to process multi-view mammograms and corresponding simple texts. Mammo-CLIP uses an early feature fusion strategy to learn multi-view relationships in four mammograms acquired from the craniocaudal (CC) and mediolateral oblique (MLO) views of the left and right breasts. To enhance learning efficiency, plug-and-play adapters are added into CLIP's image and text encoders for fine-tuning the model efficiently and limiting updates to about 1% of the parameters. For framework evaluation, we assembled two datasets retrospectively. The first dataset, comprising 470 malignant and 479 benign cases, was used for few-shot fine-tuning and internal evaluation of the proposed Mammo-CLIP via 5-fold cross-validation. The second dataset, including 60 malignant and 294 benign cases, was used to test generalizability of Mammo-CLIP. Results: Mammo-CLIP outperforms the state-of-the-art (SOTA) cross-view transformer evaluated using areas under ROC curves (AUC = 0.841 ± 0.017 vs. 0.817 ± 0.012 and 0.837 ± 0.034 vs. 0.807 ± 0.036) on both datasets. It also surpasses previous two CLIP-based methods by 20.3% and 14.3% in AUC. Conclusions: The proposed Mammo-CLIP demonstrates superior breast cancer diagnosis performance compared to SOTA methods. This study highlights the potential of applying the finetuned vision-language models for developing multi-view, image-text-based CAD schemes of breast cancer.

Cite

CITATION STYLE

APA

Chen, X., Li, Y., Hu, M., Salari, E., Chen, X., Qiu, R. L. J., … Yang, X. (2026). Leveraging pretrained vision-language model for enhanced breast cancer diagnosis with multi-view mammography. Medical Physics, 53(1). https://doi.org/10.1002/mp.70261

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free