Cross-modal attention-based multimodal deep learning framework for breast cancer diagnosis using mammography and ultrasound imaging
International Journal of Development Research
Cross-modal attention-based multimodal deep learning framework for breast cancer diagnosis using mammography and ultrasound imaging
Received 20th April, 2026 Received in revised form 18th May, 2026 Accepted 25th June, 2026 Published online 30th July, 2026
Copyright©2026, Indu et al. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Breast cancer continues to be one of the top causes of cancer-related death for women globally, highlighting the significance of prompt and precise detection. Single-modality imaging is often used by conventional computer-aided diagnosis systems, which restricts their capacity to provide additional diagnostic data. In order to diagnose breast cancer automatically, this study suggests a multimodal deep learning framework that combines ultrasound and mammography imaging. The suggested approach learns discriminative representations from heterogeneous imaging modalities by combining an ultrasonic encoder based on Swin Transformer and a mammography encoder based on ResNet101 with a cross-modal attention fusion technique. To increase lesion visibility during preprocessing, contrast enhancement utilizing Contrast Limited Adaptive Histogram Equalization (CLAHE) was used. The classification of benign and malignant conditions was done using the fused multimodal characteristics.Utilizing publicly accessible mammography and ultrasound datasets, an experimental evaluation was carried out utilizing a pseudo-paired multimodal learning approach. The suggested model demonstrated the efficacy of multimodal fusion for breast cancer diagnosis with an accuracy of 88%, an F1-score of 0.87, and an Area Under the ROC Curve (AUC) of 0.903. The findings show that cross-attention greatly enhances diagnostic performance when convolutional and transformer-based representations are combined. Future multimodal medical imaging applications and intelligent clinical decision support systems have great potential with the suggested paradigm.