CFP last date
20 August 2026
Call for Paper
September Edition
IJCA solicits high quality original research papers for the upcoming September edition of the journal. The last date of research paper submission is 20 August 2026

Submit your paper
Know more
Reseach Article

Multi-Modal Deep Learning Framework for Prediction and Early Identification of Neoplasms: A Hybrid CNN-Transformer Architecture with Explainability

by Jyotsna Patel, Ankur Pandey, Pushpendra Singh Tomar
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Number 138
Year of Publication: 2026
Authors: Jyotsna Patel, Ankur Pandey, Pushpendra Singh Tomar
10.5120/ijca69c360f8e2a2

Jyotsna Patel, Ankur Pandey, Pushpendra Singh Tomar . Multi-Modal Deep Learning Framework for Prediction and Early Identification of Neoplasms: A Hybrid CNN-Transformer Architecture with Explainability. International Journal of Computer Applications. 187, 138 ( Aug 2026), 9-15. DOI=10.5120/ijca69c360f8e2a2

@article{ 10.5120/ijca69c360f8e2a2,
author = { Jyotsna Patel, Ankur Pandey, Pushpendra Singh Tomar },
title = { Multi-Modal Deep Learning Framework for Prediction and Early Identification of Neoplasms: A Hybrid CNN-Transformer Architecture with Explainability },
journal = { International Journal of Computer Applications },
issue_date = { Aug 2026 },
volume = { 187 },
number = { 138 },
month = { Aug },
year = { 2026 },
issn = { 0975-8887 },
pages = { 9-15 },
numpages = {9},
url = { https://ijcaonline.org/archives/volume187/number138/multi-modal-deep-learning-framework-for-prediction-and-early-identification-of-neoplasms-a-hybrid-cnn-transformer-architecture-with-explainability/ },
doi = { 10.5120/ijca69c360f8e2a2 },
publisher = {Foundation of Computer Science (FCS), NY, USA},
address = {New York, USA}
}
%0 Journal Article
%1 2026-08-20T21:55:13.668348+05:30
%A Jyotsna Patel
%A Ankur Pandey
%A Pushpendra Singh Tomar
%T Multi-Modal Deep Learning Framework for Prediction and Early Identification of Neoplasms: A Hybrid CNN-Transformer Architecture with Explainability
%J International Journal of Computer Applications
%@ 0975-8887
%V 187
%N 138
%P 9-15
%D 2026
%I Foundation of Computer Science (FCS), NY, USA
Abstract

Early and accurate detection of neoplasms is a critical challenge in clinical oncology, where diagnostic delays substantially reduce patient survival rates. This paper proposes a novel Multi-Modal Hybrid CNN-Transformer (MMHCT) framework for simultaneous early prediction and identification of neoplasms across five cancer types: breast, lung, brain, prostate, and colorectal. The architecture fuses spatial feature extraction via a modified ResNet-50 backbone with long-range dependency modeling through a Vision Transformer (ViT-B/16) encoder. A Cross-Modal Attention Fusion (CMAF) mechanism integrates heterogeneous inputs, including CT, MRI, whole-slide histopathology images, and structured Electronic Health Records (EHR). A custom focal Dice loss function and federated learning protocol address class imbalance and data privacy constraints, respectively. On the TCGA-LUAD, CBIS-DDSM, BraTS-2023, and Patch Camelyon benchmarks, MMHCT achieves mean AUC = 0.974, sensitivity = 94.3%, specificity = 96.1%, and F1-score = 0.943 at stage-I detection, outperforming all nine baseline methods. GRAD-CAM++ saliency maps and SHAP feature attributions are embedded to ensure clinical interpretability. The framework complies with GDPR and HIPAA regulations through differential privacy mechanisms with privacy budget epsilon = 0.3.

References
  1. H. Sung, J. Ferlay, R. L. Siegel et al., “Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality for 36 cancers in 185 countries,” CA: A Cancer Journal for Clinicians, vol. 74, no. 3, pp. 229-263, 2024.
  2. American Cancer Society, Cancer Facts & Figures 2024. Atlanta: ACS, 2024.
  3. J. G. Elmore, G. M. Longton, P. A. Carney et al., “Variability in pathologists’ interpretations of individual breast biopsy slides,” Annals of Internal Medicine, vol. 164, no. 10, pp. 649-655, 2015.
  4. G. Litjens, T. Kooi, B. E. Bejnordi et al., “A survey on deep learning in medical image analysis,” Medical Image Analysis, vol. 42, pp. 60-88, 2017.
  5. P. Rajpurkar, J. Irvin, K. Ball et al., “CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning,” arXiv:1711.05225, 2017.
  6. A. Esteva, B. Kuprel, R. A. Novoa et al., “Dermatologist-level classification of skin cancer with deep neural networks,” Nature, vol. 542, no. 7639, pp. 115-118, 2017.
  7. A. Dosovitskiy, L. Beyer, A. Kolesnikov et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. ICLR, 2021.
  8. J. Chen, Y. Lu, Q. Yu et al., “TransUNet: Transformers make strong encoders for medical image segmentation,” arXiv:2102.04306, 2021.
  9. Z. Liu, Y. Lin, Y. Cao et al., “Swin Transformer: Hierarchical vision transformer using shifted windows,” in Proc. ICCV, pp. 10012-10022, 2021.
  10. P. Mobadersany, S. Yousefi, M. Amgad et al., “Predicting cancer outcomes from histology and genomics using convolutional networks,” Proc. Natl. Acad. Sci., vol. 115, no. 13, pp. E2970-E2979, 2018.
  11. R. J. Chen, M. Y. Lu, J. Wang et al., “Multimodal co-attention transformer for survival prediction in gigapixel whole slide images,” in Proc. ICCV, pp. 4015-4025, 2021.
  12. R. R. Selvaraju, M. Cogswell, A. Das et al., “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” in Proc. ICCV, pp. 618-626, 2017.
  13. S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proc. NeurIPS, vol. 30, 2017.
  14. S. Singh, A. Kumar, R. Sharma et al., “Explainability in deep learning for oncology: A systematic review (2020-2024),” npj Digital Medicine, vol. 7, p. 88, 2024.
  15. I. Mironov, “Rényi differential privacy,” in Proc. IEEE CSF, pp. 263-275, 2017.
  16. K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. CVPR, pp. 770-778, 2016.
  17. G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proc. CVPR, pp. 4700-4708, 2017.
  18. M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. ICML, pp. 6105-6114, 2019.
  19. M. Y. Lu, B. Chen, D. F. K. Williamson et al., “A visual-language foundation model for computational pathology,” Nature Medicine, vol. 30, no. 3, pp. 863-874, 2024.
  20. R. J. Chen, C. Ding, M. Y. Lu et al., “Towards a general-purpose foundation model for computational pathology,” Nature Medicine, vol. 30, no. 3, pp. 850-862, 2024.
  21. C. D. Lehman, R. D. Arao, B. L. Sprague et al., “National performance benchmarks for modern screening digital mammography,” Radiology, vol. 283, no. 1, pp. 49-58, 2017.
  22. E. J. Feuer, B. S. Levy, H. L. Haber et al., “CISNET: Using modeling to understand cancer control,” J. Natl. Cancer Inst. Monogr., no. 56, pp. 2-6, 2020.
Index Terms

Computer Science
Information Sciences

Keywords

Neoplasm detection deep learning Vision Transformer multi-modal fusion federated learning explainable AI oncology medical imaging