Article contents
Reliable and explainable deep learning for brain tumor diagnosis: Uncertainty-calibrated MRI classification with external validation
Abstract
Reliable brain-tumor diagnosis from magnetic resonance imaging (MRI) requires more than high classification accuracy. Deep neural networks can be overconfident under scanner, protocol, demographic, and disease-spectrum shifts, while visually plausible explanations may not faithfully identify the evidence driving a prediction. This methodological study presents a reliability-centered framework for MRI brain-tumor classification that integrates external validation, post-hoc probability calibration, predictive uncertainty estimation, out-of-distribution screening, and quantitatively evaluated explainability. The proposed pipeline uses patient-level data separation, harmonized preprocessing, a deep ensemble of independently initialized classifiers, temperature scaling on a dedicated calibration set, and entropy-based selective prediction. Grad-CAM and SHAP are combined with lesion masks to test whether model attention overlaps clinically relevant tumor regions rather than background anatomy or acquisition artifacts. Performance is assessed using macro-averaged area under the receiver operating characteristic curve, sensitivity, specificity, F1 score, Brier score, expected calibration error, negative log-likelihood, uncertainty-error association, and risk-coverage analysis. External validation is designed to hold out complete institutions or datasets, thereby testing robustness to domain shift rather than random image-level variation. A synthesis of 32 peer-reviewed studies published through March 2026 shows that brain-tumor MRI research increasingly incorporates explainability, yet calibration, uncertainty-aware referral, and independent external testing remain less consistently combined. The framework therefore treats discrimination, calibration, uncertainty, localization fidelity, and generalizability as complementary dimensions of clinical reliability. It offers a reproducible study design for developing decision-support models that can identify both confident predictions and cases that should be deferred for expert review. This design prioritizes safe failure behavior under real-world variability.
Article information
Journal
Journal of Medical and Health Studies
Volume (Issue)
7 (9)
Pages
20-211
Published
Copyright
Copyright (c) 2026 https://creativecommons.org/licenses/by/4.0/
Open access

This work is licensed under a Creative Commons Attribution 4.0 International License.

Aims & scope
Call for Papers
Article Processing Charges
Publications Ethics
Google Scholar Citations
Recruitment