Explainable Machine Learning for Cardiovascular Disease Detection: Addressing Interpretability Limitations and Evaluating Transparency for Clinical Decision Support
Keywords:
Cardiovascular disease detection; Explainable artificial intelligence; SHAP; Machine learning transparency; Clinical decision supportAbstract
Cardiovascular diseases (CVDs) remain the leading cause of mortality worldwide,
accounting for approximately 17.9 million deaths annually. Although machine learning (ML)
models have demonstrated high predictive performance for CVD detection, their
inherent black-box nature limits clinical trust, interpretability, and adoption in
healthcare practice. This study addresses three objectives: (1) to examine the limitations
of conventional ML models for CVD prediction; (2) to develop an interpretable ML
framework that enhances model transparency through the integration of SHapley
Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations
(LIME); and (3) to evaluate the predictive performance, transparency, and clinical
applicability of the resulting models. Logistic Regression, Random Forest, and Extreme
Gradient Boosting (XGBoost) classifiers were trained and evaluated using the Kaggle
Cardiovascular Disease Dataset comprising 70,000 patient records and 11 predictor
variables. Model performance was assessed using accuracy, precision, recall, F1-score,
receiver operating characteristic area under the curve (ROC-AUC), calibration error, and
a structured transparency scoring framework. XGBoost achieved the highest
discriminative performance, with an ROC-AUC of 0.794, an accuracy of 0.729, and the
lowest calibration error of 0.024, whereas Logistic Regression attained the highest
transparency score (73.8/100). Global model interpretation using SHAP identified
cholesterol, age, systolic blood pressure, and diastolic blood pressure as the most
influential predictors, consistent with established clinical evidence. Similarly, LIME
provided patient-specific explanations that highlighted the same dominant risk factors,
thereby enhancing decision interpretability at the individual level. These findings
demonstrate that high predictive accuracy and model explainability can coexist within a
unified ML framework, providing a replicable blueprint for transparent cardiovascular
risk assessment, particularly in resource-limited healthcare settings.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Pure and Applied Sciences (Science Forum)

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.


