Scalable Causal-Interpretable Machine Learning for Cancer Prescreening Using Electronic Health Records

Wait 5 sec.

Electronic health records provide a source of real-world information for disease risk prediction, but many machine learning models learn implicit associations that are difficult to inspect or use for clinical reasoning. Causal Bayesian networks (CBNs) offer a form of causal-interpretable machine learning by representing conditional dependencies, putative directional relations, and probabilistic evidence propagation in a directed acyclic graph. However, conventional CBN structure learning becomes unstable and computationally expensive when applied to large-scale EHR data. To address this limitation, we propose UPEBNL, a scalable framework for causal-interpretable cancer prescreening based on parallel CBN learning. UPEBNL integrates adaptive data slicing, quality-aware structure aggregation, and global DAG construction to learn stable and interpretable dependency structures from large-scale observational data. We evaluated UPEBNL in high-dimensional and multi-million-sample simulations and applied it to EHR-based prescreening for esophageal and colorectal cancer. In simulations, UPEBNL improved structural recovery accuracy by nearly 40% and achieved up to a 221.28-fold speedup over conventional CBN learning strategies. For cancer risk prediction, the learned CBNs provided interpretable evidence paths and achieved validation AUCs of 0.8171 for esophageal cancer and 0.784 for colorectal cancer. Calibration and decision curve analyses further supported the reliability and clinical utility of the models. These findings suggest that scalable CBN learning can support interpretable cancer prescreening from large-scale EHR data.