Timing and Context Features in Machine Learning Classification of Inter-Patient ECG Heartbeats on the MIT-BIH Benchmark

Wait 5 sec.

Explicit timing and context features may complement short electrocardiogram waveforms, but aggregate gains on imbalanced benchmarks can conceal errors in individual classes. We therefore compared Logistic Regression, Random Forest, XGBoost, and a one-dimensional convolutional neural network (CNN) in four cumulative input stages. Stage S1 used a 255-sample beat waveform alone, and stages S2 to S4 added 6, 13, and 21 engineered features. Using the inter-patient division of the MIT-BIH Arrhythmia Database, with 50,992 development and 49,686 evaluation beats in four classes, we estimated 95% intervals with a paired bootstrap over the 22 evaluation recordings. Adding six RR interval features increased macro F1 by 0.025 to 0.093, with intervals excluding zero for every model family. Further context helped Random Forest and XGBoost, whereas later CNN gains fell within resampling uncertainty. At the final stage, XGBoost and CNN reached macro F1 of 0.580 and 0.574, and the interval for their difference included zero. Despite opposite operating points, their precision for fusion beats was similar at matched recall, suggesting that neither separated fusion beats better. The most accurate pipeline, Random Forest at S2 with 94.45% accuracy, detected only 1.4% of supraventricular beats, and development and evaluation favored the same stage for only one family. The results for the minority classes were based on single recordings. Record 213 supplied 203 of the 207 correct CNN fusion detections, and elsewhere both final models detected at most 4 of 26 fusion beats. Record 232 contains 75% of the evaluation supraventricular beats, and nearly all were missed. Excluding this record nearly doubled XGBoost supraventricular F1, from 0.199 to 0.383. These findings are limited to retrospective analysis with annotated beat positions and a single training seed.