Key result
Machine learning pipeline improves real-time cardiac arrhythmia detection and decreases processing runtime.
Why the study?
Cardiac arrhythmia symptoms do not always appear on standard ECGs and physicians may misdiagnose them, necessitating continuous real-time ECG monitoring to detect arrhythmias promptly.
Does a machine learning pipeline using Apache Spark Structured Streaming improve classification performance and reduce delay in real-time cardiac arrhythmia detection?
Does a machine learning pipeline using Apache Spark Structured Streaming improve classification performance and reduce delay in real-time cardiac arrhythmia detection?
Implementing a machine learning pipeline with Apache Spark Structured Streaming and a random forest classifier enables efficient, real-time detection of cardiac arrhythmias with reduced processing delay.
Should not yet change arrhythmia monitoring practice; leaves open prospective clinical validation.
One of the major causes of death in the world is cardiac arrhythmias. In the field of healthcare, physicians use the patient's electrocardiogram (ECG) records to detect arrhythmias, which indicate the electrical activity of the patient's heart. The problem is that the symptoms do not always appear and the physician may be mistaken in the diagnosis. Therefore, patients need continuous monitoring through real-time ECG analysis to detect arrhythmias in a timely manner and prevent an eventual incident that threatens the patient's life. In this research, we used the Structured Streaming module built top on the open-source Apache Spark platform for the first time to implement a machine learning pipeline for real-time cardiac arrhythmias detection and evaluate the impact of using this new module on classification performance metrics and the rate of delay in arrhythmia detection. The ECG data collected from the MIT/BIH database for the detection of three class labels: normal beats, RBBB, and atrial fibrillation arrhythmias. We also developed three decision trees, random forest, and logistic regression multiclass classifiers for data classification where the random forest classifier showed better performance in classification than the other two classifiers. The results show previous results in performance metrics of the classification model and a significant decrease in pipeline runtime by using more class labels compared to previous studies.
No takes yet. Share an insight, caveat, or question.
Ilbeigipour et al. (2021) studied Cardiac arrhythmias. Machine learning pipeline using Apache Spark Structured Streaming vs. Other classifiers (decision trees, logistic regression) and previous studies was evaluated on Classification performance metrics and rate of delay in arrhythmia detection. A machine learning pipeline using Apache Spark Structured Streaming with a random forest classifier improved real-time cardiac arrhythmia detection and significantly decreased pipeline runtime.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: