Latest trends in the health industry suggest ever increasing amount of data accumulated digitally, making it one of the top data intensive sectors. The evolution of technology has enabled to process such large data, and accurately predict interested outcomes. In this paper, propose a scalable framework that uses healthcare data to predict heart disease based on certain attributes. Our main contribution in this work is to predict the diagnosis of heart disease with a small number of attributes. Our prediction solution uses random forest on Apache Spark, which gives massive opportunity for health care analysts to deploy this solution on ever changing, scalable big data landscape for insightful decision making. Using this approach, we show that up to 98% accuracy is achieved. We also present a comparison against Naïve-Bayes classifier, where we show the random forest approach outperforms the former by a significant margin.
No takes yet. Share an insight, caveat, or question.
Rashmi G Saboji (2017) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: