Apache Spark is one of the most widely used open source processing engines for big data, with rich language-integrated APIs and a wide range of libraries. Over the past two years, our group has worked to deploy Spark to a wide range of organizations through consulting relationships as well as our hosted service, Databricks. We describe the main challenges and requirements that appeared in taking Spark to a wide set of users, and usability and performance improvements we have made to the engine in response.
No takes yet. Share an insight, caveat, or question.
Armbrust et al. (2015) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: