Apache Spark is a widely used technology now a days for handling huge datasets in applications due to its flexibility, scalability, robustness, speed and integration with multiple programming languages like Java, Scala, Python. It provides multiple methodologies for implementation like dataframes, RDDs with these programming languages. This paper provides a deep overview of Apache spark dataframes usage for performance enhancement over Apache Spark Resilient Distributed Datasets (RDDs) and SQL based data processing. Spark dataframes are widely used in managing and processing large datasets which can be structured, non-structured or semi-structured. This paper describes the approach towards Spark dataframes for performance enhancement for large data processing in place of traditional usage of Spark RDDs with practical examples and use cases. It highlights the key points on why to use dataframes for a better performance achievement for any application where large dataset needs to be processed into a meaningful output.
No takes yet. Share an insight, caveat, or question.
Ashima Sahni (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: