PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 2021Procedia CIRP32 citationsOpen Access

Benchmarking of Data Preprocessing Methods for Machine Learning-Applications in Production

View Full Paper
MFMaik FryeJMJohannes MohrenRSRobert Schmitt

Key Points

Key points are not available for this paper at this time.

Abstract

The application of machine learning (ML) is becoming increasingly common in production. However, many ML-projects in production fail due to poor data quality. To increase the quality, data needs to be preprocessed. Hundreds of methods exist for data preprocessing (DPP) that are selected manually depending on use-case requirements. For these reasons, DPP is currently performed unstructured and accounts for 80 % of ML-projects’ duration. Thus, we introduce a structured DPP-approach, in which DPP-methods are recommended based on production use-case requirements by benchmarking identified DPP-methods according to ML-model performance on five data sets. The approach is validated through two new use-cases.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Frye et al. (2021) studied this question.

synapsesocial.com/papers/6a1bd9adc97d63156a5f0597https://doi.org/10.1016/j.procir.2021.11.009
Ask AI
Helpful
Bookmark
Share
View Full Paper