July 20, 2008

Can data transformation help in the detection of fault-prone modules?

Key Points

Key points are not available for this paper at this time.

Abstract

Data preprocessing (transformation) plays an important role in data mining and machine learning. In this study, we investigate the effect of four different preprocessing methods to fault-proneness prediction using nine datasets from NASA Metrics Data Programs (MDP) and ten classification algorithms. Our experiments indicate that log transformation rarely improves classification performance, but discretization affects the performance of many different algorithms. The impact of different transformations differs. Random forest algorithm, for example, performs better with original and log transformed data set. Boosting and NaiveBayes perform significantly better with discretized data. We conclude that no general benefit can be expected from data transformations. Instead, selected transformation techniques are recommended to boost the performance of specific classification algorithms.

اسأل الذكاء الاصطناعي

Bookmark

Cite This Study

Jiang et al. (Sun,) studied this question.

synapsesocial.com/papers/6a18254b40149b897cb4abb4 https://doi.org/https://doi.org/10.1145/1390817.1390822

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

اسأل الذكاء الاصطناعي

Bookmark