ABSTRACT This paper presents novel GPU implementation strategies that effectively exploit the available parallelism, based on the “more work per thread” approach, for three machine learning design exploration tasks: Multiple K‐means evaluation, dimensionality reduction through parallel K‐means encoding for XGBoost trees, and XGBoost tree pruning optimization. Our experimental results demonstrate significant performance improvements across all implementations. For multiple K‐means evaluations, we achieve substantial speedups of 20 to 70 compared to Nvidia's cuML library. The dimensionality reduction approach, which uses parallel K‐means encoding, achieves encoding reductions of up to two orders of magnitude while preserving classification accuracy within 1%–2% of the original performance. Additionally, we propose a Gini coefficient‐based design exploration optimization that greatly reduces the number of models to be trained during the design space optimal search, achieving a 400 speedup. Furthermore, a parallel post‐pruning evaluation framework for XGBoost demonstrates the ability to remove up to 50% of tree nodes without significant loss of accuracy.
Barros et al. (2026) studied this question.