PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 4, 2026Concurrency and Computation Practice and Experience0 citationsOpen Access

Rethinking Per‐Thread Computation for Machine Learning Design Exploration: A Work‐Efficient GPU Strategy for K‐Means and XGBoost

View Full Paper
OBOlavo BarrosIFIsabela FreitasPPPedro H. Pereira

Key Points

  • This research aims to enhance the performance of machine learning tasks using efficient GPU strategies.
  • Novel GPU implementation strategies based on a ‘more work per thread’ approach.
  • Experimental evaluation of multiple K-means, dimensionality reduction, and tree pruning for XGBoost.
  • Optimization using a Gini coefficient-based design exploration methodology.
  • Achieved speedups of 20 to 70 times compared to Nvidia's cuML library for K-means evaluations.
  • Dimensionality reduction via parallel K-means encoding improved encoding efficiency by up to two orders of magnitude while maintaining 1%-2% accuracy.
  • A Gini-based optimization resulted in a 400 times speedup in the model training process.

Abstract

ABSTRACT This paper presents novel GPU implementation strategies that effectively exploit the available parallelism, based on the “more work per thread” approach, for three machine learning design exploration tasks: Multiple K‐means evaluation, dimensionality reduction through parallel K‐means encoding for XGBoost trees, and XGBoost tree pruning optimization. Our experimental results demonstrate significant performance improvements across all implementations. For multiple K‐means evaluations, we achieve substantial speedups of 20 to 70 compared to Nvidia's cuML library. The dimensionality reduction approach, which uses parallel K‐means encoding, achieves encoding reductions of up to two orders of magnitude while preserving classification accuracy within 1%–2% of the original performance. Additionally, we propose a Gini coefficient‐based design exploration optimization that greatly reduces the number of models to be trained during the design space optimal search, achieving a 400 speedup. Furthermore, a parallel post‐pruning evaluation framework for XGBoost demonstrates the ability to remove up to 50% of tree nodes without significant loss of accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Barros et al. (2026) studied this question.

synapsesocial.com/papers/6a2117a4d499ed480b1707ddhttps://doi.org/10.1002/cpe.70744
Ask AI
Helpful
Bookmark
Share
View Full Paper