PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 18, 2024Computing and Software for Big Science12 citationsOpen Access

Optimizing High-Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

View Full Paper
CSC. SavardNMN. ManganelliBHB. Holzman

Key Points

Key points are not available for this paper at this time.

Abstract

Abstract With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high-throughput access to a HEP-specific GNN and report on the outcome.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Savard et al. (2024) studied this question.

synapsesocial.com/papers/68e5fda8b6db643587591049https://doi.org/10.1007/s41781-024-00123-2
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Learning long-term dependencies with gradient descent is difficult1994 · 8,614 citations
  2. 2The ATLAS Experiment at the CERN Large Hadron Collider2008 · 4,267 citations
  3. 3Deep Residual Learning for Image Recognition2016 · 228,344 citations
  4. 4XGBoost2016 · 52,531 citations
  5. 5Parallel processing and distributed computing2022 · 3 citations