PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 5, 20260 citations

Exploiting GPU Resources at VEGA for CMS Software Validation

View Full Paper
DSDaniele SpigaAFAdriano Di FlorioABAndrea Bocci

Key Points

  • The research aims to assess the integration of GPU resources at VEGA for validating CMS software releases.
  • Integrated VEGA as a sub-site extension for CPU and GPU resources
  • Utilized existing CMS workload management for distributed GPU job submission
  • Performed tuning for improved GPU utilization
  • Assessed error rates in comparison to traditional Grid sites
  • Successfully validated CMS software for GPU readiness
  • Achieved reasonable GPU utilization through tuning efforts
  • Error rates were assessed and compared favorably to conventional Grid sites

Abstract

In recent years, the CMS experiment has expanded the usage of HPC systems for data processing and simulation activities. These resources significantly extend the conventional pledged Grid compute capacity. Within the EuroHPC program, CMS applied for a “Benchmark Access” grant at VEGA in Slovenia, an HPC centre that is being used very successfully by the ATLAS experiment. For CMS, VEGA was integrated transparently as a sub-site extension to the Italian Tier-1 site at CNAF. In that first approach, only CPU resources were used, while all storage access was handled via CNAF through the network. Extending Grid sites with HPC resources was an established concept for CMS, however, in this project, HPC resources located in a different country from the Grid site were first integrated. CMS used the allocation primarily to validate a recent CMSSW release regarding its readiness for GPU usage. Former developments in the CMS workload management system that allow the targeting of GPU resources in the distributed infrastructure turned out to be instrumental and jobs could be submitted like any other release validation workflow. The presentation will detail aspects of the actual integration, some required tuning to achieve reasonable GPU utilisation, and an assessment of operational parameters like error rates compared to traditional Grid sites.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Spiga et al. (2025) studied this question.

synapsesocial.com/papers/698434f9f1d9ada3c1fb3bb9https://doi.org/10.1051/epjconf/202533701083/pdf
Ask AI
Helpful
Bookmark
Share
View Full Paper