PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 1, 201022 citations

Auto-tuning Dense Matrix Multiplication for GPGPU with Cache

View Full Paper
XCXiang CuiYCYifeng ChenCZChangyou Zhang

Key Points

Key points are not available for this paper at this time.

Abstract

In this paper we discuss about our experiences in improving the performance of GEMM (both single and double precision) on Fermi architecture using CUDA, and how the new features of Fermi such as cache affect performance. It is found that the addition of cache in GPU on one hand helps the processers take advantage of data locality occurred in runtime but on the other hand renders the dependency of performance on algorithmic parameters less predictable. Auto tuning then becomes a useful technique to address this issue. Our auto-tuned SGEMM and DGEMM reach 563 GFlops and 253 GFlops respectively on Tesla C2050. The design and implementation entirely use CUDA and C and have not benefited from tuning at the level of binary code.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cui et al. (2010) studied this question.

synapsesocial.com/papers/6a1c68c194dbf6307b2fbcb1https://doi.org/10.1109/icpads.2010.64
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Benchmarking GPUs to tune dense linear algebra2008 · 484 citations
  2. 2Benchmarking GPUs to tune dense linear algebra2008 · 726 citations
  3. 3Improving Performance of Matrix Multiplication and FFT on GPU2009 · 35 citations
  4. 4A Note on Auto-tuning GEMM for GPUs2009 · 186 citations
  5. 5Optimization principles and application performance evaluation of a multithreaded GPU using CUDA2008 · 909 citations