In this article, we describe the design choices behind MLPerf, a machine learning performance benchmark that has become an industry standard. The first two rounds of the MLPerf Training benchmark helped drive improvements to software-stack performance and scalability, showing a 1.3× speedup in the top 16-chip results despite higher quality targets and a 5.5× increase in system scale. The first round of MLPerf Inference received over 500 benchmark results from 14 different organizations, showing growing adoption.
No takes yet. Share an insight, caveat, or question.
Mattson et al. (2020) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: