PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 11, 2012918 citations

Workload analysis of a large-scale key-value store

View Full Paper
BABerk AtikogluYXYuehai XuEFEitan Frachtenberg

Key Points

  • Characterize real-world key-value store workloads at scale to improve system performance, scalability, and synthetic benchmark modeling.
  • Collected and analyzed production traces capturing over 284 billion requests across five distinct Memcached use cases over several days.
  • Evaluated request composition, size, arrival rates, cache efficacy, temporal patterns, and storage behaviors to develop a representative workload model.
  • Demonstrated a GET-to-SET request ratio of 30:1, substantially exceeding standard assumptions in existing literature.
  • Showed that intense key access locality does not consistently guarantee high cache hit rates, with several use cases functioning like persistent storage rather than traditional caches.
  • Identified architectural deficiencies in Memcached implementations and proposed concrete optimization strategies to improve hit rates and efficiency.

Abstract

Key-value stores are a vital component in many scale-out enterprises, including social networks, online retail, and risk analysis. Accordingly, they are receiving increased attention from the research community in an effort to improve their performance, scalability, reliability, cost, and power consumption. To be effective, such efforts require a detailed understanding of realistic key-value workloads. And yet little is known about these workloads outside of the companies that operate them. This paper aims to address this gap. To this end, we have collected detailed traces from Facebook’s Memcached deployment, arguably the world’s largest. The traces capture over 284 billion requests from five different Memcached use cases over several days. We analyze the workloads from multiple angles, including: request composition, size, and rate; cache efficacy; temporal patterns; and application use cases. We also propose a simple model of the most representative trace to enable the generation of more realistic synthetic workloads by the community. Our analysis details many characteristics of the caching workload. It also reveals a number of surprises: a GET/SET ratio of 30:1 that is higher than assumed in the literature; some applications of Memcached behave more like persistent storage than a cache; strong locality metrics, such as keys accessed many millions of times a day, do not always suffice for a high hit rate; and there is still room for efficiency and hit rate improvements in Memcached’s implementation. Toward the last point, we make several suggestions that address the exposed deficiencies.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Atikoglu et al. (2012) studied this question.

synapsesocial.com/papers/69dff6c6af3798be7f8603fbhttps://doi.org/10.1145/2254756.2254766
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A comparison of file system workloads2000 · 445 citations
  2. 2Characteristics of WWW Client-based Traces1995 · 532 citations
  3. 3Web search using mobile cores2010 · 195 citations
  4. 4Distributed caching with memcached2004 · 513 citations
  5. 524/7 Characterization of petascale I/O workloads2009 · 220 citations