Empirical comparison reveals equivalent perplexity between grouped-query and multi-head attention in small Transformers, indicating memory savings persist at small scales.
We present a controlled empirical comparison between grouped-query attention (GQA) and standard multi-head attention (MHA) in a small (~13.6M parameter) GPT-style decoder-only Transformer, trained from scratch on the TinyStories dataset. Holding architecture, random seed, data, and training budget (30,000 steps) fixed, and varying only the number of key/value heads (2 for GQA vs. 4 for MHA, with 4 query heads throughout), we find the two configurations reach nearly identical validation loss (2.2500 for GQA vs. 2.2451 for MHA; perplexity 9.49 vs. 9.44) while GQA uses exactly half the key/value projection parameters (65,536 vs. 131,072). This small, controlled result is consistent with the finding reported for GQA at much larger scale: the technique gives a real parameter and memory saving at negligible cost to model quality, and this cost remains negligible even at a scale three-plus orders of magnitude smaller than where GQA is typically studied.
No takes yet. Share an insight, caveat, or question.
Kapil (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: