The Decode Block Size Heuristic in TPU Ragged Paged Attention Wastes 28 to 69 Percent of LLM Inference Throughput | Synapse