The Decode Block Size Heuristic in TPU Ragged Paged Attention Reduces LLM Inference Throughput by 28 to 69 Percent | Synapse