Batch Scaling and Goodput of a Tuned Attention Kernel on TPU v6e: Throughput, Latency, Energy, and Cost Measurements of vLLM Serving Gemma 4 31B | Synapse