Commercial large language models bill, throttle, and budget context per token. Yet tokenizers assign more subword tokens to the same meaning in some languages than in others, so speakers of high-fertility languages pay a structural penalty before a model is ever invoked. This work measures that penalty across 16 Asian languages spanning seven language families and twelve scripts (Latin, Devanagari, Bengali, Sinhala, Tamil, Telugu, Kannada, Malayalam, Myanmar, Khmer, Lao, Thai), using parallel corpora so that the language effect is isolated from content. Contributions: - Same-content cost ratio across 16 Asian languages and 10 production tokenizers on FLORES-200 dev (n=997), with bootstrap 95% confidence intervals. Median 8. 9x on cl100kbase for the 11 non-Latin scripts, up to 11. 7x for Burmese. - Bytes-per-token (BPT) metric introduced for cross-script-fair comparison. - A 4, 000-cell needle-in-haystack recall benchmark with script-native markers, across 5 frontier OpenRouter models x 16 languages x 5 fill levels up to 131k tokens. - Headline finding: recall on non-Latin scripts collapses to 0-2/10 already at 4k tokens for 4 of 5 tested models — well inside every model's context window. gemini-2. 5 lash breaks the pattern with 84. 5% pooled recall on the same grid, indicating a vendor-level training/serving choice rather than an intrinsic limit. - Wall-clock latency penalty benchmark (800 measured trials + 240 warmup calls). Pooled Pearson r between costᵣatio and latencyᵣatio = 0. 314 — notably weaker than prior African-language findings — with per-model range -0. 19 (qwen) to +0. 74 (llama-3. 1-8b). Release artefacts: - Python package: https: //github. com/Helmo21/asia-fertility (MIT) - HuggingFace dataset: https: //huggingface. co/datasets/Helmo21/asia-fertility (CC-BY-SA 4. 0) - Manifest with SHA256 of config, pinned price + FX snapshots, tokenizer versions per run. This record bundles the preprint (PDF + LaTeX source), the three benchmark CSVs (leaderboard, NIAH v0. 3, latency), and the study manifest so any figure in the paper can be reconstructed byte-for-byte.
Antoine Pedretti (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: