Randomized trial evaluates HyperToken's effectiveness in low-resource languages, suggesting improved tokenization accuracy.
Standard subword tokenizers are typically evaluated on aggregate compression metrics, which can obscure script-specific failure modes that disproportionately affect low-resource languages. We present HyperToken, a script-aware tokenizer that combines abugida syllable clustering with unigram vocabulary pruning, and compare it against PU-Tok (PolyglotUnigramTokenizer), a general-purpose polyglot baseline, on 30 typologically diverse languages from FLORES-200. Both tokenizers are trained under identical conditions (same corpus, same 24,000-token vocabulary budget, same held-out evaluation split). HyperToken achieves lower tokens/character than PU-Tok on 26 of 30 languages, including all nine Indic abugidas evaluated, and achieves perfect round-trip reconstruction (decode(encode(text)) == text) on all 30 languages. PU-Tok, by contrast, fails round-trip on 8 of 30 languages, concentrated in scripts that use nukta diacritics to represent loanword phonology. We argue that round-trip correctness should be reported as a first-class metric alongside compression when evaluating multilingual tokenizers.This deposit contains the paper, LaTeX source, the HyperToken tokenizer package, the PU-Tok baseline implementation, benchmark scripts, and raw/parsed results needed to reproduce all reported numbers.
No takes yet. Share an insight, caveat, or question.
Prashant ARYA (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: