Measurement-validity study demonstrates that optical character recognition conceals rendering failures in text-to-image models, indicating a need for open-set visual text evaluation.
A measurement-validity study. Text-to-image (T2I) benchmarks for visual text rendering routinely score generated images with an OCR engine. We show that this practice has a structural defect: recognition-based judges are closed-set classifiers that must return some character from a known inventory, and therefore cannot report the single most informative failure mode — that the model drew a glyph that does not exist in the writing system at all. Using 300 Hangul syllable images from a commercial T2I model, exhaustively labelled by two human annotators (binary validity: Krippendorff's α = 0.942), we find that 21.7% of outputs are not valid Hangul syllables — they are Hangul-shaped glyphs that cannot be typed, e.g. a final consonant ㅅ drawn with an internal cross. Three OCR engines with independent architectures (Naver CLOVA, EasyOCR, PaddleOCR) each map 77–83% of these non-existent glyphs onto a valid syllable, and in 77% of jointly-answered cases the two engines snap to different characters — indicating the cause is the closed-set formulation itself rather than a shared dictionary. The consequences for benchmark numbers are large. Relative to human ground truth, OCR-only scoring understates absolute accuracy by 25.6–50.0 percentage points, and because the understatement grows with orthographic complexity, it exaggerates the complex-coda penalty by 2.3×–4.0×. Effect-size estimates also depend on engine choice, so results from papers using different OCR engines are not comparable. To isolate the judge property from any particular generator, we run a controlled paired experiment using no generative model at all: clean font renderings (control) versus the same glyphs with strokes programmatically added or removed (treatment, invalid by construction). EasyOCR's answer rate is statistically indistinguishable between conditions (96.4% vs 93.0%; p = 0.234; equivalence test at ±10pp margin passes), and its confidence distribution is likewise indistinguishable (0.458 vs 0.464, Mann–Whitney p = 0.841). The engine cannot tell a real character from a fabricated one, and its confidence carries no signal about the difference. The implication is not that a more accurate OCR is needed: concealment is decoupled from accuracy (the most accurate of the three engines still conceals 78.5%), because the defect lies in the closed-set output formulation rather than in recognition quality. What is needed instead is an evaluation protocol built around the limitation — reporting "not a valid character" as a distinct outcome, and quantifying rather than hiding the human fraction it requires. We release the human-labelled gold set, all code, and a verification script that recomputes every number in this note from raw data. A Korean translation (TECHNICAL_NOTE.ko.pdf) is included; the English original is the version of record.
No takes yet. Share an insight, caveat, or question.
Taewoo Lim (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: