Position paper argues for improved metrics in document extraction to address selective blindness in evaluation.
Vision-language models now beat traditional OCR on the metrics used to rankdocument extraction systems. They also introduce a failure those metrics cannotregister: fluent substitution of exactly the values that matter — a tax ID withone digit changed, a subtotal that no longer reconciles against the total. This paper argues the failure is not imprecision but selective blindness, andidentifies three measurement failures that stack: character-level scores areinsensitive to entity substitution, schema validation certifies the shape of anoutput rather than its content, and cross-field arithmetic consistency isunmeasured in every document class. It argues for calibrated coverage in placeof point accuracy, and specifies the benchmark whose absence makes accuracyclaims on tax filings unfalsifiable. Position and synthesis paper. No new experiments.
No takes yet. Share an insight, caveat, or question.
Kuluru Vineeth (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: