Translation quality estimation (QE) must support both sentence-level scoring and local error diagnosis under limited, heterogeneous supervision. This study proposes MCA-LLM-QEER, which combines leakage-controlled multi-source corpus augmentation, source-language/crosslingual/ target-language semantic evidence, explicit linguistic features, and structured LLM evidence in a joint prediction framework. On WMT'23-QE, the model achieves an average Spearman correlation of 0.687 and a Pearson correlation of 0.701, with MAE and RMSE of 0.132 and 0.178. On Domain-QE, error-span F1, error-type Macro-F1, and severity Macro-F1 reach 0.759, 0.708, and 0.724, respectively. Cross-domain and source-target mismatch tests further show that the framework improves robustness to domain shift and fluent-butunfaithful translations. The results support a unified reference-free QE design in which global quality prediction and fine-grained error diagnosis are learned from complementary bilingual evidence rather than from target-side fluency alone.
No takes yet. Share an insight, caveat, or question.
Wang et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: