PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 2026Humanities and Social Sciences Communications0 citationsOpen Access

Neural network models vs. MT evaluation metrics: a comparison between two approaches to automated assessment of information fidelity in consecutive interpreting

View Full Paper
XWXiaoman WangHeriot-Watt UniversityBWBinhua Wang

Key Points

  • The aim is to compare neural network models and machine translation evaluation metrics for assessing information fidelity in interpreting.
  • Compared automated metric scores from neural network models and MT metrics with human assessment scores at the sentence level.
  • Used three-cluster and four-cluster analysis to validate machine evaluation applicability.
  • Evaluated correlations between models like GPT, LLaMA, and traditional MT metrics.
  • Neural network models outperformed traditional machine translation evaluation metrics.
  • Moderate correlations observed with LLM embeddings from GPT (r = 0.47) and LLaMA (r = 0.46).
  • Stronger correlations found with GPT scores (r = 0.53) and MUSE embeddings (r = 0.55).
  • Cluster analysis indicates that multiple machine evaluation metrics can approximate human judgments effectively.

Abstract

Abstract Informational fidelity is the most important aspect in assessing interpreting quality, which can be assessed in two methods: human assessment and automated machine translation evaluation metrics. Despite their prevalent use in interpreter training, human assessment is time-consuming, labour-intensive, and mentally demanding. In terms of automated methods, machine translation metrics assess fidelity through inter-lingual comparisons between the interpretation and reference translations across multiple versions; and reference-free neural network models assess fidelity through cross-lingual comparison. The study proceeds to compare the automated metric scores derived from these two approaches with human scores at the sentence level, in order to explore the correlation between machine and human assessments. The applicability of machine evaluation was further substantiated through both three-cluster and four-cluster analysis. The results demonstrated that neural network models, which are trained to generate cross-lingual embeddings and LLM, outperform traditional machine translation evaluation metrics. Moderate correlations were observed with embeddings from LLM such as GPT (Pearson’s r = 0.47) and LLaMA ( r = 0.46). Meanwhile, stronger correlations were noted with GPT scores ( r = 0.53) and MUSE embeddings ( r = 0.55). The correlation is strong enough to be statistically significant and meaningful, though not so strong as to suggest near-perfect predictability. Cluster analysis reveals that aggregating multiple machine evaluation metrics can effectively approximate human judgments, highlighting distinct levels of fidelity and the necessity of a composite metric. This research posits that the adoption of pre-trained neural network models has good potential as a viable method for automated assessment of information fidelity in interpreting. This method holds promise for broader applications in low-stake interpreting assessments and can supplement human scoring, particularly in large-scale assessment tasks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2026) studied this question.

synapsesocial.com/papers/69b4b9fb18185d8a39802574https://doi.org/10.1057/s41599-026-06562-z
Ask AI
Helpful
Bookmark
Share
View Full Paper