PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 3, 2026SHILAP Revista de lepidopterología0 citationsOpen Access

Validating open-source machine translation for quantitative text analysis

HLHauke LichtRSRonja Ida SczepanskiMLMoritz Laurer

Key Points

  • This study aims to determine if open-source machine translation affects quantitative results in text analysis compared to commercial services.
  • Assessed open-source MT models against commercial services for text analysis.
  • Extended previous work by applying Transformer-based supervised text classification to multilingual corpora.
  • Negligible differences in quantitative results between open-source and commercial MT approaches, indicating strong alignment.

Abstract

Abstract Machine translation (MT) is an essential tool in many multilingual computational text analysis applications. However, relying on commercial services like Google Translate or DeepL limits reproducibility and can be expensive. This paper assesses the viability of a reproducible, transparent, and affordable alternative: open-source MT models. We ask whether using open-source MT models instead of commercial services substantially changes the measurements obtained from multilingual corpora by extending an influential study by de Vries et al. and contributing an original study focusing on Transformer-based supervised text classification. Our findings reveal negligible differences in results between the two MT approaches, suggesting that open-source MT models are highly valuable tools for multilingual text analysis.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Licht et al. (2026) studied this question.

synapsesocial.com/papers/69f6e5868071d4f1bdfc6390https://doi.org/10.1017/psrm.2026.10102
Ask AI
Helpful
Bookmark
Share
View Full Paper