Key points are not available for this paper at this time.
Background: A critical marker of high-quality systematic reviews is the identification and inclusion of all relevant, important studies. Up to 78% of systematic reviews have language restrictions; as a consequence, most reviews (93%) exclude at least 1 randomized controlled trial (RCT) (1). A 2012 study assessing Google Translate for translating nonEnglish-language studies recommended caution in using this service (2). Recently, Google updated its translation engine, reporting that it is markedly more accurate than previous versions (3). Objective: To examine the agreement between native-language and Google-translated abstractions of clinical trials published in languages other than English. Methods and Findings: We searched PubMed for RCTs published in 9 languages (Chinese, French, German, Italian, Japanese, Korean, Romanian, Russian, and Spanish) from 1 January 2000 through 15 December 2018 using the strategy randomized controlled trial Publication Type AND xx Language, where xx is the specific language. We included RCT interventions that reported outcome data and excluded publications with simultaneous English translations and those that were summaries rather than primary publications. The 5 most recent retrievable, eligible articles in each language were included to avoid arbitrariness in selecting papers. Articles were evaluated independently. Data from the original-language versions were abstracted by native-speaking physicians whose experience in conducting systematic reviews varied. The first author (J.L.J.), who speaks none of the 9 languages and is experienced in conducting systematic reviews, translated each article to English by using Google Translate (January to March 2019) and abstracted the data. Abstractors were blinded to each other's results. Data were abstracted into Excel (Microsoft) tables and included study characteristics and outcomes (both dichotomous and continuous), adverse events and withdrawals, and each abstractor's ratings of study quality on the basis of the Jadad scale (4) and Cochrane Risk of Bias Tool (5). We compared the accuracy of the translated and original data abstractions by calculating the simple percentage agreement. A priori, we specified that good agreement would be greater than 80%. Disagreements were resolved by consensus and tracked the source of any differences. In total, 6370 variables from 45 articles were abstracted, with 91% overall agreement (5791 of 6370) (Table). Agreement for the 9 languages ranged from 85% to 97%. Agreement was 96% on the Jadad quality ratings and 87% on the Cochrane Risk of Bias ratings. Only 1 disagreement in ratings was a result of translation; all other disagreements were the result of differences in interpretation of items in the quality or risk-of-bias measure. Table. Agreement in Quality Ratings and Abstraction of Outcomes, by Language* Discussion: Our results show that Google Translate is a viable tool for translating articles published in other languages into English for the purpose of abstracting data for systematic reviews, and that Google Translate is more accurate than previously reported (2). Agreement was similar for all 9 languages, approaching 90% (range, 85% to 97%). All but 1 difference in quality ratings was the result of varying interpretations of quality standards rather than differences in the translation. Differences in other abstracted variables were the result of human error and were not more likely in English- versus nonEnglish-language abstractions. Abstraction errors and disagreements are common in systematic reviews, prompting most reviewers to use dual or duplicate extraction methods. Our study had several limitations. We included a relatively small number of articles and had only 1 abstractor for each language. Some authors had no experience abstracting data or assessing study quality or bias for systematic reviews. The heterogeneous nature of the studies in our sample increased the potential for disagreement. A typical systematic review includes similar interventions and outcomes, making it easier to define the meaning of each quality or risk-of-bias measure question and to abstract outcomes. We conclude that Google Translate is a viable, accurate tool for translating nonEnglish-language trials for the purpose of conducting systematic reviews. Systematically excluding such trials from reviews may lead to substantial bias, particularly if nonEnglish-language trials tend to have a higher proportion of nonsignificant outcomes. At a minimum, we recommend that authors of systematic reviews not limit their searches to English-language studies and report how many nonEnglish-language trials were retrieved. Given the increased ability to perform sensitivity analyses and the potential for more precise estimates and generalizability of results, we recommend maximizing the number of trials available by using Google Translate to include nonEnglish-language trials in systematic reviews.
Jackson et al. (Mon,) studied this question.