PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 21, 2020470 citationsOpen Access

Beyond English-Centric Multilingual Machine Translation

View Full Paper
AFAngela FanSBShruti BhosaleHSHolger Schwenk

Key Points

Key points are not available for this paper at this time.

Abstract

Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only on data which was translated from or to English. While this is supported by large sources of training data, it does not reflect translation needs worldwide. In this work, we create a true Many-to-Many multilingual translation model that can translate directly between any pair of 100 languages. We build and open source a training dataset that covers thousands of language directions with supervised data, created through large-scale mining. Then, we explore how to effectively increase model capacity through a combination of dense scaling and language-specific sparse parameters to create high quality models. Our focus on non-English-Centric models brings gains of more than 10 BLEU when directly translating between non-English directions while performing competitively to the best single systems of WMT. We open-source our scripts so that others may reproduce the data, evaluation, and final M2M-100 model.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fan et al. (2020) studied this question.

synapsesocial.com/papers/6a0e98022c205f14b6c86f6fhttps://doi.org/10.48550/arxiv.2010.11125
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Mining the Web for bilingual text1999 · 236 citations
  2. 2Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation2020 · 82 citations
  3. 3The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali–English and Sinhala–English2019 · 91 citations
  4. 4Transformers without Tears: Improving the Normalization of Self-Attention2019 · 129 citations
  5. 5Balancing Training for Multilingual Neural Machine Translation2020 · 79 citations