Research in NLP lacks geographic diversity, and the question of how NLP can scaled to low-resourced languages has not yet been adequately solved."Low-resourced"-ness is a complex problem going beyond data availability and systemic problems in society. In this paper, we focus on the task of Translation (MT), that plays a crucial role for information and communication worldwide. Despite immense improvements in MT the past decade, MT is centered around a few high-resourced languages. As researchers cannot solve the problem of low-resourcedness alone, we propose research as a means to involve all necessary agents required in MT development process. We demonstrate the feasibility and scalability of research with a case study on MT for African languages. Its leads to a collection of novel translation datasets, MT for over 30 languages, with human evaluations for a third of them, enables participants without formal training to make a unique scientific. Benchmarks, models, data, code, and evaluation results are under https://github.com/masakhane-io/masakhane-mt.
No takes yet. Share an insight, caveat, or question.
Nekoto et al. (2020) studied this question.