We study dependency parsing for four Arabic dialects (Gulf, Levantine, Egyptian, and Maghrebi).Since no syntactically annotated data exist for Arabic dialects, we train the parser on a Modern Standard Arabic (MSA) corpus, which creates an out-of-domain setting.We investigate methods to close the gap between the source (MSA) and target data (dialects), e.g., by training on syntactically similar sentences to the test data.For testing, we manually annotate a small data set from a dialectal corpus.We focus on parsing two linguistic phenomena, which are difficult to parse: Idafa and coordination.We find that we can improve results by adding in-domain MSA data while adding dialectal embeddings only results in minor improvements.
No takes yet. Share an insight, caveat, or question.
Mokh et al. (2024) studied this question.