Key points are not available for this paper at this time.
The introduction of transformer 7 has led to the development of large language models (LLMs) like Bidirectional Encoder Representations from Transformers (BERT). However, the downstream performance of LLMs is often poor for low resourced languages (LRLs) such as Swahili because of their small dataset size.
Ndomba et al. (Sun,) studied this question.