Abstract This article examines the structural technological disadvantages faced by low-resource minority languages in India through a case study of Sindhi within contemporary Natural Language Processing and digital language infrastructures. While state-of-the-art language technologies depend on large standardized corpora, in post-Partition India, Sindhi’s territorial displacement, orthographic plurality, and declining formal language education have substantially impacted the development of digital corpus, annotation frameworks, and speech-resources suitable for the language. The study combines a dataset-level analysis of major Indian language technology initiatives with semi-structured interviews with Sindhi corpus practitioners to investigate how infrastructural design assumptions, such as script stability, Optical Character Recognition (OCR) compatibility, and broadcast-based speech pipelines, shape the language’s limited representation across digital platforms. The findings demonstrate that Sindhi’s marginal presence across digital language infrastructures reflects structural misalignment rather than simple data absence. In response, the article advances a community-integrated Digital Humanities framework that situates native linguistic expertise at the center of corpus creation. A collaborative Sindhi audio archival initiative is presented as a model for culturally grounded digitization that derives from the language’s historical and cultural context rather than adapting it to dominant technical standards. The article argues that minority language digitization must be understood as an infrastructural and cultural process, requiring language-specific strategies capable of sustaining both technological participation and long-term linguistic continuity.
Govindani et al. (Sat,) studied this question.