PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 27, 2026Journal of Open Humanities Data0 citationsOpen Access

Introducing the First Module of the Multimedia Corpus of Spoken Kazakh Language

GTGiorgia TroianiAFAndrey Filchenko

Key Points

  • To document and provide a resource for the contemporary Kazakh language as spoken in Kazakhstan and Xinjiang.
  • Collected 33 audio recordings totaling approximately 12 hours from 78 native Kazakh speakers.
  • Produced time-aligned transcriptions of the conversations for analysis.
  • Anonymized data published under a CC BY 4.0 license.
  • Dataset captures naturally occurring conversations among Kazakh speakers.
  • Intended for empirical analysis in linguistics and related fields.
  • Facilitates reuse for a variety of linguistic studies.

Abstract

The first module of the Multimedia Corpus of Spoken Kazakh Language is a dataset documenting contemporary Kazakh as spoken in Kazakhstan and Xinjiang (China). It includes 33 audio recordings (ca. 12 hours) and time-aligned transcriptions collected from 78 participants. Recordings feature naturally occurring conversation among native Kazakh speakers. The corpus is anonymized and published under a CC BY 4.0 license. The dataset is intended as a linguistic resource for the empirical analysis of Kazakh and it is suitable for reuse in a wide-range of linguistics-adjacent disciplines concerned with the analysis of naturally occurring language in use.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Troiani et al. (2026) studied this question.

synapsesocial.com/papers/6a168b160c924ddd1bd59f4fhttps://doi.org/10.5334/johd.529
Ask AI
Helpful
Bookmark
Share
View Full Paper