PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 5, 20241 citations

Development of Code-Mixed Marathi-English Dataset for Hate Speech Detection

View Full Paper
PJPrasad JoshiVPVarsha Pathak

Key Points

Key points are not available for this paper at this time.

Abstract

In India, people commonly use a mix of English and their regional languages on social media, resulting in a substantial amount of code-mixed content. As of the present date, there is only a limited number of hate speech code-mixed datasets available for Indian languages, including Hindi-English, Bengali-Hindi, Malayalam-English, and Tamil-English. However, there are no resources or datasets available for Marathi-English code-mixed data with hate speech labels. This paper introduces a novel gold standard corpus for sentiment analysis of code-mixed text in Marathi-English, annotated by voluntary annotators. The gold standard corpus achieved an impressive Krippendorff's alpha score of above 0.9 for the dataset. Utilizing this new corpus, the paper establishes a benchmark for sentiment analysis in Marathi-English code-mixed texts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Joshi et al. (2024) studied this question.

synapsesocial.com/papers/68e75a12b6db6435876d1a7dhttps://doi.org/10.1109/esci59607.2024.10497428
Ask AI
Helpful
Bookmark
Share
View Full Paper