首页 > AI前沿 > BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text

BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text

arXiv自然语言 2026-10-02 12:00 8 阅读 查看原文

Regional linguistic variation remains an important challenge for Bangla natural language processing, particularly in informal and non-standard text.

This paper introduces BanglaDial-Abuse, a balanced Bengali-script dataset developed for regional dialect identification in abusive and hostile Bangla text.

The dataset contains 1,000 sentences distributed equally across four linguistic varieties: Standard Bangla, Chattagram, Sylhet, and Barishal, with 250 samples per class.

The resource was constructed using a corpus-grounded synthetic procedure incorporating regional variation in pronouns, possessive forms, verb morphology, negation, interrogative structures, postpositions, vocabulary, and Bengali-script spelling conventions while preserving the underlying hostile or abusive meaning.

Descriptive analysis shows broadly comparable sentence-length distributions but partially distinct lexical spaces across the four classes.

Pairwise Jaccard vocabulary similarity ranges from 0.37 to 0.56.

The primary task is four-class regional dialect identification rather than binary abusive-text detection.

The dataset is publicly available through Zenodo under a Creative Commons Attribution 4.0 license.

The current version is intended as a research and prototyping corpus rather than a native-speaker-validated gold-standard linguistic resource.

Keywords: Bangla, Bengali, dialect identification, regional dialect, abusive language, low-resource NLP, Chattagram, Sylhet, Barishal, dataset