首页 > AI前沿 > ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

arXiv自然语言 2026-09-15 06:00 3 阅读 查看原文

We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian.

The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identification across seven structured target categories.

ParsHate also distinguishes explicit and implicit hate, marks explicit and implicit targets, and provides span-level rationales.

Data collection combines random and score-stratified temporal sampling to reduce keyword-driven bias while preserving natural label distributions.

Applying SOTA models for Persian hate-speech detection on ParsHate shows moderate performance (79% F1), especially with samples from earlier years, and low performance with target identification (25.5% macro-F1).

This emphasizes the diverse sampling of hate speech in ParsHate and its challenging nature that requires more advanced methods for better performance.

Dataset is made publicly available.