Mental health stigma has profoundly harmful impacts but its complexity makes it difficult to evaluate.
Stigma may involve explicit derogation, but also subtler forms of blame, fear, paternalistic pity, social distancing, structural exclusion, and discrimination.
We introduce a theory-grounded benchmark for automatic evaluation of mental health stigma in online communication
Consisting of naturally occurring online news and social media text annotated with a fine-grained taxonomy of stigma across multiple mental health conditions.
Our annotation framework comprises
- a binary stigma-detection task
- a multi-level taxonomy covering
- stigma mode
- domain
- specific components of certain forms of stigma
We apply this framework to texts mentioning six mental health conditions and evaluate large language models alongside stigma-related classifiers for detecting sentiment, toxicity, and hate speech.
Results show that mental health stigma is not well captured by models trained to detect these neighboring constructs, and that LLMs often overpredict stigma unless given explicit operational rules - mirroring the importance of decision rules in human annotation.
We release the publicly available part of benchmark, annotations, prototypical exemplars of stigma and code at: https://github.com/jemimakang/mh_stigma.