Empirical Analysis of Multi-Task Learning for Reducing Identity Bias in Toxic Comment Detection

Ameya Vaidya, Feng Mai, Yue Ning

Paper type: Full

Keywords: attention, bias, deep learning, detection, groups, identities, learning, sources, toxic, toxicity

2020-06-10 P8 (22:00-23:00 GMT) [Zoom] [Cal]

Abstract: With the recent rise of toxicity in online conversations on social media platforms, using modern machine learning algorithms for toxic comment detection has become a central focus of many online applications. Researchers and companies have developed a variety of shallow and deep learning models to identify toxicity in online conversations, reviews, or comments with mixed successes. However, these existing approaches have learned to incorrectly associate non-toxic comments that have certain trigger-words (e.g. gay, lesbian, black, muslim) as a potential source of toxicity. In this paper, we evaluate dozens of state-of-the-art models with the specific focus of reducing model bias towards these commonly-attacked identity groups. We propose a multi-task learning model with an attention layer that jointly learns to predict the toxicity of a comment as well as the identities present in the comments in order to reduce this bias. We then compare our model to an array of shallow and deep-learning models using metrics designed especially to test for unintended model bias within these identity groups.

Similar Papers

Learn2Link: Linking the Social and Academic Profiles of Researchers
Asmelash Teka Hadgu , Jayanth Kumar Reddy Gundam
Characterizing Variation in Toxic Language by Social Context
Bahar Radfar , Karthik Shivaram , Aron Culotta
The Effect of Homophily on Disparate Visibility of Minorities in People Recommender Systems
Francesco Fabbri , Francesco Bonchi , Ludovico Boratto , Carlos Castillo
Social Media Relevance Filtering Using Perplexity-Based Positive-Unlabelled Learning
Sunghwan Mac Kim , Stephen Wan , Cécile Paris , Andreas Duenser