Aggressive, Repetitive, Intentional, Visible, and Imbalanced: Refining Representations for Cyberbullying Classification

Caleb Ziems, Ymir Vigfusson, Fred Morstatter

Paper type: Full

Keywords: behaviors, cases, classification, classifiers, communities, detection, factors, large_scale, learning, linguistic, linguistic aspects, networks, performance, representations

2020-06-10 S2 (23:00-00:00 GMT) [Zoom] [Cal]

Abstract: Cyberbullying is a pervasive problem in online communities. To identify cyberbullying cases in large-scale social networks, content moderators depend on machine learning classifiers for automatic cyberbullying detection. However, existing models remain unfit for real-world applications, largely due to a shortage of publicly available training data and a lack of standard criteria for assigning ground truth labels. In this study, we address the need for reliable data using an original annotation framework. Inspired by social sciences research into bullying behavior, we characterize the nuanced problem of cyberbullying using five explicit factors to represent its social and linguistic aspects. We model this behavior using social network and language-based features, which improves classifier performance. These results demonstrate the importance of representing and modeling cyberbullying as a social phenomenon.

Similar Papers

Driving the Last Mile: Characterizing and Understanding Distracted Driving Posts on Social Networks
Hemank Lamba , Shashank Srikanth , Dheeraj Reddy Pailla , Shwetanshu Singh , Karandeep Singh Juneja , Ponnurangam Kumaraguru
Modeling and Measuring Expressed (Dis)belief in (Mis)information
Shan Jiang , Miriam Metzger , Andrew Flanagin , Christo Wilson
#MeTooMA: Multi-Aspect Annotations of Tweets Related to the MeToo Movement
Akash Gautam , Puneet Mathur , Rakesh Gosangi , Debanjan Mahata , Ramit Sawhney , Rajiv Ratn Shah