Automatic verbal aggression detection for Russian and American imageboards

Denis Gordeev
DOI: https://doi.org/10.48550/arXiv.1604.06648
2016-04-22
Computation and Language
Abstract:The problem of aggression for Internet communities is rampant. Anonymous forums usually called imageboards are notorious for their aggressive and deviant behaviour even in comparison with other Internet communities. This study is aimed at studying ways of automatic detection of verbal expression of aggression for the most popular American (4chan.org) and Russian (2ch.hk) imageboards. A set of 1,802,789 messages was used for this study. The machine learning algorithm word2vec was applied to detect the state of aggression. A decent result is obtained for English (88%), the results for Russian are yet to be improved.
What problem does this paper attempt to address?