Jay Speidell
jayspeidell.github.io › portfolio › project05-toxic-comments
Toxic Comment Classification - Natural Language Processing - Jay Speidell
The problem with this is that people will frequently write things they shouldn’t, and to maintain a positive community this toxic content and the users posting it need to be removed quickly. But they don’t have the resources to hire full-time moderators to review every comment. This problem led the Conversation AI team1, owned by Alphabet, to develop a large open dataset of labeled Wikipedia Talk Page comments, which will be the dataset used for the project.
arXiv
arxiv.org › pdf › 1903.06765 pdf
A Machine Learning Approach to Comment Toxicity Classification
any piece of text and detecting different types of toxicity like obscenity, threats, insults and identity-based hatred. The labelled Wikipedia Com- ment Dataset prepared by Jigsaw is used for the purpose.
TensorFlow
tensorflow.org › datasets › wikipedia_toxicity_subtypes
wikipedia_toxicity_subtypes | TensorFlow Datasets
The toxicity and toxicity subtype labels are binary values (0 or 1) indicating whether the majority of annotators assigned that attribute to the comment text. This config is a replica of the data released for the Jigsaw Toxic Comment Classification Challenge on Kaggle, with the test dataset joined with the test_labels released after the competition, and test data not used for scoring dropped.
Stanford University
cs229.stanford.edu › proj2019spr › report › 71.pdf pdf
Toxic Comment Detection and Classification Hao Li haoli94@stanford.edu
is the word count vector for comment i with label yi ∈ · {−1, 1} indicating whether it’s toxic or not.
Medium
johnfengphd.medium.com › classifying-toxic-comments-with-nlp-3418ee9f399
Toxic Comments Classification — NLP Project | by John Feng | Medium
February 17, 2023 - It is also important to consider the larger social implications of toxic comments. They can perpetuate harmful stereotypes, discrimination, and prejudice, and contribute to a divisive and toxic online culture. This can have a ripple effect on offline communities, leading to further social and political polarization. The dataset comes from Kaggle Toxic Comments Classification, a very popular dataset for NLP projects.
GitHub
github.com › tianqwang › Toxic-Comment-Classification-Challenge
GitHub - tianqwang/Toxic-Comment-Classification-Challenge: This repository is for the Machine Learning class project, Toxic Comment Classification Challenge in Kaggle · GitHub
The competition could be found here: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge · As a group of students with great interests in Natural Language Processing, as well as making online discussion more productive and respectful, we determined to work on this project and aim to build a model that is capable of detecting different types of toxicity like threats, obsenity, insults, and identity-based hate. The dataset we are using consists of comments from Wikipedia’s talk page edits.
Starred by 22 users
Forked by 21 users
Languages: Jupyter Notebook
GitHub
github.com › queloibtskts › Toxic-Comment-Detection
GitHub - queloibtskts/Toxic-Comment-Detection: Group project done by Xiaoyue Lin, Xiyan Zhang and Yuling Zheng in MAST30034 Applied Data Science
In this project, we aimed to classify toxic comments using Machine Learning models, which can be applied in censoring inappropriate comments on social media platforms. The dataset used for training and testing is obtained from Jigsaw Unintended Bias in Toxicity Classification competition on Kaggle: https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification/data
Author: queloibtskts
Medium
medium.com › @alaeddine.grine › toxic-comment-classification-317628632336
Toxic Comment Classification. NLP multi-label classification using… | by Ala Eddine GRINE | Medium
September 22, 2023 - Platforms struggle to identify and flag potentially harmful or offensive online comments, leading many communities to restrict or shut down user comments altogether. Kaggle issued a challenge to build a multi-label classification model that’s able to detect different types of toxicity like threats, obscenity and insults, and thus help make online discussion more productive and respectful. The data for the problem is a dataset of 159,571 comments from Wikipedia’s talk page edits.
arXiv
arxiv.org › pdf › 2304.06934 pdf
Classification of social media Toxic comments using Machine learning models
Kaggle [30], is a multi-label dataset and contains labels such · as toxic, severe_toxic, obscene, threat, insult, and · identity_hate. The non-toxic comments belong to one class, while from the other comments only those comments are · selected that have toxic labels. It means that the comments · that label severe_toxic, obscene, threat, insult, and ... TABLE 1. Example ...
Rjwave
rjwave.org › jaafr › papers › JAAFR2504003.pdf pdf
Verbiclear - Toxic Comment Detector and Remover
Comment Text – User comment. Toxicity Label – Tagged as "Toxic" or "Non-Toxic". The data was divided into 80% training and 20% validation sets with stratified sampling to maintain class balance. The dataset is mixture of Kaggle’s jigsaw dataset and a custom dataset created through surveys and research by our team on real
ACL Anthology
aclanthology.org › 2021.woah-1.17.pdf pdf
Proceedings of the Fifth Workshop on Online Abuse and Harms, pages 157–163
Table 1: Datasets currently included in the toxic comment collection (sorted by year of publication). For · this tabular presentation, we combined labels, e.g., target represents several different labels of targets. ... Source Lang. ... Nuha Albadi, Maram Kurdi, and Shivakant Mishra. 2018. Are they our brothers? analysis and detection...