🌐
Jay Speidell
jayspeidell.github.io › portfolio › project05-toxic-comments
Toxic Comment Classification - Natural Language Processing - Jay Speidell
The problem with this is that people will frequently write things they shouldn’t, and to maintain a positive community this toxic content and the users posting it need to be removed quickly. But they don’t have the resources to hire full-time moderators to review every comment. This problem led the Conversation AI team1, owned by Alphabet, to develop a large open dataset of labeled Wikipedia Talk Page comments, which will be the dataset used for the project.
🌐
Kaggle
kaggle.com › datasets › sahideseker › toxic-comment-classification-dataset
Toxic Comment Classification Dataset
April 4, 2025 - Classify comments based on whether they are toxic or not. ... This dataset contains sample social media comments labeled as toxic or non-toxic.
🌐
arXiv
arxiv.org › pdf › 1903.06765 pdf
A Machine Learning Approach to Comment Toxicity Classification
any piece of text and detecting different types of toxicity like obscenity, threats, insults and identity-based hatred. The labelled Wikipedia Com- ment Dataset prepared by Jigsaw is used for the purpose.
🌐
TensorFlow
tensorflow.org › datasets › wikipedia_toxicity_subtypes
wikipedia_toxicity_subtypes | TensorFlow Datasets
The toxicity and toxicity subtype labels are binary values (0 or 1) indicating whether the majority of annotators assigned that attribute to the comment text. This config is a replica of the data released for the Jigsaw Toxic Comment Classification Challenge on Kaggle, with the test dataset joined with the test_labels released after the competition, and test data not used for scoring dropped.
🌐
Stanford University
cs229.stanford.edu › proj2019spr › report › 71.pdf pdf
Toxic Comment Detection and Classification Hao Li haoli94@stanford.edu
is the word count vector for comment i with label yi ∈ · {−1, 1} indicating whether it’s toxic or not.
🌐
Hugging Face
huggingface.co › datasets › google › jigsaw_toxicity_pred
google/jigsaw_toxicity_pred · Datasets at Hugging Face
This dataset consists of a large number of Wikipedia comments which have been labeled by human raters for toxic behavior.
🌐
Medium
johnfengphd.medium.com › classifying-toxic-comments-with-nlp-3418ee9f399
Toxic Comments Classification — NLP Project | by John Feng | Medium
February 17, 2023 - It is also important to consider the larger social implications of toxic comments. They can perpetuate harmful stereotypes, discrimination, and prejudice, and contribute to a divisive and toxic online culture. This can have a ripple effect on offline communities, leading to further social and political polarization. The dataset comes from Kaggle Toxic Comments Classification, a very popular dataset for NLP projects.
🌐
GitHub
github.com › tianqwang › Toxic-Comment-Classification-Challenge
GitHub - tianqwang/Toxic-Comment-Classification-Challenge: This repository is for the Machine Learning class project, Toxic Comment Classification Challenge in Kaggle · GitHub
The competition could be found here: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge · As a group of students with great interests in Natural Language Processing, as well as making online discussion more productive and respectful, we determined to work on this project and aim to build a model that is capable of detecting different types of toxicity like threats, obsenity, insults, and identity-based hate. The dataset we are using consists of comments from Wikipedia’s talk page edits.
Starred by 22 users
Forked by 21 users
Languages: Jupyter Notebook
Find elsewhere
🌐
GitHub
github.com › praj2408 › Jigsaw-Toxic-Comment-Classification
GitHub - praj2408/Jigsaw-Toxic-Comment-Classification: The Toxic Comment Classification project is an application that uses deep learning to identify toxic comments as toxic, severe toxic, obscene, threat, insult, and identity hate based using various NLP algorithm · GitHub
The dataset used in this project is the Toxic Comment Classification Challenge from Kaggle. The dataset contains approximately 159,000 comments from Wikipedia talk pages that have been labeled by human annotators as toxic or non-toxic.
Author: praj2408
🌐
Towards Data Science
towardsdatascience.com › home › latest › toxic comment classification using lstm and lstm-cnn.
Toxic Comment Classification using LSTM and LSTM-CNN. | Towards Data Science
March 5, 2025 - Toxic Comment Classifier is a competition that has been organized by Jigsaw/Conversation AI and hosted on Kaggle. The data set for building the classification model was acquired from the competition site and it included the training set as well as the test set.
🌐
Papers with Code
paperswithcode.com › dataset › toxic-comment-classification-challenge
Jigsaw Toxic Comment Classification Dataset
We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks.
🌐
Medium
medium.com › @alaeddine.grine › toxic-comment-classification-317628632336
Toxic Comment Classification. NLP multi-label classification using… | by Ala Eddine GRINE | Medium
September 22, 2023 - Platforms struggle to identify and flag potentially harmful or offensive online comments, leading many communities to restrict or shut down user comments altogether. Kaggle issued a challenge to build a multi-label classification model that’s able to detect different types of toxicity like threats, obscenity and insults, and thus help make online discussion more productive and respectful. The data for the problem is a dataset of 159,571 comments from Wikipedia’s talk page edits.
🌐
Kaggle
kaggle.com › datasets › miadul › toxic-comments-detection-dataset
Social Media Toxic Comment Dataset | Kaggle
November 23, 2025 - Multilingual Toxic Comment Dataset for NLP & Social Media Moderation
🌐
arXiv
arxiv.org › pdf › 2304.06934 pdf
Classification of social media Toxic comments using Machine learning models
Kaggle [30], is a multi-label dataset and contains labels such · as toxic, severe_toxic, obscene, threat, insult, and · identity_hate. The non-toxic comments belong to one class, while from the other comments only those comments are · selected that have toxic labels. It means that the comments · that label severe_toxic, obscene, threat, insult, and ... TABLE 1. Example ...
🌐
Kaggle
kaggle.com › datasets › julian3833 › jigsaw-toxic-comment-classification-challenge
jigsaw-toxic-comment-classification-challenge
November 11, 2021 - Data from Toxic Comment Classification Challenge · Data CardCode (308)Discussion (0)Suggestions (0) info · 7.06 · CC0: Public Domain · Not specified · NLP · get_app · fullscreen · chevron_right · DetailCompactColumn · 7 of 7 columnskeyboard_arrow_down ·
🌐
Medium
medium.com › @ketakee › toxic-comment-detector-6708e2a676d6
Toxic Comment Detector. You can try this yourself by going to… | by ketakee | Medium
December 16, 2019 - This could easily be ported to your website’s comment section and you can change it to have the colour of the comment being entered change as per the toxicity of the text. We use kaggle toxic Comment classification challenge dataset to build ...
🌐
Rjwave
rjwave.org › jaafr › papers › JAAFR2504003.pdf pdf
Verbiclear - Toxic Comment Detector and Remover
Comment Text – User comment. Toxicity Label – Tagged as "Toxic" or "Non-Toxic". The data was divided into 80% training and 20% validation sets with stratified sampling to maintain class balance. The dataset is mixture of Kaggle’s jigsaw dataset and a custom dataset created through surveys and research by our team on real
🌐
ACL Anthology
aclanthology.org › 2021.woah-1.17.pdf pdf
Proceedings of the Fifth Workshop on Online Abuse and Harms, pages 157–163
Table 1: Datasets currently included in the toxic comment collection (sorted by year of publication). For · this tabular presentation, we combined labels, e.g., target represents several different labels of targets. ... Source Lang. ... Nuha Albadi, Maram Kurdi, and Shivakant Mishra. 2018. Are they our brothers? analysis and detection...
🌐
Analytics Vidhya
analyticsvidhya.com › home › predicting the toxicity of comments using text classification
Predicting the Toxicity of Comments Using Text Classification
October 15, 2024 - We will use the toxic comments dataset, which comprises a large number of Wikipedia comments labelled as toxic by human raters. There are six types of toxicities in this data: toxic, severe-toxic, obscene, threat, insult and identity-hate.