GitHub
github.com › praj2408 › Jigsaw-Toxic-Comment-classification
GitHub - praj2408/Jigsaw-Toxic-Comment-Classification: The Toxic Comment Classification project is an application that uses deep learning to identify toxic comments as toxic, severe toxic, obscene, threat, insult, and identity hate based using various NLP algorithm · GitHub
The dataset used in this project is the Toxic Comment Classification Challenge from Kaggle. The dataset contains approximately 159,000 comments from Wikipedia talk pages that have been labeled by human annotators as toxic or non-toxic.
Author: praj2408
Starred by 55 users
Forked by 31 users
Languages: Python 95.8% | Dockerfile 4.2% | Python 95.8% | Dockerfile 4.2%
GitHub
github.com › tianqwang › Toxic-Comment-Classification-Challenge
GitHub - tianqwang/Toxic-Comment-Classification-Challenge: This repository is for the Machine Learning class project, Toxic Comment Classification Challenge in Kaggle · GitHub
The competition could be found here: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge · As a group of students with great interests in Natural Language Processing, as well as making online discussion more productive and respectful, we determined to work on this project and aim to build a model that is capable of detecting different types of toxicity like threats, obsenity, insults, and identity-based hate. The dataset we are using consists of comments from Wikipedia’s talk page edits.
Starred by 22 users
Forked by 21 users
Languages: Jupyter Notebook
GitHub
github.com › topics › toxic-comment-classification
toxic-comment-classification · GitHub Topics · GitHub
The dataset consists of large number of Wikipedia comments wh ... Fine-tuning FLAN-T5 with PPO and PEFT to generate less toxic text summaries. This notebook leverages Meta AI's hate speech reward model and utilizes RLHF techniques for improved safety. nlp toxic-comment-classification hate-speech-detection toxicity-analysis ppo-pytorch dialogue-summarization generative-ai detoxification reward-model
Starred by 6 users
Forked by 3 users
Languages: XSLT 59.0% | Scala 28.7% | CSS 12.3% | XSLT 59.0% | Scala 28.7% | CSS 12.3%
Starred by 17 users
Forked by 12 users
Languages: Jupyter Notebook 100.0% | Jupyter Notebook 100.0%
GitHub
github.com › lf-data › toxic_comments_classification
GitHub - lf-data/toxic_comments_classification: This project has been developed by my collegues and I on a dataset of toxic comments. Our task was to perform a multi-label classification analysis on about 200'000 comments in order to identify which of them were to be considered as "toxic", "severe toxic", "obscene", "insult", "threat" and/or "identity hate" (6 possible classes of toxicity). · GitHub
This project has been developed by my collegues and I on a dataset of toxic comments. Our task was to perform a multi-label classification analysis on about 200'000 comments in order to identify which of them were to be considered as "toxic", "severe toxic", "obscene", "insult", "threat" and/or "identity hate" (6 possible classes of toxicity).
Author: lf-data
Jay Speidell
jayspeidell.github.io › portfolio › project05-toxic-comments
Toxic Comment Classification - Natural Language Processing - Jay Speidell
The problem with this is that people will frequently write things they shouldn’t, and to maintain a positive community this toxic content and the users posting it need to be removed quickly. But they don’t have the resources to hire full-time moderators to review every comment. This problem led the Conversation AI team1, owned by Alphabet, to develop a large open dataset of labeled Wikipedia Talk Page comments, which will be the dataset used for the project.