🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.LabelEncoder.html
LabelEncoder — scikit-learn 1.9.1 documentation
Encode categorical features as a one-hot numeric array. Examples · LabelEncoder can be used to normalize labels. >>> from sklearn.preprocessing import LabelEncoder >>> le = LabelEncoder() >>> le.fit([1, 2, 2, 6]) LabelEncoder() >>> le.classes_ array([1, 2, 6]) >>> le.transform([1, 1, 2, 6]) array([0, 0, 1, 2]...) >>> le.inverse_transform([0, 0, 1, 2]) array([1, 1, 2, 6]) It can also be used to transform non-numerical labels (as long as they are hashable and comparable) to numerical labels.
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OneHotEncoder.html
OneHotEncoder — scikit-learn 1.9.1 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to a binary one-hot encoding. >>> from sklearn.preprocessing import OneHotEncoder
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.OrdinalEncoder.html
OrdinalEncoder — scikit-learn 1.9.1 documentation
Given a dataset with two features, we let the encoder find the unique values per feature and transform the data to an ordinal encoding. >>> from sklearn.preprocessing import OrdinalEncoder >>> enc = OrdinalEncoder() >>> X = [['Male', 1], ['Female', 3], ['Female', 2]] >>> enc.fit(X) OrdinalEncoder() >>> enc.categories_ [array(['Female', 'Male'], dtype=object), array([1, 2, 3], dtype=object)] >>> enc.transform([['Female', 3], ['Male', 1]]) array([[0., 2.], [1., 0.]])
🌐
scikit-learn
scikit-learn.org › stable › modules › generated › sklearn.preprocessing.TargetEncoder.html
TargetEncoder — scikit-learn 1.9.1 documentation
This unsupervised encoding is better suited for low cardinality categorical variables as it generate one new feature per unique category. ... Micci-Barreca, Daniele. “A preprocessing scheme for high-cardinality categorical attributes in classification and prediction problems” SIGKDD Explor. Newsl. 3, 1 (July 2001), 27–32. ... >>> import numpy as np >>> from sklearn.preprocessing import TargetEncoder >>> X = np.array([["dog"] * 20 + ["cat"] * 30 + ["snake"] * 38], dtype=object).T >>> y = [90.3] * 5 + [80.1] * 15 + [20.4] * 5 + [20.1] * 25 + [21.2] * 8 + [49] * 30 >>> enc_auto = TargetEncoder(smooth="auto") >>> X_trans = enc_auto.fit_transform(X, y)
🌐
GitHub
github.com › scikit-learn-contrib › category_encoders
GitHub - scikit-learn-contrib/category_encoders: A library of sklearn compatible categorical variable encoders · GitHub
$ python setup.py install · or · pip install category_encoders · or · conda install -c conda-forge category_encoders · To install the development version, you may use: pip install --upgrade git+https://github.com/scikit-learn-contrib/category_encoders · All of the encoders are fully compatible sklearn transformers, so they can be used in pipelines or in your existing scripts.
Author: scikit-learn-contrib
🌐
GeeksforGeeks
geeksforgeeks.org › encoding-categorical-data-in-sklearn
Encoding Categorical Data in Sklearn - GeeksforGeeks
November 25, 2024 - Label Encoding is a simple and straightforward method that assigns a unique integer to each category. This method is suitable for ordinal data where the order of categories is meaningful.
🌐
Scikit-learn
contrib.scikit-learn.org › category_encoders
Category Encoders — Category Encoders 2.11.1 documentation
This can cause problems in sklearn versions prior to 1.2.0. In order to ensure full compatibility with sklearn set sklearn to also output DataFrames. This can be done by ... Pipeline( steps=[ ("preprocessor", SomePreprocessor().set_output("pandas"), ("encoder", SomeEncoder()), ] )
🌐
Scikit-learn course
inria.github.io › scikit-learn-mooc › python_scripts › 03_categorical_pipeline.html
Encoding of categorical variables — Scikit-learn course
We can encode a single feature (e.g. "education") to illustrate how the encoding works. from sklearn.preprocessing import OneHotEncoder encoder = OneHotEncoder(sparse_output=False).set_output(transform="pandas") education_encoded = encoder.fit_transform(education_column) education_encoded
🌐
scikit-learn
scikit-learn.org › dev › modules › generated › sklearn.preprocessing.LabelEncoder.html
LabelEncoder — scikit-learn 1.10.dev0 documentation
Encode categorical features as a one-hot numeric array. Examples · LabelEncoder can be used to normalize labels. >>> from sklearn.preprocessing import LabelEncoder >>> le = LabelEncoder() >>> le.fit([1, 2, 2, 6]) LabelEncoder() >>> le.classes_ array([1, 2, 6]) >>> le.transform([1, 1, 2, 6]) array([0, 0, 1, 2]...) >>> le.inverse_transform([0, 0, 1, 2]) array([1, 1, 2, 6]) It can also be used to transform non-numerical labels (as long as they are hashable and comparable) to numerical labels.
Find elsewhere
🌐
GitHub
github.com › scikit-learn › scikit-learn › blob › main › sklearn › preprocessing › _encoders.py
scikit-learn/sklearn/preprocessing/_encoders.py at main · scikit-learn/scikit-learn
f"Python string. Got {type(dry_run_combiner)} instead." ) return self.feature_name_combiner · · · class OrdinalEncoder(OneToOneFeatureMixin, _BaseEncoder): """ Encode categorical features as an integer array.
Author: scikit-learn
🌐
Snyk
snyk.io › advisor › sklearn › functions › sklearn.preprocessing.labelencoder
How to use the sklearn.preprocessing.LabelEncoder function in sklearn | Snyk
@staticmethod def encode_labels(labels): from sklearn import preprocessing intent_encoder = preprocessing.LabelEncoder() intent_encoder.fit(labels) return intent_encoder.transform(labels)
🌐
scikit-learn
scikit-learn.org › stable › auto_examples › preprocessing › plot_target_encoder.html
Comparing Target Encoder with Other Encoders — scikit-learn 1.9.1 documentation
First, we list out the encoders we will be using to preprocess the categorical features: from sklearn.compose import ColumnTransformer from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder, TargetEncoder categorical_preprocessors = [ ("drop", "drop"), ("ordinal", OrdinalEncoder(handle_unknown="use_encoded_value", unknown_value=-1)), ( "one_hot", OneHotEncoder(handle_unknown="ignore", max_categories=20, sparse_output=False), ), ("target", TargetEncoder(target_type="continuous")), ]
🌐
YouTube
youtube.com › watch
Label Encoding in Python with Scikit-Learn | Transform Data using LabelEncoder & OneHotEncoder - YouTube
In this tutorial, we dive into label encoding techniques using Python's Scikit-Learn library. Learn how to transform categorical data into numerical format w...
Published: August 29, 2024
🌐
GitHub
github.com › scikit-learn › scikit-learn › blob › f3f51f9b6 › sklearn › preprocessing › _encoders.py
scikit-learn/sklearn/preprocessing/_encoders.py at f3f51f9b611bf873bd5836748647221480071a87 · scikit-learn/scikit-learn
values per feature and transform the data to a binary one-hot encoding. · >>> from sklearn.preprocessing import OneHotEncoder · · One can discard categories not seen during `fit`: ·
Author: scikit-learn
🌐
scikit-learn
scikit-learn.org › 0.17 › modules › generated › sklearn.preprocessing.LabelEncoder.html
sklearn.preprocessing.LabelEncoder — scikit-learn 0.17.1 documentation
Encode labels with value between 0 and n_classes-1. Read more in the User Guide. ... LabelEncoder can be used to normalize labels. >>> from sklearn import preprocessing >>> le = preprocessing.LabelEncoder() >>> le.fit([1, 2, 2, 6]) LabelEncoder() >>> le.classes_ array([1, 2, 6]) >>> le.transform([1, 1, 2, 6]) array([0, 0, 1, 2]...) >>> le.inverse_transform([0, 0, 1, 2]) array([1, 1, 2, 6])
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.preprocessing.TargetEncoder.html
TargetEncoder — scikit-learn 1.5.2 documentation
This unsupervised encoding is better suited for low cardinality categorical variables as it generate one new feature per unique category. ... Micci-Barreca, Daniele. “A preprocessing scheme for high-cardinality categorical attributes in classification and prediction problems” SIGKDD Explor. Newsl. 3, 1 (July 2001), 27–32. ... >>> import numpy as np >>> from sklearn.preprocessing import TargetEncoder >>> X = np.array([["dog"] * 20 + ["cat"] * 30 + ["snake"] * 38], dtype=object).T >>> y = [90.3] * 5 + [80.1] * 15 + [20.4] * 5 + [20.1] * 25 + [21.2] * 8 + [49] * 30 >>> enc_auto = TargetEncoder(smooth="auto") >>> X_trans = enc_auto.fit_transform(X, y)
🌐
Kaggle
kaggle.com › code › residentmario › encoding-categorical-data-in-sklearn
Encoding categorical data in sklearn
July 24, 2019 - Python · Encoding categorical data in sklearnOrdinalLeave-one-outCatBoost · This Notebook has been released under the Apache 2.0 open source license. Input2 files · arrow_right_alt · Output0 files · arrow_right_alt · Logs16.6 second run - successful · arrow_right_alt ·
🌐
GitHub
github.com › scikit-learn › scikit-learn › blob › cc50648cc › sklearn › preprocessing › _encoders.py
scikit-learn/sklearn/preprocessing/_encoders.py at cc50648cc1b759b53a4edbce0f3bb6c237349448 · scikit-learn/scikit-learn
f"Python string. Got {type(dry_run_combiner)} instead." ) return self.feature_name_combiner · · · class OrdinalEncoder(OneToOneFeatureMixin, _BaseEncoder): """ Encode categorical features as an integer array.
Author: scikit-learn
🌐
GitHub
github.com › scikit-learn › scikit-learn › blob › 0fb307bf3 › sklearn › preprocessing › _encoders.py
scikit-learn/sklearn/preprocessing/_encoders.py at 0fb307bf39bbdacd6ed713c00724f8f871d60370 · scikit-learn/scikit-learn
values per feature and transform the data to a binary one-hot encoding. · >>> from sklearn.preprocessing import OneHotEncoder · · One can discard categories not seen during `fit`: ·
Author: scikit-learn
🌐
scikit-learn
scikit-learn.org › 1.5 › modules › generated › sklearn.preprocessing.LabelEncoder.html
LabelEncoder — scikit-learn 1.5.2 documentation
Encode categorical features as a one-hot numeric array. Examples · LabelEncoder can be used to normalize labels. >>> from sklearn.preprocessing import LabelEncoder >>> le = LabelEncoder() >>> le.fit([1, 2, 2, 6]) LabelEncoder() >>> le.classes_ array([1, 2, 6]) >>> le.transform([1, 1, 2, 6]) array([0, 0, 1, 2]...) >>> le.inverse_transform([0, 0, 1, 2]) array([1, 1, 2, 6]) It can also be used to transform non-numerical labels (as long as they are hashable and comparable) to numerical labels.