You can using category dtype in sklearn , it should be labelencoder
df.city=df.city.astype('category').cat.codes
df
Out[385]:
school city category capacity
0 1 0 45 23
1 2 1 12 236
2 3 2 8 63
3 4 0 7 234
Answer from BENY on Stack Overflow Top answer 1 of 3
28
You can using category dtype in sklearn , it should be labelencoder
df.city=df.city.astype('category').cat.codes
df
Out[385]:
school city category capacity
0 1 0 45 23
1 2 1 12 236
2 3 2 8 63
3 4 0 7 234
2 of 3
5
A few thousand columns is still manageable in the context of ML classifiers. Although you'd want to watch out for the curse of dimensionality.
That aside, you wouldn't want a get_dummies call to result in a memory blowout, so you could generate a SparseDataFrame instead -
v = pd.get_dummies(df.set_index('school').city, sparse=True)
v
azez6576sebd dsqozbc765aj sqdqsd12887s
school
1 1 0 0
2 0 1 0
3 0 0 1
4 1 0 0
type(v)
pandas.core.sparse.frame.SparseDataFrame
You can generate a sparse matrix using sdf.to_coo -
v.to_coo()
<4x3 sparse matrix of type '<class 'numpy.uint8'>'
with 4 stored elements in COOrdinate format>
Practical Business Python
pbpython.com › categorical-encoding.html
Guide to Encoding Categorical Values in Python - Practical Business Python
A common alternative approach is called one hot encoding (but also goes by several different names shown below). Despite the different names, the basic strategy is to convert each category value into a new column and assigns a 1 or 0 (True/False) value to the column. This has the benefit of not weighting a value improperly but does have the downside of adding more columns to the data set. Pandas supports this feature using get_dummies.
13:35
Categorical Variable Encoding Using ( One Hot Encoder & Pandas ...
05:16
Encode categorical features using OneHotEncoder or OrdinalEncoder ...
05:06
Python Tutorial: Dealing with categorical features - YouTube
27:59
How do I encode categorical features using scikit-learn? - YouTube
08:15
8. Handling Categorical Data using Python | One Hot Encoding | ...
06:15
How to convert categorical data to numerical data in python | Python ...
Scikit-learn course
inria.github.io › scikit-learn-mooc › python_scripts › 03_categorical_pipeline.html
Encoding of categorical variables — Scikit-learn course
In this notebook, we present some typical ways of dealing with categorical variables by encoding them, namely ordinal encoding and one-hot encoding. Let’s first load the entire adult dataset containing both numerical and categorical data. import pandas as pd adult_census = pd.read_csv("....
Skytowner
skytowner.com › explore › encoding_categorical_variables_in_pandas
Encoding categorical variables in Pandas
To encode categorical variables, either using one-hot encoding or dummy coding, use Pandas get_dummies(~) method.
YouTube
youtube.com › watch
How To Encode Categorical Variables Using Pandas - YouTube
Before training a machine learning model, you have to convert the texts (categorical variables) in your dataset to numbers. In this video, you will learn how...
Published: December 29, 2022
DataCamp
datacamp.com › tutorial › categorical-data
Handling Machine Learning Categorical Data with Python Tutorial | DataCamp
February 23, 2023 - It is a function in the Pandas library that can be used to perform one-hot encoding on categorical variables in a DataFrame. It takes a DataFrame and returns a new DataFrame with binary columns for each category.
Compile N Run
compilenrun.com › pandas tutorial › pandas data transformation › pandas categorical encoding
Pandas Categorical Encoding | Compile N Run
Basic Encoding: Create a dataset with at least 3 categorical features and apply label encoding to all of them.
KDnuggets
kdnuggets.com › 2023 › 07 › pandas-onehot-encode-data.html
Pandas: How to One-Hot Encode Data - KDnuggets
July 24, 2023 - df_encoded = pd.get_dummies(df, columns=['categorical_column', ]) The following commands drops the categorical_column and creates a new column for each unique value. Therefore, the single categorical column is converted into 4 new columns where only one of the 4 columns will have a 1 value, and all of the other 3 are encoded 0.
Analytics Vidhya
analyticsvidhya.com › home › what are categorical data encoding methods | binary encoding
What are Categorical Data Encoding Methods | Binary Encoding
May 1, 2025 - We use this categorical data encoding technique when the categorical feature is ordinal. In this case, retaining the order is important. Hence encoding should reflect the sequence. In Label encoding, each label is converted into an integer value. We will create a variable that contains the categories representing the education qualification of a person. import category_encoders as ce import pandas as pd train_df=pd.DataFrame({'Degree':['High school','Masters','Diploma','Bachelors','Bachelors','Masters','Phd','High school','High school']}) # create object of Ordinalencoding encoder= ce.OrdinalEncoder(cols=['Degree'],return_df=True, mapping=[{'col':'Degree', 'mapping':{'None':0,'High school':1,'Diploma':2,'Bachelors':3,'Masters':4,'phd':5}}]) #Original data print(train_df)
CodeSignal
codesignal.com › learn › courses › shaping-and-transforming-features › lessons › encoding-categorical-data-a-practical-approach
Categorical Data Encoding Techniques | CodeSignal Learn
One-hot encoding is a common technique for converting categorical data into a form that can be provided to machine learning algorithms. This is done by converting each category value into a new categorical column and assigning a 1 or 0 (True/False). You can easily perform one-hot encoding in ...
Deffro
deffro.github.io › tutorials › encoding-categorical-features
Encoding Categorical Features - Dimitris Effrosynidis
September 11, 2025 - On the other hand tree-based algorithms like Random Forest, XGBoost, LightGBM, and Naive Bayes can work with categorical data, but their accuracy might improve with encoding. import numpy as np import pandas as pd import time from sklearn.preprocessing import LabelEncoder import category_encoders as ce pd.set_option('display.max_columns', 500) pd.set_option('display.max_rows', 500) pd.set_option('max_colwidth', 100)
MachineLearningMastery
machinelearningmastery.com › home › blog › 3 ways to encode categorical variables for deep learning
3 Ways to Encode Categorical Variables for Deep Learning - MachineLearningMastery.com
August 26, 2020 - A reasonable classification accuracy ... encoding schemes. You can download the dataset and save the file as “breast-cancer.csv” in your current working directory. ... Looking at the data, we can see that all nine input variables are categorical. Specifically, all variables are quoted strings; some are ordinal and some are not. We can load this dataset into memory using the Pandas ...
CodeSignal
codesignal.com › learn › courses › cleaning-and-transforming-data-with-pandas › lessons › encoding-categorical-variables-using-python
Encoding Categorical Variables Using Python
In this lesson, we will specifically focus on using a dictionary mapping to encode a binary categorical variable, which is a form of label encoding. Be a part of our community of 1M+ users who develop and demonstrate their skills on CodeSignalStart learning today! ... Let's say we have a dataset with a Gender column containing the values Male and Female. We want to convert this column into a numerical format where Male is represented as 1 and Female is represented as 0. We can achieve this using a dictionary mapping. import pandas as pd # Creating a sample DataFrame df = pd.DataFrame({'Gender': ['Male', 'Female', 'Female', 'Male', 'Female']}) # Encoding categorical variables df['Gender_Encoded'] = df['Gender'].map({'Male': 1, 'Female': 0}) # Display the DataFrame print(df)
GeeksforGeeks
geeksforgeeks.org › machine learning › encoding-categorical-data-in-sklearn
Encoding Categorical Data in Sklearn - GeeksforGeeks
September 17, 2025 - This approach cleanly manages both ordinal and nominal encoding and fits directly into any sklearn modeling pipeline. Suitable for any supervised learning (classification/regression) with categorical inputs. ... from sklearn.compose import ColumnTransformer from sklearn.pipeline import Pipeline ordinal_features = ['safety'] ordinal_categories = [['low', 'med', 'high']] nominal_features = ['buying', 'maint', 'doors', 'persons', 'lug_boot'] preprocessor = ColumnTransformer( transformers=[ ('ord', OrdinalEncoder(categories=ordinal_categories), ordinal_features), ('nom', OneHotEncoder(sparse_output=False), nominal_features) ] ) features = ordinal_features + nominal_features X = df[features] X_prepared = preprocessor.fit_transform(X) print("Transformed shape:", X_prepared.shape)