๐ŸŒ
Dagster
dagster.io โ€บ glossary โ€บ data-discretization
What Does Discretize Mean | Dagster
Another example of discretization in Python is using the KBinsDiscretizer class from the scikit-learn library. This class allows you to specify the number of bins, the strategy for dividing the data, and whether to encode the intervals as integers ...
๐ŸŒ
Packtpub
subscription.packtpub.com โ€บ book โ€บ data โ€บ 9781838552862 โ€บ 1 โ€บ ch01lvl1sec09 โ€บ data-discretization
Introduction to Data Science and Data Pre-Processing | Data Science with Python[Instructor Edition]
The main challenge in discretization is to choose the number of intervals or bins and how to decide on their width. Here we make use of a function called pandas.cut(). This function is useful to achieve the bucketing and sorting of segmented data.
๐ŸŒ
scikit-learn
scikit-learn.org โ€บ stable โ€บ auto_examples โ€บ preprocessing โ€บ plot_discretization.html
Using KBinsDiscretizer to discretize continuous features โ€” scikit-learn 1.9.0 documentation
The example compares prediction result of linear regression (linear model) and decision tree (tree based model) with and without discretization of real-valued features. As is shown in the result be...
๐ŸŒ
DataCamp
campus.datacamp.com โ€บ courses โ€บ introduction-to-predictive-analytics-in-python โ€บ interpreting-and-explaining-models
Discretization of continuous variables | Python
Consider for instance the variable age that is discretized in 5 bins. The first bin starts at 38 and ends at 49, but it would be clearer if it would start at 40 and end at 50, and the same for the other bins. In python, you can specify the cuts that you like using the cut function.
๐ŸŒ
Towards Data Science
towardsdatascience.com โ€บ home โ€บ latest โ€บ an intro to discretization techniques for machine learning
An Intro to Discretization Techniques for Machine Learning | Towards Data Science
March 5, 2025 - In our example, suppose we wish to transform the age feature into predefined groups composing the age groups "1โ€“12", "13โ€“18", "19โ€“30", "31โ€“60", and "60โ€“80". We can achieve this in Python with the feature_engine packageโ€™s ...
๐ŸŒ
PythonProg
pythonprog.com โ€บ home โ€บ data discretization in machine learning (with python examples)
Data Discretization in Machine Learning (with Python Examples) | PythonProg
December 12, 2023 - The resulting plot shows the number of flowers in each category of the discretized column, grouped by species. Note that this is just an example and the code may need to be adapted to fit different use cases. Understanding of basic statistics concepts such as mean, median, and standard deviation ยท Knowledge of different types of data such as continuous, categorical, and ordinal ... Experience with Python programming language and related libraries such as NumPy, Pandas, and Scikit-learn
๐ŸŒ
Medium
medium.com โ€บ @gmshakil786 โ€บ mastering-discretization-in-data-science-a-step-by-step-guide-with-python-and-sklearn-examples-42d20676d458
Mastering Discretization in Data Science: A Step-by-Step Guide with Python and sklearn Examples | by Shakil Ur Rehman | Medium
September 21, 2025 - As a Data Scientist, Iโ€™ll guide you through implementing discretization using scikit-learn (sklearn) in Python, step by step. Iโ€™ll explain how to use the KBinsDiscretizer class, which is sklearnโ€™s primary tool for discretization, and provide a clear example with code.
๐ŸŒ
MachineLearningMastery
machinelearningmastery.com โ€บ home โ€บ blog โ€บ how to use discretization transforms for machine learning
How to Use Discretization Transforms for Machine Learning - MachineLearningMastery.com
August 28, 2020 - Clustered: Clusters are identified and examples are assigned to each group. The discretization transform is available in the scikit-learn Python machine learning library via the KBinsDiscretizer class.
Find elsewhere
๐ŸŒ
DataCamp
campus.datacamp.com โ€บ courses โ€บ introduction-to-predictive-analytics-in-python โ€บ interpreting-and-explaining-models
Discretization of a certain variable | Python
# Discretize the variable time_since_last_donation in 10 bins basetable["bins_recency"] = pd.qcut(____,____) # Print the group sizes of the discretized variable print(basetable.groupby("____").size())
๐ŸŒ
Medium
medium.com โ€บ @auliadafa91 โ€บ data-discretization-techniques-in-python-e7ffda574ab7
Data Discretization Techniques in Python | by Dafa Aulia | Medium
August 26, 2024 - Data discretization, or binning, is performed to simplify continuous data by converting it into discrete categories, which can improve model performance, reduce noise, and reveal underlying patterns or relationships that might not be apparent in the raw data. It also helps in creating categorical variables for models that require them, facilitates easier data visualization, and can enhance feature engineering by highlighting meaningful groupings in the data. import numpy as np import pandas as pd #Example Data data = { 'Name':['Josua','Dafa','Rio','Syahri','Krisly','Abdi'], 'House':['Kontrak','Milik','Kontrak','Kontrak','Milik','Kontrak'], 'Salary':[2100000,4000000,1400000,700000,650000,450000], 'Age':[24,23,21,28,27,22] } df = pd.DataFrame(data) df
๐ŸŒ
The Security Buddy
thesecuritybuddy.com โ€บ home โ€บ data preprocessing โ€บ how to perform equal width discretization using python pandas?
How to perform equal width discretization using Python pandas? - The Security Buddy
November 16, 2022 - We can use the pandas.cut() function to discretize a numerical variable into equal-sized buckets. For example, letโ€™s read the diamonds dataset and discretize the numerical values in the price column of the dataset. We can use the following ...
๐ŸŒ
Kaggle
kaggle.com โ€บ code โ€บ mrbisht โ€บ discretization-continuous-variables
Discretization Continuous Variables | Kaggle
September 18, 2022 - Explore and run AI code with Kaggle Notebooks | Using data from Spaceship Titanic
๐ŸŒ
Yanzhugoh
yanzhugoh.github.io โ€บ pandas โ€บ data-discretization
Data Discretization
The trick is to make use of the ... divides each age by 10 to create a float. The float is then floored, and lastly we multiply each value by 10. ages = pd.DataFrame(data={'age': [18, 27, 35, 42, 50, 69, 70, 81, 93]}) def floor_ages(age): return math.floor(age / 10) * 10 ...
๐ŸŒ
Data School
dataschool.io โ€บ discretization-for-machine-learning
Should you discretize features for Machine Learning?
June 7, 2025 - In scikit-learn, we can discretize using the KBinsDiscretizer class: When creating an instance of KBinsDiscretizer, you define the number of bins, the binning strategy, and the method used to encode the result: As an example, here's a numeric feature from the famous Titanic dataset:
Top answer
1 of 5
12

Update (Sep 2018): As of version 0.20.0, there is a function, sklearn.preprocessing.KBinsDiscretizer, which provides discretization of continuous features using a few different strategies:

  • Uniformly-sized bins
  • Bins with "equal" numbers of samples inside (as much as possible)
  • Bins based on K-means clustering

Unfortunately, at the moment, the function does not accept custom intervals (which is a bummer for me as that is what I wanted and the reason I ended up here). If you want to achieve the same, you can use Pandas function cut:

import numpy as np
import pandas as pd
n_samples = 10
a = np.random.randint(0, 10, n_samples)

# say you want to split at 1 and 3
boundaries = [1, 3]
# add min and max values of your data
boundaries = sorted({a.min(), a.max() + 1} | set(boundaries))

a_discretized_1 = pd.cut(a, bins=boundaries, right=False)
a_discretized_2 = pd.cut(a, bins=boundaries, labels=range(len(boundaries) - 1), right=False)
a_discretized_3 = pd.cut(a, bins=boundaries, labels=range(len(boundaries) - 1), right=False).astype(float)
print(a, '\n')
print(a_discretized_1, '\n', a_discretized_1.dtype, '\n')
print(a_discretized_2, '\n', a_discretized_2.dtype, '\n')
print(a_discretized_3, '\n', a_discretized_3.dtype, '\n')

which produces:

[2 2 9 7 2 9 3 0 4 0]

[[1, 3), [1, 3), [3, 10), [3, 10), [1, 3), [3, 10), [3, 10), [0, 1), [3, 10), [0, 1)]
Categories (3, interval[int64]): [[0, 1) < [1, 3) < [3, 10)]
 category

[1, 1, 2, 2, 1, 2, 2, 0, 2, 0]
Categories (3, int64): [0 < 1 < 2]
 category

[1. 1. 2. 2. 1. 2. 2. 0. 2. 0.]
 float64

Note that, by default, pd.cut returns a pd.Series object of dtype Category with elements of type interval[int64]. If you specify your own labels, the dtype of the output will still be a Category, but the elements will be of type int64. If you want the series to have a numeric dtype, you can use .astype(np.int64).

My example uses integer data, but it should work just as fine with floats.

2 of 5
10

The answer is no. There is no binning in scikit-learn. As eickenberg said, you might want to use np.histogram. Features in scikit-learn are assumed to be continuous, not discrete. The main reason why there is no binning is probably that most of sklearn is developed on text, image featuers or dataset from the scientific community. In these settings, binning is rarely helpful. Do you know of a freely available dataset where binning is really beneficial?

๐ŸŒ
scikit-learn
scikit-learn.org โ€บ stable โ€บ auto_examples โ€บ preprocessing โ€บ plot_discretization_classification.html
Feature discretization โ€” scikit-learn 1.9.0 documentation
Go to the end to download the full example code or to run this example in your browser via JupyterLite or Binder. A demonstration of feature discretization on synthetic classification datasets. Feature discretization decomposes each feature into a set of bins, here equally distributed in width.