It's probably faster and easier to use numpy.digitize():

import numpy
data = numpy.random.random(100)
bins = numpy.linspace(0, 1, 10)
digitized = numpy.digitize(data, bins)
bin_means = [data[digitized == i].mean() for i in range(1, len(bins))]

An alternative to this is to use numpy.histogram():

bin_means = (numpy.histogram(data, bins, weights=data)[0] /
             numpy.histogram(data, bins)[0])

Try for yourself which one is faster... :)

Answer from Sven Marnach on Stack Overflow
🌐
GeeksforGeeks
geeksforgeeks.org › numpy › binning-data-in-python-with-scipy-numpy
Binning Data In Python With Scipy & Numpy - GeeksforGeeks
July 23, 2025 - Binning data is a common technique in data analysis where you group continuous data into discrete intervals, or bins, to gain insights into the distribution or trends within the data.
Discussions

python - Binning a numpy array - Stack Overflow
Explore Stack Internal ... Save this question. Show activity on this post. I have a numpy array which contains time series data. I want to bin that array into equal partitions of a given length (it is fine to drop the last partition if it is not the same size) and then calculate the mean of ... More on stackoverflow.com
🌐 stackoverflow.com
python - Binning of data along one axis in numpy - Stack Overflow
I have a large two dimensional array arr which I would like to bin over the second axis using numpy. Because np.histogram flattens the array I'm currently using a for loop: import numpy as np arr... More on stackoverflow.com
🌐 stackoverflow.com
arrays - Binning values of a function in Python (numpy) - Stack Overflow
Let me expose my issue : I wrote a piece of software with Python and Numpy, it produces two numpy arrays named X and Y. This values are related as a function : Y = f(X) X values belong to the in... More on stackoverflow.com
🌐 stackoverflow.com
April 11, 2017
python - A numpy function for binning an array - Stack Overflow
5 vectorized approach to binning with numpy/scipy in Python More on stackoverflow.com
🌐 stackoverflow.com
🌐
Medium
medium.com › @heyamit10 › understanding-binning-in-numpy-02c169788d56
Understanding Binning in NumPy
March 6, 2025 - Option 2: Replace NaNs with the mean, median, or another default value before binning. Pro Tip: If you replace NaNs, choose a method that makes sense for your dataset. Median works well when dealing with outliers, while the mean is better for normally distributed data. ❓ What is the difference between numpy.histogram() and numpy.digitize()?
🌐
SciPython
scipython.com › blog › binning-a-2d-array-in-numpy
Binning a 2D array in NumPy
August 4, 2016 - This is the $2\times 3$ binned array that we wanted. Here is an illustration of the technique, based on USGS elevation data for the vicinity of Mt Ranier, which can be obtained from their download service. import os from osgeo import gdal import numpy as np import matplotlib.pyplot as plt from matplotlib import cm from mpl_toolkits.mplot3d import Axes3D imgdir = '/Users/christian/temp/mt-ranier' # Mt Ranier spans two IMG files in the USGS data set.
🌐
Statology
statology.org › home › how to bin variables in python using numpy.digitize()
How to Bin Variables in Python Using numpy.digitize()
May 24, 2022 - The following code shows how to place the values of an array into two bins: ... import numpy as np #create data data = [2, 4, 4, 7, 12, 14, 19, 20, 24, 31, 34] #place values into bins np.digitize(data, bins=[20]) array([0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1])
🌐
Llego
llego.dev › home › blog › histogramming and binning data with numpy in python
Histogramming and Binning Data with NumPy in Python - llego.dev
March 14, 2023 - Bins are sequential, non-overlapping intervals that cover the dataset’s range. Bin widths can vary but equal-width bins are commonly used. import numpy as np data = np.random.normal(size=1000) # Generate 1000 normally distributed data points num_bins = 20 # Use 20 equal-width bins counts, bin_edges = np.histogram(data, bins=num_bins) # num_bins = 'auto' can also be used to auto-determine number of bins
Find elsewhere
🌐
Arab Psychology
scales.arabpsychology.com › home › how to bin variables in python using numpy.digitize()
How To Bin Variables In Python Using Numpy.digitize()
December 24, 2025 - Binning variables, also known as discretization, is a fundamental process in data preprocessing, especially when working with statistical models that perform better with categorical inputs.
🌐
Delft Stack
delftstack.com › home › howto › python › bin data using scipy numpy and pandas in python
Bin Data Using SciPy, NumPy and Pandas in Python | Delft Stack
March 4, 2025 - The first step is to import the SciPy and NumPy libraries: ... Next, you’ll need to define the edges of the bins. It can be done using the linspace function: bin_edges = np.linspace(start, stop, num=num_bins) Where start & stop are the minimum & maximum values of the data, respectively, and num_bins is the bins’ number you want to create.
🌐
Codepointtech
codepointtech.com › home › master data digitization & binning with numpy in python
Master Data Digitization & Binning with NumPy in Python - codepointtech.com
January 18, 2026 - Binning essentially reduces the impact of minor observation errors and can help highlight overall patterns in your dataset. Instead of dealing with every unique value, you work with ranges. np.digitize() is perfectly suited for python numpy binning data.
Top answer
1 of 5
18

You could use np.apply_along_axis:

x = np.array([range(20), range(1, 21), range(2, 22)])

nbins = 2
>>> np.apply_along_axis(lambda a: np.histogram(a, bins=nbins)[0], 1, x)
array([[10, 10],
       [10, 10],
       [10, 10]])

The main advantage (if any) is that it's slightly shorter, but I wouldn't expect much of a performance gain. It's possibly marginally more efficient in the assembly of the per-row results.

2 of 5
4

For pages of many, many, many small data series I think you can do a lot faster using something like numpy.digitize (like a lot faster). Here is an example with 5000 data series, each featuring a modest 50 data points and targeting as few as 10 discrete bin locations. The speedup in this case is about ~an order of magnitude compared to the np.apply_along_axis implementation. The implementation looks like:

def histograms( data, bin_edges ):
    indices = np.digitize(data, bin_edges)
    histograms = np.zeros((data.shape[0], len(bin_edges)-1))
    for i,index in enumerate(np.unique(indices)):
        histograms[:, i]= np.sum( indices==index, axis=1 )
    return histograms

And here are some timings and verification:

data = np.random.rand(5000, 50)
bin_edges = np.linspace(0, 1, 11)

t1 = time.perf_counter()
h1 = histograms( data, bin_edges )
t2 = time.perf_counter()
print('digitize ', 1000*(t2-t1)/10., 'ms')

t1 = time.perf_counter()
h2 = np.apply_along_axis(lambda a: np.histogram(a, bins=bin_edges)[0], 1, data)
t2 = time.perf_counter()
print('numpy    ', 1000*(t2-t1)/10., 'ms')

assert np.allclose(h1, h2)

The result is something like this:

digitize  1.690 ms
numpy     15.08 ms

Cheers.

🌐
Bit Level Code
bitlevelcode.com › home › python › histogramming and binning data with numpy in python
Histogramming and Binning Data with NumPy in Python
March 16, 2025 - Each interval (bin) represents a range of values and the data points that fall within that range. Histograms are built upon binning, and NumPy makes this process seamless with its numpy.histogram() function.
🌐
Readthedocs
remu.readthedocs.io › en › latest › modules › binning › Binning.html
Binning — ReMU documentation
Subbinnings are used to get a finer binning within a given bin. The bin to be replaced by the finer binning is specified using the native bin index, i.e. the number it would have before the sub binnings are assigned. Subbinnings are inserted into the numpy arrays at the position of the original ...
Top answer
1 of 3
4

We could simply reshape to basically split into rows of such groups and hence sum each row for the desired output, like so -

np.reshape(L,(num_bins,-1)).sum(1)

For arrays with lengths not necessarily divisible by the number of bins -

def sum_groups(L, num_bins):
    n  = len(L)
    grp_len = int(np.ceil(n/float(num_bins)))
    b = int(n%num_bins!=0)
    lim = grp_len*(num_bins-b)
    p0 = np.reshape(L[:lim],(-1,grp_len)).sum(1)

    if b!=0:
        p1 = np.sum(L[lim:])
        return np.r_[p0,p1]
    else:
        return p0

Bringing in np.einsum for cases when the binned summations are within the input array dtype precision -

def sum_groups_einsum(L, num_bins):
    n  = len(L)
    grp_len = int(np.ceil(n/float(num_bins)))
    b = int(n%num_bins!=0)
    lim = grp_len*(num_bins-b)
    p0 = np.einsum('ij->i',np.reshape(L[:lim],(-1,grp_len)))

    if b!=0:
        p1 = np.einsum('i->',L[lim:])
        return np.r_[p0,p1]
    else:
        return p0

Benchmarking

Following closely the OP's timing setup -

In [404]: # Setup
     ...: np.random.seed(0)
     ...: L = np.random.randint(0,high = 6, size = 10000000)
     ...: b = 20

In [405]: %timeit sum_groups(L, num_bins=b)
     ...: %timeit sum_groups_einsum(L, num_bins=b)
     ...: %timeit np.array([t.sum() for t in np.array_split(L, b)])
     ...: %timeit np.add.reduceat(L, np.linspace(0.5, L.size+0.5, b, False, dtype=int))
100 loops, best of 3: 6.2 ms per loop
100 loops, best of 3: 6 ms per loop
100 loops, best of 3: 6.25 ms per loop # @user2699's soln
100 loops, best of 3: 6.19 ms per loop # @Paul Panzer's soln

For the case when the array length is not divisible by the number of bins, let's have few more elements in the input array to achieve the same -

In [406]: # Setup
     ...: np.random.seed(0)
     ...: L = np.random.randint(0,high = 6, size = 10000012)
     ...: b = 20

In [407]: %timeit sum_groups(L, num_bins=b)
     ...: %timeit sum_groups_einsum(L, num_bins=b)
     ...: %timeit np.array([t.sum() for t in np.array_split(L, b)])
     ...: %timeit np.add.reduceat(L, np.linspace(0.5, L.size+0.5, b, False, dtype=int))
100 loops, best of 3: 6.45 ms per loop
100 loops, best of 3: 6.05 ms per loop
100 loops, best of 3: 6.45 ms per loop
100 loops, best of 3: 6.51 ms per loop

Running those again few more times, the first one and the last two had very comparable runtimes and the second one with einsum was tiny bit faster than the rest.

2 of 3
1

The following works,

array([t.sum() for t in array_split(L, b)])

And if, as you stated, you know that b divides L evenly, you can replace array_split with the split function.

Here's some benchmarks, with b=100 and L = randint(0, 100, 1000)

%timeit sum_groups(L, b)  # Defined in Divakar's answer
8.09 µs ± 293 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)

%timeit array([t.sum() for t in array_split(L, b)])
260 µs ± 2.12 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)

%timeit np.add.reduceat(L, np.linspace(0.5, L.size+0.5, b, False, dtype=int))
15.9 µs ± 1.45 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)

and with with b=3 and L = randint(0, 100, 1000)

%timeit sum_groups(L, b)
23.2 µs ± 317 ns per loop (mean ± std. dev. of 7 runs, 10000 loops each)

%timeit array([t.sum() for t in array_split(L, b)])
16.2 µs ± 171 ns per loop (mean ± std. dev. of 7 runs, 10000 loops each)

%timeit np.add.reduceat(L, np.linspace(0.5, L.size+0.5, b, False, dtype=int))
15 µs ± 1.77 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)

Depending on your data, it looks like Divakar's answer using reshaping may be the best approach.

🌐
Kanaries
docs.kanaries.net › topics › Python › python-binning
Python Binning: Clearly Explained – Kanaries
August 17, 2023 - The most common ones include equal-width binning, equal-frequency binning, and k-means clustering. Equal-width binning divides the range of the data into N intervals of equal size. The width of the intervals is defined as (max - min) / N. The NumPy library's histogram function can be used to ...
🌐
NumPy
numpy.org › doc › 2.1 › reference › generated › numpy.bincount.html
numpy.bincount — NumPy v2.1 Manual
Count number of occurrences of each value in array of non-negative ints. The number of bins (of size 1) is one larger than the largest value in x. If minlength is specified, there will be at least this number of bins in the output array (though it will be longer if necessary, depending on the ...