It's probably faster and easier to use numpy.digitize():
import numpy
data = numpy.random.random(100)
bins = numpy.linspace(0, 1, 10)
digitized = numpy.digitize(data, bins)
bin_means = [data[digitized == i].mean() for i in range(1, len(bins))]
An alternative to this is to use numpy.histogram():
bin_means = (numpy.histogram(data, bins, weights=data)[0] /
numpy.histogram(data, bins)[0])
Try for yourself which one is faster... :)
Answer from Sven Marnach on Stack OverflowIt's probably faster and easier to use numpy.digitize():
import numpy
data = numpy.random.random(100)
bins = numpy.linspace(0, 1, 10)
digitized = numpy.digitize(data, bins)
bin_means = [data[digitized == i].mean() for i in range(1, len(bins))]
An alternative to this is to use numpy.histogram():
bin_means = (numpy.histogram(data, bins, weights=data)[0] /
numpy.histogram(data, bins)[0])
Try for yourself which one is faster... :)
The Scipy (>=0.11) function scipy.stats.binned_statistic specifically addresses the above question.
For the same example as in the previous answers, the Scipy solution would be
import numpy as np
from scipy.stats import binned_statistic
data = np.random.rand(100)
bin_means = binned_statistic(data, data, bins=10, range=(0, 1))[0]
What are the different techniques for binning data in Python?
What is Python binning?
What are the benefits of binning in Python?
Hi all.
I'm an undergrad physicist and beginner-intermediate programmer and I'm looking to sort data for a recent experiment.
I have 2 variables; x and y.
I want to sort my large array of x and y values into small width bins defined by a small range of the x values. The x values run from 0 to 30 so lets say I want to to separate the data in around 90 bins which will give me a bin width of ~0.33.
So I want to have an array of bins that contain the small range of x and the corresponding y values in each bin.
For context, I'm trying to average 10 repeat runs of an experiment together and to calculate the standard deviation of each small width bin. So in essence each bin will have around 10-20 points in it and will produce a mean value which I can then use as a data point on the new, smoothed plot, with an error bar coming from the standard deviation.
Can anyone give me any hints on how I'd do this? Most of what I've seen treat it like a histogram, so I'd have a bin with some width, and a corresponding 'frequency' value telling me how many points are in this range, but I don't want this. For a range of x values within the width of the bin I want the corresponding y values.
I'm sorry if this is poorly explained guys! Google is usually my friend but I'm struggling to find what I need.
Thanks a lot <3