It's probably faster and easier to use numpy.digitize():
import numpy
data = numpy.random.random(100)
bins = numpy.linspace(0, 1, 10)
digitized = numpy.digitize(data, bins)
bin_means = [data[digitized == i].mean() for i in range(1, len(bins))]
An alternative to this is to use numpy.histogram():
bin_means = (numpy.histogram(data, bins, weights=data)[0] /
numpy.histogram(data, bins)[0])
Try for yourself which one is faster... :)
Answer from Sven Marnach on Stack OverflowIt's probably faster and easier to use numpy.digitize():
import numpy
data = numpy.random.random(100)
bins = numpy.linspace(0, 1, 10)
digitized = numpy.digitize(data, bins)
bin_means = [data[digitized == i].mean() for i in range(1, len(bins))]
An alternative to this is to use numpy.histogram():
bin_means = (numpy.histogram(data, bins, weights=data)[0] /
numpy.histogram(data, bins)[0])
Try for yourself which one is faster... :)
The Scipy (>=0.11) function scipy.stats.binned_statistic specifically addresses the above question.
For the same example as in the previous answers, the Scipy solution would be
import numpy as np
from scipy.stats import binned_statistic
data = np.random.rand(100)
bin_means = binned_statistic(data, data, bins=10, range=(0, 1))[0]
python - Binning a numpy array - Stack Overflow
python - Binning of data along one axis in numpy - Stack Overflow
arrays - Binning values of a function in Python (numpy) - Stack Overflow
python - A numpy function for binning an array - Stack Overflow
Just use reshape and then mean(axis=1).
As the simplest possible example:
import numpy as np
data = np.array([4,2,5,6,7,5,4,3,5,7])
print data.reshape(-1, 2).mean(axis=1)
More generally, we'd need to do something like this to drop the last bin when it's not an even multiple:
import numpy as np
width=3
data = np.array([4,2,5,6,7,5,4,3,5,7])
result = data[:(data.size // width) * width].reshape(-1, width).mean(axis=1)
print result
Since you already have a numpy array, to avoid for loops, you can use reshape and consider the new dimension to be the bin:
In [33]: data.reshape(2, -1)
Out[33]:
array([[4, 2, 5, 6, 7],
[5, 4, 3, 5, 7]])
In [34]: data.reshape(2, -1).mean(0)
Out[34]: array([ 4.5, 3. , 4. , 5.5, 7. ])
Actually this will just work if the size of data is divisible by n. I'll edit a fix.
Looks like Joe Kington has an answer that handles that.
You could use np.apply_along_axis:
x = np.array([range(20), range(1, 21), range(2, 22)])
nbins = 2
>>> np.apply_along_axis(lambda a: np.histogram(a, bins=nbins)[0], 1, x)
array([[10, 10],
[10, 10],
[10, 10]])
The main advantage (if any) is that it's slightly shorter, but I wouldn't expect much of a performance gain. It's possibly marginally more efficient in the assembly of the per-row results.
For pages of many, many, many small data series I think you can do a lot faster using something like numpy.digitize (like a lot faster). Here is an example with 5000 data series, each featuring a modest 50 data points and targeting as few as 10 discrete bin locations. The speedup in this case is about ~an order of magnitude compared to the np.apply_along_axis implementation. The implementation looks like:
def histograms( data, bin_edges ):
indices = np.digitize(data, bin_edges)
histograms = np.zeros((data.shape[0], len(bin_edges)-1))
for i,index in enumerate(np.unique(indices)):
histograms[:, i]= np.sum( indices==index, axis=1 )
return histograms
And here are some timings and verification:
data = np.random.rand(5000, 50)
bin_edges = np.linspace(0, 1, 11)
t1 = time.perf_counter()
h1 = histograms( data, bin_edges )
t2 = time.perf_counter()
print('digitize ', 1000*(t2-t1)/10., 'ms')
t1 = time.perf_counter()
h2 = np.apply_along_axis(lambda a: np.histogram(a, bins=bin_edges)[0], 1, data)
t2 = time.perf_counter()
print('numpy ', 1000*(t2-t1)/10., 'ms')
assert np.allclose(h1, h2)
The result is something like this:
digitize 1.690 ms
numpy 15.08 ms
Cheers.
We could simply reshape to basically split into rows of such groups and hence sum each row for the desired output, like so -
np.reshape(L,(num_bins,-1)).sum(1)
For arrays with lengths not necessarily divisible by the number of bins -
def sum_groups(L, num_bins):
n = len(L)
grp_len = int(np.ceil(n/float(num_bins)))
b = int(n%num_bins!=0)
lim = grp_len*(num_bins-b)
p0 = np.reshape(L[:lim],(-1,grp_len)).sum(1)
if b!=0:
p1 = np.sum(L[lim:])
return np.r_[p0,p1]
else:
return p0
Bringing in np.einsum for cases when the binned summations are within the input array dtype precision -
def sum_groups_einsum(L, num_bins):
n = len(L)
grp_len = int(np.ceil(n/float(num_bins)))
b = int(n%num_bins!=0)
lim = grp_len*(num_bins-b)
p0 = np.einsum('ij->i',np.reshape(L[:lim],(-1,grp_len)))
if b!=0:
p1 = np.einsum('i->',L[lim:])
return np.r_[p0,p1]
else:
return p0
Benchmarking
Following closely the OP's timing setup -
In [404]: # Setup
...: np.random.seed(0)
...: L = np.random.randint(0,high = 6, size = 10000000)
...: b = 20
In [405]: %timeit sum_groups(L, num_bins=b)
...: %timeit sum_groups_einsum(L, num_bins=b)
...: %timeit np.array([t.sum() for t in np.array_split(L, b)])
...: %timeit np.add.reduceat(L, np.linspace(0.5, L.size+0.5, b, False, dtype=int))
100 loops, best of 3: 6.2 ms per loop
100 loops, best of 3: 6 ms per loop
100 loops, best of 3: 6.25 ms per loop # @user2699's soln
100 loops, best of 3: 6.19 ms per loop # @Paul Panzer's soln
For the case when the array length is not divisible by the number of bins, let's have few more elements in the input array to achieve the same -
In [406]: # Setup
...: np.random.seed(0)
...: L = np.random.randint(0,high = 6, size = 10000012)
...: b = 20
In [407]: %timeit sum_groups(L, num_bins=b)
...: %timeit sum_groups_einsum(L, num_bins=b)
...: %timeit np.array([t.sum() for t in np.array_split(L, b)])
...: %timeit np.add.reduceat(L, np.linspace(0.5, L.size+0.5, b, False, dtype=int))
100 loops, best of 3: 6.45 ms per loop
100 loops, best of 3: 6.05 ms per loop
100 loops, best of 3: 6.45 ms per loop
100 loops, best of 3: 6.51 ms per loop
Running those again few more times, the first one and the last two had very comparable runtimes and the second one with einsum was tiny bit faster than the rest.
The following works,
array([t.sum() for t in array_split(L, b)])
And if, as you stated, you know that b divides L evenly, you can replace array_split with the split function.
Here's some benchmarks, with b=100 and L = randint(0, 100, 1000)
%timeit sum_groups(L, b) # Defined in Divakar's answer
8.09 µs ± 293 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)
%timeit array([t.sum() for t in array_split(L, b)])
260 µs ± 2.12 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
%timeit np.add.reduceat(L, np.linspace(0.5, L.size+0.5, b, False, dtype=int))
15.9 µs ± 1.45 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)
and with with b=3 and L = randint(0, 100, 1000)
%timeit sum_groups(L, b)
23.2 µs ± 317 ns per loop (mean ± std. dev. of 7 runs, 10000 loops each)
%timeit array([t.sum() for t in array_split(L, b)])
16.2 µs ± 171 ns per loop (mean ± std. dev. of 7 runs, 10000 loops each)
%timeit np.add.reduceat(L, np.linspace(0.5, L.size+0.5, b, False, dtype=int))
15 µs ± 1.77 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)
Depending on your data, it looks like Divakar's answer using reshaping may be the best approach.