Have a look at the array module:
import array
array.array('B', [0] * 10000)
Instead of passing a list to initialize it, you can pass a generator, which is more memory efficient.
Answer from unbeknown on Stack OverflowHave a look at the array module:
import array
array.array('B', [0] * 10000)
Instead of passing a list to initialize it, you can pass a generator, which is more memory efficient.
You can pre-allocate a list with:
l = [0] * 10000
which will be slightly faster than .appending to it (as it avoids intermediate reallocations). However, this will generally allocate space for a list of pointers to integer objects, which will be larger than an array of bytes in C.
If you need memory efficiency, you could use an array object. ie:
import array, itertools
a = array.array('b', itertools.repeat(0, 10000))
Note that these may be slightly slower to use in practice, as there is an unboxing process when accessing elements (they must first be converted to a python int object).
binning data in python with scipy/numpy - Stack Overflow
Converting integer array to binary array in python - Stack Overflow
Fill array with binary numbers
arrays - Is there a way to create bins in python instead of listing all the bin numbers (as seen in code below), and maybe without having to use np.digitize? - Stack Overflow
It's probably faster and easier to use numpy.digitize():
import numpy
data = numpy.random.random(100)
bins = numpy.linspace(0, 1, 10)
digitized = numpy.digitize(data, bins)
bin_means = [data[digitized == i].mean() for i in range(1, len(bins))]
An alternative to this is to use numpy.histogram():
bin_means = (numpy.histogram(data, bins, weights=data)[0] /
numpy.histogram(data, bins)[0])
Try for yourself which one is faster... :)
The Scipy (>=0.11) function scipy.stats.binned_statistic specifically addresses the above question.
For the same example as in the previous answers, the Scipy solution would be
import numpy as np
from scipy.stats import binned_statistic
data = np.random.rand(100)
bin_means = binned_statistic(data, data, bins=10, range=(0, 1))[0]
Numpy attempts to convert your binary number to a float, except that your number contains a b which can't be interpreted; this character was added by the bin function, eg. bin(2) is 0b10. You should remove this b character before your zfill like this by using a "slice" to remove the first 2 characters:
b[i]=bin(int(a[i]))[2:].zfill(8)
bin will create a string that starts with 0b indicating that it is a binary representation. If you only want the binary representation you have to slice the first two characters before you call zfill.
Instead of doing this, you could use format like so
b[i] = '{:08b}'.format(a[i])
Basically, this will print the binary representation of a[i] padded with 0 until it has length 8.
See the Format Specification Mini-Language for further details.
Hello there. I am trying to learn python for a day. I need to create something like this:
5~ inputs
array = {}
fill array with binary 0(00000) to 31(11111) and execute some other function (not relevant here). So I need that array to be for i=0 -> array = {0,0,0,0,0}, i=1 -> array = {0,0,0,0,1}, i=2 -> array = {0,0,0,1,0} and so on, I think it can be also reversed, like i=1 -> array = {1,0,0,0,0} etc.
I managed to do a for loop to get my binary numbers 00000, 00001 etc, how can I split these numbers and put them into each array position?
I cant find the original author of a different SO post where I got this from using Pandas but maybe try something like this below that I thru together really fast for an idea to try. The data frame is just numpy random range to generate the fake data in the ranges you are looking for.
import pandas as pd
import numpy as np
#create bins & categories for data ranges
cats = ['4100000_4155303',
'4155304_4210608',
'4210608_4321215',
'4321216_4542431',
'4542432_4984864',
'4984865_5327532',
'5327533_5670200',
'5670201_5746216',
'5746217_5873108',
'5873109_6000000']
bins = [0,
4100000,
4210608,
4321215,
4542431,
4984864,
5327532,
5670200,
5746216,
5873108,
6000000]
def binn(df):
df = (df.groupby([df.index, pd.cut(df['A'], bins, labels=cats)])
.size()
.unstack(fill_value=0)
.reindex(columns=cats, fill_value=0))
return df
rng = np.random.default_rng()
df = pd.DataFrame(rng.integers(4155304, 6000000, size=(1000, 1)), columns=list('A'))
dfBinned = binn(df)
print('All data binned in column A of the df')
print(dfBinned.sum(axis = 0))
This prints:
All data binned in column A of the df
A
4100000_4155303 0
4155304_4210608 35
4210608_4321215 42
4321216_4542431 130
4542432_4984864 239
4984865_5327532 174
5327533_5670200 205
5670201_5746216 37
5746217_5873108 63
5873109_6000000 75
dtype: int64
Simply use the numpy.arange method:
bins = np.arange(4100000, 6000000, 55304)
bins
Output
array([4100000, 4155304, 4210608, 4265912, 4321216, 4376520, 4431824,
4487128, 4542432, 4597736, 4653040, 4708344, 4763648, 4818952,
4874256, 4929560, 4984864, 5040168, 5095472, 5150776, 5206080,
5261384, 5316688, 5371992, 5427296, 5482600, 5537904, 5593208,
5648512, 5703816, 5759120, 5814424, 5869728, 5925032, 5980336])
Cheers