Numpy attempts to convert your binary number to a float, except that your number contains a b which can't be interpreted; this character was added by the bin function, eg. bin(2) is 0b10. You should remove this b character before your zfill like this by using a "slice" to remove the first 2 characters:
b[i]=bin(int(a[i]))[2:].zfill(8)
Answer from Aaron Christiansen on Stack OverflowNumpy attempts to convert your binary number to a float, except that your number contains a b which can't be interpreted; this character was added by the bin function, eg. bin(2) is 0b10. You should remove this b character before your zfill like this by using a "slice" to remove the first 2 characters:
b[i]=bin(int(a[i]))[2:].zfill(8)
bin will create a string that starts with 0b indicating that it is a binary representation. If you only want the binary representation you have to slice the first two characters before you call zfill.
Instead of doing this, you could use format like so
b[i] = '{:08b}'.format(a[i])
Basically, this will print the binary representation of a[i] padded with 0 until it has length 8.
See the Format Specification Mini-Language for further details.
You can use bin built-in function to convert the integers into their binary representations:
>>> l = [116, 101, 115, 116]
>>> [bin(i) for i in l]
['0b1110100', '0b1100101', '0b1110011', '0b1110100']
If you do not want the 0b prefix, use string formatting with binary integer representation:
>>> l = [116, 101, 115, 116]
>>> ["{0:b}".format(i) for i in l]
['1110100', '1100101', '1110011', '1110100']
You can use the str.format method with b as the presentation type:
print('[{}]'.format(', '.join(map('{:b}'.format, s_acii))))
This outputs:
[1110100, 1100101, 1110011, 1110100]
Hello there. I am trying to learn python for a day. I need to create something like this:
5~ inputs
array = {}
fill array with binary 0(00000) to 31(11111) and execute some other function (not relevant here). So I need that array to be for i=0 -> array = {0,0,0,0,0}, i=1 -> array = {0,0,0,0,1}, i=2 -> array = {0,0,0,1,0} and so on, I think it can be also reversed, like i=1 -> array = {1,0,0,0,0} etc.
I managed to do a for loop to get my binary numbers 00000, 00001 etc, how can I split these numbers and put them into each array position?
You should be able to vectorize this, something like
>>> d = np.array([1,2,3,4,5])
>>> m = 8
>>> (((d[:,None] & (1 << np.arange(m)))) > 0).astype(int)
array([[1, 0, 0, 0, 0, 0, 0, 0],
[0, 1, 0, 0, 0, 0, 0, 0],
[1, 1, 0, 0, 0, 0, 0, 0],
[0, 0, 1, 0, 0, 0, 0, 0],
[1, 0, 1, 0, 0, 0, 0, 0]])
which just gets the appropriate bit weights and then takes the bitwise and:
>>> (1 << np.arange(m))
array([ 1, 2, 4, 8, 16, 32, 64, 128])
>>> d[:,None] & (1 << np.arange(m))
array([[1, 0, 0, 0, 0, 0, 0, 0],
[0, 2, 0, 0, 0, 0, 0, 0],
[1, 2, 0, 0, 0, 0, 0, 0],
[0, 0, 4, 0, 0, 0, 0, 0],
[1, 0, 4, 0, 0, 0, 0, 0]])
There are lots of ways to convert this to 1s wherever it's non-zero (> 0)*1, .astype(bool).astype(int), etc. I chose one basically at random.
One-line version, taking advantage of the fast path in numpy.binary_repr:
def bin_array(num, m):
"""Convert a positive integer num into an m-bit bit vector"""
return np.array(list(np.binary_repr(num).zfill(m))).astype(np.int8)
Example:
In [1]: bin_array(15, 6)
Out[1]: array([0, 0, 1, 1, 1, 1], dtype=int8)
Vectorized version for expanding an entire numpy array of ints at once:
def vec_bin_array(arr, m):
"""
Arguments:
arr: Numpy array of positive integers
m: Number of bits of each integer to retain
Returns a copy of arr with every element replaced with a bit vector.
Bits encoded as int8's.
"""
to_str_func = np.vectorize(lambda x: np.binary_repr(x).zfill(m))
strs = to_str_func(arr)
ret = np.zeros(list(arr.shape) + [m], dtype=np.int8)
for bit_ix in range(0, m):
fetch_bit_func = np.vectorize(lambda x: x[bit_ix] == '1')
ret[...,bit_ix] = fetch_bit_func(strs).astype("int8")
return ret
Example:
In [1]: vec_bin_array(np.array([[100, 42], [2, 5]]), 8)
Out[1]: array([[[0, 1, 1, 0, 0, 1, 0, 0],
[0, 0, 1, 0, 1, 0, 1, 0]],
[[0, 0, 0, 0, 0, 0, 1, 0],
[0, 0, 0, 0, 0, 1, 0, 1]]], dtype=int8)
Have a look at the array module:
import array
array.array('B', [0] * 10000)
Instead of passing a list to initialize it, you can pass a generator, which is more memory efficient.
You can pre-allocate a list with:
l = [0] * 10000
which will be slightly faster than .appending to it (as it avoids intermediate reallocations). However, this will generally allocate space for a list of pointers to integer objects, which will be larger than an array of bytes in C.
If you need memory efficiency, you could use an array object. ie:
import array, itertools
a = array.array('b', itertools.repeat(0, 10000))
Note that these may be slightly slower to use in practice, as there is an unboxing process when accessing elements (they must first be converted to a python int object).
Hello,
I have a problem using NumPy arrays. Basically, I have a NumPy array consisting of 0's and 1's. An array might look something like: np.array([1,1,0,1,1,0,1,1,0,1). However, I want to take that array and cut it into 'n' sections so that I could convert the sections into decimal form.
For example, I want to take the array presented earlier and cut it into two sections. The first section will contain [1,1,0,1,1]. Then I want to convert that into decimal form. I will do the same to the second section as well.
My question is: how do I go about sectioning my array? I think I know how to convert them to decimal But the sectioning is where I'm having a lot of difficulties. Thanks!
Things I tried:
> I tried doing array indexing and taking the first section and converting it to decimal, but I can't seem to get the second section.
> My next idea would be to put it in a for loop, but I don't know how to go about implementing that.
Thanks!
This is the fastest I came up with. A slight variation of your initial solution:
digits = ['0', '1']
int("".join([ digits[y] for y in x ]), 2)
%timeit int("".join([digits[y] for y in x]),2)
100000 loops, best of 3: 6.15 us per loop
%timeit int("".join(map(str, x)),2)
100000 loops, best of 3: 7.49 us per loop
(Btw, it seems that in this case, using a list comprehension is faster than using a generator expression.)
EDIT:
Also, I hate being a smartass, but you can always trade memory for speed:
# one time precalculation
cache_N = 16 # or much bigger?!
cache = {
tuple(x): int("".join([digits[y] for y in x]),2)
for x in itertools.product((0,1), repeat=cache_N)
}
Then:
res = cache[tuple(x)]
Way faster. Of course, this is only feasible up to a point...
EDIT2:
I now see you say your lists have 32 elements. In this case the caching solution is probably infeasible, BUT we have more ways to trade speed for memory. E.g., with cache_N=16, which is surely feasible, you can access it twice:
c = 2 ** cache_N # compute once
xx = tuple(x)
cache[xx[:16]] * c + cache[xx[16:]]
%timeit cache[xx[:16]] * c + cache[xx[16:]]
1000000 loops, best of 3: 1.23 us per loop # YES!
I decided to create a script to trial 4 different methods of doing this task.
import time
trials = range(1000000)
list1 = [1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,0,0]
list0 = [0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]
listmix = [1,0,1,0,1,0,1,0,1,0,1,0,1,0,1,0,1,0,1,0,1,0,1,0,1,0,1,0,1,0,1,0]
def test1(l):
start = time.time()
for trial in trials:
tot = 0
n = 1
for i in reversed(l):
if i:
tot += 2**n
n += 1
print 'Time taken:', str(time.time() - start)
def test2(l):
start = time.time()
for trial in trials:
int("".join(map(str, l)),2)
print 'Time taken:', str(time.time() - start)
def test3(l):
start = time.time()
for trial in trials:
sum(x << i for i, x in enumerate(reversed(l)))
print 'Time taken:', str(time.time() - start)
def test4(l):
start = time.time()
for trial in trials:
int("".join([str(i) for i in l]),2)
print 'Time taken:', str(time.time() - start)
test1(list1)
test2(list1)
test3(list1)
test4(list1)
print '.'
test1(list0)
test2(list0)
test3(list0)
test4(list0)
print '.'
test1(listmix)
test2(listmix)
test3(listmix)
test4(listmix)
My results:
Time taken: 7.14670491219
Time taken: 5.4076821804
Time taken: 4.7349550724
Time taken: 7.24234819412
.
Time taken: 2.29213285446
Time taken: 5.38784003258
Time taken: 4.70707392693
Time taken: 7.27936697006
.
Time taken: 4.78960323334
Time taken: 5.36612486839
Time taken: 4.70103287697
Time taken: 7.22436404228
Conclusion: @goncalopp's solution is probably the best one. It is consistently fast. On the other hand, if you're likely to have more zeros than ones, stepping through the list and manually multiplying powers of two and adding them will be fastest.
EDIT: I re-wrote my script to use timeit, the source code is at http://pastebin.com/m6sSmmR6
My output result:
7.78366303444
2.79321694374
5.29976511002
.
5.72017598152
5.70349907875
5.66881299019
.
5.25683712959
5.17318511009
5.20052909851
.
8.23388290405
8.24193501472
8.15649604797
.
3.94102287292
3.95323395729
3.9201271534
My method of stepping through the list backwards adding powers of two is still faster if you have all zeros, but otherwise, @sxh2's method is definitely the fastest, and my implementation didn't even include his caching optimization.