You could use np.array(list(result.items()), dtype=dtype):
import numpy as np
result = {0: 1.1181753789488595, 1: 0.5566080288678394, 2: 0.4718269778030734, 3: 0.48716683119447185, 4: 1.0, 5: 0.1395076201641266, 6: 0.20941558441558442}
names = ['id','data']
formats = ['f8','f8']
dtype = dict(names = names, formats=formats)
array = np.array(list(result.items()), dtype=dtype)
print(repr(array))
yields
array([(0.0, 1.1181753789488595), (1.0, 0.5566080288678394),
(2.0, 0.4718269778030734), (3.0, 0.48716683119447185), (4.0, 1.0),
(5.0, 0.1395076201641266), (6.0, 0.20941558441558442)],
dtype=[('id', '<f8'), ('data', '<f8')])
If you don't want to create the intermediate list of tuples, list(result.items()), then you could instead use np.fromiter:
In Python2:
array = np.fromiter(result.iteritems(), dtype=dtype, count=len(result))
In Python3:
array = np.fromiter(result.items(), dtype=dtype, count=len(result))
Why using the list [key,val] does not work:
By the way, your attempt,
numpy.array([[key,val] for (key,val) in result.iteritems()],dtype)
was very close to working. If you change the list [key, val] to the tuple (key, val), then it would have worked. Of course,
numpy.array([(key,val) for (key,val) in result.iteritems()], dtype)
is the same thing as
numpy.array(result.items(), dtype)
in Python2, or
numpy.array(list(result.items()), dtype)
in Python3.
np.array treats lists differently than tuples: Robert Kern explains:
As a rule, tuples are considered "scalar" records and lists are recursed upon. This rule helps numpy.array() figure out which sequences are records and which are other sequences to be recursed upon; i.e. which sequences create another dimension and which are the atomic elements.
Since (0.0, 1.1181753789488595) is considered one of those atomic elements, it should be a tuple, not a list.
You could use np.array(list(result.items()), dtype=dtype):
import numpy as np
result = {0: 1.1181753789488595, 1: 0.5566080288678394, 2: 0.4718269778030734, 3: 0.48716683119447185, 4: 1.0, 5: 0.1395076201641266, 6: 0.20941558441558442}
names = ['id','data']
formats = ['f8','f8']
dtype = dict(names = names, formats=formats)
array = np.array(list(result.items()), dtype=dtype)
print(repr(array))
yields
array([(0.0, 1.1181753789488595), (1.0, 0.5566080288678394),
(2.0, 0.4718269778030734), (3.0, 0.48716683119447185), (4.0, 1.0),
(5.0, 0.1395076201641266), (6.0, 0.20941558441558442)],
dtype=[('id', '<f8'), ('data', '<f8')])
If you don't want to create the intermediate list of tuples, list(result.items()), then you could instead use np.fromiter:
In Python2:
array = np.fromiter(result.iteritems(), dtype=dtype, count=len(result))
In Python3:
array = np.fromiter(result.items(), dtype=dtype, count=len(result))
Why using the list [key,val] does not work:
By the way, your attempt,
numpy.array([[key,val] for (key,val) in result.iteritems()],dtype)
was very close to working. If you change the list [key, val] to the tuple (key, val), then it would have worked. Of course,
numpy.array([(key,val) for (key,val) in result.iteritems()], dtype)
is the same thing as
numpy.array(result.items(), dtype)
in Python2, or
numpy.array(list(result.items()), dtype)
in Python3.
np.array treats lists differently than tuples: Robert Kern explains:
As a rule, tuples are considered "scalar" records and lists are recursed upon. This rule helps numpy.array() figure out which sequences are records and which are other sequences to be recursed upon; i.e. which sequences create another dimension and which are the atomic elements.
Since (0.0, 1.1181753789488595) is considered one of those atomic elements, it should be a tuple, not a list.
Similarly to the approved answer. If you want to create an array from dictionary keys:
np.array( tuple(dict.keys()) )
If you want to create an array from dictionary values:
np.array( tuple(dict.values()) )
You have a 0-dimensional array of object dtype. Making this array at all is probably a mistake, but if you want to use it anyway, you can extract the dictionary by indexing the array with a tuple of no indices:
x[()]
or by calling the array's item method:
x.item()
If you add square brackets to the array assignment you will have a 1-dimensional array:
x = np.array([{'x': 2, 'y': 5}])
then you could use:
x[0]['y']
I believe it would make more sense.
I have about 200,000 records to iterate about and a list of group names whose corresponding values need to be pulled for each record.
import numpy as np
#Vectorized function to use array elements as keys to pull from dict
pull_dict = np.vectorize(lambda e,g: g[e])
#Group names to pull under each record
groups= []
groups.append(["A","B"])
groups.append(["A","C"])
groups.append(["C","D"])
groups = np.asarray(groups)
#Example of record_dicts for each record
# record_dict = {"A":1,"B":2,"C":3,"D":4} #Record 1
# record_dict = {"A":2,"B":3,"C":1,"D":4} #Record 2
# .
# .
# n (where n= 200000)
#For each record object
all_values = []
for record in records:
#Pull dictionary stored inside record object
record_dict = record["dict"]
#Pull values
values = pull_dict(groups,record_dict)
all_values.append(values)
#Example Input/Output for Record #1 (0th index in loop)
#-----------------------------------------------------------
# #Input ndarray
# [["A","B"]
# ["A","C"]
# ["C","D"]]
# #Output ndarray (Ex: For record #1)
# [[1,2]
# [1,3]
# [3,4]]There's about 200,000 records and the above code is just a sample. It takes about 8 seconds to go through this portion of code. Trying to optimize it to be as fast as possible. I am currently using a numpy vectorized function to treat each element as a key used to pull values from the dictionary. The vectorized function returns an ndarray of equal size to "groups". Each record has a unique dict that ties the group name to an integer.
Is there a more direct way to use an ndarray as a bunch of keys for pulling values from a dictionary? Looking for further speed reductions.