A jagged array in Python is pretty much a list of lists as you mentioned.
I would use a dictionary to store the counselors activity information, where the key is the name of the counselor, and the value is the list of activities the counselor will be in charge of e.g.
counselors_activities = {"Adam": ["archery", "canoeing"],
"Bob": ["frisbee", "golf", "painting", "trampoline"],
"Carol": ["tennis", "dance", "skating"],
"Denise": ["cycling"],
"Eddie": ["horseback", "fencing", "soccer"],
"Fiona": ["painting"],
"George": ["basketball", "football"]}
And access each counselor in the dictionary as such:
counselors_activites["Adam"] # when printed will display the result => ['archery', 'canoeing']
In regards to the question, I would store the list of activities available in a list, and anytime an activity is chosen, remove it from the list and add it to the counselor in the dictionary as such:
list_of_available_activities.remove("archery")
counselors_activities["Adam"].append("archery")
And if a counselor no longer was in charge of the activity, remove it from them and add it back to the list of available activities.
Update: I have provided a more fully fledged solution below based on your requirements from your comments.
Text file, activites.txt:
Adam: archery, canoeing
Bob: frisbee, golf, painting, trampoline
Carol: tennis, dance, skating
Denise: cycling
Eddie: horseback, fencing, soccer
Fiona: painting
George: basketball, football
Code:
#Set of activities available for counselors to choose from
set_of_activities = {"archery",
"canoeing",
"frisbee",
"golf",
"painting",
"trampoline",
"tennis",
"dance",
"skating",
"cycling",
"horseback",
"fencing",
"soccer",
"painting",
"basketball",
"football"}
with open('activities.txt', 'r') as f:
for line in f:
# Iterate over the file and pull out the counselor's names
# and insert their activities into a list
counselor_and_activities = line.split(':')
counselor = counselor_and_activities[0]
activities = counselor_and_activities[1].strip().split(', ')
# Iterate over the list of activities chosen by the counselor and
# see if that activity is free to choose from and if the activity
# is free to choose, remove it from the set of available activities
# and if it is not free remove it from the counselor's activity list
for activity in activities:
if activity in set_of_activities:
set_of_activities.remove(activity)
else:
activities.remove(activity)
# Insert the counselor and their chosen activities into the dictionary
counselors_activities[counselor] = activities
# print(counselors_activities)
I have made one assumption with this new example, which is that you will already have a set of activities that can be chosen from already available:
I made the text file the same format of the counselors and their activities listed in the question, but the logic can be applied to other methods of storage.
As a side note and a correction from my second example previously, I have used a set to represent the list of activities instead of a list in this example. This set will only be used to verify that no counselor will be in charge of an activity that has already been assigned to someone else; i.e., removing an activity from the set will be faster than removing an activity from the list in worst case.
The counselors can be inserted into the dictionary from the notepad file without having to insert them into a list.
When the dictionary is printed it will yield the result:
{"Adam": ["archery", "canoeing"],
"Bob": ["frisbee", "golf", "painting", "trampoline"],
"Carol": ["tennis", "dance", "skating"],
"Denise": ["cycling"],
"Eddie": ["horseback", "fencing", "soccer"],
"Fiona": [], # Empty activity list as the painting activity was already chosen by Bob
"George": ["basketball", "football"]}
Answer from user8605709 on Stack OverflowA jagged array in Python is pretty much a list of lists as you mentioned.
I would use a dictionary to store the counselors activity information, where the key is the name of the counselor, and the value is the list of activities the counselor will be in charge of e.g.
counselors_activities = {"Adam": ["archery", "canoeing"],
"Bob": ["frisbee", "golf", "painting", "trampoline"],
"Carol": ["tennis", "dance", "skating"],
"Denise": ["cycling"],
"Eddie": ["horseback", "fencing", "soccer"],
"Fiona": ["painting"],
"George": ["basketball", "football"]}
And access each counselor in the dictionary as such:
counselors_activites["Adam"] # when printed will display the result => ['archery', 'canoeing']
In regards to the question, I would store the list of activities available in a list, and anytime an activity is chosen, remove it from the list and add it to the counselor in the dictionary as such:
list_of_available_activities.remove("archery")
counselors_activities["Adam"].append("archery")
And if a counselor no longer was in charge of the activity, remove it from them and add it back to the list of available activities.
Update: I have provided a more fully fledged solution below based on your requirements from your comments.
Text file, activites.txt:
Adam: archery, canoeing
Bob: frisbee, golf, painting, trampoline
Carol: tennis, dance, skating
Denise: cycling
Eddie: horseback, fencing, soccer
Fiona: painting
George: basketball, football
Code:
#Set of activities available for counselors to choose from
set_of_activities = {"archery",
"canoeing",
"frisbee",
"golf",
"painting",
"trampoline",
"tennis",
"dance",
"skating",
"cycling",
"horseback",
"fencing",
"soccer",
"painting",
"basketball",
"football"}
with open('activities.txt', 'r') as f:
for line in f:
# Iterate over the file and pull out the counselor's names
# and insert their activities into a list
counselor_and_activities = line.split(':')
counselor = counselor_and_activities[0]
activities = counselor_and_activities[1].strip().split(', ')
# Iterate over the list of activities chosen by the counselor and
# see if that activity is free to choose from and if the activity
# is free to choose, remove it from the set of available activities
# and if it is not free remove it from the counselor's activity list
for activity in activities:
if activity in set_of_activities:
set_of_activities.remove(activity)
else:
activities.remove(activity)
# Insert the counselor and their chosen activities into the dictionary
counselors_activities[counselor] = activities
# print(counselors_activities)
I have made one assumption with this new example, which is that you will already have a set of activities that can be chosen from already available:
I made the text file the same format of the counselors and their activities listed in the question, but the logic can be applied to other methods of storage.
As a side note and a correction from my second example previously, I have used a set to represent the list of activities instead of a list in this example. This set will only be used to verify that no counselor will be in charge of an activity that has already been assigned to someone else; i.e., removing an activity from the set will be faster than removing an activity from the list in worst case.
The counselors can be inserted into the dictionary from the notepad file without having to insert them into a list.
When the dictionary is printed it will yield the result:
{"Adam": ["archery", "canoeing"],
"Bob": ["frisbee", "golf", "painting", "trampoline"],
"Carol": ["tennis", "dance", "skating"],
"Denise": ["cycling"],
"Eddie": ["horseback", "fencing", "soccer"],
"Fiona": [], # Empty activity list as the painting activity was already chosen by Bob
"George": ["basketball", "football"]}
Answer from user8605709 on Stack OverflowHow to make a jagged array neat in Python? - Stack Overflow
Iterate through jagged array values and indices in Python - Stack Overflow
python jagged array operation efficiency - Stack Overflow
python - How to make 2D jagged array using NumPy - Stack Overflow
Just have a 2d array and fill it with zeros
jaggedArray = [[] for row in range(3)]
'''
above line same as
jaggedArray = []
for row in range(3):
jaggedArray.append([])
'''
jaggedArray[0] = [0]*5
jaggedArray[1] = [0]*4
jaggedArray[2] = [0]*2
print(jaggedArray)
You can do something like this:
jaggedArray=list()
jaggedArray.append([0]*5)
jaggedArray.append([0]*4)
jaggedArray.append([0]*2)
This will create similar array that you have created in sample code.
Unless I misunderstand the question, you just want the product of the sub-lists, although you have to wrap any single elements into lists first.
>>> from itertools import product
>>> arr = ['a', ['e', 'r', 't'], ['c', 'd']]
>>> listified = [x if isinstance(x, list) else [x] for x in arr]
>>> listified
[['a'], ['e', 'r', 't'], ['c', 'd']]
>>> list(product(*listified))
[('a', 'e', 'c'),
('a', 'e', 'd'),
('a', 'r', 'c'),
('a', 'r', 'd'),
('a', 't', 'c'),
('a', 't', 'd')]
I have a recursive solution:
inlist1 = ['ab', ['e', 'r', 't'], ['c', 'd']]
inlist2 = [['a', 'b'], ['e', 'r', 't'], ['c', 'd']]
inlist3 = [['a', 'b'], 'e', ['c', 'd']]
def jagged(inlist):
a = [None] * len(inlist)
def _jagged(index):
if index == 0:
print(a)
return
v = inlist[index - 1]
if isinstance(v, list):
for i in v:
a[index - 1] = i
_jagged(index - 1, )
else:
a[index - 1] = v
_jagged(index - 1)
_jagged(len(inlist))
jagged(inlist3)
Use enumerate.
jagged = [[1], [2, 3], [4, 5, 6]]
for i, sub_list in enumerate(jagged):
for j, value in enumerate(sub_list):
print 'a[{}][{}] = {}'.format(i, j, value)
# a[0][0] = 1
# a[1][0] = 2
# a[1][1] = 3
# a[2][0] = 4
# a[2][1] = 5
# a[2][2] = 6
The first code snippet fails because you reset i in every outer iteration:
for value in arr:
i = 0 # reset the counter?
j = 0
if isinstance(value, list):
for other in value:
print("a[%s][%s] = %s" % (i, j, other))
i += 1 # increment in the inner loop?
j += 1
The second code snippet fails because you increment too early:
i = 0
for value in arr1:
i += 1 # increment *before* the inner loop?
for j in range(0, len(value)):
print("a[%s][%s] = %s" % (i, j, value[j]))
Nevertheless you can simply use enumerate(..) and make things easier:
for i,value in enumerate(arr1):
for j,item in enumerate(value):
print("a[%s][%s] = %s" % (i, j, item))
enumerate(..) takes as input an iterable and generates tuples containing the index (as first item of the tuple) and the element. So enumerate([1,'a',2,5.0]), it generates tuples (0,1), (1,'a'), (2,2) and (3,5.0).
Your array is 2x2:
In [298]: A
Out[298]:
array([[array([1, 2, 3]), array([4, 5])],
[array([6, 7, 8, 9]), array([10])]], dtype=object)
While A+A works, boolean tests have not been implemented for this kind of array:
In [299]: A>4
...
ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()
I'm going to flatten A because it makes it easier to compare with list operations:
In [301]: A1=A.flatten()
In [303]: A1+A1
Out[303]:
array([array([2, 4, 6]), array([ 8, 10]), array([12, 14, 16, 18]),
array([20])], dtype=object)
In [304]: [a+a for a in A1]
Out[304]: [array([2, 4, 6]), array([ 8, 10]), array([12, 14, 16, 18]), array([20])]
In [305]: timeit A1+A1
100000 loops, best of 3: 6.85 µs per loop
In [306]: timeit [a+a for a in A1]
100000 loops, best of 3: 9.09 µs per loop
The array operation is a bit faster than a list comprehension. But if I first turn the array into a list:
In [307]: A1l=A1.tolist()
In [308]: A1l
Out[308]: [array([1, 2, 3]), array([4, 5]), array([6, 7, 8, 9]), array([10])]
In [309]: timeit [a+a for a in A1l]
100000 loops, best of 3: 5.2 µs per loop
times improve. This is a good indication that the A1+A1 (or even A+A) is using a similar sort of iteration.
So the straight forward way of performing your A,B calculation is
In [310]: A2=[a[a>4] for a in A1]
In [311]: B=[a+a for a in A2]
In [312]: B
Out[312]: [array([], dtype=int32), array([10]), array([12, 14, 16, 18]), array([20])]
(we can convert to/from arrays and lists as needed).
A numpy array stores its data a flat databuffer, and uses the shape and strides attributes to quickly calculate the location of any element, regardless of the dimensions. The fast array operations use compiled code that rapidly steps though the databuffers of arguments, performing the operations element by element (or some other combination).
A dtype object array also has the flat databuffer, but the elements are pointers to lists or arrays elsewhere. So while it can index individual elements quickly, it still has to perform a Python call(s) to access the arrays. So especially when the array is 1d, it is virtually the same as a flat list with the same pointers.
Multidimensional object arrays are nicer than nested lists. You can reshape them, access elements (A[1,3] v Al[1][3]), transpose them, etc. But when it comes to iterating through all the subarrays they don't offer much of a benefit.
Looking again at your 2d array:
In [315]: timeit A+A
100000 loops, best of 3: 6.93 µs per loop # 6.85 for A1+A1 (above)
In [316]: timeit [[j+j for j in i] for i in A]
100000 loops, best of 3: 17.1 µs per loop
In [317]: Al = A.tolist()
In [318]: timeit [[j+j for j in i] for i in Al]
100000 loops, best of 3: 7.01 µs per loop # 5.2 for A1l flat list
Basically the same time for summing the array and iterating through the equivalent nested list.
The performance of numpy jagged array may not be optimal, but there are enough reasons to believe that it should be much better than using python nested list. As explained in your earlier post:
On principle you should have some performance bonus because every element is a numpy array. So you just need a 2 dimensional loop rather than a 3D loop (if you store every number in nested lists). Also it always saves you lots of memory allocation time to avoid using python list.
Here is a simple test:
import time,sys,random
import numpy as np
rand = np.random.rand
L = np.array([[rand(100), rand(200)],[rand(400), rand(300)]], dtype=object)
L1 = [random.random() for i in range(1000)]
arrFunc = np.vectorize(lambda x:x[x>0.3],otypes=[np.ndarray])
start = time.time()
if sys.argv[1]=='np':
for i in range(100000):
B=i*L
else:
for i in range(100000):
B=[i*x for x in L1]
end = time.time()
print ('Arithmetic Op: ', end-start)
start = time.time()
if sys.argv[1]=='np':
for i in range(100000):
B=arrFunc(L)
else:
for i in range(100000):
B=[x for x in L1 if x<0.3]
end = time.time()
print ('Indexing ', end-start)
Result:
> python testNpJarray.py np
Arithmetic Op: 3.9719998836517334
Indexing 8.079999923706055
> python testNpJarray.py list
Arithmetic Op: 53.289000034332275
Indexing 52.10899996757507
This test may not be quite fare because the outter numpy array is quite small, you are welcome to change the size to fit into your application and tell us the results.
Based on this StackOverflow answer:
NumPy does not support jagged arrays natively. gives an array that may or may not behave as you expect.
A workaround using masked arrays can be as follows:
import numpy as np
import numpy.ma as ma
a = np.array([0, 1])
b = np.array([2, 3, 4, 5])
c = np.array([6, 7, 8, 9, 10, 11])
jagged_array = ma.vstack(
[
ma.array(np.resize(a, c.shape[0]), mask=[False, False, True, True, True, True]),
ma.array(
np.resize(b, c.shape[0]), mask=[False, False, False, False, True, True]
),
c,
]
)
print(jagged_array)
print(jagged_array.ndim)
print(jagged_array.shape)
Your output would look like:
❯ python3 sample.py
[[0 1 -- -- -- --]
[2 3 4 5 -- --]
[6 7 8 9 10 11]]
2
(3, 6)
def ndim(arr):
return len(arr)-1
jagged_array = np.array([[None, None], [None, None, None, None], [None, None, None,None, None, None]])
print(jagged_array)
print(ndim(jagged_array))
print(jagged_array.shape)
Perhaps not the most efficient but it works nicely in numpy. and will short circuit as soon as one of the conditions is False. If the first three conditions are True, we have no choice but to iterate through the rows.
Thankfull, all will shortcircuit as soon as one of the iterations is False so it won't check all the rows if it doesn't have to.
def jagged(x):
x = np.asarray(x)
return (
x.dtype == "object"
and x.ndim == 1
and isinstance(x[0], list)
and not all(len(row) == len(x[0]) for row in x)
)
If you wan't to squeeze more efficieny out, it's actually performing the len(x[0]) every iteration of the all part but this is probably inconsequential and this is a lot more legible than the alaternative which would have you write out the whole if statement.
First what you show are lists, not arrays (but more on that later):
In [305]: alist1 = [[1, 2], [3, 4, 5]]
In [306]: alist2 = [[1, 2], [3, 4], [5, 6], [[7], [8]]]
Mixed len at the first level is a simple and obvious test
In [307]: [len(i) for i in alist1]
Out[307]: [2, 3]
but it's not enough with the 2nd example:
In [308]: [len(i) for i in alist2]
Out[308]: [2, 2, 2, 2]
Making an array from list1 produces a 1d object dtype:
In [310]: np.array(alist1)
Out[310]: array([list([1, 2]), list([3, 4, 5])], dtype=object)
list2 is 2d, but still object dtype:
In [311]: np.array(alist2)
Out[311]:
array([[1, 2],
[3, 4],
[5, 6],
[list([7]), list([8])]], dtype=object)
np.array is not the most efficient tool; while compiled, it does have to evaluate the nest list at least down to the level where it finds the discrepency.
If the list isn't ragged, at any level, the result is a numeric dtype:
In [321]: alist3 = [[1, 2], [3, 4], [5, 6], [7, 8]]
In [322]: np.array(alist3)
Out[322]:
array([[1, 2],
[3, 4],
[5, 6],
[7, 8]])
If the list elements are arrays, there can be a further result - a broadcasting error. This is results when the first dimensions match, but the differences are in the lower level(s).
In sum, if it is already a numpy array, then object is a good indicator, especially if you were expecting a numeric dtype. If the lowest level elements might themselves be objects (other than lists) this won't help. In both the list1 and list2 cases, some or all of the lowest level elements are objects - lists.
If it's a list of lists, then recursive evaluation of the len is probably the way to go. But only time tests can prove that this is better than np.array(alist).