Based on this StackOverflow answer:

NumPy does not support jagged arrays natively. gives an array that may or may not behave as you expect.

A workaround using masked arrays can be as follows:

import numpy as np
import numpy.ma as ma

a = np.array([0, 1])
b = np.array([2, 3, 4, 5])
c = np.array([6, 7, 8, 9, 10, 11])

jagged_array = ma.vstack(
    [
        ma.array(np.resize(a, c.shape[0]), mask=[False, False, True, True, True, True]),
        ma.array(
            np.resize(b, c.shape[0]), mask=[False, False, False, False, True, True]
        ),
        c,
    ]
)
print(jagged_array)
print(jagged_array.ndim)
print(jagged_array.shape)

Your output would look like:

❯ python3 sample.py
[[0 1 -- -- -- --]
 [2 3 4 5 -- --]
 [6 7 8 9 10 11]]
2
(3, 6)
Answer from user4109800 on Stack Overflow
🌐
Reddit
reddit.com › r/learnpython › numpy stack jagged arrays – can i make this code cleaner?
r/learnpython on Reddit: NumPy stack jagged arrays – can I make this code cleaner?
September 4, 2015 - Subreddit for posting questions and asking for general advice about your python code. ... I have 3 NumPy arrays of different lengths and want to combine them into a matrix, filling in 0s to make them equal length. I've used a rather dirty for-loop solution – is there a better way to do this? #this matrix may be jagged.
Discussions

What is the Python equivalent of a jagged array? - Stack Overflow
After years of using Excel and learning VBA, I am now trying to learn Python. Here's the scenario: I asked 7 summer camp counselors which activities they would like to be in charge of. Each stud... More on stackoverflow.com
🌐 stackoverflow.com
python jagged array operation efficiency - Stack Overflow
I am new to Python and I am looking for the most efficient way to do operations with a jagged array. ... Apparently python is very efficient for doing operations like this with numpy arrays, but unfortunetely I need to do this for jagged arrays and I havent found such an object in Python. More on stackoverflow.com
🌐 stackoverflow.com
July 26, 2016
How to make a jagged array neat in Python? - Stack Overflow
0 How to apply a function on jagged Numpy arrays (unequal row lengths) without using np.apply_along_axis()? More on stackoverflow.com
🌐 stackoverflow.com
python - Convert jagged lists into numpy array - Stack Overflow
I have list consisting of different lengths (see below) [(880), (880, 1080), (880, 1080, 1080), (470, 470, 470, 1250)] I want to convert it to same looking numpy.array, even if I have to fill b... More on stackoverflow.com
🌐 stackoverflow.com
May 11, 2021
🌐
GitHub
github.com › scikit-hep › awkward-0.x › issues › 13
Conversion of JaggedArray to numpy array broken · Issue #13 · scikit-hep/awkward-0.x
October 24, 2018 - import numpy as np from awkward import * from awkward.type import * a = JaggedArray([0, 3, 3, 5], [3, 3, 5, 10], [0.0, 1.1, 2.2, 3.3, 4.4, 5.5, 6.6, 7.7, 8.8, 9.9]) np.asarray(a) ... <snip> File "/home/phxlk/.local/lib/python2.7/site-packages/awkward/array/base.py", line 39, in __array__ return ...
Author: scikit-hep
🌐
GitHub
github.com › scikit-hep › awkward-0.x
GitHub - scikit-hep/awkward-0.x: Manipulate arrays of complex data structures as easily as Numpy. · GitHub
In some cases, that may be what you want, but in many, especially any cases involving jagged arrays, it will be a major performance loss and a loss of functionality: jagged arrays turn into Numpy dtype=object arrays containing Numpy arrays, which ...
Starred by 214 users
Forked by 38 users
Languages: Python 63.7% | Jupyter Notebook 36.3%
🌐
GitHub
github.com › topics › jagged-array
jagged-array · GitHub Topics · GitHub
December 2, 2022 - python data-science data-structure ... jagged-array ... A Python library for numpy arrays that persist on disk in a format that is simple, self-documented and tool-independent, and maximizes universal readability....
🌐
Readthedocs
landlab.readthedocs.io › en › latest › _modules › landlab › utils › jaggedarray.html
landlab.utils.jaggedarray - landlab
[docs] def flatten_jagged_array(jagged, dtype=None): """Flatten a list of lists. Parameters ---------- jagged : array_like of array_like An array of arrays of unequal length. Returns ------- (data, offset) : (ndarray, ndarray of int) A tuple the data, as a flat numpy array, and offsets into that array for every item of the original list.
🌐
Frank Sauerburger
frank.sauerburger.io › 2020 › 03 › 11 › awkward-and-numba.html
Awkward arrays and numba | Frank Sauerburger
March 11, 2020 - We can construct a JITed wrapper taking these arrays as input, which then slices the content array and passes the slices to the signum defined above. We can even go one step further and package all of this in a decorator. from functools import wraps import numba import numpy as np def jagged_loop(func): """ Function decorator.
Find elsewhere
Top answer
1 of 1
4

A jagged array in Python is pretty much a list of lists as you mentioned.

I would use a dictionary to store the counselors activity information, where the key is the name of the counselor, and the value is the list of activities the counselor will be in charge of e.g.

counselors_activities = {"Adam": ["archery", "canoeing"],
                      "Bob": ["frisbee", "golf", "painting", "trampoline"],
                      "Carol": ["tennis", "dance", "skating"],
                      "Denise": ["cycling"],
                      "Eddie": ["horseback", "fencing", "soccer"],
                      "Fiona": ["painting"],
                      "George": ["basketball", "football"]}

And access each counselor in the dictionary as such:

counselors_activites["Adam"] # when printed will display the result => ['archery', 'canoeing']

In regards to the question, I would store the list of activities available in a list, and anytime an activity is chosen, remove it from the list and add it to the counselor in the dictionary as such:

list_of_available_activities.remove("archery")
counselors_activities["Adam"].append("archery")

And if a counselor no longer was in charge of the activity, remove it from them and add it back to the list of available activities.

Update: I have provided a more fully fledged solution below based on your requirements from your comments.

Text file, activites.txt:

Adam: archery, canoeing
Bob: frisbee, golf, painting, trampoline
Carol: tennis, dance, skating
Denise: cycling
Eddie: horseback, fencing, soccer
Fiona: painting
George: basketball, football

Code:

#Set of activities available for counselors to choose from

set_of_activities = {"archery",
                  "canoeing",
                  "frisbee",
                  "golf",
                  "painting",
                  "trampoline",
                  "tennis",
                  "dance",
                  "skating",
                  "cycling",
                  "horseback",
                  "fencing",
                  "soccer",
                  "painting",
                  "basketball",
                  "football"}

with open('activities.txt', 'r') as f:
    for line in f:

        # Iterate over the file and pull out the counselor's names
        # and insert their activities into a list

        counselor_and_activities = line.split(':')
        counselor = counselor_and_activities[0]
        activities = counselor_and_activities[1].strip().split(', ')

    # Iterate over the list of activities chosen by the counselor and
    # see if that activity is free to choose from and if the activity
    # is free to choose, remove it from the set of available activities
    # and if it is not free remove it from the counselor's activity list

    for activity in activities:
        if activity in set_of_activities:
            set_of_activities.remove(activity)
        else:
            activities.remove(activity)

    # Insert the counselor and their chosen activities into the dictionary

    counselors_activities[counselor] = activities

# print(counselors_activities)

I have made one assumption with this new example, which is that you will already have a set of activities that can be chosen from already available:

I made the text file the same format of the counselors and their activities listed in the question, but the logic can be applied to other methods of storage.

As a side note and a correction from my second example previously, I have used a set to represent the list of activities instead of a list in this example. This set will only be used to verify that no counselor will be in charge of an activity that has already been assigned to someone else; i.e., removing an activity from the set will be faster than removing an activity from the list in worst case.

The counselors can be inserted into the dictionary from the notepad file without having to insert them into a list.

When the dictionary is printed it will yield the result:

{"Adam": ["archery", "canoeing"],
 "Bob": ["frisbee", "golf", "painting", "trampoline"],
 "Carol": ["tennis", "dance", "skating"],
 "Denise": ["cycling"],
 "Eddie": ["horseback", "fencing", "soccer"],
 "Fiona": [], # Empty activity list as the painting activity was already chosen by Bob
 "George": ["basketball", "football"]}
🌐
GitHub
github.com › scikit-hep › awkward-0.x › blob › master › docs › classes.adoc
awkward-0.x/docs/classes.adoc at master · scikit-hep/awkward-0.x
June 21, 2022 - If jagged arrays are passed into a Numpy ufunc (or equivalent mapped kernel), they are computed elementwise at the deepest level of jaggedness, adjusting for different starts/stops/content representations of the same logical structure, and broadcasting scalars and non-jagged values to the jagged structure.
Author: scikit-hep
Top answer
1 of 2
3

Your array is 2x2:

In [298]: A
Out[298]: 
array([[array([1, 2, 3]), array([4, 5])],
       [array([6, 7, 8, 9]), array([10])]], dtype=object)

While A+A works, boolean tests have not been implemented for this kind of array:

In [299]: A>4
...
ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()

I'm going to flatten A because it makes it easier to compare with list operations:

In [301]: A1=A.flatten()

In [303]: A1+A1
Out[303]: 
array([array([2, 4, 6]), array([ 8, 10]), array([12, 14, 16, 18]),
       array([20])], dtype=object)

In [304]: [a+a for a in A1]
Out[304]: [array([2, 4, 6]), array([ 8, 10]), array([12, 14, 16, 18]), array([20])]

In [305]: timeit A1+A1
100000 loops, best of 3: 6.85 µs per loop

In [306]: timeit [a+a for a in A1]
100000 loops, best of 3: 9.09 µs per loop

The array operation is a bit faster than a list comprehension. But if I first turn the array into a list:

In [307]: A1l=A1.tolist()

In [308]: A1l
Out[308]: [array([1, 2, 3]), array([4, 5]), array([6, 7, 8, 9]), array([10])]

In [309]: timeit [a+a for a in A1l]
100000 loops, best of 3: 5.2 µs per loop

times improve. This is a good indication that the A1+A1 (or even A+A) is using a similar sort of iteration.

So the straight forward way of performing your A,B calculation is

In [310]: A2=[a[a>4] for a in A1]
In [311]: B=[a+a for a in A2]
In [312]: B
Out[312]: [array([], dtype=int32), array([10]), array([12, 14, 16, 18]), array([20])]

(we can convert to/from arrays and lists as needed).

A numpy array stores its data a flat databuffer, and uses the shape and strides attributes to quickly calculate the location of any element, regardless of the dimensions. The fast array operations use compiled code that rapidly steps though the databuffers of arguments, performing the operations element by element (or some other combination).

A dtype object array also has the flat databuffer, but the elements are pointers to lists or arrays elsewhere. So while it can index individual elements quickly, it still has to perform a Python call(s) to access the arrays. So especially when the array is 1d, it is virtually the same as a flat list with the same pointers.

Multidimensional object arrays are nicer than nested lists. You can reshape them, access elements (A[1,3] v Al[1][3]), transpose them, etc. But when it comes to iterating through all the subarrays they don't offer much of a benefit.

Looking again at your 2d array:

In [315]: timeit A+A
100000 loops, best of 3: 6.93 µs per loop  # 6.85 for A1+A1 (above)

In [316]: timeit [[j+j for j in i] for i in A]
100000 loops, best of 3: 17.1 µs per loop

In [317]: Al = A.tolist()

In [318]: timeit [[j+j for j in i] for i in Al]
100000 loops, best of 3: 7.01 µs per loop    # 5.2 for A1l flat list

Basically the same time for summing the array and iterating through the equivalent nested list.

2 of 2
0

The performance of numpy jagged array may not be optimal, but there are enough reasons to believe that it should be much better than using python nested list. As explained in your earlier post:

On principle you should have some performance bonus because every element is a numpy array. So you just need a 2 dimensional loop rather than a 3D loop (if you store every number in nested lists). Also it always saves you lots of memory allocation time to avoid using python list.

Here is a simple test:

import time,sys,random
import numpy as np
rand = np.random.rand
L = np.array([[rand(100), rand(200)],[rand(400), rand(300)]], dtype=object)
L1 = [random.random() for i in range(1000)]
arrFunc = np.vectorize(lambda x:x[x>0.3],otypes=[np.ndarray])

start = time.time()
if sys.argv[1]=='np':
  for i in range(100000):
    B=i*L
else:
  for i in range(100000):
    B=[i*x for x in L1]

end = time.time()
print ('Arithmetic Op: ', end-start)


start = time.time()
if sys.argv[1]=='np':
  for i in range(100000):
    B=arrFunc(L)
else:
  for i in range(100000):
    B=[x for x in L1 if x<0.3]
end = time.time()
print ('Indexing       ', end-start)

Result:

> python testNpJarray.py np
Arithmetic Op:  3.9719998836517334
Indexing        8.079999923706055

> python testNpJarray.py list
Arithmetic Op:  53.289000034332275
Indexing        52.10899996757507

This test may not be quite fare because the outter numpy array is quite small, you are welcome to change the size to fit into your application and tell us the results.

Top answer
1 of 1
1

You might think, your input is a list of tuples. However, it is a list of integers and tuples. (880) will be interpreted as an integer, but not as a tuple. So you have to deal with both datatypes.

First of all I suggest converting your input data to a list of lists. Each of the lists contained in that list should have the same length, because an array supports constant dimensions only. Therefore, I would convert the elements into a list and fill missing values with zeros (to make all elements equal in length).

If we do this for all of the elements given in the input list, we create a new list containing lists of equal length which can be converted into an array.

A very basic (and error-prone) approach would look like this:

import numpy as np


original_list = [
    (880),
    (880, 1080),
    (880, 1080, 1080),
    (470, 470, 470, 1250),
]


def get_len(item):
    try:
        return len(item)
    except TypeError:
        # `(880)` will be interpreted as an int instead of a tuple
        # so we need to handle tuples and integers
        # as integers do not support len(), a TypeError will be raised
        return 1


def to_list(item):
    try:
        return list(item)
    except TypeError:
        # `(880)` will be interpreted as an int instead of a tuple
        # so we need to handle tuples and integers
        # as integers do not support __iter__(), a TypeError will be raised
        return [item]


def fill_zeros(item, max_len):
    item_len = get_len(item)
    to_fill = [0] * (max_len - item_len)
    as_list = to_list(item) + to_fill
    return as_list


max_len = max([get_len(item) for item in original_list])
filled = [fill_zeros(item, max_len) for item in original_list]

arr = np.array(filled)
print(arr)

Printing:

[[ 880    0    0    0]
[ 880 1080    0    0]
[ 880 1080 1080    0]
[ 470  470  470 1250]]
🌐
CERN
root-forum.cern.ch › t › rdataframe-asnumpy-as-jagged-arrays › 48835
RDataFrame -> AsNumpy as jagged arrays - ROOT - ROOT Forum
February 17, 2022 - Hello I would like to use something like scikit-hep jagged array to retrieve information from the ROOT trees → dictionary-of-flat numpy-arrays. E.g.: event entry with dynamic array tracks with parameter attributes track entry with array of clusters (position charge) In many use cases we need ...
Top answer
1 of 6
43

Short answer: you can't. NumPy does not support jagged arrays natively.

Long answer:

>>> a = ones((3,))
>>> b = ones((2,))
>>> c = array([a, b])
>>> c
array([[ 1.  1.  1.], [ 1.  1.]], dtype=object)

gives an array that may or may not behave as you expect. E.g. it doesn't support basic methods like sum or reshape, and you should treat this much as you'd treat the ordinary Python list [a, b] (iterate over it to perform operations instead of using vectorized idioms).

Several possible workarounds exist; the easiest is to coerce a and b to a common length, perhaps using masked arrays or NaN to signal that some indices are invalid in some rows. E.g. here's b as a masked array:

>>> ma.array(np.resize(b, a.shape[0]), mask=[False, False, True])
masked_array(data = [1.0 1.0 --],
             mask = [False False  True],
       fill_value = 1e+20)

This can be stacked with a as follows:

>>> ma.vstack([a, ma.array(np.resize(b, a.shape[0]), mask=[False, False, True])])
masked_array(data =
 [[1.0 1.0 1.0]
 [1.0 1.0 --]],
             mask =
 [[False False False]
 [False False  True]],
       fill_value = 1e+20)

(For some purposes, scipy.sparse may also be interesting.)

2 of 6
7

In general, there is an ambiguity in putting together arrays of different length because alignment of data might matter. Pandas has different advanced solutions to deal with that, e.g. to merge series into dataFrames.

If you just want to populate columns starting from first element, what I usually do is build a matrix and populate columns. Of course you need to fill the empty spaces in the matrix with a null value (in this case np.nan)

a = ones((3,))
b = ones((2,))
arraylist=[a,b]

outarr=np.ones((np.max([len(ps) for ps in arraylist]),len(arraylist)))*np.nan #define empty array
for i,c in enumerate(arraylist):  #populate columns
    outarr[:len(c),i]=c

In [108]: outarr
Out[108]: 
array([[  1.,   1.],
       [  1.,   1.],
       [  1.,  nan]])
🌐
Vlad Feinberg
vladfeinberg.com › 2021 › 01 › 07 › vectorizing-ragged-arrays.html
Vectorizing Ragged Arrays (Numpy Gems, Part 4)
January 7, 2021 - Luckily, notice that our main reduction (np.mean) over the ragged arrays is a composition of two operations: sum / count. Extracting the reduction operation (the sum) into its own step will let us use our numpy gem, np.cumsum + np.diff, to aggregate across ragged arrays.
🌐
Wikipedia
en.wikipedia.org › wiki › Jagged_array
Jagged array - Wikipedia
March 5, 2026 - Jagged array can be implemented with Iliffe vector data structure in languages such as Java, PHP, Python (multidimensional lists), Ruby, C#.NET, Visual Basic.NET, Perl, JavaScript, Objective-C, Swift, and Atlas Autocode.
🌐
GitHub
github.com › scikit-hep › awkward-0.x › blob › master › README.rst
awkward-0.x/README.rst at master · scikit-hep/awkward-0.x
In some cases, that may be what you want, but in many, especially any cases involving jagged arrays, it will be a major performance loss and a loss of functionality: jagged arrays turn into Numpy dtype=object arrays containing Numpy arrays, which ...
Author: scikit-hep
🌐
GitHub
github.com › scikit-hep › awkward-0.x › issues › 17
Saving a jagged array · Issue #17 · scikit-hep/awkward-0.x
October 25, 2018 - And I actually have several arrays so I don't want to save three differently named arrays per jagged arrays by hand if possible: from awkward import JaggedArray import numpy as np ja = JaggedArray([0,4],[4,6],[1,2,3,4,5,6]) from tempfile import TemporaryFile outfile = TemporaryFile() np.savez(outfile, ja=[ja.starts, ja.stops, ja.content]) outfile.seek(0) f = np.load(outfile) JaggedArray(*f['ja']) If I directly try to save the jaggedarray, Python oddly hangs up instead of giving an error or working...
Author: scikit-hep