Based on this StackOverflow answer:

NumPy does not support jagged arrays natively. gives an array that may or may not behave as you expect.

A workaround using masked arrays can be as follows:

import numpy as np
import numpy.ma as ma

a = np.array([0, 1])
b = np.array([2, 3, 4, 5])
c = np.array([6, 7, 8, 9, 10, 11])

jagged_array = ma.vstack(
    [
        ma.array(np.resize(a, c.shape[0]), mask=[False, False, True, True, True, True]),
        ma.array(
            np.resize(b, c.shape[0]), mask=[False, False, False, False, True, True]
        ),
        c,
    ]
)
print(jagged_array)
print(jagged_array.ndim)
print(jagged_array.shape)

Your output would look like:

❯ python3 sample.py
[[0 1 -- -- -- --]
 [2 3 4 5 -- --]
 [6 7 8 9 10 11]]
2
(3, 6)
Answer from user4109800 on Stack Overflow
🌐
Reddit
reddit.com › r/learnpython › numpy stack jagged arrays – can i make this code cleaner?
r/learnpython on Reddit: NumPy stack jagged arrays – can I make this code cleaner?
September 4, 2015 - Subreddit for posting questions and asking for general advice about your python code. ... I have 3 NumPy arrays of different lengths and want to combine them into a matrix, filling in 0s to make them equal length. I've used a rather dirty for-loop solution – is there a better way to do this? #this matrix may be jagged.
Discussions

python jagged array operation efficiency - Stack Overflow
I am new to Python and I am looking for the most efficient way to do operations with a jagged array. ... Apparently python is very efficient for doing operations like this with numpy arrays, but unfortunetely I need to do this for jagged arrays and I havent found such an object in Python. More on stackoverflow.com
🌐 stackoverflow.com
July 26, 2016
How to make a jagged array neat in Python? - Stack Overflow
0 How to apply a function on jagged Numpy arrays (unequal row lengths) without using np.apply_along_axis()? More on stackoverflow.com
🌐 stackoverflow.com
python - Convert jagged lists into numpy array - Stack Overflow
I have list consisting of different lengths (see below) [(880), (880, 1080), (880, 1080, 1080), (470, 470, 470, 1250)] I want to convert it to same looking numpy.array, even if I have to fill b... More on stackoverflow.com
🌐 stackoverflow.com
May 11, 2021
RDataFrame -> AsNumpy as jagged arrays
Hello I would like to use something like scikit-hep jagged array to retrieve information from the ROOT trees → dictionary-of-flat numpy-arrays. E.g.: event entry with dynamic array tracks with parameter attributes track entry with array of clusters (position charge) In many use cases we need ... More on root-forum.cern.ch
🌐 root-forum.cern.ch
0
0
February 17, 2022
🌐
GitHub
github.com › scikit-hep › awkward-0.x › issues › 13
Conversion of JaggedArray to numpy array broken · Issue #13 · scikit-hep/awkward-0.x
October 24, 2018 - import numpy as np from awkward import * from awkward.type import * a = JaggedArray([0, 3, 3, 5], [3, 3, 5, 10], [0.0, 1.1, 2.2, 3.3, 4.4, 5.5, 6.6, 7.7, 8.8, 9.9]) np.asarray(a) ... <snip> File "/home/phxlk/.local/lib/python2.7/site-packages/awkward/array/base.py", line 39, in __array__ return ...
Author: scikit-hep
🌐
GitHub
github.com › scikit-hep › awkward-0.x
GitHub - scikit-hep/awkward-0.x: Manipulate arrays of complex data structures as easily as Numpy. · GitHub
In some cases, that may be what you want, but in many, especially any cases involving jagged arrays, it will be a major performance loss and a loss of functionality: jagged arrays turn into Numpy dtype=object arrays containing Numpy arrays, which ...
Starred by 214 users
Forked by 38 users
Languages: Python 63.7% | Jupyter Notebook 36.3%
🌐
GitHub
github.com › topics › jagged-array
jagged-array · GitHub Topics · GitHub
December 2, 2022 - python data-science data-structure ... jagged-array ... A Python library for numpy arrays that persist on disk in a format that is simple, self-documented and tool-independent, and maximizes universal readability....
🌐
Readthedocs
landlab.readthedocs.io › en › latest › _modules › landlab › utils › jaggedarray.html
landlab.utils.jaggedarray - landlab
[docs] def flatten_jagged_array(jagged, dtype=None): """Flatten a list of lists. Parameters ---------- jagged : array_like of array_like An array of arrays of unequal length. Returns ------- (data, offset) : (ndarray, ndarray of int) A tuple the data, as a flat numpy array, and offsets into that array for every item of the original list.
🌐
Frank Sauerburger
frank.sauerburger.io › 2020 › 03 › 11 › awkward-and-numba.html
Awkward arrays and numba | Frank Sauerburger
March 11, 2020 - We can construct a JITed wrapper taking these arrays as input, which then slices the content array and passes the slices to the signum defined above. We can even go one step further and package all of this in a decorator. from functools import wraps import numba import numpy as np def jagged_loop(func): """ Function decorator.
Find elsewhere
🌐
GitHub
github.com › scikit-hep › awkward-0.x › blob › master › docs › classes.adoc
awkward-0.x/docs/classes.adoc at master · scikit-hep/awkward-0.x
June 21, 2022 - If jagged arrays are passed into a Numpy ufunc (or equivalent mapped kernel), they are computed elementwise at the deepest level of jaggedness, adjusting for different starts/stops/content representations of the same logical structure, and broadcasting scalars and non-jagged values to the jagged structure.
Author: scikit-hep
Top answer
1 of 2
3

Your array is 2x2:

In [298]: A
Out[298]: 
array([[array([1, 2, 3]), array([4, 5])],
       [array([6, 7, 8, 9]), array([10])]], dtype=object)

While A+A works, boolean tests have not been implemented for this kind of array:

In [299]: A>4
...
ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()

I'm going to flatten A because it makes it easier to compare with list operations:

In [301]: A1=A.flatten()

In [303]: A1+A1
Out[303]: 
array([array([2, 4, 6]), array([ 8, 10]), array([12, 14, 16, 18]),
       array([20])], dtype=object)

In [304]: [a+a for a in A1]
Out[304]: [array([2, 4, 6]), array([ 8, 10]), array([12, 14, 16, 18]), array([20])]

In [305]: timeit A1+A1
100000 loops, best of 3: 6.85 µs per loop

In [306]: timeit [a+a for a in A1]
100000 loops, best of 3: 9.09 µs per loop

The array operation is a bit faster than a list comprehension. But if I first turn the array into a list:

In [307]: A1l=A1.tolist()

In [308]: A1l
Out[308]: [array([1, 2, 3]), array([4, 5]), array([6, 7, 8, 9]), array([10])]

In [309]: timeit [a+a for a in A1l]
100000 loops, best of 3: 5.2 µs per loop

times improve. This is a good indication that the A1+A1 (or even A+A) is using a similar sort of iteration.

So the straight forward way of performing your A,B calculation is

In [310]: A2=[a[a>4] for a in A1]
In [311]: B=[a+a for a in A2]
In [312]: B
Out[312]: [array([], dtype=int32), array([10]), array([12, 14, 16, 18]), array([20])]

(we can convert to/from arrays and lists as needed).

A numpy array stores its data a flat databuffer, and uses the shape and strides attributes to quickly calculate the location of any element, regardless of the dimensions. The fast array operations use compiled code that rapidly steps though the databuffers of arguments, performing the operations element by element (or some other combination).

A dtype object array also has the flat databuffer, but the elements are pointers to lists or arrays elsewhere. So while it can index individual elements quickly, it still has to perform a Python call(s) to access the arrays. So especially when the array is 1d, it is virtually the same as a flat list with the same pointers.

Multidimensional object arrays are nicer than nested lists. You can reshape them, access elements (A[1,3] v Al[1][3]), transpose them, etc. But when it comes to iterating through all the subarrays they don't offer much of a benefit.

Looking again at your 2d array:

In [315]: timeit A+A
100000 loops, best of 3: 6.93 µs per loop  # 6.85 for A1+A1 (above)

In [316]: timeit [[j+j for j in i] for i in A]
100000 loops, best of 3: 17.1 µs per loop

In [317]: Al = A.tolist()

In [318]: timeit [[j+j for j in i] for i in Al]
100000 loops, best of 3: 7.01 µs per loop    # 5.2 for A1l flat list

Basically the same time for summing the array and iterating through the equivalent nested list.

2 of 2
0

The performance of numpy jagged array may not be optimal, but there are enough reasons to believe that it should be much better than using python nested list. As explained in your earlier post:

On principle you should have some performance bonus because every element is a numpy array. So you just need a 2 dimensional loop rather than a 3D loop (if you store every number in nested lists). Also it always saves you lots of memory allocation time to avoid using python list.

Here is a simple test:

import time,sys,random
import numpy as np
rand = np.random.rand
L = np.array([[rand(100), rand(200)],[rand(400), rand(300)]], dtype=object)
L1 = [random.random() for i in range(1000)]
arrFunc = np.vectorize(lambda x:x[x>0.3],otypes=[np.ndarray])

start = time.time()
if sys.argv[1]=='np':
  for i in range(100000):
    B=i*L
else:
  for i in range(100000):
    B=[i*x for x in L1]

end = time.time()
print ('Arithmetic Op: ', end-start)


start = time.time()
if sys.argv[1]=='np':
  for i in range(100000):
    B=arrFunc(L)
else:
  for i in range(100000):
    B=[x for x in L1 if x<0.3]
end = time.time()
print ('Indexing       ', end-start)

Result:

> python testNpJarray.py np
Arithmetic Op:  3.9719998836517334
Indexing        8.079999923706055

> python testNpJarray.py list
Arithmetic Op:  53.289000034332275
Indexing        52.10899996757507

This test may not be quite fare because the outter numpy array is quite small, you are welcome to change the size to fit into your application and tell us the results.

Top answer
1 of 1
1

You might think, your input is a list of tuples. However, it is a list of integers and tuples. (880) will be interpreted as an integer, but not as a tuple. So you have to deal with both datatypes.

First of all I suggest converting your input data to a list of lists. Each of the lists contained in that list should have the same length, because an array supports constant dimensions only. Therefore, I would convert the elements into a list and fill missing values with zeros (to make all elements equal in length).

If we do this for all of the elements given in the input list, we create a new list containing lists of equal length which can be converted into an array.

A very basic (and error-prone) approach would look like this:

import numpy as np


original_list = [
    (880),
    (880, 1080),
    (880, 1080, 1080),
    (470, 470, 470, 1250),
]


def get_len(item):
    try:
        return len(item)
    except TypeError:
        # `(880)` will be interpreted as an int instead of a tuple
        # so we need to handle tuples and integers
        # as integers do not support len(), a TypeError will be raised
        return 1


def to_list(item):
    try:
        return list(item)
    except TypeError:
        # `(880)` will be interpreted as an int instead of a tuple
        # so we need to handle tuples and integers
        # as integers do not support __iter__(), a TypeError will be raised
        return [item]


def fill_zeros(item, max_len):
    item_len = get_len(item)
    to_fill = [0] * (max_len - item_len)
    as_list = to_list(item) + to_fill
    return as_list


max_len = max([get_len(item) for item in original_list])
filled = [fill_zeros(item, max_len) for item in original_list]

arr = np.array(filled)
print(arr)

Printing:

[[ 880    0    0    0]
[ 880 1080    0    0]
[ 880 1080 1080    0]
[ 470  470  470 1250]]
🌐
CERN
root-forum.cern.ch › t › rdataframe-asnumpy-as-jagged-arrays › 48835
RDataFrame -> AsNumpy as jagged arrays - ROOT - ROOT Forum
February 17, 2022 - Hello I would like to use something like scikit-hep jagged array to retrieve information from the ROOT trees → dictionary-of-flat numpy-arrays. E.g.: event entry with dynamic array tracks with parameter attributes track entry with array of clusters (position charge) In many use cases we need ...
🌐
Vlad Feinberg
vladfeinberg.com › 2021 › 01 › 07 › vectorizing-ragged-arrays.html
Vectorizing Ragged Arrays (Numpy Gems, Part 4)
January 7, 2021 - Luckily, notice that our main reduction (np.mean) over the ragged arrays is a composition of two operations: sum / count. Extracting the reduction operation (the sum) into its own step will let us use our numpy gem, np.cumsum + np.diff, to aggregate across ragged arrays.
Top answer
1 of 6
43

Short answer: you can't. NumPy does not support jagged arrays natively.

Long answer:

>>> a = ones((3,))
>>> b = ones((2,))
>>> c = array([a, b])
>>> c
array([[ 1.  1.  1.], [ 1.  1.]], dtype=object)

gives an array that may or may not behave as you expect. E.g. it doesn't support basic methods like sum or reshape, and you should treat this much as you'd treat the ordinary Python list [a, b] (iterate over it to perform operations instead of using vectorized idioms).

Several possible workarounds exist; the easiest is to coerce a and b to a common length, perhaps using masked arrays or NaN to signal that some indices are invalid in some rows. E.g. here's b as a masked array:

>>> ma.array(np.resize(b, a.shape[0]), mask=[False, False, True])
masked_array(data = [1.0 1.0 --],
             mask = [False False  True],
       fill_value = 1e+20)

This can be stacked with a as follows:

>>> ma.vstack([a, ma.array(np.resize(b, a.shape[0]), mask=[False, False, True])])
masked_array(data =
 [[1.0 1.0 1.0]
 [1.0 1.0 --]],
             mask =
 [[False False False]
 [False False  True]],
       fill_value = 1e+20)

(For some purposes, scipy.sparse may also be interesting.)

2 of 6
7

In general, there is an ambiguity in putting together arrays of different length because alignment of data might matter. Pandas has different advanced solutions to deal with that, e.g. to merge series into dataFrames.

If you just want to populate columns starting from first element, what I usually do is build a matrix and populate columns. Of course you need to fill the empty spaces in the matrix with a null value (in this case np.nan)

a = ones((3,))
b = ones((2,))
arraylist=[a,b]

outarr=np.ones((np.max([len(ps) for ps in arraylist]),len(arraylist)))*np.nan #define empty array
for i,c in enumerate(arraylist):  #populate columns
    outarr[:len(c),i]=c

In [108]: outarr
Out[108]: 
array([[  1.,   1.],
       [  1.,   1.],
       [  1.,  nan]])
🌐
Wikipedia
en.wikipedia.org › wiki › Jagged_array
Jagged array - Wikipedia
March 5, 2026 - Jagged array can be implemented with Iliffe vector data structure in languages such as Java, PHP, Python (multidimensional lists), Ruby, C#.NET, Visual Basic.NET, Perl, JavaScript, Objective-C, Swift, and Atlas Autocode.
🌐
GitHub
github.com › scikit-hep › awkward-0.x › issues › 17
Saving a jagged array · Issue #17 · scikit-hep/awkward-0.x
October 25, 2018 - And I actually have several arrays so I don't want to save three differently named arrays per jagged arrays by hand if possible: from awkward import JaggedArray import numpy as np ja = JaggedArray([0,4],[4,6],[1,2,3,4,5,6]) from tempfile import TemporaryFile outfile = TemporaryFile() np.savez(outfile, ja=[ja.starts, ja.stops, ja.content]) outfile.seek(0) f = np.load(outfile) JaggedArray(*f['ja']) If I directly try to save the jaggedarray, Python oddly hangs up instead of giving an error or working...
Author: scikit-hep
🌐
GitHub
github.com › scikit-hep › awkward-0.x › blob › master › README.rst
awkward-0.x/README.rst at master · scikit-hep/awkward-0.x
In some cases, that may be what you want, but in many, especially any cases involving jagged arrays, it will be a major performance loss and a loss of functionality: jagged arrays turn into Numpy dtype=object arrays containing Numpy arrays, which ...
Author: scikit-hep
🌐
James D. McCaffrey
jamesmccaffreyblog.com › home › loading a jagged numeric matrix from text file using python
Loading a Jagged Numeric Matrix From Text File Using Python - James D. McCaffreyJames D. McCaffrey
July 8, 2019 - Without fromstring() I’d have to split each line using split() then create a NumPy numeric array with the correct number of cells for the line using np.zeros(), convert each value from object to int, add each value into the NumPy array, then append the array to the result list. Actually, my problem was to read data from a text file into training and test jagged PyTorch tensors.