Based on this StackOverflow answer:

NumPy does not support jagged arrays natively. gives an array that may or may not behave as you expect.

A workaround using masked arrays can be as follows:

import numpy as np
import numpy.ma as ma

a = np.array([0, 1])
b = np.array([2, 3, 4, 5])
c = np.array([6, 7, 8, 9, 10, 11])

jagged_array = ma.vstack(
    [
        ma.array(np.resize(a, c.shape[0]), mask=[False, False, True, True, True, True]),
        ma.array(
            np.resize(b, c.shape[0]), mask=[False, False, False, False, True, True]
        ),
        c,
    ]
)
print(jagged_array)
print(jagged_array.ndim)
print(jagged_array.shape)

Your output would look like:

❯ python3 sample.py
[[0 1 -- -- -- --]
 [2 3 4 5 -- --]
 [6 7 8 9 10 11]]
2
(3, 6)
Answer from user4109800 on Stack Overflow
🌐
Awkward-array
awkward-array.org › doc › main › getting-started › jagged-ragged-awkward-arrays.html
Jagged, Ragged, Awkward Arrays! — Awkward Array 2.9.1 documentation
import numpy as np mass = np.sqrt( 2 * mu1.pt * mu2.pt * (np.cosh(mu1.eta - mu2.eta) - np.cos(mu1.phi - mu2.phi)) ) Quick quiz: how many masses do we have in each event? How does this compare with muons, mu1, and mu2? Since this mass is a jagged array, it can’t be directly histogrammed.
Discussions

RDataFrame -> AsNumpy as jagged arrays
Hello I would like to use something like scikit-hep jagged array to retrieve information from the ROOT trees → dictionary-of-flat numpy-arrays. E.g.: event entry with dynamic array tracks with parameter attributes track entry with array of clusters (position charge) In many use cases we need ... More on root-forum.cern.ch
🌐 root-forum.cern.ch
0
0
February 17, 2022
Awkward: Nested, jagged, differentiable, mixed type, GPU-enabled, JIT'd NumPy
Just a note, I noticed that there is an unfortunate error in the page describing the bike route calculations. Right in the box where it's supposed to show how much faster things are after JIT, instead a traceback is displayed ending with: · I don't think this is intentional. More on news.ycombinator.com
🌐 news.ycombinator.com
44
144
December 20, 2021
c++ - Jagged Numpy Arrays in Boost Numpy - Stack Overflow
I need to efficiently share data between Python and C++ and use Boost Python and Boost Numpy for this. It works very well for "cartesian" arrays. For jagged arrays I am not sure if a direct indexing is possible or not. Here is the example I that shows how I extract the second array from a jagged ... More on stackoverflow.com
🌐 stackoverflow.com
June 1, 2022
python - Convert jagged lists into numpy array - Stack Overflow
I have list consisting of different lengths (see below) [(880), (880, 1080), (880, 1080, 1080), (470, 470, 470, 1250)] I want to convert it to same looking numpy.array, even if I have to fill b... More on stackoverflow.com
🌐 stackoverflow.com
May 11, 2021
🌐
Reddit
reddit.com › r/learnpython › numpy stack jagged arrays – can i make this code cleaner?
r/learnpython on Reddit: NumPy stack jagged arrays – can I make this code cleaner?
September 4, 2015 - I have 3 NumPy arrays of different lengths and want to combine them into a matrix, filling in 0s to make them equal length. I've used a rather dirty for-loop solution – is there a better way to do this? #this matrix may be jagged.
🌐
Readthedocs
landlab.readthedocs.io › en › latest › _modules › landlab › utils › jaggedarray.html
landlab.utils.jaggedarray - landlab
Parameters ---------- jagged : array_like of array_like An array of arrays of unequal length. Returns ------- (data, offset) : (ndarray, ndarray of int) A tuple the data, as a flat numpy array, and offsets into that array for every item of the original list. Examples -------- >>> from landlab.utils.jaggedarray import flatten_jagged_array >>> data, offset = flatten_jagged_array([[1, 2], [], [3, 4, 5]], dtype=int) >>> data array([1, 2, 3, 4, 5]) >>> offset array([0, 2, 2, 5]) """ data = np.concatenate(jagged).astype(dtype=dtype) # if len(jagged) > 1: # data = np.concatenate(jagged).astype(dtype=dtype) # else: # data = np.array(jagged[0]).astype(dtype=dtype) items_per_block = np.array([len(block) for block in jagged], dtype=int) offset = np.empty(len(items_per_block) + 1, dtype=int) offset[0] = 0 offset[1:] = np.cumsum(items_per_block) return data, offset
🌐
CERN
root-forum.cern.ch › t › rdataframe-asnumpy-as-jagged-arrays › 48835
RDataFrame -> AsNumpy as jagged arrays - ROOT - ROOT Forum
February 17, 2022 - Hello I would like to use something like scikit-hep jagged array to retrieve information from the ROOT trees → dictionary-of-flat numpy-arrays. E.g.: event entry with dynamic array tracks with parameter attributes track entry with array of clusters (position charge) In many use cases we need ...
🌐
Hacker News
news.ycombinator.com › item
Awkward: Nested, jagged, differentiable, mixed type, GPU-enabled, JIT'd NumPy | Hacker News
December 20, 2021 - Just a note, I noticed that there is an unfortunate error in the page describing the bike route calculations. Right in the box where it's supposed to show how much faster things are after JIT, instead a traceback is displayed ending with: · I don't think this is intentional.
Top answer
1 of 1
1

Arrays of arrays are generally inefficient because the sub array are object causing Numpy to fallback on a slow path interacting with the interpreter (due to reference counting, checks, indirections, etc.) or to implicitly convert them into a big array internally (which is generally not possible with jagged arrays). AFAIK, your Boost code should also interact with the interpreter internally in this case. In fact An array of arrays is generally slower than a list of array because Numpy arrays are not built-ins type so CPython needs to calls Numpy functions doing many checks over and over. Additionally, Numpy does not support jagged array natively (see this related post).

Question 1: Is this the proper way of handling jagged arrays?

This is the usual way but clearly not a fast way. An efficient way to encore jagged array is to concatenate them in a big 1D array and use an additional array to store the start/stop indices (or offset/size informations). That way enable you to still use some basic Numpy vectorized methods on all arrays of sub-slices. Numba can be used to speed up the iteration over sub-arrays. The same thing applies for C++ with Boost.

Question 2: I create an object np::ndarray row to index into the second dimensions. This seems a potential performance bottleneck. Can that somehow be avoided? Can I use somehow the length of the 2nd arrays and index directly into the data buffer? I am not sure if there is padding or if there is a contiguous block of memory behind a jagged array? I assume it is not.

AFAIK, the performance bottleneck is due to the interaction with the CPython interpreter (via the CPython API) so to deal with CPython object. The above solution solve this problem since you only need to read an integer from the slicing array. The jagged array can be seen as a custom type or as a simple tuple of two arrays (possibly three if you want to split the start/stop or offsets/size). This representation is a bit similar to sparse matrices.

Note that arrays of arrays objects are indeed not stored contiguously in memory (each array object is stored are a different location in memory independently of other arrays that is dependant of the underlying CPython allocator). Arrays of arrays objects are typically stored as a pointer to an object structure containing a pointer to a memory buffer that contains pointers to objects containing each a pointer to other memory buffers. This cause a lot of pointer indirections and thus bad performance.

🌐
Frank Sauerburger
frank.sauerburger.io › 2020 › 03 › 11 › awkward-and-numba.html
Awkward arrays and numba | Frank Sauerburger
March 11, 2020 - We can construct a JITed wrapper taking these arrays as input, which then slices the content array and passes the slices to the signum defined above. We can even go one step further and package all of this in a decorator. from functools import wraps import numba import numpy as np def jagged_loop(func): """ Function decorator.
Find elsewhere
🌐
James D. McCaffrey
jamesmccaffreyblog.com › home › loading a jagged numeric matrix from text file using python
Loading a Jagged Numeric Matrix From Text File Using Python - James D. McCaffreyJames D. McCaffrey
July 8, 2019 - For example: m = np.loadtxt(".\thefile.txt", delimiter=" ", usecols=[0,1,2,3], dtype=np.int) But if the data is jagged, loadtxt() won’t work. Suppose the data in the file looks like: 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 · The ...
Top answer
1 of 1
1

You might think, your input is a list of tuples. However, it is a list of integers and tuples. (880) will be interpreted as an integer, but not as a tuple. So you have to deal with both datatypes.

First of all I suggest converting your input data to a list of lists. Each of the lists contained in that list should have the same length, because an array supports constant dimensions only. Therefore, I would convert the elements into a list and fill missing values with zeros (to make all elements equal in length).

If we do this for all of the elements given in the input list, we create a new list containing lists of equal length which can be converted into an array.

A very basic (and error-prone) approach would look like this:

import numpy as np


original_list = [
    (880),
    (880, 1080),
    (880, 1080, 1080),
    (470, 470, 470, 1250),
]


def get_len(item):
    try:
        return len(item)
    except TypeError:
        # `(880)` will be interpreted as an int instead of a tuple
        # so we need to handle tuples and integers
        # as integers do not support len(), a TypeError will be raised
        return 1


def to_list(item):
    try:
        return list(item)
    except TypeError:
        # `(880)` will be interpreted as an int instead of a tuple
        # so we need to handle tuples and integers
        # as integers do not support __iter__(), a TypeError will be raised
        return [item]


def fill_zeros(item, max_len):
    item_len = get_len(item)
    to_fill = [0] * (max_len - item_len)
    as_list = to_list(item) + to_fill
    return as_list


max_len = max([get_len(item) for item in original_list])
filled = [fill_zeros(item, max_len) for item in original_list]

arr = np.array(filled)
print(arr)

Printing:

[[ 880    0    0    0]
[ 880 1080    0    0]
[ 880 1080 1080    0]
[ 470  470  470 1250]]
🌐
Tonysyu
tonysyu.github.io › ragged-arrays.html
Ragged arrays - Tony S. Yu
I often need to save a series of arrays in which one dimension varies in length---sometimes called a ragged array [1]. For example, I'm running particle tracking experiments, and I need to save the 2D coordinates of all particles in each video frame. The number of particles in each frame will ...
🌐
GitHub
github.com › scikit-hep › awkward-0.x › blob › master › docs › classes.adoc
awkward-0.x/docs/classes.adoc at master · scikit-hep/awkward-0.x
June 21, 2022 - Jagged arrays have a logical structure that is independent of how they are represented in memory, but since Awkward Array defines this structure in terms of a basic array library (Numpy), the structure we choose is a visible part of the Awkward Array specification.
Author: scikit-hep
Top answer
1 of 6
43

Short answer: you can't. NumPy does not support jagged arrays natively.

Long answer:

>>> a = ones((3,))
>>> b = ones((2,))
>>> c = array([a, b])
>>> c
array([[ 1.  1.  1.], [ 1.  1.]], dtype=object)

gives an array that may or may not behave as you expect. E.g. it doesn't support basic methods like sum or reshape, and you should treat this much as you'd treat the ordinary Python list [a, b] (iterate over it to perform operations instead of using vectorized idioms).

Several possible workarounds exist; the easiest is to coerce a and b to a common length, perhaps using masked arrays or NaN to signal that some indices are invalid in some rows. E.g. here's b as a masked array:

>>> ma.array(np.resize(b, a.shape[0]), mask=[False, False, True])
masked_array(data = [1.0 1.0 --],
             mask = [False False  True],
       fill_value = 1e+20)

This can be stacked with a as follows:

>>> ma.vstack([a, ma.array(np.resize(b, a.shape[0]), mask=[False, False, True])])
masked_array(data =
 [[1.0 1.0 1.0]
 [1.0 1.0 --]],
             mask =
 [[False False False]
 [False False  True]],
       fill_value = 1e+20)

(For some purposes, scipy.sparse may also be interesting.)

2 of 6
7

In general, there is an ambiguity in putting together arrays of different length because alignment of data might matter. Pandas has different advanced solutions to deal with that, e.g. to merge series into dataFrames.

If you just want to populate columns starting from first element, what I usually do is build a matrix and populate columns. Of course you need to fill the empty spaces in the matrix with a null value (in this case np.nan)

a = ones((3,))
b = ones((2,))
arraylist=[a,b]

outarr=np.ones((np.max([len(ps) for ps in arraylist]),len(arraylist)))*np.nan #define empty array
for i,c in enumerate(arraylist):  #populate columns
    outarr[:len(c),i]=c

In [108]: outarr
Out[108]: 
array([[  1.,   1.],
       [  1.,   1.],
       [  1.,  nan]])
🌐
GitHub
github.com › scikit-hep › awkward-0.x
GitHub - scikit-hep/awkward-0.x: Manipulate arrays of complex data structures as easily as Numpy. · GitHub
The star attributes ("name", "ra" or right ascension in degrees, "dec" or declination in degrees, "dist" or distance in parsecs, "mass" in multiples of the sun's mass, and "radius" in multiples of the sun's radius) are plain Numpy arrays and the planet attributes ("name", "orbit" or orbital distance in AU, "eccen" or eccentricity, "period" or periodicity in days, "mass" in multiples of Jupyter's mass, and "radius" in multiples of Jupiter's radius) are jagged because each star may have a different number of planets.
Starred by 214 users
Forked by 38 users
Languages: Python 63.7% | Jupyter Notebook 36.3%
🌐
GitHub
github.com › scikit-hep › awkward-0.x › issues › 59
numpy functions on jagged arrays don't produce compatible arrays · Issue #59 · scikit-hep/awkward-0.x
December 12, 2018 - Now when computing the np.sin function on the selected subarray, the jaggedness (starts/stops structure) of the subarray changes. mu_phi_sel.starts[:10] array([ 0, 2, 4, 9, 11, 13, 16, 18, 20, 22]) np.sin(mu_phi_sel).starts[:10] array([ 0, 2, ...
Author: scikit-hep
🌐
GitHub
github.com › scikit-hep › ragged
GitHub - scikit-hep/ragged: Manipulating ragged arrays in an Array API compliant way. · GitHub
For example, this is a ragged/jagged array: >>> import ragged >>> a = ragged.array([[[1.1, 2.2, 3.3], []], [[4.4]], [], [[5.5, 6.6, 7.7, 8.8], [9.9]]]) >>> a ragged.array([ [[1.1, 2.2, 3.3], []], [[4.4]], [], [[5.5, 6.6, 7.7, 8.8], [9.9]] ]) ...
Author: scikit-hep
Top answer
1 of 2
3

Your array is 2x2:

In [298]: A
Out[298]: 
array([[array([1, 2, 3]), array([4, 5])],
       [array([6, 7, 8, 9]), array([10])]], dtype=object)

While A+A works, boolean tests have not been implemented for this kind of array:

In [299]: A>4
...
ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()

I'm going to flatten A because it makes it easier to compare with list operations:

In [301]: A1=A.flatten()

In [303]: A1+A1
Out[303]: 
array([array([2, 4, 6]), array([ 8, 10]), array([12, 14, 16, 18]),
       array([20])], dtype=object)

In [304]: [a+a for a in A1]
Out[304]: [array([2, 4, 6]), array([ 8, 10]), array([12, 14, 16, 18]), array([20])]

In [305]: timeit A1+A1
100000 loops, best of 3: 6.85 µs per loop

In [306]: timeit [a+a for a in A1]
100000 loops, best of 3: 9.09 µs per loop

The array operation is a bit faster than a list comprehension. But if I first turn the array into a list:

In [307]: A1l=A1.tolist()

In [308]: A1l
Out[308]: [array([1, 2, 3]), array([4, 5]), array([6, 7, 8, 9]), array([10])]

In [309]: timeit [a+a for a in A1l]
100000 loops, best of 3: 5.2 µs per loop

times improve. This is a good indication that the A1+A1 (or even A+A) is using a similar sort of iteration.

So the straight forward way of performing your A,B calculation is

In [310]: A2=[a[a>4] for a in A1]
In [311]: B=[a+a for a in A2]
In [312]: B
Out[312]: [array([], dtype=int32), array([10]), array([12, 14, 16, 18]), array([20])]

(we can convert to/from arrays and lists as needed).

A numpy array stores its data a flat databuffer, and uses the shape and strides attributes to quickly calculate the location of any element, regardless of the dimensions. The fast array operations use compiled code that rapidly steps though the databuffers of arguments, performing the operations element by element (or some other combination).

A dtype object array also has the flat databuffer, but the elements are pointers to lists or arrays elsewhere. So while it can index individual elements quickly, it still has to perform a Python call(s) to access the arrays. So especially when the array is 1d, it is virtually the same as a flat list with the same pointers.

Multidimensional object arrays are nicer than nested lists. You can reshape them, access elements (A[1,3] v Al[1][3]), transpose them, etc. But when it comes to iterating through all the subarrays they don't offer much of a benefit.

Looking again at your 2d array:

In [315]: timeit A+A
100000 loops, best of 3: 6.93 µs per loop  # 6.85 for A1+A1 (above)

In [316]: timeit [[j+j for j in i] for i in A]
100000 loops, best of 3: 17.1 µs per loop

In [317]: Al = A.tolist()

In [318]: timeit [[j+j for j in i] for i in Al]
100000 loops, best of 3: 7.01 µs per loop    # 5.2 for A1l flat list

Basically the same time for summing the array and iterating through the equivalent nested list.

2 of 2
0

The performance of numpy jagged array may not be optimal, but there are enough reasons to believe that it should be much better than using python nested list. As explained in your earlier post:

On principle you should have some performance bonus because every element is a numpy array. So you just need a 2 dimensional loop rather than a 3D loop (if you store every number in nested lists). Also it always saves you lots of memory allocation time to avoid using python list.

Here is a simple test:

import time,sys,random
import numpy as np
rand = np.random.rand
L = np.array([[rand(100), rand(200)],[rand(400), rand(300)]], dtype=object)
L1 = [random.random() for i in range(1000)]
arrFunc = np.vectorize(lambda x:x[x>0.3],otypes=[np.ndarray])

start = time.time()
if sys.argv[1]=='np':
  for i in range(100000):
    B=i*L
else:
  for i in range(100000):
    B=[i*x for x in L1]

end = time.time()
print ('Arithmetic Op: ', end-start)


start = time.time()
if sys.argv[1]=='np':
  for i in range(100000):
    B=arrFunc(L)
else:
  for i in range(100000):
    B=[x for x in L1 if x<0.3]
end = time.time()
print ('Indexing       ', end-start)

Result:

> python testNpJarray.py np
Arithmetic Op:  3.9719998836517334
Indexing        8.079999923706055

> python testNpJarray.py list
Arithmetic Op:  53.289000034332275
Indexing        52.10899996757507

This test may not be quite fare because the outter numpy array is quite small, you are welcome to change the size to fit into your application and tell us the results.

🌐
Annasguidetopython
annasguidetopython.com › python3 › data structures › arrays-creating-a-jagged-array
Arrays - Creating a jagged array in Python
May 12, 2023 - number_of_rows = 5 jagged_array = [[] for _ in range(number_of_rows)] In this example, we’ve created an empty jagged array with five rows.
🌐
GitHub
github.com › topics › jagged-array
jagged-array · GitHub Topics · GitHub
December 2, 2022 - python data-science data-structure data-analysis dask columnar-format ragged-array jagged-array ... A Python library for numpy arrays that persist on disk in a format that is simple, self-documented and tool-independent, and maximizes universal readability.
Top answer
1 of 2
1

You can transform your initial array of iterable-objects to ndarray by padding them with zeros in a vectorized manner:

import numpy as np

a = np.array([[1, 2, 3, 4], 
              [2, 3, 4], 
              [4, 5, 6]])
max_len = len(max(a, key = lambda x: len(x))) # max length of iterable-objects contained in array
cust_func = np.vectorize(pyfunc=lambda x: np.pad(array=x, 
                                                 pad_width=(0,max_len), 
                                                 mode='constant', 
                                                 constant_values=(0,0))[:max_len], otypes=[list])
a_pad = np.stack(cust_func(a))

output:

array([[1, 2, 3, 4],
       [2, 3, 4, 0],
       [4, 5, 6, 0]])
2 of 2
1

It depends. Do you know the size of the vectors before or are you appending to a list?

see e.g. http://stackoverflow.com/a/58085045/7919597

You could for example pad the arrays

import numpy as np

a1 = [1, 2, 3, 4]
a2 = [2, 3, 4, np.nan] # pad with nan
a3 = [4, 5, 6, np.nan] # pad with nan

b = np.stack([a1, a2, a3], axis=0)

print(b)

# you can apply the normal numpy operations on 
# arrays with nan, they usually just result in a nan
# in a resulting array
c = np.diff(b, axis=-1)

print(c)

Afterwards you can apply a moving window on each row over the columns.

Have a look at https://stackoverflow.com/a/22621523/7919597 which is only 1d, but can give you an idea of how it could work.

It is possible to use a 2d array with only one row as kernel (shape e.g. (1, 3)) with scipy.signal.convolve2d and use the idea above. This is a workaround to get a "row-wise 1D convolution":

from scipy import signal

krnl = np.array([[0, 1, 0]])

d = signal.convolve2d(c, krnl, mode='same')
print(d)