The following works (numba 0.37):

@nb.njit
def desired_fn(thing):
    thing.blah[:] += np.arange(len(thing))
    # or
    # thing['blah'][:] += np.arange(len(thing))

If you are operating primarily on columns of your data instead of rows, you might consider using a different data container. A numpy structured array is laid out like a vector of structs rather than a struct of arrays. This means that when you want to update blah, you are moving through non-contiguous memory space as you traverse the array.

Also, with any code optimizations, it's aways worth it to use timeit or some other timing harness (that removes the time required to jit the code) to see what is the actual performance. You might find with numba that explicit looping while more verbose could actually be faster than your vectorized code.

Answer from JoshAdel on Stack Overflow
🌐
Numba
numba.readthedocs.io › en › stable › reference › numpysupported.html
Supported NumPy features — Numba 0+untagged.1117.g3190b91.dirty documentation
... Because of the way Numba logic ... be aware of this when using these features. Numba presently supports accessing fields of individual elements in structured arrays by attribute as well as by getting and setting....
Top answer
1 of 2
2

The following works (numba 0.37):

@nb.njit
def desired_fn(thing):
    thing.blah[:] += np.arange(len(thing))
    # or
    # thing['blah'][:] += np.arange(len(thing))

If you are operating primarily on columns of your data instead of rows, you might consider using a different data container. A numpy structured array is laid out like a vector of structs rather than a struct of arrays. This means that when you want to update blah, you are moving through non-contiguous memory space as you traverse the array.

Also, with any code optimizations, it's aways worth it to use timeit or some other timing harness (that removes the time required to jit the code) to see what is the actual performance. You might find with numba that explicit looping while more verbose could actually be faster than your vectorized code.

2 of 2
1

Without numba, accessing field values is no slower than accessing columns of a 2d array:

In [1]: arr2 = np.zeros((10000), dtype='i,i')
In [2]: arr2.dtype
Out[2]: dtype([('f0', '<i4'), ('f1', '<i4')])

Modifying a field:

In [4]: %%timeit x = arr2.copy()
   ...: x['f0'] += 1
   ...: 
16.2 µs ± 13.7 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)

Similar time if I assign the field to a new variable:

In [5]: %%timeit x = arr2.copy()['f0']
   ...: x += 1
   ...: 
15.2 µs ± 14.2 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)

Much faster if I construct a 1d array of the same size:

In [6]: %%timeit x = np.zeros(arr2.shape, int)
   ...: x += 1
   ...: 
8.01 µs ± 15.1 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)

But similar time when accessing the column of a 2d array:

In [7]: %%timeit x = np.zeros((arr2.shape[0],2), int)
   ...: x[:,0] += 1
   ...: 
17.3 µs ± 23.8 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)
🌐
GitHub
github.com › numba › numba › issues › 5753
Declaring Numpy structured array dtypes throws TypeError · Issue #5753 · numba/numba
May 26, 2020 - @njit def my_test(): struct_dtype = np.dtype([ ('one', np.float64), ('two', np.float64)]) struct_array = np.array([ (1.0, 1.1), (2.0, 2.1) ], dtype=struct_dtype) return struct_array result = my_test() print(result) A TypeError is thrown: TypeError: unknown dtype descriptor: list(Tuple(unicode_type, class(float64))) ... I am using the latest released version of Numba (most recent is visible in the change log (https://github.com/numba/numba/blob/master/CHANGE_LOG).
Author: numba
🌐
GitHub
github.com › numba › numba › issues › 878
Numba semantics for structured arrays differs to Numpy semantics for structured arrays · Issue #878 · numba/numba
December 1, 2014 - Numba allows access to fields of individual elements in structured arrays by attribute, whereas Numpy does not. Consider the following as an example: from numba import jit, numpy_support import num...
Author: numba
🌐
Numba Discussion
numba.discourse.group › support: what is this error message?
Structured arrays with nd-fields - Support
February 3, 2021 - Hi 🙂 Not sure if I am misunderstanding the docs, or if I am doing something wrong, so bear with me please. Scalar types Numba supports the following Numpy scalar types: [...] Structured scalars: structured scalars made of any of the types above and arrays of the types above The following scalar types and features are not supported: [...] Nested structured scalars the fields of structured scalars may not contain other structured scalars (from Supported NumPy features — Numba ...
🌐
Numba Discussion
numba.discourse.group › support: how do i do ...?
Numba dictionary(list) and numpy structured arrays? - Support: How do I do ...? - Numba Discussion
October 19, 2023 - Does numba dictionaries recognize numpy structured arrays? Here is a simple example, pointDtype = np.dtype([('x', 'f4'),('y', 'f4'),('z', 'f4')]) pointNBtype = nb.from_dtype(pointDtype) p0 = np.array((0,0,0), dtype= pointNBtype) dct = nb.typed.Dict.empty(nb.i8, pointNBtype) Up to here all goes okay.
🌐
Numba
numba.pydata.org › numba-doc › dev › reference › types.html
Types and signatures — Numba 0.52.0.dev0+274.g626b40e-py3.7-linux-x86_64.egg documentation
Instead of using typeof(), non-trivial scalars such as structured types can also be constructed programmatically. ... >>> struct_dtype = np.dtype([('row', np.float64), ('col', np.float64)]) >>> ty = numba.from_dtype(struct_dtype) >>> ty Record([('row', '<f8'), ('col', '<f8')]) >>> ty[:, :] unaligned array(Record([('row', '<f8'), ('col', '<f8')]), 2d, A) ... Create a Numba type for Numpy datetimes of the given unit.
🌐
Numba Discussion
numba.discourse.group › support: what is this error message?
Accessing structured array scalars by index - Numba Discussion
January 20, 2021 - This code works in plain-python mode, but borks in jit mode. I don’t believe this is the same thing as another similar post but I wouldn’t be surprised if there’s some relationship under the covers. Numpy docs indicate this should be legal, and Numba supported features don’t seem to mention it. from numba import njit import numpy as np arr = np.array([(1, 2., 3.)], dtype='i, f, f') # @njit def get_elem(arr, idx, member_idx): return arr[idx][member_idx] print(get_elem(arr, 0, 1)) return...
🌐
Stack Overflow
stackoverflow.com › questions › 69713623 › how-to-create-and-fill-structured-array-with-numba
numpy ndarray - How to create and fill structured array with numba? - Stack Overflow
Python 3.9.7, Numba 0.54.0 I wrote this simple code, that works well: import numpy as np import numba from collections import namedtuple Deal = namedtuple('Deal' , ['DateTime', 'Price', 'Quantity'])
Find elsewhere
🌐
Numba
numba.pydata.org › numba-doc › dev › reference › numpysupported.html
Supported NumPy features — Numba 0.52.0.dev0+274.g626b40e-py3.7-linux-x86_64.egg documentation
For numeric dtypes, Numba follows Numpy’s behavior. The real attribute returns a view of the real part of the complex array and it behaves as an identity function for other numeric dtypes. The imag attribute returns a view of the imaginary part of the complex array and it returns a zero array with the same shape and dtype for other numeric dtypes. For non-numeric dtypes, including all structured/record dtypes, using these attributes will result in a compile-time (TypingError) error.
🌐
GitHub
github.com › numba › numba › issues › 6473
Array assignment to a structure's nested array · Issue #6473 · numba/numba
November 12, 2020 - Impossible to assign an array to the nested array in the structured array. Without numba all the options below are working, but with numba only the last option works. The first option falls with an error "Can only insert i8 at [0] in [1024 x i8]: got i8*". The second option falls with "Buffer dtype cannot be buffer, have dtype: nestedarray(uint8, (1024,))" The third option falls with "Buffer dtype cannot be buffer, have dtype: nestedarray(uint8, (1024,))" import numba as nb import numpy as np BLOCK = np.dtype([("id", np.uint64), ("props", (np.uint8, 1024))]) @nb.njit() def set_props1(_source,
Author: numba
🌐
GitHub
github.com › numba › numba › issues › 5798
Initilizing a Numpy structured array inside a jitclass constructor throws an error · Issue #5798 · numba/numba
June 2, 2020 - This works: from numba.experimental import jitclass from numba import from_dtype import numpy as np values_dtype = np.dtype([ ('one', 'U10'), ('two', 'f8') ]) class My_test: def __init__(self): self.values = values = np.empty(2, dtype=va...
Author: numba
🌐
Stack Overflow
stackoverflow.com › questions › 52409479 › python-numba-accessing-structured-numpy-array-elements-as-fast-as-possible
Python & Numba: Accessing structured numpy array elements as fast as possible - Stack Overflow
September 19, 2018 - What you get with my_array['field1'] and my_array[0]. As for the last, I suspect numba has no indication that my_array is a structured array or that the [v] indexing step requires something different from an ordinary 2d column access.
🌐
Numba
numba.pydata.org › numba-doc › 0.24.0 › reference › numpysupported.html
2.6. Supported NumPy features — Numba 0.24.0-py2.7-macosx-10.5-x86_64.egg documentation
Structured scalars support attribute getting and setting, as well as member lookup using constant strings. ... Numpy scalars reference. Numpy arrays of any of the scalar types above are supported, regardless of the shape or layout.
🌐
Numba
numba.pydata.org › numba-doc › 0.17.0 › reference › types.html
2.1. Numba Types — Numba 0.17.0-py2.7-linux-x86_64.egg documentation
Non-trivial scalars such as structured types need to be constructed programmatically. ... >>> struct_dtype = np.dtype([('row', np.float64), ('col', np.float64)]) >>> numba.from_dtype(struct_dtype) Record([('row', '<f8'), ('col', '<f8')]) ... Create a Numba type for Numpy datetimes of the given unit.
🌐
Numba
numba.pydata.org › numba-doc › 0.17.0 › reference › numpysupported.html
2.5. Supported Numpy features — Numba 0.17.0-py2.7-linux-x86_64.egg documentation
Structured scalars support attribute getting and setting. ... Numpy scalars reference. Arrays of any of the scalar types above are supported, regardless of the shape or layout.
Top answer
1 of 1
1

Consider the following functions.

from numba import njit


@njit
def func():
    i = 777
    t = i
    return t


@njit
def func2():
    i = 776
    t = i + 1
    return t

You can check how each variable's type is inferred using the following method.

func()
func.inspect_types()

This is the key lines:

    #   i = const(int, 777)  :: Literalint
    #   t = i  :: Literalint

The part after :: is the type of the variable. This indicates that both i and t are of integer literal type.

Next, for func2:

func2()
func2.inspect_types()
    #   i = const(int, 776)  :: Literalint
    #   t = i + $const10.2  :: int64

Compared to func, you can see that t is inferred as int64 rather than an integer literal type. This means, numba performs type inference on the code before optimization.

This is a reasonable choice. Typed code is required for optimization, but type inference is required to generate typed code. So first type inference is performed on Python bytecode, and then optimization is performed based on the inferred types. For more accurate and detailed information on this flow, please refer to the official documentation.

In summary, you need a constant variable at the Python bytecode phase.


As an additional note, numba does not support indexing records with non-literal variables. However, it is somehow possible by explicitly defining the mapping via overloading.

from operator import setitem

import numpy as np
from numba import njit, types
from numba.core.extending import overload

a_dtype = np.dtype([("id", "i4"), ("qtrnm0", "S4"), ("qtr0", "f4")])


@overload(setitem)
def setitem_overload_for_a(a, index, value):
    if getattr(a, "dtype", None) != a_dtype:
        return None

    if isinstance(value, (types.Integer, types.Float)):
        def numeric_impl(a, index, value):
            # You need to map these indexes correctly according to the dtype.
            if index == 0:
                a[0] = value
            elif index == 2:
                a[2] = value
            else:
                raise ValueError()

        return numeric_impl
    elif isinstance(value, (types.Bytes, types.CharSeq)):
        def bytes_impl(a, index, value):
            if index == 1:
                a[1] = value
            else:
                raise ValueError()

        return bytes_impl
    else:
        raise TypeError(f"Unsupported type: {index=}, {value=}, {a.dtype=}")


@njit
def upsert_numba(a, sid, qtrnm, val):
    i = 0
    a[i + 1] = qtrnm
    a[i + 2] = val
    return a


x = (1, b"24q2", 3.0)
a = np.array([(1, b"24q1", 1.0)], dtype=a_dtype)
print(upsert_numba(a[0].copy(), *x))  # (1, b'24q2', 3.)

Note that this is an ad hoc strategy that requires you to hardcode setitem for each record type, and may not work in some cases. That said, it should work unless you're doing something very tricky.

Author: numba
🌐
Stack Overflow
stackoverflow.com › questions › 60118008 › accessing-structured-data-types-in-numba-vs-numpy
python - Accessing structured data types in numba vs. numpy - Stack Overflow
I think recarray is a something relic from the past, currently coded as a thin subclass layer on top of the more basic structured array. numba is newer, and appears to implement that record attribute functionality without adding the subclassing layer. ... The main area of development in numpy structured arrays is in the handling of multi-field indexing.
🌐
Numba Discussion
numba.discourse.group › support: how do i do ...?
Structured array field index to field name - Support: How do I do ...? - Numba Discussion
January 8, 2021 - Is there a better way to do this? I’m trying to convert a index into the fields of a structured array into the corresponding field name. My workaround is to explicitly pull the field names into a separate tuple, but I was hoping there was something more elegant. from numba import njit, from_dtype import numpy as np x = np.array([('Rex', 9, 81.0), ('Fido', 3, 27.0)], dtype=[('name', 'U10'), ('age', 'i4'), ('weight', 'f4')]) numba_type = from_dtype(x.dtype) print("x.dtype.names", x.dtype.names...