You can use metaclasses (and updated to Python3 code):
class PostInitCaller(type):
def __call__(cls, *args, **kwargs):
obj = type.__call__(cls, *args, **kwargs)
obj.__post_init__()
return obj
class BaseClass(metaclass=PostInitCaller):
def __init__(self):
print('base __init__')
self.common1()
def common1(self):
print('common 1')
def finalizeInitialization(self):
print('finalizeInitialization [common2]')
def __post_init__(self): # this is called at the end of __init__
self.finalizeInitialization()
class Subclass1(BaseClass):
def __init__(self):
super().__init__()
self.specific()
def specific(self):
print('specific')
s = Subclass1()
base __init__
common 1
specific
finalizeInitialization [common2]
Answer from Cam.Davidson.Pilon on Stack OverflowYou can use metaclasses (and updated to Python3 code):
class PostInitCaller(type):
def __call__(cls, *args, **kwargs):
obj = type.__call__(cls, *args, **kwargs)
obj.__post_init__()
return obj
class BaseClass(metaclass=PostInitCaller):
def __init__(self):
print('base __init__')
self.common1()
def common1(self):
print('common 1')
def finalizeInitialization(self):
print('finalizeInitialization [common2]')
def __post_init__(self): # this is called at the end of __init__
self.finalizeInitialization()
class Subclass1(BaseClass):
def __init__(self):
super().__init__()
self.specific()
def specific(self):
print('specific')
s = Subclass1()
base __init__
common 1
specific
finalizeInitialization [common2]
Template Method Design Pattern to the rescue:
class BaseClass:
def __init__(self, specifics=None):
print 'base __init__'
self.common1()
if specifics is not None:
specifics()
self.finalizeInitialization()
def common1(self):
print 'common 1'
def finalizeInitialization(self):
print 'finalizeInitialization [common2]'
class Subclass1(BaseClass):
def __init__(self):
BaseClass.__init__(self, self.specific)
def specific(self):
print 'specific'
When should I implement __post_init__?
Add a __post__ method, equivalent to the __new__method, but called after __init__? - Ideas - Discussions on Python.org
Default __post_init__ Implementation in Dataclasses - Ideas - Discussions on Python.org
Clean Code Writing: Dataclasses __post_init__ question
The __post_init__ method is specific to the dataclasses library, because the __init__ method on dataclass classes is generated and overriding it would entirely defeat the purpose of generating it in the first place.
SQLAlchemy, on the other hand, provides an __init__ implementation on the base model class (generated for you with declarative_base()). You can safely re-use that method yourself after setting up default values, via super().__init__(). Take into account that the SQLAlchemy-provided __init__ method only takes keyword arguments:
def __init__(self, useragent, profile):
"""specify the main information"""
id = generate_new_id(self)
super().__init__(id=id, useragent=useragent, profile=profile)
If you need to wait for the other columns to be given updated values first (because perhaps they define Python functions as a default), then you can also run functions after calling super().__init__(), and just assign to self:
def __init__(self, useragent, profile):
"""specify the main information"""
super().__init__(useragent=useragent, profile=profile)
self.id = generate_new_id(self)
Note: you do not want to use the built-in id() function to generate ids for SQL-inserted data, the values that the function returns are not guaranteed to be unique. They are only unique for the set of all active Python objects only, and only in the current process. The next time you run Python, or when objects are deleted from memory, values can and will be reused, and you can't control what values it'll generate next time, or in a different process altogether.
If you were looking to only ever create rows with unique combinations of the useragent and profile columns, then you need to define a UniqueConstraint in the table arguments. Don't try to detect uniqueness at the Python level, as you can't guarantee that another process will not make the same check at the same time. The database is in a much better position to determine if you have duplicate values, without risking race conditions:
class Worker(Base):
__tablename__='worker'
id = Column(Integer, primary_key=True, autoincrement=True)
profile = Column(String(100), nullable=False)
useragent = Column(String(100), nullable=False)
__table_args__ = (
UniqueConstraint("profile", "useragent"),
)
or you could use a composite primary key based on the two columns; primary keys (composite or otherwise) must always be unique:
class Worker(Base):
__tablename__='worker'
profile = Column(String(100), primary_key=True, nullable=False)
useragent = Column(String(100), primary_key=True, nullable=False)
I implemented the similar behavior using the __init_subclass__ method:
class Parent:
def __init_subclass__(cls, **kwargs):
def init_decorator(previous_init):
def new_init(self, *args, **kwargs):
previous_init(self, *args, **kwargs)
if type(self) == cls:
self.__post_init__()
return new_init
cls.__init__ = init_decorator(cls.__init__)
def __post_init__(self):
pass
class Child(Parent):
def __init__(self):
print('Child __init__')
def __post_init__(self):
print('Child __post_init__')
class GrandChild(Child):
def __init__(self):
print('Before calling Child __init__')
Child.__init__(self)
print('After calling Child __init__')
def __post_init__(self):
print('GrandChild __post_init__')
child = Child()
# output:
# Child __init__
# Child __post_init__
grand_child = GrandChild()
# output:
# Before calling Child __init__
# Child __init__
# After calling Child __init__
# GrandChild __post_init__
I'm having a hard time wrapping my head around when to use __post_init__ in general. I'm building some stuff using the @dataclass decorator, but I don't really see the point in __post_init__ if the init argument is already set to true, by default? Like at that point, what would the __post_init__ being doing that the __init__ hasn't already done? Like dataclass is going to do its own thing and also define its own repr as well, so I guess the same could be questionable for why define a __repr__ for a dataclass?
Maybe its just for customization purposes that both of those are optional. But at that point, what would be the point of a dataclass over a regular class. Like assume I do something like this
@dataclass(init=False, repr=False)
class Thing:
def __init__(self):
...
def __repr__(self):
...
# what else is @dataclass doing if both of these I have to implement
# ik there are more magic / dunder methods to each class,
# is it making this type 'Thing' more operable with others that share those features?I guess what I'm getting at is: What would the dataclass be doing for me that a regular class wouldn't?
Idk maybe that didn't make sense. I'm confused haha, maybe I just don't know. Maybe I'm using it wrong, that probably is the case lol. HALP!!! lol
Hello,
I have a question about the best way to initialize my instance variables for a data class in python. Some of the instance variables depend on some of the fields of the data class in python, which are inputs to a webscraping method. This means I need a __post_init__ method to retrieve the values from the webscrape. For the __post_init__ method, I would have way more than 3 variables being scraped from the website, so getting the key variable from data seems really inefficient. I know there are fields you can add to dataclasses, but I am not sure if that would help me here. Is there anyway I can simplify this? Here is my code (This is not the actual code, just the general structure of the dataclass):
from dataclasses import dataclass
from external_scrape_module import run
@dataclass
class Scrape:
path: int
criteria1: str
criteria2: str
criteria3: str
def __post_init__(self) -> None:
self.data: dict = self.scrape_website()
self.scraped_info1: str = self.data['scraped_info1']
self.scraped_info2: str = self.data['scraped_info2']
self.scraped_info3: str = self.data['scraped_info3']
def scrape_website(self) -> dict:
return run(self.path, self.criteria1, self.criteria2, self.criteria3)Much help would be appreciated, as I am fairly new to dataclasses. Thanks!
This is easiest to use the actual examples. I'm making data classes for a baseball game.
There is a parent method called player. This incorporates all the base running information (pitchers are runners too).
There are two children classes, hitter and pitcher, these incorporate the hitting and defense for the hitter and the pitching information for the pitcher.
Then there is (Shohei Ohtani) the rare two way player for this class which is a child of both hitter and pitcher.
For all of these classes there is a __post__init__ method, for the hitters and pitchers I want to run the player __post__init__ method as well as the specific hitter or pitcher one. For the two way player I want to be able to run the player, hitter and pitcher __post__init__ methods.
I think for the single class inheritance I could use super(), but I don't know how either use super to specify which or both methods to call or another way I cold call all those post__init__ methods.
If I can call a parent classes __post_init__ method from within the __post_init__ of the child that would work, I just don't know how I would do it.
If not using actual code is a problem, please let me know and I will edit my question so that there is actual code.
Indeed - that is tricky.
Having a decorator on the __init__ of base class that would freeze the instance after __init__, as you mentioned have the problem of state - it is possible to add other state variables, or count the __init__ depth, and use a metaclass (or __init_subclass__) to decorate all __init__ method in the subclasses, so that it would do the freezing only when exiting the outermost __init__.
But there is an easier way using metaclasses: the metaclass' __call__ method is what calls a class __new__ and then __init__ when creating a new instance. So, it suffices to put code to calling a __post_init__ on a custom metaclass' __call__ (or freeze the instance directly from it)
class PostInitMeta(type):
def __call__(cls, *args, **kw):
instance = super().__call__(*args, **kw) # < runs __new__ and __init__
instance.__post_init__()
return instance
class Freezing(metaclass=PostInitMeta):
_frozen = False
def __post_init__(self):
self._frozen = True
def __setattr__(self, name, value):
if self._frozen:
raise AttributeError() # or do nothing, and return as you prefer
super().__setattr__(name, value)
class A(Freezing):
def __init__(self, a):
self.a = a
class B(A):
def __init__(self, a, b):
super().__init__(a)
self.b = b
And testing this in interactive mode:
In [18]: b = B(2, 3)
In [19]: b.a
Out[19]: 2
In [20]: b.b
Out[20]: 3
In [21]: b.b = 5
---------------------------------------------------------------------------
AttributeError Traceback (most recent call last)
Cell In [21], line 1
----> 1 b.b = 5
There is a simpler solution using __slots__.
__slots__ basically freezes the instance (-specific) attributes.
Note: There are some use-cases which may prohibit the use of __slots__, check Using Slots.
class Foo:
__slots__ = ['a', 'b']
def __init__(self, a, b):
self.a = a
self.b = b
foo = Foo(1,2)
# foo.c = 3 # => attribute error - 'Foo' object has no attribute 'c'
class Bar(Foo):
__slots__ = ['c']
def __init__(self, a, b, c):
super().__init__(a, b)
self.c = 3
bar = Bar(1,2,3)
# bar.d = 4 # => attribute error - 'Bar' object has no attribute 'd'
Explanation
Without __slots__, __dict__ is used by default which is of type dict.
__dict__ contains the data stored in the program's memory for a specific object.
Since it's possible to add keys to a dict simply by doing __dict__[key] = value, this allows for adding instance attributes whenever we want (by writing foo.c = 3 which translates to foo.__dict__['c'] = 3).
By using __slots__, and using a non-dict for it, we stop that from happening.
On top of that, __slots__ also provides memory optimization.