What is a variable? A variable is a name that is bound to a non-null pointer to an object (a PyObject*, in CPython). Critically, None is also a non-null PyObject*. What is an object? An object is a PyObject whose data has been allocated on the heap. Crucially, a PyObject might exist for some time after nothing is pointing to it; it takes time for the garbage collector to step in and free the memory. What is a value? A value is the collective state of an object's members. In other words, in C, a PyObject is a struct, and the value of the object is the data inside all of the various members that it could contain. Now, let's look at your example. You are allocating memory for a new PyLongObject on the heap, and x now points to that object. Calling id(x) returns the contents of that pointer (a location in memory). Next, you allocate a new integer PyLongObject on the heap, and x now points to that object. The old object no longer has any references pointing to it, and will just sit there orphaned until the garbage collector frees it. What's the value? The value is the collective state of the members of that PyLongObject, whatever the implementation actually to define the data inside the object. In CPython, it's an array of uint32_ts.
So at this point, the idea of a variable as being fundamentally linked to a physical memory location is kind of stuck in my head. Yeah, I think you should unlink this idea. This is not a universally true statement, since different languages will have fundamentally different "concepts" or "models" of memory. This in turn means that the definition of a "variable" will change from language to language. I think this is easiest to explain by comparing and contrasting C with Python Lower-level languages like C fundamentally conceptualizes memory as a giant array of bytes split into two parts: the stack and the heap. Whenever you create a variable, you are reserving a small chunk of bytes on this "stack". You can store small quantities of data here -- things like ints, floats... If you want to store larger quantities of data, you typically do so on the heap. You ask your runtime (e.g. your operating system) permission to use a certain range of bytes (e.g. via malloc) and stick your data there. But in order to actually use this data, you now need to keep track of which specific index within this giant array your data starts at. We call these "memory addresses". And where can we store these indexes? Well, typically within one of your variables on the stack. So when we are doing a pointer dereference (e.g. *my_variable), what we are doing is: Looking at the specific chunk of bytes on the stack associated with my_variable. Reading the data located within those bytes Interpreting those bytes as an index within this giant bytes array. Grabbing the data starting at that index. Of course, everything I described above is a giant over-simplification and a lie. For example: Most modern-day operating systems don't literally represent memory as a giant array of bytes and instead split them up into discrete chunks called "pages". This enables them to do things like swap pages in and out of disk to allow programs to use more RAM then is physically present on the machine. This in turn means that the "memory address" is also a bit of an illusion. You are typically being given a "virtual memory address" which your operating system will under the hood map to a "physical memory address". This mapping will change as data is paged in and out, since where your data is literally stored on your physical RAM can change over time. Your variable does not necessarily need to correspond to something on the stack. Your compiler may opt to not bother storing any data there, and instead use your CPU's registers instead. This is done for efficiency purposes, since reading/writing to a register is faster then reading/writing to RAM. Most modern-day CPUs do a lot of clever caching whenever you read data from RAM. Instead of literally fetching data byte-for-byte from RAM, it pre-emptively fetches a whole bunch of bytes at once and stores it within a cache within your CPU. This helps speed up code that tries reading bytes in sequential order -- and can lead to unexpected slowness in code that doesn't. C's model by itself wouldn't give you any way of predicting this behavior. That said, you can mostly get away with ignoring all of the above implementation details. C's model and abstractions are consistent and self-contained and is sufficient for understanding how to write well-formed C programs. Higher-level programming languages like Python have a very different model. Python does still sort of have the concept of a stack and a heap, but that's about where the similarities with C end. In particular, unlike C Python does not have any concept of memory being a giant array nor any concept of memory having an address. Instead, it conceptualizes the stack as being literally the stack data structure (with push/pop operations), where each entry in the stack is a key-value map (e.g. a hashmap or dictionary). A variable is just a key within this dict. You can actually modify this stack frame dict directly. For example, try running the following from the Python IDE. >>> my_variable = 3 >>> stack_frame = locals() >>> stack_frame["my_variable"] = "surprise" >>> print(my_variable) surprise The locals() builtin function will return a reference to the current stack frame dict. We can then modify that dict directly to change out what our variable happens to be referring to. (Does this mean we can dynamically create new variables? Yes -- try doing stack_frame["brand_new_variable"] = 101; print(brand_new_variable)) Python's concept of a "heap" is even more abstract. In Python, the heap is just a space where objects can be stored. These objects are not stored at a particular memory address nor are in any particular order: the heap is basically one giant bag of stuff. These objects are given a unique integer id. You can imagine that's how the stack frame dict is keeping track of each object under the hood: a stack frame is a mapping of strings (variable names) to object ids, and the heap is a mapping of object ids to the underlying object. This is the mechanism by which a Python variable refers to an object. But I prefer to conceptualize this more visually. To me, a variable in Python is an arrow with a name that points to some object floating in the aether. Thinking about what's happening in terms of these "object ids" is honestly a bit too cumbersome to be useful in practice. And what exactly is an object? Well, either: A glorified wrapper around a dict, for user-defined classes Or special "primitive" object that the Python interpreter creates on your behalf Just like with C, this model is a gross oversimplification of what's actually happening under the hood. And like C, this model is internally consistent and self-contained, so you mostly don't need to worry about the implementation details. Granted, Python's model does break down a bit more frequently then C's does. For example, you do need to mentally consider how everything maps to the underlying physical memory if you care about performance. You also need to start caring about C's model if you want to write Python wrappers around C code. I know a Python variable's ID (obtained via the id function) IS just the memory address (that's what the CPython documentation says at least) This is just an implementation detail. It's true today, but CPython is under no obligation to continue using the memory address as the unique ID in the future. Similarly, other implementations of Python may choose to use something completely different. (That said, it's unlikely in practice that CPython will stop using the memory address as the unique id any time soon. Switching to something else would probably make it harder to implement the model we discussed above.) My friend said that's because the id doesn't really belong to the variable itself, but to the object. Yup, this is absolutely correct. So is the first ID number that would be printed by the above code by then a reference to the ID of the object 5? No: what's printed is literally the id of the object 5. If we want to be more precise, here's what's happening when we do print(id(x)): Python evaluates the expression x. To do this, it looks at the current stack frame dict, looks up the key "x", then grabs the object id -- a reference to an object representing the integer 5. Let's say for the sake of argument that this object id is 3017463234928. The x variable is evaluated to this object reference. The id(...) function accepts this object reference -- the 3017463234928 object id. It does some Magic™ and creates a new int object that represents the number 3017463234928. (Why does it bother creating an object? Well, because Python's model states that every value the user interacts with must be represented as an object). Anyways, the 3017463234928 int object is stored on the heap. Let's say this new int object has an id of 4999584263006. The id(x) expression evaluates to this object reference. The print(...) function accepts the 4999584263006 object id. It uses that to look up the underlying object, reads it, and prints out 3017463234928 to stdout. Visually, I guess this ends up looking sort of like this: +--------------------------+ | object id: 3017463234928 | x -------------> | type: int | | value: 5 | +--------------------------+ +--------------------------+ | object id: 4999584263006 | id(x) ---------> | type: int | | value: 3017463234928 | +--------------------------+ It's a bit weird to have id(x) point to something -- it's not a variable, after all. But hopefully you get the idea. But then, if 5 is an object, that what's the value? This is a bit easier to explain using custom objects instead of Python's builtin ones. Consider the following: >>> class Url: ... def __init__(self, domain, path): ... self.domain = domain ... self.path = path ... >>> >>> a = Url("reddit.com", "/r/learnprogramming") >>> b = Url("reddit.com", "/r/learnprogramming") >>> c = Url("reddit.com", "/r/aww") This is creating three separate Url objects: two representing reddit.com/r/learnprogramming and one representing reddit.com/r/aww. But how many distinct "entities" or "values" do we have? Arguably, just two. The a and b variables may be pointing to two distinct objects, but what they fundamentally represent is the same. So, we say those two objects have the same "value": the same high-level "meaning" or "data". But then, if 5 is an object, that what's the value? The "value" is the integer 5. Python will represent this integer using an int object.