I'm restricting this answer, perhaps unnecessarily, to IEEE754 floating point.

DBL_MIN is not allowed to be a subnormal number.

But std::nextafter is allowed to return a subnormal number.

Hence the return value of the latter could be less than DBL_MIN.

For more details see https://en.wikipedia.org/wiki/Denormal_number

Answer from Bathsheba on Stack Overflow
Top answer
1 of 6
226

The IEEE 754 format has one bit reserved for the sign and the remaining bits representing the magnitude. This means that it is "symmetrical" around origo (as opposed to the Integer values, which have one more negative value). Thus the minimum value is simply the same as the maximum value, with the sign-bit flipped, so yes, -Double.MAX_VALUE is the lowest actual number you can represent with a double.

I suppose the Double.MAX_VALUE should be seen as maximum magnitude, in which case it actually makes sense to simply write -Double.MAX_VALUE. It also explains why Double.MIN_VALUE is the least positive value (since that represents the least possible magnitude).

But sure, I agree that the naming is a bit misleading. Being used to the meaning Integer.MIN_VALUE, I too was a bit surprised when I read that Double.MIN_VALUE was the smallest absolute value that could be represented. Perhaps they thought it was superfluous to have a constant representing the least possible value as it is simply a - away from MAX_VALUE :-)

(Note, there is also Double.NEGATIVE_INFINITY but I'm disregarding from this, as it is to be seen as a "special case" and does not in fact represent any actual number.)

Here is a good text on the subject.

2 of 6
14

These constants have nothing to do with sign. This makes more sense if you consider a double as a composite of three parts: Sign, Exponent and Mantissa. Double.MIN_VALUE is actually the smallest value Mantissa can assume when the Exponent is at minimun value before a flush to zero occurs. Likewise MAX_VALUE can be understood as the largest value Mantissa can assume when the Exponent is at maximum value before a flush to infinity occurs.

A more descriptive name for these two could be Largest Absolute (add non-zero for verbositiy) and Smallest Absolute value (add non-infinity for verbositiy).

Check out the IEEE 754 (1985) standard for details. There is a revised (2008) version, but that only introduces more formats which aren't even supported by java (strictly speaking java even lacks support for some mandatory features of IEEE 754 1985, like many other high level languages).

Top answer
1 of 2
8

Okay, let's imagine it this way, sticking with decimal numbers. Suppose you have a floating decimal point type which allows you to represent 5 decimal digits, and a number between 0 and 3 for the exponent, to multiple the result by 1, 10, 100 or 1000.

So the smallest non-zero value is just 1 (i.e. mantissa=00001, exponent=0). The largest value is 99999000 (mantissa=99999, exponent=3).

Now, what happens when you add 1 to 50000000? You can't represent 50000001...the next representable number after 500000000 is 50001000. So if you try to add them together, the result is just going to be the closest value to the "true" result - which is still 500000000. That's like adding Double.MIN_VALUE to a large double.

My version (converting to bits, incrementing and then converting back) is like taking that 50000000, splitting into mantissa and exponent (m=50000, e=3) then incrementing it the smallest amount, to (m=50001, e=3) and then reassembling to 50001000.

Do you see how they're different?


Now here's a concrete example:

public class Test{
    public static void main(String[] args) {
        double before = 100000000000000d;
        double after = before + Double.MIN_VALUE;
        System.out.println(before == after);

        long bits = Double.doubleToLongBits(before);
        bits++;
        double afterBits = Double.longBitsToDouble(bits);
        System.out.println(before == afterBits);
        System.out.println(afterBits - before);
    }
}

This tries both approaches with a large number. The output is:

true
false
0.015625

Going through the output, that means:

  • Adding Double.MIN_VALUE didn't have any effect
  • Incrementing the bit did have an effect
  • The difference between afterBits and before is 0.015625, which is much bigger than Double.MIN_VALUE. No wonder the simple addition had no effect!
2 of 2
3

It's exactly as Jon said:

"If you try to add a very little number to a very big number, the difference may well be so small that the closest result is the same as the original."

For example:

// True:
(Double.MAX_VALUE + Double.MIN_VALUE) == Double.MAX_VALUE
// False:
Double.longBitsToDouble(Double.doubleToLongBits(Double.MAX_VALUE) + 1) == Double.MAX_VALUE)

MIN_VALUE is the smallest representable positive double, but that certainly does not imply that adding it to an arbitrary double results in a unequal one.

In contrast, adding 1 to the underlying bits results in a new bit pattern, and thus does result in a unequal double.

Top answer
1 of 2
23

.Machine$double.xmin gives the value of the smallest positive number whose representation meets the requirements of IEEE 754 technical standard for floating point computation. As is mentioned in the Wikipedia article on double-precision floating point numbers, that standard requires that:

If a decimal string with at most 15 significant digits is converted to IEEE 754 double precision representation and then converted back to a string with the same number of significant digits, then the final string should match the original. If an IEEE 754 double precision is converted to a decimal string with at least 17 significant digits and then converted back to double, then the final number must match the original.

The same article goes on to note that, by compromising precision, even smaller positive numbers (which do not meet the standards' precision requirements) can be represented:

The 11 bit width of the exponent allows the representation of numbers between 10-308 and 10308, with full 15–17 decimal digits precision. By compromising precision, the subnormal representation allows even smaller values up to about 5 × 10-324.

R's doubles behave in exactly this way, as is noted in the Details section of ?.Machine:

Note that on most platforms smaller positive values than ‘.Machine$double.xmin’ can occur. On a typical R platform the smallest positive double is about ‘5e-324’.

To confirm that that is the smallest positive value that can be represented using R's doubles and to see the cost in loss of precision, try out a few operations like this:

5e-324
# [1] 4.940656e-324
2e-324
# [1] 0
1.4 * 5e-324
# [1] 4.940656e-324
1.6 * 5e-324
# [1] 9.881313e-324
2 of 2
1

Here are some representations using SAS, IEEE 754 Big Endian?

data _null_;
    y=constant('big');
    put y hex16.;
    put y E21.3;
run;quit;

Biggest

7FEFFFFFFFFFFFFF 1.79769313486230E+308

data _null_;
    y=constant('small');
    put y hex16.;
    put y E21.3;
run;quit;

Smallest

0010000000000000 2.22507385850720E-308

I am not sure the smallest because SAS may set aside some values for missings.

Find elsewhere
Top answer
1 of 2
3

I doubt there is a standard library function for finding the number you need, however, it is not too hard to find the needed value using a simple binary search:

#include <iostream>
#include <cmath>
#include <typeinfo>

template<typename T> T find_magic_epsilon(T from, T to) {
    if (to == std::nextafter(from, to)) return to;
    T mid = (from + to)/2;
    if (std::isinf((T)1.0/mid)) return find_magic_epsilon<T>(mid, to);
    else return find_magic_epsilon<T>(from, mid);
}

template<typename T> T meps() {
    return find_magic_epsilon<T>(0.0, 0.1);
}

template<typename T> T test_meps() {
    T value = meps<T>();
    std::cout << typeid(T).name() << ": MEPS=" << value
              << " 1/MEPS=" << (T)1.0/value << " 1/(MEPS--)="
              << (T)1.0/std::nextafter(value,(T)0.0) << std::endl;
}

int main() {
    test_meps<float>();
    test_meps<double>();
    test_meps<long double>();
    return 0;
}

The output of the script above:

f: MEPS=2.93874e-39 1/MEPS=3.40282e+38 1/(MEPS--)=inf
d: MEPS=5.56268e-309 1/MEPS=1.79769e+308 1/(MEPS--)=inf
e: MEPS=8.40526e-4933 1/MEPS=1.18973e+4932 1/(MEPS--)=inf
2 of 2
2

Here is a starting point to answer your question. I have simply divided by 2. Once you get to dividing by 2^^148, you come pretty close. You could then iteratively move closer and print out the hex representation of the number to see what the compiler is doing:

#include <iostream>
#include <stdio.h>


#include <limits>//is it here?

int main() {
    double d_eps,invd_eps;
    float f_eps,invf_eps;
    invf_eps = 1.f/f_eps;//invf_eps should be finite
    float last_seed = 0;
    float seed = 1.0;
    for(int i = 0; i < 1000000; i++) {
        last_seed = seed;
        seed = seed/2;
        if(seed/2 == invd_eps) {
            printf("Breaking at i = %d\n", i);
            printf("Seed:  %g, last seed: %g\n", seed, last_seed);
            break;
        }
    }

    printf("%f, %lf, %f, %lf\n\n", f_eps, d_eps, invf_eps, invd_eps);
    return 0;
}

Output:

Breaking at i = 148
Seed:  1.4013e-45, last seed: 2.8026e-45
0.000000, 0.000000, inf, 0.000000


Process finished with exit code 0
🌐
Xona
xona.com › 2006 › 07 › 26.html
Xona Games - Smallest Positive Floating Point Values
The exponent = -126..127 for 32-bit float, and -1022..1023 for 64-bit double. The mantissa (also known as the significand) is >= 1.0 and < 2.0. Thus the smallest possible positive float and double are such that the exponent and mantissa contain their smallest values:
Top answer
1 of 4
4

Actual zero is zero. The result can become zero in different ways. A double has an value range of +/-10^+/-308 (roughly). A number smaller than the smallest number will be considered zero. Using #include <limits>, you can get numeric_limits<double>::denorm_min(), which is the smallest value that can be represented in a double.

But you can get "the effect of zero" in other ways. Say you have a fairly large number, 10 million, and you add (or subtract - read add as add or subtract in the rest of this paragraph) a very small number, say 1/10 million, then the addition will have no effect, because it is outside the actual value bits of the mantissa of the floating point number - that is, 53 bits in the case of double - then the effect will be the same as adding zero. In other words, even if you have a number that is not zero, using it to add to another number is not always going to change the other number.

See IEEE-754 on Wikipedia (other floating point formats do exist, but they are unusual).

2 of 4
2

You could try:

#include <limits>
std::numeric_limits<double>::denorm_min();

Doc for denormal (aka subnormal) numbers (here).

If this number is divided by e.g. by 2 the result is 0.

To check this values on a specific platform the following code can be used:

#include <iostream>
#include <limits>
using std::cout;
using std::endl;

int main() {
    typedef double real;
    union dbl {
        real d;
        unsigned char c[sizeof(d)];

        dbl(const dbl &n = 0.0) : d(n.d) {}
        dbl(double n) : d(n) {}

        void pr(const char *txt = 0) const {
            if (txt) cout << txt << ": ";
            cout << d << ":";
            for (int i = sizeof(d) -1; i >= 0; --i)
                cout << std::hex << " " << (int)c[i];
            cout << endl;
        }
    };

    dbl n = 1.0;
    for (; n.d > 0.0; n.d /= 2.0)
        n.pr();
    n.pr("zero");
    n.d = std::numeric_limits<real>::min();
    n.pr("min");
    n.d = std::numeric_limits<real>::denorm_min();
    n.pr("denorm_min");
}

Output on 32 bit linux (intel cpu) (doc about double format):

1: 3f f0 0 0 0 0 0 0
0.5: 3f e0 0 0 0 0 0 0
0.25: 3f d0 0 0 0 0 0 0
0.125: 3f c0 0 0 0 0 0 0
0.0625: 3f b0 0 0 0 0 0 0
...
8.9003e-308: 0 30 0 0 0 0 0 0
4.45015e-308: 0 20 0 0 0 0 0 0
2.22507e-308: 0 10 0 0 0 0 0 0
1.11254e-308: 0 8 0 0 0 0 0 0
5.56268e-309: 0 4 0 0 0 0 0 0
...
7.90505e-323: 0 0 0 0 0 0 0 10
3.95253e-323: 0 0 0 0 0 0 0 8
1.97626e-323: 0 0 0 0 0 0 0 4
9.88131e-324: 0 0 0 0 0 0 0 2
4.94066e-324: 0 0 0 0 0 0 0 1
zero: 0: 0 0 0 0 0 0 0 0
min: 2.22507e-308: 0 10 0 0 0 0 0 0
denorm_min: 4.94066e-324: 0 0 0 0 0 0 0 1

If real is defined as long double the output is:

1: 0 0 3f ff 80 0 0 0 0 0 0 0
0.5: 0 0 3f fe 80 0 0 0 0 0 0 0
0.25: 0 0 3f fd 80 0 0 0 0 0 0 0
0.125: 0 0 3f fc 80 0 0 0 0 0 0 0
0.0625: 0 0 3f fb 80 0 0 0 0 0 0 0
...
5.83232e-4950: 0 0 0 0 0 0 0 0 0 0 0 10
2.91616e-4950: 0 0 0 0 0 0 0 0 0 0 0 8
1.45808e-4950: 0 0 0 0 0 0 0 0 0 0 0 4
7.2904e-4951: 0 0 0 0 0 0 0 0 0 0 0 2
3.6452e-4951: 0 0 0 0 0 0 0 0 0 0 0 1
zero: 0: 0 0 0 0 0 0 0 0 0 0 0 0
min: 3.3621e-4932: 0 0 0 1 80 0 0 0 0 0 0 0
denorm_min: 3.6452e-4951: 0 0 0 0 0 0 0 0 0 0 0 1

Or for float:

1: 3f 80 0 0
0.5: 3f 0 0 0
0.25: 3e 80 0 0
0.125: 3e 0 0 0
0.0625: 3d 80 0 0
...
2.24208e-44: 0 0 0 10
1.12104e-44: 0 0 0 8
5.60519e-45: 0 0 0 4
2.8026e-45: 0 0 0 2
1.4013e-45: 0 0 0 1
zero: 0: 0 0 0 0
min: 1.17549e-38: 0 80 0 0
denorm_min: 1.4013e-45: 0 0 0 1
🌐
Baeldung
baeldung.com › home › java › core java › overflow and underflow in java
Overflow and Underflow in Java | Baeldung
January 8, 2024 - The minimum exponent for the binary representation of a double is given as -1074. That means the smallest positive value a double can have is Math.pow(2, -1074), which is equal to 4.9e-324.
🌐
Cplusplus
cplusplus.com › forum › general › 53760
smallest double value greater zero - C++ Forum
October 30, 2011 - Floating point representation is ... exponent we could have was 10^-100 and we had 5 digit representation. The the smallest positive number is 1.0000x10^-100, the next number above this is 1.0001x10^-100 but the next smallest one would be 0....