I'm restricting this answer, perhaps unnecessarily, to IEEE754 floating point.
DBL_MIN is not allowed to be a subnormal number.
But std::nextafter is allowed to return a subnormal number.
Hence the return value of the latter could be less than DBL_MIN.
For more details see https://en.wikipedia.org/wiki/Denormal_number
Answer from Bathsheba on Stack OverflowI'm restricting this answer, perhaps unnecessarily, to IEEE754 floating point.
DBL_MIN is not allowed to be a subnormal number.
But std::nextafter is allowed to return a subnormal number.
Hence the return value of the latter could be less than DBL_MIN.
For more details see https://en.wikipedia.org/wiki/Denormal_number
Is
DBL_MINthe smallest positive double?
Not certainly.
DBL_MIN is the smallest positive normal double.
DBL_TRUE_MIN is the smallest positive double (since C++17). It will be smaller than DBL_MIN when double supports subnormals.
-DBL_MAX in ANSI C, which is defined in float.h.
Floating point numbers (IEEE 754) are symmetrical, so if you can represent the greatest value (DBL_MAX or numeric_limits<double>::max()), just prepend a minus sign.
And then is the cool way:
double f;
(*((uint64_t*)&f))= ~(1LL<<52);
Double.NEGATIVE_INFINITY is not a finite value, so it can't be the smallest finite value of type double.
Double.MIN_VALUE is also not the smallest finite value of type double, since that's the
smallest positive nonzero value of type double
The smallest finite value of type double is -Double.MAX_VALUE, whose value is -1.79769313348623157E308, or -(2-2-52)*21023 or Double.longBitsToDouble(0xffefffffffffffffL).
This can be calculated based on JLS 4.2.3. Floating-Point Types, Formats, and Values:
The finite nonzero values of any floating-point value set can all be expressed in the form s ⋅ m ⋅ 2(e - N + 1), where s is +1 or -1, m is a positive integer less than 2N, and e is an integer between Emin = -(2K-1-2) and Emax = 2K-1-1, inclusive.
For double, Emax==1023 and N==53.
When you assign in s ⋅ m ⋅ 2(e - N + 1) the values s=-1, m=2N-1 and e = Emax, you get the minimum finite value: -1 * (253-1)*2(1023 - 53 + 1) = - (2 - 2-52) * 21023.
The finite range of double is 4.9E-324 to 1.7e+038. Double.NEGATIVE_INFINITY is just a contant the is defined by the ieee floating point specification
public static final double NEGATIVE_INFINITY
A constant holding the negative infinity of type double. It is equal to the value returned by Double.longBitsToDouble(0xfff0000000000000L)
The IEEE 754 format has one bit reserved for the sign and the remaining bits representing the magnitude. This means that it is "symmetrical" around origo (as opposed to the Integer values, which have one more negative value). Thus the minimum value is simply the same as the maximum value, with the sign-bit flipped, so yes, -Double.MAX_VALUE is the lowest actual number you can represent with a double.
I suppose the Double.MAX_VALUE should be seen as maximum magnitude, in which case it actually makes sense to simply write -Double.MAX_VALUE. It also explains why Double.MIN_VALUE is the least positive value (since that represents the least possible magnitude).
But sure, I agree that the naming is a bit misleading. Being used to the meaning Integer.MIN_VALUE, I too was a bit surprised when I read that Double.MIN_VALUE was the smallest absolute value that could be represented. Perhaps they thought it was superfluous to have a constant representing the least possible value as it is simply a - away from MAX_VALUE :-)
(Note, there is also Double.NEGATIVE_INFINITY but I'm disregarding from this, as it is to be seen as a "special case" and does not in fact represent any actual number.)
Here is a good text on the subject.
These constants have nothing to do with sign. This makes more sense if you consider a double as a composite of three parts: Sign, Exponent and Mantissa. Double.MIN_VALUE is actually the smallest value Mantissa can assume when the Exponent is at minimun value before a flush to zero occurs. Likewise MAX_VALUE can be understood as the largest value Mantissa can assume when the Exponent is at maximum value before a flush to infinity occurs.
A more descriptive name for these two could be Largest Absolute (add non-zero for verbositiy) and Smallest Absolute value (add non-infinity for verbositiy).
Check out the IEEE 754 (1985) standard for details. There is a revised (2008) version, but that only introduces more formats which aren't even supported by java (strictly speaking java even lacks support for some mandatory features of IEEE 754 1985, like many other high level languages).
Okay, let's imagine it this way, sticking with decimal numbers. Suppose you have a floating decimal point type which allows you to represent 5 decimal digits, and a number between 0 and 3 for the exponent, to multiple the result by 1, 10, 100 or 1000.
So the smallest non-zero value is just 1 (i.e. mantissa=00001, exponent=0). The largest value is 99999000 (mantissa=99999, exponent=3).
Now, what happens when you add 1 to 50000000? You can't represent 50000001...the next representable number after 500000000 is 50001000. So if you try to add them together, the result is just going to be the closest value to the "true" result - which is still 500000000. That's like adding Double.MIN_VALUE to a large double.
My version (converting to bits, incrementing and then converting back) is like taking that 50000000, splitting into mantissa and exponent (m=50000, e=3) then incrementing it the smallest amount, to (m=50001, e=3) and then reassembling to 50001000.
Do you see how they're different?
Now here's a concrete example:
public class Test{
public static void main(String[] args) {
double before = 100000000000000d;
double after = before + Double.MIN_VALUE;
System.out.println(before == after);
long bits = Double.doubleToLongBits(before);
bits++;
double afterBits = Double.longBitsToDouble(bits);
System.out.println(before == afterBits);
System.out.println(afterBits - before);
}
}
This tries both approaches with a large number. The output is:
true
false
0.015625
Going through the output, that means:
- Adding
Double.MIN_VALUEdidn't have any effect - Incrementing the bit did have an effect
- The difference between
afterBitsandbeforeis 0.015625, which is much bigger thanDouble.MIN_VALUE. No wonder the simple addition had no effect!
It's exactly as Jon said:
"If you try to add a very little number to a very big number, the difference may well be so small that the closest result is the same as the original."
For example:
// True:
(Double.MAX_VALUE + Double.MIN_VALUE) == Double.MAX_VALUE
// False:
Double.longBitsToDouble(Double.doubleToLongBits(Double.MAX_VALUE) + 1) == Double.MAX_VALUE)
MIN_VALUE is the smallest representable positive double, but that certainly does not imply that adding it to an arbitrary double results in a unequal one.
In contrast, adding 1 to the underlying bits results in a new bit pattern, and thus does result in a unequal double.
The documentation for sys.float_info says that this is the way to get the smallest possible float:
>>> math.ulp(0.0)
5e-324
The value of sys.float_info.min (part of a previous answer that I deleted) gives the smallest normalized float, a much bigger value.
Another way to approach this problem is to look into the details of how floats work, e.g. on Wikipedia. Acknowledging that python uses IEEE double precision floats on most platforms, and slogging through the relevant bits of https://en.wikipedia.org/wiki/Double-precision_floating-point_format#Double-precision_examples:
+2−1022 × 2−52 = 2−1074 ≈ 4.9406564584124654 × 10−324 (Min. subnormal positive double)
If your numbers are finite, you can use a couple of convenient methods in the BitConverter class:
long bits = BitConverter.DoubleToInt64Bits(value);
if (value > 0)
return BitConverter.Int64BitsToDouble(bits - 1);
else if (value < 0)
return BitConverter.Int64BitsToDouble(bits + 1);
else
return -double.Epsilon;
IEEE-754 formats were designed so that the bits that make up the exponent and mantissa together form an integer that has the same ordering as the floating-point numbers. So, to get the largest smaller number, you can subtract one from this number if the value is positive, and you can add one if the value is negative.
The key reason why this works is that the leading bit of the mantissa is not stored. If your mantissa is all zeros, then your number is a power of two. If you subtract 1 from the exponent/mantissa combination, you get all ones and you'll have to borrow from the exponent bits. In other words: you have to decrement the exponent, which is exactly what we want.
In .NET Core 3.0 you can use Math.BitIncrement/Math.BitDecrement. No need to do manual bit manipulation anymore
Returns the smallest value that compares greater than a specified value.
Returns the largest value that compares less than a specified value.
Since .NET 7.0 there are also Double.BitIncrement/Double.BitDecrement and the generic versions in IFloatingPointIeee754<TSelf>: BitDecrement/BitIncrement
.Machine$double.xmin gives the value of the smallest positive number whose representation meets the requirements of IEEE 754 technical standard for floating point computation. As is mentioned in the Wikipedia article on double-precision floating point numbers, that standard requires that:
If a decimal string with at most 15 significant digits is converted to IEEE 754 double precision representation and then converted back to a string with the same number of significant digits, then the final string should match the original. If an IEEE 754 double precision is converted to a decimal string with at least 17 significant digits and then converted back to double, then the final number must match the original.
The same article goes on to note that, by compromising precision, even smaller positive numbers (which do not meet the standards' precision requirements) can be represented:
The 11 bit width of the exponent allows the representation of numbers between 10-308 and 10308, with full 15–17 decimal digits precision. By compromising precision, the subnormal representation allows even smaller values up to about 5 × 10-324.
R's doubles behave in exactly this way, as is noted in the Details section of ?.Machine:
Note that on most platforms smaller positive values than ‘.Machine$double.xmin’ can occur. On a typical R platform the smallest positive double is about ‘5e-324’.
To confirm that that is the smallest positive value that can be represented using R's doubles and to see the cost in loss of precision, try out a few operations like this:
5e-324
# [1] 4.940656e-324
2e-324
# [1] 0
1.4 * 5e-324
# [1] 4.940656e-324
1.6 * 5e-324
# [1] 9.881313e-324
Here are some representations using SAS, IEEE 754 Big Endian?
data _null_;
y=constant('big');
put y hex16.;
put y E21.3;
run;quit;
Biggest
7FEFFFFFFFFFFFFF 1.79769313486230E+308
data _null_;
y=constant('small');
put y hex16.;
put y E21.3;
run;quit;
Smallest
0010000000000000 2.22507385850720E-308
I am not sure the smallest because SAS may set aside some values for missings.
I doubt there is a standard library function for finding the number you need, however, it is not too hard to find the needed value using a simple binary search:
#include <iostream>
#include <cmath>
#include <typeinfo>
template<typename T> T find_magic_epsilon(T from, T to) {
if (to == std::nextafter(from, to)) return to;
T mid = (from + to)/2;
if (std::isinf((T)1.0/mid)) return find_magic_epsilon<T>(mid, to);
else return find_magic_epsilon<T>(from, mid);
}
template<typename T> T meps() {
return find_magic_epsilon<T>(0.0, 0.1);
}
template<typename T> T test_meps() {
T value = meps<T>();
std::cout << typeid(T).name() << ": MEPS=" << value
<< " 1/MEPS=" << (T)1.0/value << " 1/(MEPS--)="
<< (T)1.0/std::nextafter(value,(T)0.0) << std::endl;
}
int main() {
test_meps<float>();
test_meps<double>();
test_meps<long double>();
return 0;
}
The output of the script above:
f: MEPS=2.93874e-39 1/MEPS=3.40282e+38 1/(MEPS--)=inf
d: MEPS=5.56268e-309 1/MEPS=1.79769e+308 1/(MEPS--)=inf
e: MEPS=8.40526e-4933 1/MEPS=1.18973e+4932 1/(MEPS--)=inf
Here is a starting point to answer your question. I have simply divided by 2. Once you get to dividing by 2^^148, you come pretty close. You could then iteratively move closer and print out the hex representation of the number to see what the compiler is doing:
#include <iostream>
#include <stdio.h>
#include <limits>//is it here?
int main() {
double d_eps,invd_eps;
float f_eps,invf_eps;
invf_eps = 1.f/f_eps;//invf_eps should be finite
float last_seed = 0;
float seed = 1.0;
for(int i = 0; i < 1000000; i++) {
last_seed = seed;
seed = seed/2;
if(seed/2 == invd_eps) {
printf("Breaking at i = %d\n", i);
printf("Seed: %g, last seed: %g\n", seed, last_seed);
break;
}
}
printf("%f, %lf, %f, %lf\n\n", f_eps, d_eps, invf_eps, invd_eps);
return 0;
}
Output:
Breaking at i = 148
Seed: 1.4013e-45, last seed: 2.8026e-45
0.000000, 0.000000, inf, 0.000000
Process finished with exit code 0
By definition, for floating types, min returns the smallest positive value the type can encode, not the lowest.
If you want the lowest value, use numeric_limits::lowest instead.
Documentation: http://en.cppreference.com/w/cpp/types/numeric_limits/min
As for why it is this way, I can only speculate that the Standard committee needed to have a way to represent all forms of extreme values for all different native types. In the case of integral types, there's only two types of extreme: max positive and max negative. For floats there is another: smallest possible.
If you think the semantics are a bit muddled, I agree. The semantics of the related #defines in the C standard are muddled in much the same way.
It's unfortunate, but behind similar names completely different meaning lies. It was kinda carried over from C, where DBL_MIN and INT_MIN has the very same "problem".
As not much can be done, just remember what means what.
I'm guessing at what you want from your code: You want to start with largest possible negative value and increment it toward positive infinity in the smallest possible steps until the value is no longer negative.
I think the function you want is nextafter().
int main() {
double value(-std::numeric_limits<double>::max());
while(value < 0) {
std::cout << value << '\n';
value = std::nextafter(value,std::numeric_limits<double>::infinity());
}
}
Firstly,
while(i < 0); // <--- remove this semicolon
{
cout << i << endl;
i += min ;
}
Then, std::numeric_limits<double>::min() is a positive value, so i < 0 will never be true. If you need the most negative value, you'll need
double min = -std::numeric_limits<double>::max();
but I don't know what your i += min line is supposed to do. Adding two most negative number will just yield −∞, and the loop will never finish. If you want to add a number, you'll need another variable, like
double most_negative = -std::numeric_limits<double>::max();
double most_positive = std::numeric_limits<double>::max();
double i = most_negative;
while (i < 0)
{
std::cout << i << std::endl;
i += most_positive;
}
Of course this will just print the most negative number (-1.8e+308), and then i becomes 0 and the loop will exit.
Have you looked at double.PositiveInfinity? (aka System.Double.PositiveInfinity). The are similar values for negative infinity, NaN, epsilon (the smallest positive double) and min/max values.
It's a constant in the double struct. If that doesn't do what you want it to, please clarify.
Note that to test for infinity, you can use double.IsPositiveInfinity (and IsNegativeInfinity, and just IsInfinity).
Use the Double.PositiveInfinity constant. Example:
double huge = Double.PositiveInfinity;
Take a look at http://www.cplusplus.com/reference/clibrary/cfloat/ the FLT_EPSILON can be used to compare a float variable to 0 (zero), or nearly zero in that case.
There's an other article here: http://www.cygnus-software.com/papers/comparingfloats/Comparing%20floating%20point%20numbers.htm
EDIT: what you need is this q&a: https://stackoverflow.com/questions/9380108/get-machine-epsilon-in-microsoft-excel it describes the epsilon value in excel.
I think you're probably remembering FLT_EPSILON, DBL_EPSILON, and so on.
They aren't quite what you're describing though: they're not the smallest number greater than 0. Rather, they're the smallest number greater than 1. For better or worse, however, DBL_EPSILON-1 won't be even close to the smallest number greater than 0. Epsilon is really intended to be scaled by multiplication, not addition or subtraction.
This is not normally used to avoid problems with division by zero, but to give some idea of the smallest possible difference between two numbers, and whether two numbers are "substantially different", or any inequality between them is really due to rounding errors (and such).
I'd also note that replacing x/y with x/(y+EPSILON) could quite easily cause a division by zero rather than preventing it (i.e., if y == -EPSILON).
Think about it in these terms. Take a 2-bit number with a preceding sign:
000 = 0
001 = 1
010 = 2
011 = 3
Now let's have some negatives:
111 = -1
110 = -2
101 = -3
Wait, we also have
100 ...
It has to be negative, because the sign-bit is 1. So, logically, it must be -4.
(Edit: As WorldEngineer rightly points out, not all numbering systems work this way -- but the ones you're asking about do.)
Because there are not two classes of numbers in the integer range, but three: negative numbers, zero, and positive numbers. Zero has to take up a slot (would be rather impractical not to be able to represent zero...), so either the positive or the negative class has to give up a slot. The fact that it's usually the positive range that has to make that sacrifice is to a certain extent arbitrary, but on the level of bit manipulations there are some things that this decision makes more convenient.
Actual zero is zero. The result can become zero in different ways. A double has an value range of +/-10^+/-308 (roughly). A number smaller than the smallest number will be considered zero. Using #include <limits>, you can get numeric_limits<double>::denorm_min(), which is the smallest value that can be represented in a double.
But you can get "the effect of zero" in other ways. Say you have a fairly large number, 10 million, and you add (or subtract - read add as add or subtract in the rest of this paragraph) a very small number, say 1/10 million, then the addition will have no effect, because it is outside the actual value bits of the mantissa of the floating point number - that is, 53 bits in the case of double - then the effect will be the same as adding zero. In other words, even if you have a number that is not zero, using it to add to another number is not always going to change the other number.
See IEEE-754 on Wikipedia (other floating point formats do exist, but they are unusual).
You could try:
#include <limits>
std::numeric_limits<double>::denorm_min();
Doc for denormal (aka subnormal) numbers (here).
If this number is divided by e.g. by 2 the result is 0.
To check this values on a specific platform the following code can be used:
#include <iostream>
#include <limits>
using std::cout;
using std::endl;
int main() {
typedef double real;
union dbl {
real d;
unsigned char c[sizeof(d)];
dbl(const dbl &n = 0.0) : d(n.d) {}
dbl(double n) : d(n) {}
void pr(const char *txt = 0) const {
if (txt) cout << txt << ": ";
cout << d << ":";
for (int i = sizeof(d) -1; i >= 0; --i)
cout << std::hex << " " << (int)c[i];
cout << endl;
}
};
dbl n = 1.0;
for (; n.d > 0.0; n.d /= 2.0)
n.pr();
n.pr("zero");
n.d = std::numeric_limits<real>::min();
n.pr("min");
n.d = std::numeric_limits<real>::denorm_min();
n.pr("denorm_min");
}
Output on 32 bit linux (intel cpu) (doc about double format):
1: 3f f0 0 0 0 0 0 0
0.5: 3f e0 0 0 0 0 0 0
0.25: 3f d0 0 0 0 0 0 0
0.125: 3f c0 0 0 0 0 0 0
0.0625: 3f b0 0 0 0 0 0 0
...
8.9003e-308: 0 30 0 0 0 0 0 0
4.45015e-308: 0 20 0 0 0 0 0 0
2.22507e-308: 0 10 0 0 0 0 0 0
1.11254e-308: 0 8 0 0 0 0 0 0
5.56268e-309: 0 4 0 0 0 0 0 0
...
7.90505e-323: 0 0 0 0 0 0 0 10
3.95253e-323: 0 0 0 0 0 0 0 8
1.97626e-323: 0 0 0 0 0 0 0 4
9.88131e-324: 0 0 0 0 0 0 0 2
4.94066e-324: 0 0 0 0 0 0 0 1
zero: 0: 0 0 0 0 0 0 0 0
min: 2.22507e-308: 0 10 0 0 0 0 0 0
denorm_min: 4.94066e-324: 0 0 0 0 0 0 0 1
If real is defined as long double the output is:
1: 0 0 3f ff 80 0 0 0 0 0 0 0
0.5: 0 0 3f fe 80 0 0 0 0 0 0 0
0.25: 0 0 3f fd 80 0 0 0 0 0 0 0
0.125: 0 0 3f fc 80 0 0 0 0 0 0 0
0.0625: 0 0 3f fb 80 0 0 0 0 0 0 0
...
5.83232e-4950: 0 0 0 0 0 0 0 0 0 0 0 10
2.91616e-4950: 0 0 0 0 0 0 0 0 0 0 0 8
1.45808e-4950: 0 0 0 0 0 0 0 0 0 0 0 4
7.2904e-4951: 0 0 0 0 0 0 0 0 0 0 0 2
3.6452e-4951: 0 0 0 0 0 0 0 0 0 0 0 1
zero: 0: 0 0 0 0 0 0 0 0 0 0 0 0
min: 3.3621e-4932: 0 0 0 1 80 0 0 0 0 0 0 0
denorm_min: 3.6452e-4951: 0 0 0 0 0 0 0 0 0 0 0 1
Or for float:
1: 3f 80 0 0
0.5: 3f 0 0 0
0.25: 3e 80 0 0
0.125: 3e 0 0 0
0.0625: 3d 80 0 0
...
2.24208e-44: 0 0 0 10
1.12104e-44: 0 0 0 8
5.60519e-45: 0 0 0 4
2.8026e-45: 0 0 0 2
1.4013e-45: 0 0 0 1
zero: 0: 0 0 0 0
min: 1.17549e-38: 0 80 0 0
denorm_min: 1.4013e-45: 0 0 0 1
According to the javadoc for Double.MIN_VALUE, MIN_VALUE is:
A constant holding the smallest positive nonzero value of type double
So Double.MIN_VALUE is not negative, it's the positive value that's as close as a Double can get to zero (without being zero).
Double.MIN_VALUE is the smallest positive non-zero value which can be represented by a Java double (see the JavaDoc at http://download.oracle.com/javase/8/docs/api/java/lang/Double.html).