The compiler is allowed to evaluate float expressions in any bigger precision it likes, so it looks like the first expression is evaluated in long double precision. In the second expression you enforce scaling the result down to float again.

In answer to some of your additional questions and the discussion below: you are basically looking for the smallest non-zero difference with 1 of some floating point type. Depending on the setting of FLT_EVAL_METHOD a compiler may decide to evaluate all floating point expressions in a higher precision than the types involved. On a Pentium traditionally the internal registers of the floating point unit are 80 bits and it is convenient to use that precision for all the smaller floating point types. So in the end your test depends on the precision of your compare !=. In the absence of an explicit cast the precision of this comparison is determined by your compiler not by your code. With the explicit cast you scale the comparison down to the type you desire.

As you confirmed your compiler has set FLT_EVAL_METHOD to 2 so it uses the highest precision for any floating point calculation.

As a conclusion to the discussion below we are confident to say that there is a bug relating to implementation of the FLT_EVAL_METHOD=2 case in gcc prior to version 4.5 and that is fixed from of at least version 4.6. If the integer constant 2 is used in the expression instead of the floating point constant 2.0, the cast to float is omitted in the generated assembly. It is also worth noticing that from of optimization level -O1 the right results are produced on these older compilers, but the generated assembly is quite different and contains only few floating point operations.

Answer from Bryan Olivier on Stack Overflow
๐ŸŒ
GNU
gnu.org โ€บ software โ€บ c-intro-and-ref โ€บ manual โ€บ html_node โ€บ Machine-Epsilon.html
Machine Epsilon (GNU C Language Manual)
static const float epsf = 0x1p-23; /* about 1.192e-07 */ static const double eps = 0x1p-52; /* about 2.220e-16 */ static const long double epsl = 0x1p-63; /* about 1.084e-19 */ Instead of the hexadecimal constants, we could also have used the Standard C macros, FLT_EPSILON, DBL_EPSILON, and LDBL_EPSILON.
๐ŸŒ
cppreference.com
en.cppreference.com โ€บ cpp โ€บ types โ€บ numeric_limits โ€บ epsilon
std::numeric_limits<T>::epsilon - cppreference.com
October 23, 2023 - Taking the bigger // one is also reasonable, I guess. const T m = std::min(std::fabs(x), std::fabs(y)); // Subnormal numbers have fixed exponent, which is `min_exponent - 1`. const int exp = m < std::numeric_limits<T>::min() ? std::numeric_limits<T>::min_exponent - 1 : std::ilogb(m); // We consider `x` and `y` equal if the difference between them is // within `n` ULPs. return std::fabs(x - y) <= n * std::ldexp(std::numeric_limits<T>::epsilon(), exp); } int main() { double x = 0.3; double y = 0.1 + 0.2; std::cout << std::hexfloat; std::cout << "x = " << x << '\n'; std::cout << "y = " << y << '\n'; std::cout << (x == y ?
Discussions

c - Floating point arithmetic and machine epsilon - Stack Overflow
I'm trying to compute an approximation of the epsilon value for the float type (and I know it's already in the standard library). The epsilon values on this machine are (printed with some approximation): FLT_EPSILON = 1.192093e-07 DBL_EPSILON = 2.220446e-16 LDBL_EPSILON = 1.084202e-19 ยท FLT_EVAL_METHOD is 2 so everything is done in long double ... More on stackoverflow.com
๐ŸŒ stackoverflow.com
c++ - Compare double to zero using epsilon - Stack Overflow
Today, I was looking through some C++ code (written by somebody else) and found this section: double someValue = ... if (someValue ::epsilon() && More on stackoverflow.com
๐ŸŒ stackoverflow.com
When double.Epsilon can equal 0
Having tracked down similar bugs before (the joys of trying to do deterministic floating point calculations), I immediately suspected someone messing with the FPU control word (turns out that was pretty close). Other fun stuff I've read about (can't find the articles anymore though): Printer driver messing with the control word (good look finding this bug when it only occurs for customers with a certain printer!) Plugins/COM components changing it (sometimes just because they've used a certain compiler) DirectX setting float precision More on reddit.com
๐ŸŒ r/programming
42
282
September 18, 2020
machine epsilon - C++ Forum
Hi all, I need high precision for my numerical calculations. As a test, I compute machine precision ("machine epsilon" (?)) for "float", "double" and "long double". (For thus purpose, I apply the C-program as you can find at the bottom this page). More on cplusplus.com
๐ŸŒ cplusplus.com
๐ŸŒ
"The" Book of C
thebookofc.com โ€บ home โ€บ floating point โ€บ comparing floats using epsilon
Comparing Floats Using Epsilon - "The" Book of C
April 15, 2019 - C standard guarantees minimum epsilon for float to be 1E-5, i.e. 1 X 10-5 or 0.00001. Hence any value that is greater than 0.99999 and less than 1.00001 is effectively 1.0. Epsilon value for double and long double is defined to be at least 1E-9.

The compiler is allowed to evaluate float expressions in any bigger precision it likes, so it looks like the first expression is evaluated in long double precision. In the second expression you enforce scaling the result down to float again.

In answer to some of your additional questions and the discussion below: you are basically looking for the smallest non-zero difference with 1 of some floating point type. Depending on the setting of FLT_EVAL_METHOD a compiler may decide to evaluate all floating point expressions in a higher precision than the types involved. On a Pentium traditionally the internal registers of the floating point unit are 80 bits and it is convenient to use that precision for all the smaller floating point types. So in the end your test depends on the precision of your compare !=. In the absence of an explicit cast the precision of this comparison is determined by your compiler not by your code. With the explicit cast you scale the comparison down to the type you desire.

As you confirmed your compiler has set FLT_EVAL_METHOD to 2 so it uses the highest precision for any floating point calculation.

As a conclusion to the discussion below we are confident to say that there is a bug relating to implementation of the FLT_EVAL_METHOD=2 case in gcc prior to version 4.5 and that is fixed from of at least version 4.6. If the integer constant 2 is used in the expression instead of the floating point constant 2.0, the cast to float is omitted in the generated assembly. It is also worth noticing that from of optimization level -O1 the right results are produced on these older compilers, but the generated assembly is quite different and contains only few floating point operations.

Answer from Bryan Olivier on Stack Overflow
Top answer
1 of 2
7

The compiler is allowed to evaluate float expressions in any bigger precision it likes, so it looks like the first expression is evaluated in long double precision. In the second expression you enforce scaling the result down to float again.

In answer to some of your additional questions and the discussion below: you are basically looking for the smallest non-zero difference with 1 of some floating point type. Depending on the setting of FLT_EVAL_METHOD a compiler may decide to evaluate all floating point expressions in a higher precision than the types involved. On a Pentium traditionally the internal registers of the floating point unit are 80 bits and it is convenient to use that precision for all the smaller floating point types. So in the end your test depends on the precision of your compare !=. In the absence of an explicit cast the precision of this comparison is determined by your compiler not by your code. With the explicit cast you scale the comparison down to the type you desire.

As you confirmed your compiler has set FLT_EVAL_METHOD to 2 so it uses the highest precision for any floating point calculation.

As a conclusion to the discussion below we are confident to say that there is a bug relating to implementation of the FLT_EVAL_METHOD=2 case in gcc prior to version 4.5 and that is fixed from of at least version 4.6. If the integer constant 2 is used in the expression instead of the floating point constant 2.0, the cast to float is omitted in the generated assembly. It is also worth noticing that from of optimization level -O1 the right results are produced on these older compilers, but the generated assembly is quite different and contains only few floating point operations.

2 of 2
3

A C99 C compiler can evaluate floating-point expressions as if they were of a more precise floating-point type than their actual type.

The macro FLT_EVAL_METHOD is set by the compiler to indicate the strategy:

-1 indeterminable;

0 evaluate all operations and constants just to the range and precision of the type;

1 evaluate operations and constants of type float and double to the range and precision of the double type, evaluate long double operations and constants to the range and precision of the long double type;

2 evaluate all operations and constants to the range and precision of the long double type.

For historical reasons, two common choices when targeting the x86 processors are 0 and 2.

File m.c is your first program. If I compile it, using my compiler, thus, I obtain:

$ gcc -std=c99 -mfpmath=387 m.c
$ ./a.out 
float eps = 1.084202e-19
 ./a.out 
float eps = 1.192093e-07

If I compile this other program below, the compiler sets the macro according to what it does:

#include <stdio.h>
#include <float.h>

int main(){
  printf("%d\n", FLT_EVAL_METHOD);
}

Results:

$ gcc -std=c99 -mfpmath=387 t.c
 gcc -std=c99 t.c
$ ./a.out 
0
๐ŸŒ
Wikipedia
en.wikipedia.org โ€บ wiki โ€บ Machine_epsilon
Machine epsilon - Wikipedia
1 month ago - The following examples compute interval machine epsilon in the sense of the spacing of the floating point numbers at 1 rather than in the sense of the unit roundoff. Note that results depend on the particular floating-point format used, such as float, double, long double, or similar as supported ...
๐ŸŒ
GeeksforGeeks
geeksforgeeks.org โ€บ c++ โ€บ cpp-program-to-find-machine-epsilon
C++ program to find Machine Epsilon - GeeksforGeeks
January 23, 2018 - Mathematically, for each floating ... In C, machine epsilon is specified in the standard header with the names FLT_EPSILON, DBL_EPSILON, and LDBL_EPSILON....
Top answer
1 of 12
206

Assuming 64-bit IEEE double, there is a 52-bit mantissa and 11-bit exponent. Let's break it to bits:

1.0000 00000000 00000000 00000000 00000000 00000000 00000000 ร— 2^0 = 1

The smallest representable number greater than 1:

1.0000 00000000 00000000 00000000 00000000 00000000 00000001 ร— 2^0 = 1 + 2^-52

Therefore:

epsilon = (1 + 2^-52) - 1 = 2^-52

Are there any numbers between 0 and epsilon? Plenty... E.g. the minimal positive representable (normal) number is:

1.0000 00000000 00000000 00000000 00000000 00000000 00000000 ร— 2^-1022 = 2^-1022

In fact there are (1022 - 52 + 1)ร—2^52 = 4372995238176751616 numbers between 0 and epsilon, which is 47% of all the positive representable numbers...

2 of 12
20

The test certainly is not the same as someValue == 0. The whole idea of floating-point numbers is that they store an exponent and a significand. They therefore represent a value with a certain number of binary significant figures of precision (53 in the case of an IEEE double). The representable values are much more densely packed near 0 than they are near 1.

To use a more familiar decimal system, suppose you store a decimal value "to 4 significant figures" with exponent. Then the next representable value greater than 1 is 1.001 * 10^0, and epsilon is 1.000 * 10^-3. But 1.000 * 10^-4 is also representable, assuming that the exponent can store -4. You can take my word for it that an IEEE double can store exponents less than the exponent of epsilon.

You can't tell from this code alone whether it makes sense or not to use epsilon specifically as the bound, you need to look at the context. It may be that epsilon is a reasonable estimate of the error in the calculation that produced someValue, and it may be that it isn't.

Find elsewhere
๐ŸŒ
Reddit
reddit.com โ€บ r/programming โ€บ when double.epsilon can equal 0
r/programming on Reddit: When double.Epsilon can equal 0
September 18, 2020 - My fuzzy number of choice is sqrt(numeric_limits<T>::epsilon()). Unlike the one quoted in the article, the C float/double epsilon isn't denormalized (it's the smallest number e such that 1+e is representable), and using sqrt basically means ...
๐ŸŒ
Hacker News
news.ycombinator.com โ€บ item
That's why you don't compare floating point numbers with the = operator. You use... | Hacker News
December 12, 2012 - You use something like: ยท fabs( a-b ) < epsilon where epsilon is usually defined by the language as a very small quantity (example values in C are 1E-8 for a float, 1E-15 for double), or you can just use a hard coded value appropriate to the values you are comparing
๐ŸŒ
GeeksforGeeks
geeksforgeeks.org โ€บ c++ โ€บ compare-float-and-double-while-accounting-for-precision-loss
How to Compare Float and Double While Accounting for Precision Loss? - GeeksforGeeks
January 29, 2024 - There are two methods to compare two float or double values while accounting for precision loss. ... It is advised to use an epsilon, or tiny number, to provide a tolerance threshold for comparisons when comparing float and double values.
๐ŸŒ
Cplusplus
cplusplus.com โ€บ forum โ€บ general โ€บ 39951
machine epsilon - C++ Forum
#include <stdio.h> int main( int argc, char **argv ) { //float machEps = 1.0f; double machEps = 1.0f; printf( "current Epsilon, 1 + current Epsilon\n" ); do { printf( "%G\t%.20f\n", machEps, (1.0f + machEps) ); machEps /= 2.0f; // If next epsilon yields 1, then break, because current // epsilon is the machine epsilon.
Top answer
1 of 1
3

C provides a function for this, in the <math.h> header. nextafterf(x, INFINITY) is the next representable value after x, in the direction toward INFINITY.

However, if you'd prefer to do it yourself:

The following returns the epsilon you seek, for single precision (float), assuming IEEE 754. See notes at the bottom about using library routines.

#include <float.h>
#include <math.h>


/*  Return the ULP of q.

    This was inspired by Algorithm 3.5 in Siegfried M. Rump, Takeshi Ogita, and
    Shin'ichi Oishi, "Accurate Floating-Point Summation", _Technical Report
    05.12_, Faculty for Information and Communication Sciences, Hamburg
    University of Technology, November 13, 2005.
*/
float ULP(float q)
{
    // SmallestPositive is the smallest positive floating-point number.
    static const float SmallestPositive = FLT_EPSILON * FLT_MIN;

    /*  Scale is .75 ULP, so multiplying it by any significand in [1, 2) yields
        something in [.75 ULP, 1.5 ULP) (even with rounding).
    */
    static const float Scale = 0.75f * FLT_EPSILON;

    q = fabsf(q);

    /*  In fmaf(q, -Scale, q), we subtract q*Scale from q, and q*Scale is
        something more than .5 ULP but less than 1.5 ULP.  That must produce q
        - 1 ULP.  Then we subtract that from q, so we get 1 ULP.

        The significand 1 is of particular interest.  We subtract .75 ULP from
        q, which is midway between the greatest two floating-point numbers less
        than q.  Since we round to even, the lesser one is selected, which is
        less than q by 1 ULP of q, although 2 ULP of itself.
    */
    return fmaxf(SmallestPositive, q - fmaf(q, -Scale, q));
}

The following returns the next value representable in float after the value it is passed (treating โˆ’0 and +0 as the same).

#include <float.h>
#include <math.h>


/*  Return the next floating-point value after the finite value q.

    This was inspired by Algorithm 3.5 in Siegfried M. Rump, Takeshi Ogita, and
    Shin'ichi Oishi, "Accurate Floating-Point Summation", _Technical Report
    05.12_, Faculty for Information and Communication Sciences, Hamburg
    University of Technology, November 13, 2005.
*/
float NextAfterf(float q)
{
    /*  Scale is .625 ULP, so multiplying it by any significand in [1, 2)
        yields something in [.625 ULP, 1.25 ULP].
    */
    static const float Scale = 0.625f * FLT_EPSILON;

    /*  Either of the following may be used, according to preference and
        performance characteristics.  In either case, use a fused multiply-add
        (fmaf) to add to q a number that is in [.625 ULP, 1.25 ULP].  When this
        is rounded to the floating-point format, it must produce the next
        number after q.
    */
#if 0
    // SmallestPositive is the smallest positive floating-point number.
    static const float SmallestPositive = FLT_EPSILON * FLT_MIN;

    if (fabsf(q) < 2*FLT_MIN)
        return q + SmallestPositive;

    return fmaf(fabsf(q), Scale, q);
#else
    return fmaf(fmaxf(fabsf(q), FLT_MIN), Scale, q);
#endif
}

Library routines are used, but fmaxf (maximum of its arguments) and fabsf (absolute value) are easily replaced. fmaf should compile to a hardware instruction on architectures with fused multiply-add. Failing that, fmaf(a, b, c) in this use can be replaced by (double) a * b + c. (IEEE-754 binary64 has sufficient range and precision to replaced fmaf. Other choices for double might not.)

Another alternative to the fused-multiply add would be to add some tests for cases where q * Scale would be subnormal and handle those separately. For other cases, the multiplication and addition can be performed separately with ordinary * and + operators.

๐ŸŒ
Bdach
bdach.github.io โ€บ debugging โ€บ 2020 โ€บ 09 โ€บ 18 โ€บ when-double-epsilon-can-equal-zero.html
When double.Epsilon can equal 0 - Lunar Transmissions
September 18, 2020 - Therefore, the largest denormal ...-----mantissa/significand----------------| Because double.Epsilon is essentially a (uint64_t)0x1, it definitely is a denormal number....
๐ŸŒ
G. Samaras
gsamaras.wordpress.com โ€บ code โ€บ difference-between-two-doubles-is-wrong-c
Difference between two doubles is wrong (C++) | G. Samaras
June 24, 2013 - In order to understand why this code provides not the result the programmer expected, we should take a look at this code. #include <iostream> #include <iomanip> using namespace std; int main() { double epsilon=0.001; double d1=2.334; double d2=2.335; cout<<"epsilon is: "<< setprecision(20) << epsilon<<endl; cout<<"d2-d1 is: "<< setprecision(20) << d2-d1 <<endl; if ((d2-d1)==epsilon) cout<<"Equal!"; else cout<<"Not equal!"; }
๐ŸŒ
Bnikolic
bnikolic.co.uk โ€บ blog โ€บ cpp-rr-epsilon
RR: Small numerical values in C++ - Bojan Nikolic
This is the value if a small perturbation to a number is necessary, e.g., f(x+std::numeric_limits<double>::epsilon())-f(x)
๐ŸŒ
Reddit
reddit.com โ€บ r/c_programming โ€บ comparing doubles in c
r/C_Programming on Reddit: Comparing doubles in C
July 18, 2016 -

Is it OK to compare doubles in C? Like a == b. Will it give me correct results everytime? If not, workarounds?

What happens when doubles overflow? Does the C state machine go into an undefined state and then your program behavior is undefined/non deterministic? Or does something else happen? How does one prevent this from happening?

๐ŸŒ
CodeRivers
coderivers.org โ€บ c โ€บ c-basic โ€บ c-double
Understanding C double: Fundamental Concepts, Usage, and Best Practices - CodeRivers
January 19, 2025 - In this example, the areApproximatelyEqual function uses the fabs function from the <math.h> library to calculate the absolute difference between two double values and checks if it is less than a small tolerance value (EPSILON).
๐ŸŒ
Reddit
reddit.com โ€บ r/cpp_questions โ€บ how do i compare floats/doubles?
r/cpp_questions on Reddit: How do I compare floats/doubles?
June 16, 2021 -

I thought the below is how I am suppose to compare float values. A) How am I suppose to compare doubles? B) What's the point of DBL_EPSILON?

#include <math.h>
#include <float.h>
#include <stdio.h>

int main(int argc, char *argv[])
{
    int value = fabs((30.1L/5.0L) - 6.02L) <= DBL_EPSILON;
    printf("Is same? %d\n", value);
    return value;
}