The limits for floating point types are defined in float.h not limits.h
Answer from Clifford on Stack OverflowJust started learning C. so forigve me if this is a dumb question.
For int and long, the maximum number we can assign them is 2^32 and 2^64 respectively. how about for float and double? I know that double has higher precision of about 15 decimal points while float is 7.
i also know that float is 32 bit while double is 64 bit. So does that mean the highest we can store them (2^32)and (2^64) but with decimal points in between their numbers? i aslo learnt that these data type cant have unsigned value so does that value doubles?
Alright. Using what I learned from here (thanks everyone) and the other parts of the web I wrote a neat little summary of the two just in case I run into another issue like this.
In C++ there are two ways to represent/store decimal values.
Floats and Doubles
A float can store values from:
- -340282346638528859811704183484516925440.0000000000000000 Float lowest
- 340282346638528859811704183484516925440.0000000000000000 Float max
A double can store values from:
-179769313486231570814527423731704356798070567525844996598917476803157260780028538760589558632766878171540458953514382464234321326889464182768467546703537516986049910576551282076245490090389328944075868508455133942304583236903222948165808559332123348274797826204144723168738177180919299881250404026184124858368.0000000000000000 Double lowest
179769313486231570814527423731704356798070567525844996598917476803157260780028538760589558632766878171540458953514382464234321326889464182768467546703537516986049910576551282076245490090389328944075868508455133942304583236903222948165808559332123348274797826204144723168738177180919299881250404026184124858368.0000000000000000 Double max
Float's precision allows it to store a value of up to 9 digits (7 real digits, +2 from decimal to binary conversion)
Double, like the name suggests can store twice as much precision as a float. It can store up to 17 digits. (15 real digits, +2 from decimal to binary conversion)
e.g.
float x = 1.426;
double y = 8.739437;
Decimals & Math
Due to a float being able to carry 7 real decimals, and a double being able to carry 15 real decimals, to print them out when performing calculations a proper method must be used.
e.g
include
typedef std::numeric_limits<double> dbl;
cout.precision(dbl::max_digits10-2); // sets the precision to the *proper* amount of digits.
cout << dbl::max_digits10 <<endl; // prints 17.
double x = 12345678.312;
double a = 12345678.244;
// these calculations won't perform correctly be printed correctly without setting the precision.
cout << endl << x+a <<endl;
example 2:
typedef std::numeric_limits< float> flt;
cout.precision(flt::max_digits10-2);
cout << flt::max_digits10 <<endl;
float x = 54.122111;
float a = 11.323111;
cout << endl << x+a <<endl; /* without setting precison this outputs a different value, as well as making sure we're *limited* to 7 digits. If we were to enter another digit before the decimal point, the digits on the right would be one less, as there can only be 7. Doubles work in the same way */
Roughly how accurate is this description? Can it be used as a standard when confused?
The std::numerics_limits class in the <limits> header provides information about the characteristics of numeric types.
For a floating-point type T, here are the greatest and least values representable in the type, in various senses of “greatest” and “least.” I also include the values for the common IEEE 754 64-bit binary type, which is called double in this answer. These are in decreasing order:
std::numeric_limits<T>::infinity()is the largest representable value, ifTsupports infinity. It is, of course, infinity. Whether the typeTsupports infinity is indicated bystd::numeric_limits<T>::has_infinity.std::numeric_limits<T>::max()is the largest finite value. Fordouble, this is 21024−2971, approximately 1.79769•10308.std::numeric_limits<T>::min()is the smallest positive normal value. Floating-point formats often have an interval where the exponent cannot get any smaller, but the significand (fraction portion of the number) is allowed to get smaller until it reaches zero. This comes at the expense of precision but has some desirable mathematical-computing properties.min()is the point where this precision loss starts. Fordouble, this is 2−1022, approximately 2.22507•10−308.std::numeric_limits<T>::denorm_min()is the smallest positive value. In types which have subnormal values, it is subnormal. Otherwise, it equalsstd::numeric_limits<T>::min(). Fordouble, this is 2−1074, approximately 4.94066•10−324.std::numeric_limits<T>::lowest()is the least finite value. It is usually a negative number large in magnitude. Fordouble, this is −(21024−2971), approximately −1.79769•10308.If
std::numeric_limits<T>::has_infinityandstd::numeric_limits<T>::is_signedare true, then-std::numeric_limits<T>::infinity()is the least value. It is, of course, negative infinity.
Another characteristic you may be interested in is:
std::numeric_limits<T>::digits10is the greatest number of decimal digits such that converting any decimal number with that many digits toTand then converting back to the same number of decimal digits will yield the original number. Fordouble, this is 15.
Use the same format you'd use for any other values of those types:
#include <float.h>
#include <stdio.h>
int main(void) {
printf("FLT_MAX = %g\n", FLT_MAX);
printf("DBL_MAX = %g\n", DBL_MAX);
printf("LDBL_MAX = %Lg\n", LDBL_MAX);
}
Arguments of type float are promoted to double for variadic functions like printf, which is why you use the same format for both.
%f prints a floating-point value using decimal notation with no exponent, which will give you a very long string of (mostly insignificant) digits for very large values.
%e forces the use of an exponent.
%g uses either %f or %e, depending on the magnitude of the number being printed.
On my system, the above prints the following:
FLT_MAX = 3.40282e+38
DBL_MAX = 1.79769e+308
LDBL_MAX = 1.18973e+4932
As Eric Postpischil points out in a comment, the above prints only approximations of the values. You can print more digits by specifying a precision (the number of digits you'll need depends on the precision of the types); for example, you can replace %g by %.20g.
Or, if your implementation supports it, C99 added the ability to print floating-point values in hexadecimal with as much precision as necessary:
printf("FLT_MAX = %a\n", FLT_MAX);
printf("DBL_MAX = %a\n", DBL_MAX);
printf("LDBL_MAX = %La\n", LDBL_MAX);
But the result is not as easily human-readable as the usual decimal format:
FLT_MAX = 0x1.fffffep+127
DBL_MAX = 0x1.fffffffffffffp+1023
LDBL_MAX = 0xf.fffffffffffffffp+16380
(Note: main() is an obsolescent definition; use int main(void) instead.)
To print approximations of the maximums with enough digits to represent the actual values (the result of converting the printed value back to floating-point should be the original value), you can use:
#include <float.h>
#include <stdio.h>
int main(void)
{
printf("%.*g\n", DECIMAL_DIG, FLT_MAX);
printf("%.*g\n", DECIMAL_DIG, DBL_MAX);
printf("%.*Lg\n", DECIMAL_DIG, LDBL_MAX);
return 0;
}
In C 2011, you can use the more specific FLT_DECIMAL_DIG, DBL_DECIMAL_DIG, and LDBL_DECIMAL_DIG in place of DECIMAL_DIG.
To print the exact values, instead of approximations, you need to specify more precision. (int) (log10(x)+1) digits should be enough.
Approximations of the minimums and the epsilons can be printed with sufficient accuracy in the same way. However, calculating the numbers of digits needed for exact values may be more complicated than for the maximums. (Technically, it may be impossible in exotic C implementations. E.g., a base-three floating-point system would have a minimum not representable in any finite number of decimal digits. I am not aware of any such implementations in use.)