Floating point types work differently than integer types.
Integer types directly map their min-max range values to all the binary permutations from all 0s to all 1s. So they can directly represent a value, if its in range.
Floating point types are essentially 2 numbers: fraction (mantissa) and exponent.
So number 1,000,000 could be represented by fraction 1 and exponent 6 (10^6).
So a floating point number can represent a massive range of numbers (range is limited by exponent range) with variable accuracy (accuracy is limited by fraction range). As numbers get larger, due to the limited digits of fraction, smaller numbers will get less and less accurate. Eg if you reach 10,000 , adding 0.0001 will result in closest number to 10,000.0001 that can be represented, like 10,000.025786 or something.
As per IEE765, a double has 11 bits of exponent and 52 bits for fraction.
So when you pair a 52 bit long number, with a 11 bit long exponent, you get quite a large number. But that doesnt mean every number in that smallest to largest range can be accurately represented, unlike an integer type.
Answer from pacukluka on Stack OverflowFloating point types work differently than integer types.
Integer types directly map their min-max range values to all the binary permutations from all 0s to all 1s. So they can directly represent a value, if its in range.
Floating point types are essentially 2 numbers: fraction (mantissa) and exponent.
So number 1,000,000 could be represented by fraction 1 and exponent 6 (10^6).
So a floating point number can represent a massive range of numbers (range is limited by exponent range) with variable accuracy (accuracy is limited by fraction range). As numbers get larger, due to the limited digits of fraction, smaller numbers will get less and less accurate. Eg if you reach 10,000 , adding 0.0001 will result in closest number to 10,000.0001 that can be represented, like 10,000.025786 or something.
As per IEE765, a double has 11 bits of exponent and 52 bits for fraction.
So when you pair a 52 bit long number, with a 11 bit long exponent, you get quite a large number. But that doesnt mean every number in that smallest to largest range can be accurately represented, unlike an integer type.
C++ why max of 64bit double has 308 digits?
In my environment (Win10 64bit
Because your system uses hardware that conforms to the IEEE 754 specification (as does most hardware). That document specifies the the largest finite representable value of 64 bit binary floating point to be 21024 which is a bit less than 1.8 E308.
sizeof(double) is 8 bytes, the max value should be 1.84e19
There are about 1.84e19 positive integers that are representable by 8 bytes. Are you assuming that double is an unsigned integer type? It is not, so your assumption is misplaced.
The biggest/largest integer that can be stored in a double without losing precision is the same as the largest possible value of a double. That is, DBL_MAX or approximately 1.8 × 10308 (if your double is an IEEE 754 64-bit double). It's an integer, and it's represented exactly.
What you might want to know instead is what the largest integer is, such that it and all smaller integers can be stored in IEEE 64-bit doubles without losing precision. An IEEE 64-bit double has 52 bits of mantissa, so it's 253 (and -253 on the negative side):
- 253 + 1 cannot be stored, because the 1 at the start and the 1 at the end have too many zeros in between.
- Anything less than 253 can be stored, with 52 bits explicitly stored in the mantissa, and then the exponent in effect giving you another one.
- 253 obviously can be stored, since it's a small power of 2.
Or another way of looking at it: once the bias has been taken off the exponent, and ignoring the sign bit as irrelevant to the question, the value stored by a double is a power of 2, plus a 52-bit integer multiplied by 2exponent − 52. So with exponent 52 you can store all values from 252 through to 253 − 1. Then with exponent 53, the next number you can store after 253 is 253 + 1 × 253 − 52. So loss of precision first occurs with 253 + 1.
9007199254740992 (that's 9,007,199,254,740,992 or 2^53) with no guarantees :)
Program
#include <math.h>
#include <stdio.h>
int main(void) {
double dbl = 0; /* I started with 9007199254000000, a little less than 2^53 */
while (dbl + 1 != dbl) dbl++;
printf("%.0f\n", dbl - 1);
printf("%.0f\n", dbl);
printf("%.0f\n", dbl + 1);
return 0;
}
Result
9007199254740991 9007199254740992 9007199254740992
confused about the max limit we can store for float and double
Questions about 64-bit stuff
Is the maximum supported digits for a double type in c is 15 digits or more
c - Double precision - Max value - Stack Overflow
Just started learning C. so forigve me if this is a dumb question.
For int and long, the maximum number we can assign them is 2^32 and 2^64 respectively. how about for float and double? I know that double has higher precision of about 15 decimal points while float is 7.
i also know that float is 32 bit while double is 64 bit. So does that mean the highest we can store them (2^32)and (2^64) but with decimal points in between their numbers? i aslo learnt that these data type cant have unsigned value so does that value doubles?
Floating point types work differently than integer types.
Integer types directly map their min-max range values to all the binary permutations from all 0s to all 1s. So they can directly represent a value, if its in range.
Floating point types are essentially 2 numbers: fraction (mantissa) and exponent.
So number 1,000,000 could be represented by fraction 1 and exponent 6 (10^6).
So a floating point number can represent a massive range of numbers (range is limited by exponent range) with variable accuracy (accuracy is limited by fraction range). As numbers get larger, due to the limited digits of fraction, smaller numbers will get less and less accurate. Eg if you reach 10,000 , adding 0.0001 will result in closest number to 10,000.0001 that can be represented, like 10,000.025786 or something.
As per IEE765, a double has 11 bits of exponent and 52 bits for fraction.
So when you pair a 52 bit long number, with a 11 bit long exponent, you get quite a large number. But that doesnt mean every number in that smallest to largest range can be accurately represented, unlike an integer type.
Answer from pacukluka on Stack OverflowHi Debojit Acharjee,
The significand of the double type is approximately 15 to 17 decimal digits for most platforms. In most cases, a variable of type double can accurately represent 15 to 17 decimal digits. Numbers outside this range may lose precision or be rounded.
When I defined a 18 digits number, the result lost precision.
When using floating-point numbers, you should choose the appropriate data type according to your specific needs and precision requirements, and avoid using values beyond its representation range for calculations.
Regarding the double type, this documentation states:
Microsoft Specific The double type contains 64 bits: 1 for sign, 11 for the exponent, and 52 for the mantissa. Its range is +/-1.7E308 with at least 15 digits of precision.
You could also refer to this document for the float type.
Best regards,
Elya Yao
If the answer is the right solution, please click "Accept Answer" and kindly upvote it. If you have extra questions about this answer, please click "Comment".
Note: Please follow the steps in our documentation to enable e-mail notifications if you want to receive the related email notification for this thread.
A double is stored in base 2 not decimal. It’s stored in 64 bits. The mantissa is 52 bits, or max 179769313486232 in decimal. The exponent is 11 bits or max of 2047 in decimal. The final bit is the sign bit.
See:
https://en.wikipedia.org/wiki/Computer_number_format#:~:text=an%2011%2Dbit%20binary%20exponent,gives%20the%20actual%20signed%20value
It will not hold the 308 digits of the 10^308 number. The double precision number holds the exponent and a limited number of digits.
See https://en.wikipedia.org/wiki/IEEE_floating_point (english) http://fr.wikipedia.org/wiki/IEEE_754 (french) for a detailed description of floating points encoding in memory.
According to C standard, there are three floating point types: float, double, and long double, and the value representation of all floating-point types are implementation defined.
Most compilers however do follow the binary64 format, as specified by the IEEE 754 standard.
This format has:
- 1 sign bit
- 11 bits for exponent
- 52 bits for mantissa
To find the largest value double can hold, you should check the DBL_MAX defined in the header <float.h>. It will be approximately 1.8 × 10308 for implementations using binary64 IEEE 754 standard.
It's usually not a good idea to test the equality of floating-point numbers. The behavior of binary floating-point numbers can differ drastically from what you may expect from base-10 decimals. Consider the example:
>> isequal(0.1, 0.3/3)
ans =
0
Ultimately, you have 53 bits of precision. This means that integers can be represented exactly (with no loss in accuracy) up to the number 253 (which is a little over 9 x 1015). After that, well:
>> (2^53 + 1) - 2^53
ans =
0
>> 2^53 + (1 - 2^53)
ans =
1
For non-integers, you are almost never going to be representing them exactly, even for simple-looking decimals such as 0.1 (as shown in that first example). However, it still guarantees you at least 15 significant figures of precision.
This means that if you take any number and round it to the nearest number representable as a double-precision floating point, then this new number will match your original number at least up to the first 15 digits (regardless of where these digits are with respect to the decimal point).
You might want to use variable precision arithmetics (VPA) in matlab. It computes expressions exactly up to a given digit count, which may be quite large. See here.
If you know the number of exponent bits and mantissa bits, then based on the IEEE-754 format, one can establish that the maximum absolute representable value is:
2^(2^(E-1)-1)) * (1 + (2^M-1)/2^M)
The minimum absolute value (not including zero or denormals) is:
2^(2-2^(E-1))
- For single-precision,
Eis 8,Mis 23. - For double-precision,
Eis 11,Mis 52. - For extended-precision, I'm not sure. If you're referring to the 80-bit precision of the x87 FPU, then so far as I can tell, it's not's IEEE-754 compliant...
The answer (if you're on a machine with IEEE floating point) is
in float.h. FLT_MAX, DBL_MAX and LDBL_MAX. On a system
with full IEEE support, something around 3.4e+38, 1.8E+308 and
1.2E4932. (The exact values may vary, and may be expressed
differently, depending on how the compiler does its input and
rounding. g++, for example, defines them to be compiler
built-ins.)
EDIT:
WRT your question (since neither I nor the other responders
actually answered it): the range of representable values is
[-type_MAX...type], where
type is one of FLT, DBL, or LDBL.