Floating point types work differently than integer types.

Integer types directly map their min-max range values to all the binary permutations from all 0s to all 1s. So they can directly represent a value, if its in range.

Floating point types are essentially 2 numbers: fraction (mantissa) and exponent.

So number 1,000,000 could be represented by fraction 1 and exponent 6 (10^6).

So a floating point number can represent a massive range of numbers (range is limited by exponent range) with variable accuracy (accuracy is limited by fraction range). As numbers get larger, due to the limited digits of fraction, smaller numbers will get less and less accurate. Eg if you reach 10,000 , adding 0.0001 will result in closest number to 10,000.0001 that can be represented, like 10,000.025786 or something.

As per IEE765, a double has 11 bits of exponent and 52 bits for fraction.

So when you pair a 52 bit long number, with a 11 bit long exponent, you get quite a large number. But that doesnt mean every number in that smallest to largest range can be accurately represented, unlike an integer type.

Answer from pacukluka on Stack Overflow
Top answer
1 of 3
4

Floating point types work differently than integer types.

Integer types directly map their min-max range values to all the binary permutations from all 0s to all 1s. So they can directly represent a value, if its in range.

Floating point types are essentially 2 numbers: fraction (mantissa) and exponent.

So number 1,000,000 could be represented by fraction 1 and exponent 6 (10^6).

So a floating point number can represent a massive range of numbers (range is limited by exponent range) with variable accuracy (accuracy is limited by fraction range). As numbers get larger, due to the limited digits of fraction, smaller numbers will get less and less accurate. Eg if you reach 10,000 , adding 0.0001 will result in closest number to 10,000.0001 that can be represented, like 10,000.025786 or something.

As per IEE765, a double has 11 bits of exponent and 52 bits for fraction.

So when you pair a 52 bit long number, with a 11 bit long exponent, you get quite a large number. But that doesnt mean every number in that smallest to largest range can be accurately represented, unlike an integer type.

2 of 3
2

C++ why max of 64bit double has 308 digits?

In my environment (Win10 64bit

Because your system uses hardware that conforms to the IEEE 754 specification (as does most hardware). That document specifies the the largest finite representable value of 64 bit binary floating point to be 21024 which is a bit less than 1.8 E308.


sizeof(double) is 8 bytes, the max value should be 1.84e19

There are about 1.84e19 positive integers that are representable by 8 bytes. Are you assuming that double is an unsigned integer type? It is not, so your assumption is misplaced.

Top answer
1 of 11
733

The biggest/largest integer that can be stored in a double without losing precision is the same as the largest possible value of a double. That is, DBL_MAX or approximately 1.8 × 10308 (if your double is an IEEE 754 64-bit double). It's an integer, and it's represented exactly.

What you might want to know instead is what the largest integer is, such that it and all smaller integers can be stored in IEEE 64-bit doubles without losing precision. An IEEE 64-bit double has 52 bits of mantissa, so it's 253 (and -253 on the negative side):

  • 253 + 1 cannot be stored, because the 1 at the start and the 1 at the end have too many zeros in between.
  • Anything less than 253 can be stored, with 52 bits explicitly stored in the mantissa, and then the exponent in effect giving you another one.
  • 253 obviously can be stored, since it's a small power of 2.

Or another way of looking at it: once the bias has been taken off the exponent, and ignoring the sign bit as irrelevant to the question, the value stored by a double is a power of 2, plus a 52-bit integer multiplied by 2exponent − 52. So with exponent 52 you can store all values from 252 through to 253 − 1. Then with exponent 53, the next number you can store after 253 is 253 + 1 × 253 − 52. So loss of precision first occurs with 253 + 1.

2 of 11
113

9007199254740992 (that's 9,007,199,254,740,992 or 2^53) with no guarantees :)

Program

#include <math.h>
#include <stdio.h>

int main(void) {
  double dbl = 0; /* I started with 9007199254000000, a little less than 2^53 */
  while (dbl + 1 != dbl) dbl++;
  printf("%.0f\n", dbl - 1);
  printf("%.0f\n", dbl);
  printf("%.0f\n", dbl + 1);
  return 0;
}

Result

9007199254740991
9007199254740992
9007199254740992
Discussions

confused about the max limit we can store for float and double
Interestingly, a double can exactly represent any integer up to about 2^53. Here's more info on the max limits: https://en.wikipedia.org/wiki/IEEE_754 More on reddit.com
🌐 r/C_Programming
8
3
February 1, 2024
Questions about 64-bit stuff
The smallest positive integer that a 64 bit double in IEEE 754 format cannot represent properly is 2^53 + 1. Numbers represented in double are restricted to about 16 digits in accuracy. Once the values get above about 10^16 then the distance between adjacent representable numbers becomes larger ... More on mathworks.com
🌐 mathworks.com
4
1
July 27, 2015
Is the maximum supported digits for a double type in c is 15 digits or more
In your code, the variable a is assigned the value 12345678901234567, which is a 17-bit number. Due to the precision limitation of double, it is difficult to accurately represent such a large number. Therefore, rounding or approximation occurs on output, causing the last digit to become 8 instead of the original 7. ... A double is stored in base 2 not decimal. It’s stored in 64 bits. The mantissa is 52 bits, or max ... More on learn.microsoft.com
🌐 learn.microsoft.com
2
0
May 28, 2023
c - Double precision - Max value - Stack Overflow
I have a pretty silly question on double precision. I have read that a double (in C for example) is represented on 64 bits, but I have also read that the maximum value that can be represented by a double is approximately 10^308. More on stackoverflow.com
🌐 stackoverflow.com
🌐
Reddit
reddit.com › r/c_programming › confused about the max limit we can store for float and double
r/C_Programming on Reddit: confused about the max limit we can store for float and double
February 1, 2024 -

Just started learning C. so forigve me if this is a dumb question.

For int and long, the maximum number we can assign them is 2^32 and 2^64 respectively. how about for float and double? I know that double has higher precision of about 15 decimal points while float is 7.

i also know that float is 32 bit while double is 64 bit. So does that mean the highest we can store them (2^32)and (2^64) but with decimal points in between their numbers? i aslo learnt that these data type cant have unsigned value so does that value doubles?

Floating point types work differently than integer types.

Integer types directly map their min-max range values to all the binary permutations from all 0s to all 1s. So they can directly represent a value, if its in range.

Floating point types are essentially 2 numbers: fraction (mantissa) and exponent.

So number 1,000,000 could be represented by fraction 1 and exponent 6 (10^6).

So a floating point number can represent a massive range of numbers (range is limited by exponent range) with variable accuracy (accuracy is limited by fraction range). As numbers get larger, due to the limited digits of fraction, smaller numbers will get less and less accurate. Eg if you reach 10,000 , adding 0.0001 will result in closest number to 10,000.0001 that can be represented, like 10,000.025786 or something.

As per IEE765, a double has 11 bits of exponent and 52 bits for fraction.

So when you pair a 52 bit long number, with a 11 bit long exponent, you get quite a large number. But that doesnt mean every number in that smallest to largest range can be accurately represented, unlike an integer type.

Answer from pacukluka on Stack Overflow
🌐
John Tromp
tromp.github.io › blog › 2023 › 11 › 24 › largest-number
The largest number representable in 64 bits
November 24, 2023 - That is indeed the maximum possible value of 64 bit unsigned integers, available as datatype uint64_t in C or u64 in Rust. We can easily surpass this with floating point numbers. The 64-bit double floating point format has a largest (finite) representable value of 21024(1-2-53) ~ 1.8*10308.
🌐
Wikipedia
en.wikipedia.org › wiki › Double-precision_floating-point_format
Double-precision floating-point format - Wikipedia
3 days ago - Double-precision floating-point format (sometimes called FP64 or float64) is a floating-point number format, usually occupying 64 bits in computer memory; it represents a wide range of numeric values by using a floating radix point.
🌐
Quora
quora.com › What-is-the-largest-64-bit-number
What is the largest 64 bit number? - Quora
Answer (1 of 3): Depends on the use. Unsigned - i.e. representation without negative numbers - it will be 2^64–1 18,446,744,073,709,551,615 However, if the representation needed includes negative numbers, the leading bit is used as a sign bit. The range is -2^63 to (2^63)-1. i.e. -9,223,372,...
Find elsewhere
🌐
MathWorks
mathworks.com › matlabcentral › answers › 231253-questions-about-64-bit-stuff
Questions about 64-bit stuff - MATLAB Answers - MATLAB Central
July 27, 2015 - The smallest positive integer that a 64 bit double in IEEE 754 format cannot represent properly is 2^53 + 1. Numbers represented in double are restricted to about 16 digits in accuracy.
🌐
Jazz Community Site
jazz.net › dxl › html › 4150 - Max value that can be held in Real Variable.html
Max value that can be held in Real Variable
Hi We are trying to store the maximum number that can be held in 64 bits in a real variable as follows: real maxVal = 18446744073709551615.0 We want to test other real variables against this limit. However, we have noticed in maxVal that it doesn't hold this number instead it rounds the number ...
🌐
Fandom
googology.fandom.com › wiki › User_blog:JohnTromp › The_largest_number_representable_in_64_bits
User blog:JohnTromp/The largest number representable in 64 bits | Googology Wiki | Fandom
We can reach quite a bit further with floating point values. The 64-bit double floating point format finds its largest (finite) representable value in the 309 digit number 21024(1-2-53) = 17976931...24858368.
🌐
Note.nkmk.me
note.nkmk.me › home › python
Maximum and Minimum float Values in Python | note.nkmk.me
August 11, 2023 - In Python, the float type is a 64-bit double-precision floating-point number, equivalent to double in languages like C. This article explains how to get and check the range (maximum and minimum values) that float can represent in Python. In many environments, the representable range for float ...
Top answer
1 of 2
1

Hi Debojit Acharjee,

The significand of the double type is approximately 15 to 17 decimal digits for most platforms. In most cases, a variable of type double can accurately represent 15 to 17 decimal digits. Numbers outside this range may lose precision or be rounded.

When I defined a 18 digits number, the result lost precision.

When using floating-point numbers, you should choose the appropriate data type according to your specific needs and precision requirements, and avoid using values beyond its representation range for calculations.

Regarding the double type, this documentation states:

Microsoft Specific The double type contains 64 bits: 1 for sign, 11 for the exponent, and 52 for the mantissa. Its range is +/-1.7E308 with at least 15 digits of precision.

You could also refer to this document for the float type.

Best regards,

Elya Yao


If the answer is the right solution, please click "Accept Answer" and kindly upvote it. If you have extra questions about this answer, please click "Comment".

Note: Please follow the steps in our documentation to enable e-mail notifications if you want to receive the related email notification for this thread.

2 of 2
0

A double is stored in base 2 not decimal. It’s stored in 64 bits. The mantissa is 52 bits, or max 179769313486232 in decimal. The exponent is 11 bits or max of 2047 in decimal. The final bit is the sign bit.

See:

https://en.wikipedia.org/wiki/Computer_number_format#:~:text=an%2011%2Dbit%20binary%20exponent,gives%20the%20actual%20signed%20value

🌐
Autodesk
help.autodesk.com › cloudhelp › 2016 › ENU › MAXScript-Help › files › GUID-20409BF0-2192-4551-AB01-EA46E1DE51EE.htm
64 Bit Values - Double, Integer64, IntegerPtr
When a Double value is printed, ... as 0.0. In versions prior to 3ds Max 2011, the float value of 1.192092896e-07 was being used for Doubles too. ... In 3ds Max 9 and higher, returns true if running in a 64 bit build of 3ds Max, false otherwise....
🌐
Microsoft Learn
learn.microsoft.com › en-us › dotnet › api › system.double.maxvalue
Double.MaxValue Field (System) | Microsoft Learn
January 25, 2023 - Represents the largest possible value of a Double. This field is constant. public: double MaxValue = 1.7976931348623157E+308;
🌐
IBM
ibm.com › docs › en › db2-big-sql › 6.0.0
IBM Documentation
The numeric data types are integer, decimal, floating-point, and decimal floating-point.
🌐
javaspring
javaspring.net › blog › java-double-max-value
Java Double Max Value: A Comprehensive Guide — javaspring.net
This blog post will explore the ... numbers. It is a 64-bit IEEE 754 floating-point number, which means it can represent a very large range of values....
🌐
Cplusplus
cplusplus.com › forum › beginner › 111796
Values larger than 2^64 - C++ Forum
I don't understand how the range of a 64 bit unsigned integer is 2^64 - 1 but for a double it's < 1.8e301 . Checked wikipedia but didn't really understand it's explanation. Say a data type such as a double is 8 bytes , why is its range not the same as an 8 byte integer.