Data Types and Their Sizes
Chapter Nine
Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"
Pages 39 to 44 of 222
In one line
A data type tells the compiler how many bytes an object occupies and how to read the bits in them, and C's basic types are char, int, float and double, widened by the modifiers short, long, signed and unsigned.
Why a type is necessary at all
A byte in memory is eight bits and nothing more. The bits 01000001 are the number 65, and they are also the letter A, and as part of a larger group they are part of some floating-point value. Nothing in the memory says which. The type is where that information lives, and it lives in the compiler, not in the machine.
That single fact explains most of C. The type decides how much space is set aside, how the bits are interpreted, which operations are allowed, and what happens when two different types meet in one expression. It also explains why C is fast and why it will let you do something disastrous without complaint: once the program is compiled, the types are gone, and the machine does what the instructions say.
Ritchie's stated reason for inventing C out of B was precisely this: B was typeless, and a language in which every value is one machine word cannot conveniently talk about a character. Chapter 3 has the history.
The basic types
char holds one character, and is exactly one byte by definition. sizeof(char) is 1, always, on every implementation. A byte need not be eight bits in the standard's terms, but on every machine you will meet it is.
int holds a whole number, and is meant to be the natural size for the machine.
float holds a number with a fractional part, at single precision.
double holds a number with a fractional part, at double precision. It is the one to use.
void holds nothing, and is not a type you can make a variable of. It appears in three places: as the return type of a function that returns nothing, as (void) for a function that takes no parameters, and as void *, the pointer that can point at anything, which is chapter 39.
The four modifiers
Two change the size:
shortasks for a smaller integer thanint.longasks for a larger one.long long, added in C99, asks for larger still.
Two change the meaning of the bits:
signedallows negative values. It is the default forint.unsigneddoes not, and spends the bit that would have carried the sign on magnitude instead, so the largest value roughly doubles.
They combine, and the spelling is flexible: unsigned long int, long unsigned, unsigned long all name the same type. int may be left out whenever a modifier is present, so short means short int and unsigned means unsigned int.
Data Types and Their Sizes
What the standard actually guarantees
This is the table that is true everywhere. It is the standard's own required minimum magnitudes.
| Type | At least this range | Which needs |
|---|---|---|
signed char | -127 to +127 | 8 bits |
unsigned char | 0 to 255 | 8 bits |
short | -32767 to +32767 | 16 bits |
unsigned short | 0 to 65535 | 16 bits |
int | -32767 to +32767 | 16 bits |
unsigned int | 0 to 65535 | 16 bits |
long | -2147483647 to +2147483647 | 32 bits |
unsigned long | 0 to 4294967295 | 32 bits |
long long | -9223372036854775807 to +9223372036854775807 | 64 bits |
unsigned long long | 0 to 18446744073709551615 | 64 bits |
Two further guarantees are worth as much as the table:
sizeof(char)is 1. Every other size is measured in units of it.- The sizes are ordered and the ranges nest:
charno larger thanshort, no larger thanint, no larger thanlong, no larger thanlong long. So ashortvalue always fits in anint, whatever the machine.
Notice what is not guaranteed. int is not four bytes. float is not four bytes. Nothing says int and long differ. On a small embedded processor int really is two bytes and the minimum above really is the range.
What this machine actually has
Ask the machine. That is what sizeof is for: an operator, not a function, that gives the size of a type or of an object in bytes.
#include <stdio.h>
int main(void)
{
printf("type bytes\n");
printf("char %5zu\n", sizeof(char));
printf("short %5zu\n", sizeof(short));
printf("int %5zu\n", sizeof(int));
printf("long %5zu\n", sizeof(long));
printf("long long %5zu\n", sizeof(long long));
printf("float %5zu\n", sizeof(float));
printf("double %5zu\n", sizeof(double));
printf("long double %5zu\n", sizeof(long double));
printf("void * %5zu\n", sizeof(void *));
return 0;
}type bytes
char 1
short 2
int 4
long 8
long long 8
float 4
double 8
long double 16
void * 8Those are the sizes on the 64-bit Linux machine that compiled this book's programs. Four of them differ elsewhere, and these are the four worth knowing about:
longis 4 bytes on 64-bit Windows, not 8. Code that assumes alongcan hold a large count is not portable between Linux and Windows.long longis 8 everywhere that has it.long doubleis 16 bytes here, 8 on 64-bit Windows with MSVC, and 12 on 32-bit Linux. It is the least portable type in C.void *is 8 bytes on any 64-bit machine and 4 on a 32-bit one. It is the size of an address, so it tracks the machine, not the standard.intis 4 bytes on every desktop machine and 2 on many microcontrollers.
Data Types and Their Sizes
%zu is the conversion for a value of type size_t, which is what sizeof produces. Using %d for it is a mistake that -Wall catches and that a great many textbooks make.
Signed against unsigned, and the trap in it
An unsigned type cannot hold a negative value. What it does instead is the thing to learn, because it is defined behaviour and it surprises everybody once.
#include <stdio.h>
int main(void)
{
unsigned int u = 0;
int i = 0;
printf("as unsigned, 0 - 1 is %u\n", u - 1);
printf("as signed, 0 - 1 is %d\n", i - 1);
return 0;
}as unsigned, 0 - 1 is 4294967295
as signed, 0 - 1 is -1Subtracting one from an unsigned zero does not give minus one and it is not an error. Unsigned arithmetic wraps round: the result is reduced modulo one more than the largest value the type can hold, so it comes out as the largest value. This is exactly why a loop written as
for (unsigned int i = n; i >= 0; i--)never ends. An unsigned i is always at least zero, so the condition is always true.
Use int for counting unless you have a reason not to. unsigned is for bit patterns, for sizes returned by the library, and for the one case where you genuinely need the extra range.
The floating-point types, and why double
float and double hold approximations. They are stored as a sign, a set of significant digits and an exponent, in the manner of scientific notation, which means a fixed number of significant digits and not a fixed number of decimal places.
The number of decimal digits you can rely on is in <float.h> as FLT_DIG and DBL_DIG. On the machine that ran this book's programs those are 6 and 15, and those two values are typical of every desktop machine, because both types follow the same international format for floating-point arithmetic.
Six digits is not many. Here is what that costs, in a program that is not buggy:
#include <stdio.h>
int main(void)
{
float f = 123456789.0f;
double d = 123456789.0;
printf("as a float : %.1f\n", (double) f);
printf("as a double : %.1f\n", d);
printf("0.1 + 0.2 as a double, to 20 places: %.20f\n", 0.1 + 0.2);
return 0;
}as a float : 123456792.0
as a double : 123456789.0
0.1 + 0.2 as a double, to 20 places: 0.30000000000000004441The float did not store 123456789. It stored the nearest value it is able to represent, and the last two digits are not the ones that were written. Nothing went wrong: a float has about seven significant decimal digits and was asked for nine.
Data Types and Their Sizes
Try the same program with 1234567.0f and the float prints it back perfectly. Seven digits happens to be inside what a float can do, which is exactly why this trap is hard to spot: it appears only once the numbers grow, and by then the program is finished and trusted.
And 0.1 + 0.2 is not 0.3, in any language, on any machine that uses binary floating point, because one tenth cannot be written exactly in binary any more than one third can be written exactly in decimal.
The rule that follows is absolute and it is worth marks: never compare two floating-point values with ==. Compare the size of their difference against a small tolerance instead. Chapter 16 gives the form.
Use double, not float. A float saves four bytes, which no program in this course will notice, and costs nine significant digits. float exists for very large arrays and for hardware that is faster at it.
The complete picture
| Declaration | Typical size here | What it is for |
|---|---|---|
char | 1 | One character, or a very small integer |
signed char | 1 | A small integer, definitely signed |
unsigned char | 1 | A byte, 0 to 255 |
short | 2 | A small integer |
unsigned short | 2 | A small non-negative integer |
int | 4 | The default whole number |
unsigned int | 4 | Non-negative, or a bit pattern |
long | 8 here, 4 on Windows | A larger integer |
long long | 8 | A large integer, everywhere |
float | 4 | Fractions, about 6 digits |
double | 8 | Fractions, about 15 digits. The default |
long double | 16 here | Fractions, more digits, least portable |
void | none | Nothing: no value, no parameters, or an untyped pointer |
The conversion to print each one
Getting this wrong is undefined behaviour, not a small mistake, because printf reads its arguments according to what the format string says is there. -Wall catches nearly all of it, which is the best argument for compiling with it.
| Type | Conversion |
|---|---|
char as a character | %c |
char as a number | %d |
int | %d |
unsigned int | %u |
short | %hd |
long | %ld |
long long | %lld |
unsigned long | %lu |
float | %f, and the value is promoted to double |
double | %f, or %g, or %e |
long double | %Lf |
size_t, from sizeof | %zu |
| a pointer | %p |
There is no conversion for float in printf, and it is not an omission: a float argument is automatically promoted to double when passed, so %f is correct for both. In scanf, which takes addresses and promotes nothing, %f is for a float and %lf is for a double . Confusing the two in scanf is a real and common bug.
Data Types and Their Sizes
What this does NOT mean
int is not four bytes. It is four bytes on the machines you will use. The language says at least two. Write sizeof(int) when you need the number and never the literal 4.
sizeof is not a function. It is an operator and it is evaluated by the compiler, not at run time. sizeof x without brackets is legal for an object; brackets are required for a type name.
char is not "the character type" and nothing else. It is a one-byte integer, and whether it is signed is up to the implementation. When you want a small number use signed char or unsigned char and say which. When you want a character, use char.
unsigned does not mean positive. It means non-negative, and it means arithmetic that wraps rather than going negative.
float is not "less accurate for big numbers only". It has about six significant digits at every magnitude. float can hold 3.4 times ten to the thirty-eighth; it just cannot tell that number from its neighbours.
A type is not a promise about the value. Declaring int age does not stop anyone storing -500 in it. Checking the value is your job, and it is part of what integrity meant in chapter 6.
Quick revision
- Basic types:
char,int,float,double, plusvoid. - Modifiers:
short,long,signed,unsigned. sizeof(char)is 1 by definition; everything else is measured in those units.- The standard fixes minimum ranges, not sizes:
intat least -32767 to +32767. - Sizes never shrink going
char,short,int,long,long long. - Here: char 1, short 2, int 4, long 8, long long 8, float 4, double 8, pointer 8.
longis 4 bytes on 64-bit Windows.long longis 8 everywhere.- Unsigned arithmetic wraps modulo the range;
0u - 1is the maximum value. floatabout 6 significant digits,doubleabout 15. Usedouble.- Never compare floating-point values with
==. sizeofgives asize_t; print it with%zu.
Test yourself
1. What is the size of int in C?
The standard does not say. It requires a range of at least -32767 to +32767, so at least two bytes, and it is four bytes on every desktop machine. sizeof(int) is the only reliable answer for a given machine.
2. Which is bigger, int or long?
long is never smaller than int, and on many machines they are the same size. sizeof(long) >= sizeof(int) is guaranteed; > is not.
3. What does this print, and why?
unsigned int u = 5;
printf("%u\n", u - 10);A very large number, the maximum for unsigned int minus four, which is 4294967291 where unsigned int is 32 bits. Unsigned arithmetic is reduced modulo one more than the maximum, so it cannot produce a negative result.
Data Types and Their Sizes
4. Why should you not write if (x == 0.1) when x is a double?
Because 0.1 has no exact binary representation, so the value stored is the nearest one that does, and any arithmetic that produced x may have landed on a different nearest value. Test if (fabs(x - 0.1) < 1e-9) instead.
5. How many bytes does char occupy, and how do you know?
One, by definition. sizeof(char) is 1 in every conforming implementation, because the byte is the unit sizeof measures in.
6. Give the right printf conversion for sizeof(int), and say why %d is wrong.
%zu. sizeof produces a size_t, which is an unsigned type that may be wider than int, so %d tells printf to read the wrong number of bytes as the wrong kind of value.
What can be asked on this, and how to answer it
"Explain the data types in C with their sizes and ranges." Give the four basic types and the four modifiers, then the table of typical sizes, and then the sentence that earns the extra marks: the standard fixes minimum ranges rather than sizes, so the correct way to get a size is sizeof. Give int as 2 bytes minimum and 4 in practice, and you have answered both versions of the question.
"What is the difference between signed and unsigned?" A signed type spends one bit on the sign and can hold negative values; an unsigned type of the same width holds only non-negative values and reaches roughly twice as far. Add that unsigned arithmetic wraps modulo the range rather than going negative, and give 0u - 1 as the example.
"Differentiate between float and double." Both hold approximations. float is typically 4 bytes with about 6 significant decimal digits; double is typically 8 with about 15. double is the default for floating-point constants and the type to prefer. float is used to save space in large arrays.
"What is the use of the void data type?" Three uses: the return type of a function that returns no value, the parameter list of a function that takes no arguments, and void *, a pointer to an object of unspecified type.
"What is sizeof? Write a program to find the size of each data type." sizeof is a compile-time operator giving the size in bytes of a type or object, with sizeof(char) equal to 1 by definition. Then give the program in this chapter, and remember %zu.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.