munotes®

Data Types and Their Sizes

Chapter Nine

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

Pages 39 to 44 of 222

In one line

A data type tells the compiler how many bytes an object occupies and how to read the bits in them, and C's basic types are char, int, float and double, widened by the modifiers short, long, signed and unsigned.

Why a type is necessary at all

A byte in memory is eight bits and nothing more. The bits 01000001 are the number 65, and they are also the letter A, and as part of a larger group they are part of some floating-point value. Nothing in the memory says which. The type is where that information lives, and it lives in the compiler, not in the machine.

That single fact explains most of C. The type decides how much space is set aside, how the bits are interpreted, which operations are allowed, and what happens when two different types meet in one expression. It also explains why C is fast and why it will let you do something disastrous without complaint: once the program is compiled, the types are gone, and the machine does what the instructions say.

Ritchie's stated reason for inventing C out of B was precisely this: B was typeless, and a language in which every value is one machine word cannot conveniently talk about a character. Chapter 3 has the history.

The basic types

char holds one character, and is exactly one byte by definition. sizeof(char) is 1, always, on every implementation. A byte need not be eight bits in the standard's terms, but on every machine you will meet it is.

int holds a whole number, and is meant to be the natural size for the machine.

float holds a number with a fractional part, at single precision.

double holds a number with a fractional part, at double precision. It is the one to use.

void holds nothing, and is not a type you can make a variable of. It appears in three places: as the return type of a function that returns nothing, as (void) for a function that takes no parameters, and as void *, the pointer that can point at anything, which is chapter 39.

The four modifiers

Two change the size:

  • short asks for a smaller integer than int.
  • long asks for a larger one. long long, added in C99, asks for larger still.

Two change the meaning of the bits:

  • signed allows negative values. It is the default for int.
  • unsigned does not, and spends the bit that would have carried the sign on magnitude instead, so the largest value roughly doubles.

They combine, and the spelling is flexible: unsigned long int, long unsigned, unsigned long all name the same type. int may be left out whenever a modifier is present, so short means short int and unsigned means unsigned int.

munotes.in39

Data Types and Their Sizes

What the standard actually guarantees

This is the table that is true everywhere. It is the standard's own required minimum magnitudes.

TypeAt least this rangeWhich needs
signed char-127 to +1278 bits
unsigned char0 to 2558 bits
short-32767 to +3276716 bits
unsigned short0 to 6553516 bits
int-32767 to +3276716 bits
unsigned int0 to 6553516 bits
long-2147483647 to +214748364732 bits
unsigned long0 to 429496729532 bits
long long-9223372036854775807 to +922337203685477580764 bits
unsigned long long0 to 1844674407370955161564 bits

Two further guarantees are worth as much as the table:

  1. sizeof(char) is 1. Every other size is measured in units of it.
  2. The sizes are ordered and the ranges nest: char no larger than short, no larger than int, no larger than long, no larger than long long. So a short value always fits in an int, whatever the machine.

Notice what is not guaranteed. int is not four bytes. float is not four bytes. Nothing says int and long differ. On a small embedded processor int really is two bytes and the minimum above really is the range.

What this machine actually has

Ask the machine. That is what sizeof is for: an operator, not a function, that gives the size of a type or of an object in bytes.

#include <stdio.h>

int main(void)
{
    printf("type          bytes\n");
    printf("char          %5zu\n", sizeof(char));
    printf("short         %5zu\n", sizeof(short));
    printf("int           %5zu\n", sizeof(int));
    printf("long          %5zu\n", sizeof(long));
    printf("long long     %5zu\n", sizeof(long long));
    printf("float         %5zu\n", sizeof(float));
    printf("double        %5zu\n", sizeof(double));
    printf("long double   %5zu\n", sizeof(long double));
    printf("void *        %5zu\n", sizeof(void *));
    return 0;
}
type          bytes
char              1
short             2
int               4
long              8
long long         8
float             4
double            8
long double      16
void *            8

Those are the sizes on the 64-bit Linux machine that compiled this book's programs. Four of them differ elsewhere, and these are the four worth knowing about:

  • long is 4 bytes on 64-bit Windows, not 8. Code that assumes a long can hold a large count is not portable between Linux and Windows. long long is 8 everywhere that has it.
  • long double is 16 bytes here, 8 on 64-bit Windows with MSVC, and 12 on 32-bit Linux. It is the least portable type in C.
  • void * is 8 bytes on any 64-bit machine and 4 on a 32-bit one. It is the size of an address, so it tracks the machine, not the standard.
  • int is 4 bytes on every desktop machine and 2 on many microcontrollers.
munotes.in40

Data Types and Their Sizes

%zu is the conversion for a value of type size_t, which is what sizeof produces. Using %d for it is a mistake that -Wall catches and that a great many textbooks make.

Signed against unsigned, and the trap in it

An unsigned type cannot hold a negative value. What it does instead is the thing to learn, because it is defined behaviour and it surprises everybody once.

#include <stdio.h>

int main(void)
{
    unsigned int u = 0;
    int i = 0;

    printf("as unsigned, 0 - 1 is %u\n", u - 1);
    printf("as signed,   0 - 1 is %d\n", i - 1);
    return 0;
}
as unsigned, 0 - 1 is 4294967295
as signed,   0 - 1 is -1

Subtracting one from an unsigned zero does not give minus one and it is not an error. Unsigned arithmetic wraps round: the result is reduced modulo one more than the largest value the type can hold, so it comes out as the largest value. This is exactly why a loop written as

for (unsigned int i = n; i >= 0; i--)

never ends. An unsigned i is always at least zero, so the condition is always true.

Use int for counting unless you have a reason not to. unsigned is for bit patterns, for sizes returned by the library, and for the one case where you genuinely need the extra range.

The floating-point types, and why double

float and double hold approximations. They are stored as a sign, a set of significant digits and an exponent, in the manner of scientific notation, which means a fixed number of significant digits and not a fixed number of decimal places.

The number of decimal digits you can rely on is in <float.h> as FLT_DIG and DBL_DIG. On the machine that ran this book's programs those are 6 and 15, and those two values are typical of every desktop machine, because both types follow the same international format for floating-point arithmetic.

Six digits is not many. Here is what that costs, in a program that is not buggy:

#include <stdio.h>

int main(void)
{
    float f = 123456789.0f;
    double d = 123456789.0;

    printf("as a float  : %.1f\n", (double) f);
    printf("as a double : %.1f\n", d);
    printf("0.1 + 0.2 as a double, to 20 places: %.20f\n", 0.1 + 0.2);
    return 0;
}
as a float  : 123456792.0
as a double : 123456789.0
0.1 + 0.2 as a double, to 20 places: 0.30000000000000004441

The float did not store 123456789. It stored the nearest value it is able to represent, and the last two digits are not the ones that were written. Nothing went wrong: a float has about seven significant decimal digits and was asked for nine.

munotes.in41

Data Types and Their Sizes

Try the same program with 1234567.0f and the float prints it back perfectly. Seven digits happens to be inside what a float can do, which is exactly why this trap is hard to spot: it appears only once the numbers grow, and by then the program is finished and trusted.

And 0.1 + 0.2 is not 0.3, in any language, on any machine that uses binary floating point, because one tenth cannot be written exactly in binary any more than one third can be written exactly in decimal.

The rule that follows is absolute and it is worth marks: never compare two floating-point values with ==. Compare the size of their difference against a small tolerance instead. Chapter 16 gives the form.

Use double, not float. A float saves four bytes, which no program in this course will notice, and costs nine significant digits. float exists for very large arrays and for hardware that is faster at it.

The complete picture

DeclarationTypical size hereWhat it is for
char1One character, or a very small integer
signed char1A small integer, definitely signed
unsigned char1A byte, 0 to 255
short2A small integer
unsigned short2A small non-negative integer
int4The default whole number
unsigned int4Non-negative, or a bit pattern
long8 here, 4 on WindowsA larger integer
long long8A large integer, everywhere
float4Fractions, about 6 digits
double8Fractions, about 15 digits. The default
long double16 hereFractions, more digits, least portable
voidnoneNothing: no value, no parameters, or an untyped pointer

The conversion to print each one

Getting this wrong is undefined behaviour, not a small mistake, because printf reads its arguments according to what the format string says is there. -Wall catches nearly all of it, which is the best argument for compiling with it.

TypeConversion
char as a character%c
char as a number%d
int%d
unsigned int%u
short%hd
long%ld
long long%lld
unsigned long%lu
float%f, and the value is promoted to double
double%f, or %g, or %e
long double%Lf
size_t, from sizeof%zu
a pointer%p

There is no conversion for float in printf, and it is not an omission: a float argument is automatically promoted to double when passed, so %f is correct for both. In scanf, which takes addresses and promotes nothing, %f is for a float and %lf is for a double . Confusing the two in scanf is a real and common bug.

munotes.in42

Data Types and Their Sizes

What this does NOT mean

int is not four bytes. It is four bytes on the machines you will use. The language says at least two. Write sizeof(int) when you need the number and never the literal 4.

sizeof is not a function. It is an operator and it is evaluated by the compiler, not at run time. sizeof x without brackets is legal for an object; brackets are required for a type name.

char is not "the character type" and nothing else. It is a one-byte integer, and whether it is signed is up to the implementation. When you want a small number use signed char or unsigned char and say which. When you want a character, use char.

unsigned does not mean positive. It means non-negative, and it means arithmetic that wraps rather than going negative.

float is not "less accurate for big numbers only". It has about six significant digits at every magnitude. float can hold 3.4 times ten to the thirty-eighth; it just cannot tell that number from its neighbours.

A type is not a promise about the value. Declaring int age does not stop anyone storing -500 in it. Checking the value is your job, and it is part of what integrity meant in chapter 6.

Quick revision

  • Basic types: char, int, float, double, plus void.
  • Modifiers: short, long, signed, unsigned.
  • sizeof(char) is 1 by definition; everything else is measured in those units.
  • The standard fixes minimum ranges, not sizes: int at least -32767 to +32767.
  • Sizes never shrink going char, short, int, long, long long.
  • Here: char 1, short 2, int 4, long 8, long long 8, float 4, double 8, pointer 8.
  • long is 4 bytes on 64-bit Windows. long long is 8 everywhere.
  • Unsigned arithmetic wraps modulo the range; 0u - 1 is the maximum value.
  • float about 6 significant digits, double about 15. Use double.
  • Never compare floating-point values with ==.
  • sizeof gives a size_t; print it with %zu.

Test yourself

1. What is the size of int in C?

The standard does not say. It requires a range of at least -32767 to +32767, so at least two bytes, and it is four bytes on every desktop machine. sizeof(int) is the only reliable answer for a given machine.

2. Which is bigger, int or long?

long is never smaller than int, and on many machines they are the same size. sizeof(long) >= sizeof(int) is guaranteed; > is not.

3. What does this print, and why?

unsigned int u = 5;
printf("%u\n", u - 10);

A very large number, the maximum for unsigned int minus four, which is 4294967291 where unsigned int is 32 bits. Unsigned arithmetic is reduced modulo one more than the maximum, so it cannot produce a negative result.

munotes.in43

Data Types and Their Sizes

4. Why should you not write if (x == 0.1) when x is a double?

Because 0.1 has no exact binary representation, so the value stored is the nearest one that does, and any arithmetic that produced x may have landed on a different nearest value. Test if (fabs(x - 0.1) < 1e-9) instead.

5. How many bytes does char occupy, and how do you know?

One, by definition. sizeof(char) is 1 in every conforming implementation, because the byte is the unit sizeof measures in.

6. Give the right printf conversion for sizeof(int), and say why %d is wrong.

%zu. sizeof produces a size_t, which is an unsigned type that may be wider than int, so %d tells printf to read the wrong number of bytes as the wrong kind of value.

What can be asked on this, and how to answer it

"Explain the data types in C with their sizes and ranges." Give the four basic types and the four modifiers, then the table of typical sizes, and then the sentence that earns the extra marks: the standard fixes minimum ranges rather than sizes, so the correct way to get a size is sizeof. Give int as 2 bytes minimum and 4 in practice, and you have answered both versions of the question.

"What is the difference between signed and unsigned?" A signed type spends one bit on the sign and can hold negative values; an unsigned type of the same width holds only non-negative values and reaches roughly twice as far. Add that unsigned arithmetic wraps modulo the range rather than going negative, and give 0u - 1 as the example.

"Differentiate between float and double." Both hold approximations. float is typically 4 bytes with about 6 significant decimal digits; double is typically 8 with about 15. double is the default for floating-point constants and the type to prefer. float is used to save space in large arrays.

"What is the use of the void data type?" Three uses: the return type of a function that returns no value, the parameter list of a function that takes no arguments, and void *, a pointer to an object of unspecified type.

"What is sizeof? Write a program to find the size of each data type." sizeof is a compile-time operator giving the size in bytes of a type or object, with sizeof(char) equal to 1 by definition. Then give the program in this chapter, and remember %zu.

munotes.in44

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Report or request
Done!