munotes®

Programming with C Notes | B.Sc. (Information Technology) Semester 1 | Mumbai University | munotes

Official Notes munotes.in

Programming with C

B.SC. (INFORMATION TECHNOLOGY) · SEMESTER 1

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Information Technology)

For B.Sc. (Information Technology) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in First Year

Programming with C

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I Introduction and Type of operators

  1. What a Program Is, and What an Algorithm Is 1
  2. Pseudocode Statements and Flowchart Symbols 6
  3. The History of C, and Which C This Book Teaches 11
  4. The Structure of a C Program 14
  5. From Source to Running Program: Preprocessor, Compiler, Linker 19
  6. Program Characteristics, and Which Ones Are Desirable 25
  7. The C Character Set 30
  8. Identifiers and Keywords 35
  9. Data Types and Their Sizes 39
  10. Constants and Their Types 45
  11. Variables: Declaration, Definition and Initialisation 50
  12. Characters and Character Strings 55
  13. typedef 60
  14. Type Conversion and Typecasting 64
  15. Arithmetic Operators 69
  16. Relational and Logical Operators 74
  17. Increment and Decrement Operators 80
  18. Assignment Operators and Expressions 85
  19. The Conditional Operator 90
  20. Precedence and Order of Evaluation 94
  21. Block Structure and Initialization 99
  22. The C Preprocessor 104

Module II Control Flow, Functions, Pointers and User-defined data types

  1. Statements and Blocks 109
  2. if and if-else 113
  3. else-if Ladders 117
  4. switch 122
  5. while Loops 127
  6. for Loops 132
  7. do-while 137
  8. break and continue 141
  9. goto and Labels 146
  10. What a Function Is: Definition, Call, Arguments, Return 150
  11. User-Defined Functions 155
  12. Library Functions 159
  13. Recursion 164
  14. Arrays 169
  15. Strings as Arrays, and the String Library 175
  16. Two-Dimensional Arrays and Matrices 182
  17. Pointers and Addresses 188
  18. Pointers as Function Arguments: Call by Value and Call by Reference 194
  19. Pointers and Arrays 199
  20. Structures 205
  21. Unions 211
  22. Putting It Together: A Menu-Driven Program 216
munotes.in

Module I

Introduction and Type of operators

munotes.in

Chapter One

What a Program Is, and What an Algorithm Is

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics."

In one line

An algorithm is a finite list of unambiguous steps that turns its inputs, if it has any, into the wanted output, and a program is that algorithm written in a language a computer can be made to follow.

Why you start here and not at the keyboard

A beginner's instinct is to open the editor and start typing. It is the wrong instinct, and the reason is worth understanding rather than being told.

A computer does exactly what it is told, in the order it is told, and it has no idea what you meant. Every part of a problem you have not thought through is a part the machine will get wrong at full speed. So the thinking is done first, in a form you can check with a pencil, and only then translated into C.

That form is the algorithm. It is not a lesser version of the program. It is the part where the problem is actually solved. The C is the part where the solution is written down for a machine.

Your University asks for it in exactly this way. All three parts of the very first practical in Major Practical 1 ask for more than a program. Each one ends "Write algorithm & draw flowchart for the same." The algorithm carries marks of its own.

What makes a list of steps an algorithm

Not every list of instructions qualifies. Five properties are required, and each one rules out a specific way of getting it wrong.

1. Finiteness. The algorithm must stop, after a finite number of steps, for every input it accepts.

A procedure that says "keep adding 1 to n and print it" is not an algorithm. It never ends. This is the property a beginner breaks most often, and it has a name in the wild: the infinite loop.

2. Definiteness. Every step must mean exactly one thing.

"Take a suitable value of interest" is not a step. Suitable to whom? "Set the rate to 7.5" is a step. If two competent people can read your step and do different things, it is not definite.

3. Input. The algorithm takes zero or more inputs, and where they come from is stated.

Zero is allowed. An algorithm that prints the first ten natural numbers needs no input at all. What is not allowed is a step that silently assumes a value nobody supplied.

4. Output. It produces at least one output, and that output is the answer to the problem.

An algorithm that computes the interest correctly and never displays it has failed. This sounds pedantic until the first time a student loses marks for a program that calculates and does not print.

munotes.in1

What a Program Is, and What an Algorithm Is

5. Effectiveness. Every step must be basic enough that a person could carry it out exactly, with a pencil, in a finite time.

"Set X to the exact decimal value of one third" is not such a step. It is perfectly definite, and you know exactly what is meant, but the digits never end, so nobody can finish performing it. "Set X to 0.3333, correct to four decimal places" is a step. The test is not whether the instruction can be described; it is whether it can be performed.

A useful way to remember the five: an algorithm must stop, must be unambiguous, must know what it is given, must say what it found, and must be made of steps a human could actually perform.

Writing one out: simple interest

This is the first program on your practical list, so it is the first algorithm here.

To calculate simple interest taking principal, rate of interest and number of years as input from user.

The formula is the easy part. Simple interest is SI = (P R T) / 100, where P is the principal, R is the annual rate as a percentage, and T is the time in years.

The algorithm is written as numbered steps, with a Start and a Stop, and nothing else:

Step 1: Start
Step 2: Read P, R, T from the user
Step 3: Set SI = (P * R * T) / 100
Step 4: Display SI
Step 5: Stop

Read it against the five properties. It stops at step 5, so it is finite. Every step means one thing. Step 2 names its inputs. Step 4 produces the output. Every step could be done by hand. It is an algorithm.

Notice what is not in it: no mention of C, no printf, no data types, no semicolons. An algorithm is language independent, and that is the point of writing one. The same five steps become a C program, a Python program or a set of instructions to a clerk with a calculator.

Tracing it

An algorithm is checked by tracing: choosing values, walking the steps in order, and writing down what every quantity holds after each one. A trace table is how you prove to yourself, and to an examiner, that the thing works before a compiler is involved.

Take P as Rs 12,000, R as 8.5 per cent and T as 3 years.

StepPRTSIWhat is displayed
1. Start----
2. Read P, R, T120008.53-
3. SI = (P R T) / 100120008.533060
4. Display SI120008.5330603060
5. Stop
munotes.in2

What a Program Is, and What an Algorithm Is

The arithmetic: 12000 times 8.5 is 1,02,000, times 3 is 3,06,000, divided by 100 is Rs 3,060. A trace table with one row per step, and one column per quantity, is a habit worth forming now. In the loop chapters it stops being a formality and becomes the only reliable way to find a bug.

A second one, where a decision appears: the greatest of three numbers

Write a program to find greatest of three numbers using conditional operator.

Her practical names the C construct she wants, and the chapter on the conditional operator writes that program. The algorithm below names no construct at all, because an algorithm never does, and it is the same algorithm whichever one you finally reach for.

Simple interest had no decision in it. Most problems do. The moment a step depends on a comparison, the algorithm branches.

Step 1: Start
Step 2: Read A, B, C
Step 3: If A > B and A > C, then set MAX = A
Step 4: Otherwise, if B > C, then set MAX = B
Step 5: Otherwise, set MAX = C
Step 6: Display MAX
Step 7: Stop

Steps 3, 4 and 5 are one decision with three outcomes, not three separate decisions. Exactly one of them runs, and step 5 can afford to carry no condition at all precisely because steps 3 and 4 have already failed by the time it is reached.

Write them instead as three independent tests and that stops being true. Every one of them would then have to carry a condition of its own, including the last, or MAX would be overwritten with C on every run whatever steps 3 and 4 decided. All three would also be evaluated even after the answer was known.

Trace it with A as 14, B as 27 and C as 9. Step 3 asks whether 14 is greater than both 27 and 9; it is not, so MAX is not set to A. Step 4 asks whether 27 is greater than 9; it is, so MAX becomes 27. Step 5 is skipped because step 4 ran. Step 6 displays 27, which is correct.

Now trace it with all three equal, say 5, 5 and 5. Step 3 fails, because 5 is not greater than 5. Step 4 fails for the same reason. Step 5 runs and MAX becomes C, which is 5. Correct, but only just: the algorithm reaches the right answer through its last resort. Testing the equal case, the all-negative case and the two-equal case is what separates an algorithm that works from one that has only been tried once.

munotes.in3

What a Program Is, and What an Algorithm Is

A third one, where the rule itself is the difficulty: the leap year

Write a program to check if the year entered is leap year or not.

Here the coding is trivial and the rule is the whole problem. Most students remember "divisible by 4" and stop. The full rule, as fixed by the Gregorian calendar, has three parts:

  • a year divisible by 4 is a leap year,
  • except that a year divisible by 100 is not,
  • except that a year divisible by 400 is.

So 2024 is a leap year, 1900 was not, and 2000 was. Any algorithm that gets 1900 and 2000 both right has the rule; one that gets either wrong does not.

Step 1: Start
Step 2: Read Y
Step 3: If Y is divisible by 400, then set R = "Leap year"
Step 4: Otherwise, if Y is divisible by 100, then set R = "Not a leap year"
Step 5: Otherwise, if Y is divisible by 4, then set R = "Leap year"
Step 6: Otherwise, set R = "Not a leap year"
Step 7: Display R
Step 8: Stop

This has the same shape as the greatest of three: one decision with four outcomes, exactly one of which runs, and then a single Display that every path reaches. An algorithm with one way out is easier to check and easier to draw than one that finishes in four different places.

The order of the tests is the algorithm. Test 400 first, then 100, then 4, and each test can be simple because the ones before it have already removed the exceptions. Reverse the order and every test needs conditions bolted onto it.

Check it on the years that settle whether a rule is right:

YDivisible by 400?Divisible by 100?Divisible by 4?Displayed
2000yesnot reachednot reachedLeap year
1900noyesnot reachedNot a leap year
2024nonoyesLeap year
2023nononoNot a leap year

From algorithm to program

Once the algorithm is right, writing the C is mechanical, and the rest of this book is how. The order never changes:

  1. Understand the problem, and write down what is given and what is wanted.
  2. Write the algorithm, in steps, in ordinary words.
  3. Trace it by hand on values you have chosen to be awkward.
  4. Only then, translate it into C.
  5. Compile it, run it, and check it against the trace you already did.

Step 3 is the one that gets skipped, and skipping it is why a program that compiles cleanly still prints the wrong number. The compiler checks your grammar. Nothing but a trace checks your thinking.

munotes.in4

What a Program Is, and What an Algorithm Is

What an examiner can ask on this, and how to answer it

"Define an algorithm and state its characteristics." Give the one-line definition, then the five properties with a word of explanation each. A bare list of five nouns is a weak answer; each property with the failure it prevents is a full one.

"Write an algorithm to ..." Number the steps, open with Start and close with Stop, name your inputs in a Read step, and finish with a Display step. Keep the steps in ordinary words: an algorithm written in C statements is no longer language independent, which is the one thing it exists to demonstrate.

"What is the difference between an algorithm and a program?" An algorithm is a finite sequence of unambiguous steps solving a problem, written in ordinary language, independent of any programming language and not executable by a machine. A program is that algorithm expressed in a particular programming language, following that language's grammar, and executable once translated. One is the solution; the other is the solution written down for a computer.

A trace question. You may be given a short algorithm and asked what it displays for stated inputs. Build the trace table. Do not read the algorithm and guess.

Contents This chapter on its own page

munotes.in5

Chapter Two

Pseudocode Statements and Flowchart Symbols

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics."

In one line

Pseudocode writes an algorithm in structured, language-free statements that look like code without being code, and a flowchart draws the same algorithm as boxes joined by arrows, so that its path through every decision can be seen at a glance.

Why the same algorithm gets written three ways

Chapter 1 wrote algorithms as numbered steps in ordinary English. That is the loosest of the three notations and the easiest to write. It has one weakness: as soon as an algorithm has decisions inside loops, plain numbered steps stop showing you the shape of the thing, and you have to hold the structure in your head.

Pseudocode fixes that by borrowing the structure of a programming language, and nothing else. It has an IF that visibly ends, a loop whose body is visibly indented, and no semicolons, no data types and no library to remember.

A flowchart fixes it a different way, by making the structure visible instead of readable. Every path the program can take is a path your finger can follow.

Neither replaces the other, and your University asks for both. All three parts of her first practical end with the same instruction:

Write algorithm & draw flowchart for the same.

So the working order for this subject is: steps in English, then pseudocode when the logic gets complicated, then the flowchart, then the C.

Pseudocode

There is no official pseudocode, and that is the point

No standards body defines it. That sounds like a weakness and is actually the whole idea: pseudocode exists so that a human reader understands the algorithm without knowing any particular language. What matters is that it is consistent, indented, and unambiguous.

What follows is the ordinary academic convention, and it is the one to use in an answer book.

The statements

Input and output. READ or INPUT takes a value from the user. PRINT or DISPLAY shows one.

READ P, R, T
PRINT "Simple interest is", SI

Assignment. A value is put into a name. Write it as an arrow or as a plain equals sign, and then do not mix the two in one answer.

SET SI = (P * R * T) / 100

Decision. IF, with an optional ELSE, and always a closing ENDIF. The closing keyword is what makes the extent of the branch visible.

IF marks >= 40 THEN
    PRINT "Pass"
ELSE
    PRINT "Fail"
ENDIF

Multi-way decision. ELSE IF chains, closed once.

IF A > B AND A > C THEN
    SET MAX = A
ELSE IF B > C THEN
    SET MAX = B
ELSE
    SET MAX = C
ENDIF

Loops. WHILE tests before the body, REPEAT tests after it, and FOR counts. Two operators appear here that pseudocode spells out in words: MOD is the remainder after division, and DIV is division that throws the fraction away, so 27 MOD 10 is 7 and 27 DIV 10 is 2.

munotes.in6

Pseudocode Statements and Flowchart Symbols

WHILE n > 0 DO
    SET digit = n MOD 10
    SET n = n DIV 10
ENDWHILE

FOR i = 1 TO 10 DO
    PRINT i
ENDFOR

Modules. A named piece of work called from elsewhere.

CALL swap(x, y)

Two conventions carry most of the weight. Keywords go in capitals so the structure can be seen without reading the words. The body of every IF and every loop is indented, and the matching ENDIF or ENDWHILE sits at the same indentation as the keyword that opened it.

The same problem, in steps and in pseudocode

Chapter 1 wrote the simple interest algorithm as five numbered steps. Here it is as pseudocode:

BEGIN
    READ P, R, T
    SET SI = (P * R * T) / 100
    PRINT SI
END

For a problem this small the two notations are equally clear, and the numbered steps are shorter. The difference appears the moment there is a decision, because pseudocode shows you where the branch closes and numbered steps only tell you.

Flowcharts

The symbols

A flowchart is drawn with a fixed set of shapes, and the shape carries the meaning. Using a rectangle where a diamond belongs is a mistake even when the words inside it are right.

SymbolShapeWhat it meansExample
TerminalOval, or a rectangle with rounded endsWhere the flowchart starts and where it stopsSTART, STOP
Input / OutputParallelogramA value is read from the user, or shown to themREAD P, R, T
ProcessRectangleA calculation, or a value put into a nameSI = (PRT)/100
DecisionDiamondA question with two answers, and two arrows outIs A > B?
Predefined processRectangle with a double bar down each sideWork done by a module defined elsewhereCALL swap(x, y)
ConnectorSmall circle with a letter in itJoins two points on the same page without a long lineA
Off-page connectorFive-sided tagContinues the chart on another page1
Flow lineArrowThe order the boxes are carried out in

Two shapes do the work in nearly every chart a first-year student draws: the rectangle for doing something, and the diamond for asking something.

The rules

  1. Every flowchart begins with exactly one START terminal and ends with a STOP terminal.
  2. Arrows carry the flow, and every arrow has a head. A line without an arrowhead does not say which way the work goes.
  3. A decision diamond has one arrow in and exactly two arrows out, and both are labelled, normally Yes and No.
  4. Every other symbol has one arrow in and one arrow out.
  5. The chart flows top to bottom and left to right, unless a loop takes it back up.
  6. Where two paths finish, they join before the chart carries on, so that the flow after them is drawn once and not twice.
munotes.in7

Pseudocode Statements and Flowchart Symbols

Rule 3 is the one that is broken most often. A diamond with three arrows out is not a decision, it is two decisions that have been drawn on top of each other.

How these are drawn here

The charts below are drawn in text so they read on a phone and in a printout. The shapes stand in for the real ones like this, and in your journal and your answer book you draw the shapes from the table above, not these:

Drawn here asMeans
( ... )Terminal, an oval
/ ... /Input or output, a parallelogram
[ ... ]Process, a rectangle
< ... >Decision, a diamond
A vertical bar, and -->Flow lines

Flowchart 1: simple interest

Her practical, part (a):

To calculate simple interest taking principal, rate of interest and number of years as input from user.

        (  START  )
             |
             v
     / READ P, R, T /
             |
             v
   [ SI = (P * R * T) / 100 ]
             |
             v
      / DISPLAY SI /
             |
             v
        (  STOP  )

There is no decision in it, so the chart is a straight line. Read it against the rules: one START, one STOP, every symbol has one arrow in and one out, and the two parallelograms are input and output rather than processes.

Flowchart 2: the greatest of three numbers

Her practical, part (b):

Write a program to find greatest of three numbers using conditional operator.

Now there are decisions, and the chart earns its keep.

              (  START  )
                   |
                   v
           / READ A, B, C /
                   |
                   v
        < A > B AND A > C ? >--- Yes -->[ MAX = A ]
                   |                          |
                   No                         |
                   |                          |
                   v                          |
           < B > C ? >--- Yes -->[ MAX = B ]  |
                   |                    |     |
                   No                   |     |
                   |                    |     |
                   v                    |     |
             [ MAX = C ]                |     |
                   |                    |     |
                   +<-------------------+<----+
                   |
                   v
           / DISPLAY MAX /
                   |
                   v
              (  STOP  )

The three paths set MAX and then join, and the join is the part beginners leave out. DISPLAY MAX is drawn once, not three times, because whichever branch ran, the work after it is the same.

Follow A as 14, B as 27, C as 9 with your finger. The first diamond asks whether 14 is greater than both 27 and 9, and the answer is No, so you go down. The second asks whether 27 is greater than 9, and the answer is Yes, so you go right into MAX = B. Then you come back to the join, display 27, and stop.

munotes.in8

Pseudocode Statements and Flowchart Symbols

Flowchart 3: the leap year

Her practical, part (c):

Write a program to check if the year entered is leap year or not.

           (  START  )
                |
                v
           / READ Y /
                |
                v
  < Y divisible by 400 ? >-- Yes -->[ R = "Leap year" ]
                |                            |
                No                           |
                |                            |
                v                            |
  < Y divisible by 100 ? >-- Yes -->[ R = "Not a leap year" ]
                |                            |
                No                           |
                |                            |
                v                            |
  < Y divisible by 4 ? >---- Yes -->[ R = "Leap year" ]
                |                            |
                No                           |
                |                            |
                v                            |
     [ R = "Not a leap year" ]               |
                |                            |
                +<---------------------------+
                |
                v
            / DISPLAY R /
                |
                v
           (  STOP  )

Three diamonds in a chain, each with its two labelled exits, and one join at the end. Every process box holds an action, R = "Leap year", and not a bare piece of text. A box containing only the words Leap year would not say what the flowchart is supposed to do with them. Chapter 1 showed why the order has to be 400, then 100, then 4; the chart shows the same fact as a shape, because each No arrow leads into a test that no longer has to worry about the case above it.

The mistakes that cost marks

  • A decision with one exit, or three. Two exits, both labelled.
  • Unlabelled arrows out of a diamond. The reader cannot tell which branch is which, and neither can you a week later.
  • No arrowheads. A flowchart without arrowheads is a picture of some boxes.
  • A rectangle used for input. READ and PRINT are parallelograms. This is the commonest shape error.
  • Branches that never join. Two copies of the same ending drawn under two branches is a sign the join was forgotten.
  • C code inside the boxes. printf("%d", si); in a process box defeats the purpose. Write DISPLAY SI.
  • Pseudocode with no closing keyword. An IF without an ENDIF leaves the extent of the branch to the reader's guess.

What can be asked on this, and how to answer it

"Draw the flowchart symbols and state their use." Draw each shape, name it, and give one example of what goes inside it. A named shape with no example is half an answer.

munotes.in9

Pseudocode Statements and Flowchart Symbols

"Write algorithm & draw flowchart for the same." This is her own wording, and both halves are wanted. They must also agree with each other. Write the algorithm first, then draw the chart from it, then check that every step appears in the chart and every box appears in the algorithm.

"What is pseudocode? How does it differ from an algorithm and from a program?" Pseudocode is a structured, language-free way of writing an algorithm, using keywords and indentation but no particular language's grammar. An algorithm may be written in plain English or as pseudocode, so pseudocode is one form an algorithm can take. A program is the algorithm in a real language, which a compiler can translate and a machine can run.

A trace question on a flowchart. You may be given a chart and asked what it prints. Put your finger at START and walk it, writing down each quantity as it changes, exactly as chapter 1 traced a table.

Contents This chapter on its own page

munotes.in10

Chapter Three

The History of C, and Which C This Book Teaches

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics."

In one line

C was written at Bell Labs between 1969 and 1973 by Dennis Ritchie, out of an older language called B, to write the Unix operating system in, and it has been an international standard since 1990, revised four times since.

Where it came from

C did not arrive from nowhere. It is the third language in a short family line, and each step in that line was a response to a shortage.

BCPL. Martin Richards designed BCPL in the mid-1960s while visiting MIT. It was typeless: every value was one machine word, and what that word meant was the programmer's business.

B. Ken Thompson created B in 1969 and 1970, derived directly from Richards's BCPL, for the small PDP-7 machine that the first Unix ran on. B kept BCPL's typelessness, and on a machine that addressed bytes rather than words, that became a real problem. A language in which everything is a word cannot conveniently talk about a character.

C. In Ritchie's own words, C "came into being in the years 1969-1973, in parallel with the early development of the Unix operating system", and "the most creative period occurred during 1972". What C added to B is the thing this whole course rests on: a type structure. Ritchie describes C as derived from the typeless BCPL and having "evolved a type structure". By early 1973, he writes, "the essentials of modern C were complete".

The purpose is worth holding on to, because it explains almost every design decision in the language you are about to learn. C was built to write an operating system. It had to be close enough to the machine to control it, and portable enough to move that operating system to a different machine. That is why C gives you pointers and direct access to memory, and also why it does so little for you automatically.

How it became a standard

For about a decade, C was defined by a book rather than by a standard. The first widely available description was The C Programming Language by Brian Kernighan and Dennis Ritchie, published in 1978, and universally called K&R. For years, "what K&R says" was the definition of C.

That stopped being enough as C spread to machines and compilers Bell Labs had nothing to do with. In Ritchie's account, "in the middle 1980s, the language was officially standardized by the ANSI X3J11 committee, which made further changes." The ANSI standard was then adopted internationally, and C has been maintained as an ISO standard ever since.

Every revision since is recorded in the standard's own change history, which identifies each edition by the value the compiler must give the macro __STDC_VERSION__:

munotes.in11

The History of C, and Which C This Book Teaches

Edition__STDC_VERSION__Commonly called
First edition, Amendment 1199409LC95
Second edition199901LC99
Third edition201112LC11
Fourth edition201710LC17
Fifth edition202311LC23

The current published standard is the fifth edition, ISO/IEC 9899:2024, which in the standard's own words "cancels and replaces the fourth edition (ISO/IEC 9899:2018), which has been technically revised."

The names in the last column are what everybody actually says. They are the year, not the edition number, and the fourth edition is called C17 although it was published as 9899:2018, because 2017 was the year its content was fixed.

Which C this book teaches, and why it matters to you

This book teaches C17, the fourth edition, and every program in it is compiled with warnings turned on:

cc -std=c17 -Wall -Wextra program.c -o program

Your prescribed first text is Kernighan and Ritchie, and you should read it, but you should know what it is. The book predates the first standard. Its second edition was rewritten for ANSI C, which is the first edition of the standard, and even that is now four revisions old. Almost everything in it is still exactly right. A few things are not, and they are the things that will confuse you if nobody warns you:

  • Old-style function definitions, where the parameter types are listed on separate lines after the parameter names, are gone. The fifth edition removed them outright. Write the modern form, which chapter 32 gives.
  • main() with an empty parameter list is not the same as main(void) in older C. Write main(void) when the program takes no arguments.
  • Declaring your variables at the top of a function was required by the first edition and has not been since the second. Declare a variable where you first need it.
  • // comments were not in the first edition. They are in every edition since.

You can ask your own compiler which standard it is using, because the standard requires it to say:

#include <stdio.h>

int main(void)
{
    printf("__STDC_VERSION__ is %ld\n", __STDC_VERSION__);
    return 0;
}

Compiled with -std=c17, this prints:

__STDC_VERSION__ is 201710

Notice that the printed number has no L on the end, although the table above writes 201710L. Both are right. The macro is defined as the constant 201710L, where the L tells the compiler to treat it as a long, and the value of that constant is the number 201710. The suffix is part of how the constant is written in the source, not part of the number. Constants and their suffixes are chapter 10.

Compile the same file with -std=c99 and it prints 199901 instead. The program has not changed; the rules the compiler is applying to it have.

munotes.in12

The History of C, and Which C This Book Teaches

Why C is still taught first

It would be fair to ask why a language designed for a 1970s minicomputer is the first one on your degree. Three reasons, and none of them are sentimental.

It is small. The whole language fits in a book. There is very little to memorise, and what there is, you have met by the end of this semester.

It hides nothing. A variable is a place in memory, an array is a run of them, and a pointer is the address of one. Languages that came later hide all of that, which is a kindness until something goes wrong. If you learn C first, the later languages feel like conveniences rather than magic.

Nearly everything is built on it. Operating systems, databases, language runtimes and embedded firmware are largely written in C or in languages that borrowed its syntax wholesale. The for loop and the curly brace you are about to learn are the same ones you will meet in C++, Java, C#, JavaScript, PHP and Go.

What can be asked on this, and how to answer it

"Write a short note on the history of C." Give the family line in order, with dates and people: BCPL by Martin Richards in the mid-1960s, B by Ken Thompson in 1969 to 1970, and C by Dennis Ritchie at Bell Labs between 1969 and 1973, developed alongside Unix. Then the standardisation: K&R in 1978, ANSI X3J11 in the mid-1980s, and the ISO revisions since. Say what C added to B, which is the type structure, because that is the part that answers "why a new language at all".

"Why was C developed?" To write an operating system in. It had to be low-level enough to control the machine and portable enough to move Unix to another one, and both of those requirements are visible in the language today.

"What is the difference between C and its predecessors?" B and BCPL were typeless, treating every value as one machine word. C introduced data types, which let the language describe characters, integers of several sizes and floating-point numbers, and let the compiler check what you are doing with them.

Contents This chapter on its own page

munotes.in13

Chapter Four

The Structure of a C Program

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

A C program is a file of text made of preprocessor directives and function definitions, exactly one of which must be a function called main, and execution starts at the first statement inside main and stops when main returns.

Why a program has a fixed shape at all

A compiler is not a reader. It cannot tell from the sense of what you wrote where your program begins, so the language fixes that in advance: the starting point is always the function named main. Nothing else will do, and no other name works, however sensible.

The same is true of everything else in the shape. A name must be declared before it is used, because the compiler reads your file once from top to bottom and cannot know what is coming. A statement ends with a semicolon, because a newline in C means nothing at all and something has to mark the end. Every rule in this chapter exists because a program has to be read by a machine that does not guess.

The smallest complete C program that does something

#include <stdio.h>

int main(void)
{
    printf("This program did one thing and stopped.\n");
    return 0;
}

Compiled and run, it prints:

This program did one thing and stopped.

Six lines, and every part of the language's required shape is in them. Read them one at a time.

#include <stdio.h> is a preprocessor directive. A line beginning with # is not a C statement at all. It is an instruction to a program that runs before the compiler, and this one says: find the file stdio.h and paste its contents in here. stdio.h is the standard input and output header, and it is what tells the compiler that a function called printf exists and what kind of arguments it takes. Without it the program does not compile cleanly. Chapter 22 is the preprocessor in full.

int main(void) is a function definition header, and it says three things. main is the name, and it is the one name the language reserves for the starting point. (void) says the function takes no arguments. int says it gives back an integer when it finishes, and the operating system reads that integer to find out whether the program worked.

{ and } are braces, and together they make a block: the body of the function. Everything the function does sits between them.

printf("This program did one thing and stopped.\n"); is a statement, which is one complete instruction. It calls a library function to write text to the screen. The \n inside the quotation marks is one character, the newline, and it is what moves the cursor to the next line. The semicolon at the end is part of the statement, not decoration.

munotes.in14

The Structure of a C Program

return 0; ends main and hands the integer 0 back. By long convention 0 means the program succeeded and anything else means it failed.

The six sections MU's answer wants

Written out in full, a C program can have six sections in this order. The smallest program above has only two of them, which is the point: four of the six are optional and two are not.

SectionWhat goes in itRequired?
Documentation sectionComments saying what the file is, who wrote it and whenNo
Link section#include directives, which bring in headersIn practice, yes
Definition section#define directives, which name constantsNo
Global declaration sectionVariables and function declarations visible to the whole fileNo
The main functionThe block execution starts inYes, always
Subprogram sectionThe other functions the program is made ofNo

Only main is compulsory as a matter of language. The link section is listed as required in practice because a program that prints anything needs stdio.h, and a program that needs nothing from the library needs no header at all.

A program with all six sections in it

This is the same simple interest calculation the lab asks for in its very first practical, written so that each of the six sections is present and labelled.

/*
 * simple-interest.c
 * Simple interest on a deposit, with the calculation in its own function.
 * Documentation section: this comment, and nothing else.
 */

#include <stdio.h>             /* link section */

#define PERCENT 100.0          /* definition section */

double interest(double principal, double rate, double years);
int calls = 0;                 /* global declaration section */

int main(void)                 /* the main function */
{
    double si = interest(15000.0, 8.5, 2.0);

    printf("Simple interest: %.2f\n", si);
    printf("Function calls : %d\n", calls);
    return 0;
}

double interest(double principal, double rate, double years)
{                              /* subprogram section */
    calls = calls + 1;
    return principal * rate * years / PERCENT;
}
Simple interest: 2550.00
Function calls : 1

Four things in that file are worth naming now and learning later.

  • / ... / is a comment. The compiler deletes it. It is for the person reading the file, and a comment may run over as many lines as you like.
  • #define PERCENT 100.0 makes the preprocessor replace the word PERCENT with 100.0 everywhere below it. It is not a variable. Chapter 22.
  • double interest(double principal, double rate, double years); with a semicolon and no body is a function declaration, also called a prototype. It promises the compiler that such a function exists further down, so that main can call it before the compiler has read it. Chapter 32.
  • int calls = 0; outside any function is a global variable, reachable from every function in the file. Chapter 21 says why you should nearly always prefer one that is not.
munotes.in15

The Structure of a C Program

Move the definition of interest above main and the declaration becomes unnecessary, because the compiler will have read the real thing by the time it reaches the call. Both arrangements are correct. The declaration is the one that keeps working when the program grows into more than one file.

The rules that are not negotiable

  1. There is exactly one main. No main and the linker fails with an undefined reference. Two and it fails with a duplicate symbol.
  2. C is case sensitive. main, Main and MAIN are three different names, and only the first is the starting point.
  3. Every statement ends with a semicolon. A function definition header and a preprocessor directive do not, and those are the two places beginners put one.
  4. Braces come in pairs, and each { needs its }.
  5. A name must be declared before it is used. The compiler reads the file once, downwards.
  6. A preprocessor directive occupies its own line and begins with #.

What the structure does NOT mean

A newline does not end a statement. C ignores where you press Return. The whole of the first program can be written on one line and it compiles to exactly the same thing. Indentation and line breaks are for the reader, and they are worth a great deal, but the compiler is not reading them.

main is not called by your program. It is called for you when the program starts, by code the linker puts in. This is why nothing in the file appears to call it.

#include <stdio.h> does not make the program bigger by including a library. The header is a text file of declarations. It tells the compiler what printf looks like; the machine code for printf is joined on later by the linker, which is chapter 5.

return 0; is not what prints the output. It is the value handed back to whatever started the program. Leaving it out of main is legal in C99 and later, where reaching the closing brace of main returns 0 anyway, but write it: it says plainly what you meant, and in every function other than main leaving it out is a real bug.

The six sections are not six blocks of a form. They are the places things may go, in order. A real file usually has three of them.

Quick revision

  • A C program is preprocessor directives plus function definitions.
  • Exactly one function must be called main. Execution starts there.
  • int main(void) returns an int and takes no arguments; 0 means success.
  • #include is a directive to the preprocessor, not a C statement, and takes no semicolon.
  • A header such as stdio.h supplies declarations, not machine code.
  • Six sections, in order: documentation, link, definition, global declaration, main, subprogram. Only main is compulsory.
  • Statements end with a semicolon. Braces delimit a block. C is case sensitive.
  • A name must be declared before it is used, which is why prototypes exist.
munotes.in16

The Structure of a C Program

Test yourself

1. Write the smallest C program that compiles, runs and prints nothing.

int main(void)
{
    return 0;
}

It prints nothing at all, so the output block beside it is empty:

It needs no header, because it uses nothing from the library. It compiles clean under -Wall -Wextra and exits with status 0.

2. Name the six sections of a C program in order, and say which are compulsory.

Documentation, link, definition, global declaration, main, subprogram. Only main is compulsory; in practice the link section is present in any program that does input or output.

3. Why must printf be declared before it is used, and where does that declaration come from?

Because the compiler reads the file once and must know the function's return type and parameter types to compile the call correctly. The declaration comes from stdio.h, pasted in by the preprocessor.

4. What is wrong with this?

#include <stdio.h>;

void main(void)
{
    printf("hello\n")
}

Three errors. The directive has a semicolon it must not have. main is declared void, and the standard requires int. The printf statement has no semicolon.

5. Does a comment cost anything at run time?

No. The preprocessor removes comments before the compiler sees the file, so they are not in the finished program at all.

What can be asked on this, and how to answer it

"Explain the structure of a C program." Give the six sections in order with one line each, then draw or write out a short program with the sections labelled, exactly as this chapter does. Say which are compulsory. An answer that lists the six and shows no program is half an answer, and this is the commonest form of the question on this paper.

"What is the purpose of main()?" It is the function where execution begins, and the one function every C program must have. Its return value is reported to the operating system, and 0 conventionally means success.

"What is the use of #include <stdio.h>?" It tells the preprocessor to insert the standard input and output header, which declares the library functions for input and output, printf and scanf among them, so the compiler can check your calls to them.

munotes.in17

The Structure of a C Program

"Distinguish between a declaration and a definition of a function." A declaration gives the name, the return type and the parameter types and ends in a semicolon; it promises the function exists. A definition gives all of that and the body in braces; it is the function. A file may declare a function many times and must define it once.

"Is void main() correct?" Not in standard C. The standard gives main a return type of int. Some older compilers accept void main(), which is exactly why students write it, and it should not be written.

Contents This chapter on its own page

munotes.in18

Chapter Five

From Source to Running Program: Preprocessor, Compiler, Linker

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

Turning a .c file into a program you can run takes four steps in a fixed order, the preprocessor, the compiler, the assembler and the linker, and cc program.c -o program is one command that quietly runs all four.

Why it is four steps and not one

Each step exists because it does a job the others cannot.

The preprocessor works on text and understands no C at all. It pastes headers in, replaces the names you defined, and strips comments. Its whole reason for being is that a program is built out of pieces kept in different files, and something has to assemble the pieces into one stream of text before anybody tries to make sense of it.

The compiler is the part that understands C. It reads the text, checks that it is a legal program, and translates it into instructions for one particular kind of processor.

The assembler turns those instructions from text into the numbers a processor actually reads.

The linker exists because your file is not the whole program. You called printf, and you did not write printf. Somebody compiled it years ago and it sits in a library on your machine. The linker's job is to find every name your file uses but does not contain, fetch the machine code for it, and join everything into one file the operating system can load.

Splitting the work this way is what lets a large program be compiled one file at a time, and it is why a mistake gets caught by a different tool depending on what kind of mistake it is. That last point is worth more marks than any other in this chapter.

The program we will push through all four stages

#include <stdio.h>

#define SIDE 7

int main(void)
{
    printf("Area of a square of side %d is %d\n", SIDE, SIDE * SIDE);
    return 0;
}
Area of a square of side 7 is 49

Nine lines, saved as area.c. Now watch what each stage does to it.

Stage 1: the preprocessor

cc -std=c17 -E area.c > area.i

-E means stop after preprocessing. The result, by convention a .i file, is still C text you can read. Its size is the first surprise:

$ wc -l < area.c
9
$ wc -l < area.i
554

Nine lines became 554. Almost all of that is stdio.h, pasted in whole, bringing the declarations of every standard input and output function with it. The end of the file is your own code, and it has changed:

$ tail -6 area.i
# 5 "area.c"
int main(void)
{
    printf("Area of a square of side %d is %d\n", 7, 7 * 7);
    return 0;
}
munotes.in19

From Source to Running Program: Preprocessor, Compiler, Linker

Three things happened, and each one tells you what the preprocessor is.

  • SIDE became 7. Both times. That is all #define does: replace a name with a piece of text.
  • SIDE SIDE became 7 7, and not 49. The preprocessor does not do arithmetic. It moved text around and stopped. The multiplication is the compiler's job, and in fact the compiler does it before the program ever runs, because both operands are constants.
  • #include and #define are gone, along with the blank line structure, replaced by line markers like # 5 "area.c" that let the compiler report errors against your file rather than against the 554-line stream.

If a comment had been in the file, it would be gone too. The preprocessor removes comments, which is why a comment costs nothing at run time.

Stage 2: the compiler

cc -std=c17 -S area.c -o area.s

-S means stop after compiling, and leave the result as assembly language: text, one instruction per line, for one specific processor. The nine-line program becomes 35 lines of it, beginning like this:

$ sed -n '1,11p' area.s
	.arch armv8-a
	.file	"area.c"
	.text
	.section	.rodata
	.align	3
.LC0:
	.string	"Area of a square of side %d is %d\n"
	.text
	.global	main
	.type	main, %function
main:

You are not expected to read assembly on this paper. Three things in it are worth seeing once.

.arch armv8-a names the processor family. This listing was produced on arm64, and the same program compiled in a lab with an Intel processor produces a different listing that does the same thing. Assembly is not portable; C is. That is the whole argument for a compiled high-level language.

.string "Area of a square of side %d is %d\n" is your text, stored as data. The format string was not turned into instructions, because it is not an instruction. It is a run of characters the program will hand to printf.

main: is a label, and it marks the place where your function's instructions begin. The linker will later look for exactly that name.

Stage 3: the assembler

cc -std=c17 -c area.c -o area.o

-c means compile and assemble, but do not link. The result is an object file, and it is the first thing in this chain you cannot read as text: it holds machine code. What you can read is its table of names, with nm:

$ nm area.o
0000000000000000 T main
                 U printf

Two lines, and they are the whole idea of linking.

  • T main means this file defines a function called main, in its text (code) section, at offset 0. T is for text.
  • U printf means this file uses a function called printf and does not define it. U is for undefined. There is a hole in the machine code where a call to printf should go, and this file cannot fill it.
munotes.in20

From Source to Running Program: Preprocessor, Compiler, Linker

The object file is 1,648 bytes on the machine that produced these numbers. Your own will differ, and the number does not matter. What matters is that it is small, because it holds your nine lines and nothing else.

Stage 4: the linker

cc -std=c17 area.o -o area

Now the hole gets filled. The linker takes your object file, finds printf in the standard C library, copies in the machine code for it and for everything printf itself needs, adds the small piece of start-up code that calls main for you, and writes one executable file.

$ stat -c %s area
70312
$ ./area
Area of a square of side 7 is 49

1,648 bytes of object file became a 70,312 byte program. Your nine lines are a tiny fraction of what you are running, and the rest is library and start-up code you did not write and did not have to.

All four in one command

cc -std=c17 -Wall -Wextra area.c -o area

This is what you will actually type. It runs the preprocessor, the compiler, the assembler and the linker in order, keeps none of the intermediate files, and stops at the first stage that fails. -o area names the output; without it the program is called a.out, which is a habit from the earliest Unix and has stuck for fifty years.

Then run it. The ./ is not decoration: it means "in this directory", and without it the shell looks for area in the standard places and does not find it.

./area

Which stage caught your mistake

This is the useful half of the chapter. An error message tells you which tool is complaining, and therefore what kind of thing is wrong. All five below are real messages from gcc 13.3.

The preprocessor cannot find a header. Misspell stdio.h and the chain stops before any C is looked at:

$ cc -std=c17 e1.c -o e1
e1.c:1:10: fatal error: stdioo.h: No such file or directory
    1 | #include <stdioo.h>
      |          ^~~~~~~~~~

The compiler finds a broken statement. Leave out a semicolon and the compiler complains, and notice where: it names line 5, although the missing semicolon is on line 4. The compiler discovered the problem when it read the next thing that could not follow.

$ cc -std=c17 e2.c -o e2
e2.c: In function 'main':
e2.c:5:5: error: expected ',' or ';' before 'return'
    5 |     return a;
      |     ^~~~~~

The linker cannot find a definition. Declare a function, call it, and never write it. Nothing is wrong with your C, so the compiler is satisfied, and the linker is the one that objects:

munotes.in21

From Source to Running Program: Preprocessor, Compiler, Linker

$ cc -std=c17 e3.c -o e3
/usr/bin/ld: in function `main':
e3.c:(.text+0xc): undefined reference to `twice'
collect2: error: ld returned 1 exit status

The linker cannot find main. Compile a file with functions but no main, and the start-up code is the thing left with a hole:

$ cc -std=c17 e4.c -o e4
(.text+0x1c): undefined reference to `main'
collect2: error: ld returned 1 exit status

A warning is not an error, and matters more. Delete #include <stdio.h> and the program still links and still runs, and gcc tells you that you are calling a function it knows nothing about:

$ cc -std=c17 -Wall -Wextra nohdr.c -o nohdr
nohdr.c:6:5: warning: implicit declaration of function 'printf' [-Wimplicit-function-declaration]
nohdr.c:1:1: note: include '<stdio.h>' or provide a declaration of 'printf'

That is the class of mistake -Wall -Wextra exists to surface. Compile with them always. A program that runs today with a warning is a program that breaks on somebody else's compiler.

SymptomStageTypical cause
No such file or directory on a headerPreprocessorHeader name misspelled or not installed
expected ';', undeclared identifier, type errorsCompilerGrammar or type mistake in your C
undefined reference to 'x'LinkerDeclared or called but never defined, or a library not named
undefined reference to 'main'LinkerNo main in anything being linked
implicit declaration of functionCompiler, as a warningA missing #include
Nothing at all, wrong answerNoneYour logic. No tool can catch this

What this does NOT mean

The compiler does not produce a runnable program. It produces an object file with holes in it. Only the linker produces something you can run. Saying "the compiler turns C into an executable" is the single commonest error on this topic.

A header is not a library. stdio.h is text: declarations telling the compiler what printf looks like. The machine code for printf lives in a compiled library and is joined on by the linker. This is why a missing header gives a compiler warning while a missing library gives a linker error.

The preprocessor is not part of the C language. It has its own grammar, does not understand types, expressions or scope, and would happily replace a name inside something you never meant. Chapter 22 shows what that costs.

An interpreter is not doing this. In an interpreted language there is no separate translation step and no executable. C compiles ahead of time, which is why a C program starts instantly and why you must compile again after every edit.

cc and gcc are not necessarily different programs. On the machine these transcripts came from, cc is gcc 13.3 reached through another name. On a Mac, cc is usually clang. cc is the portable name, which is why this book uses it.

munotes.in22

From Source to Running Program: Preprocessor, Compiler, Linker

Quick revision

  • Four stages in order: preprocessor, compiler, assembler, linker.
  • Preprocessor: text only. Pastes #include, substitutes #define, strips comments. No arithmetic, no type checking. Stop after it with -E.
  • Compiler: checks the C and emits assembly for one processor. Stop after it with -S.
  • Assembler: assembly text into machine code, giving an object file. Stop after it with -c.
  • Linker: joins object files and libraries, resolves every undefined name, adds start-up code, writes the executable.
  • nm on an object file shows T name for defined and U name for used but undefined.
  • undefined reference is always the linker. expected ';' is always the compiler.
  • cc -std=c17 -Wall -Wextra file.c -o file then ./file.

Test yourself

1. Put these in order and say what each produces: linker, assembler, preprocessor, compiler.

Preprocessor, giving expanded C text. Compiler, giving assembly. Assembler, giving an object file. Linker, giving an executable.

2. You get undefined reference to 'average'. Which stage failed, and what are the two likely causes?

The linker. Either you declared or called average and never defined it, or its definition is in another file or library that you did not include in the link command.

3. A program prints the wrong answer. Which stage will find the bug?

None. Every stage succeeded; the program you described is the program you wrote. This is what tracing by hand and testing are for.

4. What does -E do, and why is the output so much longer than the input?

It stops after preprocessing. The output is longer because #include <stdio.h> pastes the whole header in, which is several hundred lines of declarations.

5. After #define N 5, does N * N reach the compiler as 25?

No. It reaches the compiler as 5 * 5. The preprocessor substitutes text and does no arithmetic. The compiler then works out 25 while compiling, because both operands are constants.

6. Why does ./area need the ./?

Because the shell searches only a fixed list of directories for a command name, and the current directory is not normally on that list. ./area says explicitly which file to run.

What can be asked on this, and how to answer it

"Explain the compilation and execution of a C program." Name the four stages in order, say what each takes in and gives out, and name the file at each boundary: .c, then expanded text, then .s, then .o, then the executable. Finish with the run step. A diagram of five boxes and four arrows earns the marks quickly; label the arrows with the tool and the boxes with the file.

munotes.in23

From Source to Running Program: Preprocessor, Compiler, Linker

"What is the role of the linker?" To resolve names. It joins your object files with the library code for every function you used but did not write, adds the start-up code that calls main, and produces a single executable. Give undefined reference as the error it reports when it cannot.

"Distinguish between the compiler and the interpreter." A compiler translates the whole program once, ahead of time, into machine code, which then runs on its own and runs fast; errors in the whole file are reported before anything runs. An interpreter translates and executes statement by statement every time the program runs, needs the interpreter present to run at all, and reports an error only when execution reaches it.

"What is the purpose of the preprocessor?" To process the file as text before compilation: insert headers, substitute defined names, select code conditionally, and remove comments. Say plainly that it does not understand C, because that is the part examiners are testing.

"A header file and a library file, distinguish." A header is source text holding declarations, read by the compiler so it can check your calls. A library is compiled machine code holding definitions, read by the linker so your calls have something to reach. stdio.h against the C standard library is the example to give.

Contents This chapter on its own page

munotes.in24

Chapter Six

Program Characteristics, and Which Ones Are Desirable

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

Every program has characteristics you can judge it by, and the ones worth aiming at are integrity, clarity, simplicity, efficiency, modularity and generality, in that order of importance, because a fast program that gives the wrong answer is worth nothing.

Why this is on the syllabus at all

Two programs can produce identical output and be very different things. One can be read by a stranger in a minute, changed safely, and reused; the other can only be understood by the person who wrote it, and only for about a week. In a first semester it feels like fussiness. It stops feeling like fussiness the first time you have to change a program you wrote a month ago.

The list is also an examination topic in its own right, and it is asked as bookwork, so learn the six words. But the six words are worth very little without an example each, and an answer that gives one line of reasoning per word is worth roughly twice one that gives the words alone.

The two programs

Both of these read nothing, compute the same thing, and print the same numbers. The task is the one from the lab's first practical, extended: print a small table of simple interest for three deposits.

Program A. It works.

#include <stdio.h>

int main(void)
{
    int i;
    double a[3] = {15000.0, 20000.0, 32000.0};
    double b[3] = {8.5, 7.25, 9.0};
    double c[3] = {2.0, 3.0, 1.5};
    for (i = 0; i < 3; i++) {
        printf("%.2f %.2f %.2f %.2f\n", a[i], b[i], c[i],
               a[i] * b[i] * c[i] / 100.0);
    }
    return 0;
}
15000.00 8.50 2.00 2550.00
20000.00 7.25 3.00 4350.00
32000.00 9.00 1.50 4320.00

Program B. It works too, and produces the same four numbers per row.

/*
 * interest-table.c
 * Prints the simple interest earned on each of several deposits.
 */

#include <stdio.h>

#define PERCENT      100.0
#define DEPOSITS     3

/* Simple interest = P * R * T / 100, with R a percentage per year. */
double simple_interest(double principal, double rate_percent, double years)
{
    return principal * rate_percent * years / PERCENT;
}

int main(void)
{
    double principal[DEPOSITS] = {15000.0, 20000.0, 32000.0};
    double rate_percent[DEPOSITS] = {8.5, 7.25, 9.0};
    double years[DEPOSITS] = {2.0, 3.0, 1.5};

    printf("%12s %8s %7s %12s\n", "Principal", "Rate", "Years", "Interest");
    for (int i = 0; i < DEPOSITS; i++) {
        printf("%12.2f %8.2f %7.2f %12.2f\n",
               principal[i], rate_percent[i], years[i],
               simple_interest(principal[i], rate_percent[i], years[i]));
    }
    return 0;
}
   Principal     Rate   Years     Interest
    15000.00     8.50    2.00      2550.00
    20000.00     7.25    3.00      4350.00
    32000.00     9.00    1.50      4320.00

Now the six characteristics, each read off the difference between those two files.

1. Integrity

Integrity is the accuracy of the result. It is first because nothing else can compensate for its absence.

munotes.in25

Program Characteristics, and Which Ones Are Desirable

Both programs have it here, and it is worth seeing how easily one could lose it. Write the formula as principal rate_percent years / 100 with an integer 100 and the answer is still right, because one operand is already a double. Write it as principal (rate_percent years / 100) and it is still right. Write it as (int) principal rate_percent years / PERCENT and it is right only while the principal happens to be a whole number. Integrity is not preserved by good intentions; it is preserved by knowing the rules, which is why chapter 14 on conversions exists.

Integrity also covers what happens on bad input. A program that divides by a number the user typed has no integrity until it checks whether that number is zero.

2. Clarity

Clarity is how easily a reader who did not write it can tell what the program does.

Program A calls its arrays a, b and c. Nothing in the file says which is the rate and which is the number of years, so the expression a[i] b[i] c[i] / 100.0 cannot be checked by reading: you have to go back to the initialisers and count. Program B calls them principal, rate_percent and years, and the same expression becomes a formula you can compare against the one in your textbook.

Clarity in C comes from four cheap things, all visible above:

  1. Names that say what the thing is. rate_percent rather than b, and rate_percent rather than rate, because the unit was the thing a reader would have had to guess.
  2. A comment saying why, not what. / Simple interest = P R T / 100 / is worth writing. i++; / add one to i / is not.
  3. Consistent indentation. Everything inside a block moves right by the same amount. The compiler does not care, and every reader does.
  4. One idea to a line. Program B's printf is split across lines so each argument can be seen.

3. Simplicity

Simplicity is doing the job without more machinery than the job needs.

This is the one that cuts both ways, and Program B is longer than Program A. Length is not complexity. Program B has one idea per construct: a function that computes one interest figure, a loop that walks the deposits, a header row. Program A has fewer lines and puts the formula inside the printf call, where it is mixed up with the formatting.

Simplicity is also what stops you writing a[i] b[i] c[i] / 100.0 inside a printf in the first place. A reader now has to hold the formatting and the arithmetic in mind at once, and the two have nothing to do with each other.

munotes.in26

Program Characteristics, and Which Ones Are Desirable

4. Efficiency

Efficiency is the time and the memory the program uses.

Here the two programs are effectively identical, and that is the honest answer: at three deposits, nothing you do to this program can be measured. Efficiency matters when the size of the input grows, and the thing that decides it is almost always the algorithm rather than the code. A sort that compares every element with every other takes a hundred times as long on a hundred times as much data, whatever style it is written in.

Do not write an unclear program in the belief that it is a fast one. Program B calls a function three times, which Program A avoids, and on any modern compiler at any optimisation level that call disappears. Choosing bad names to save typing buys nothing at all.

5. Modularity

Modularity is dividing the program into pieces that each do one job and can be understood, tested and changed on their own.

Program A has one piece. Program B has two, and simple_interest is the useful one: it takes three numbers and returns one, it does not print anything, it does not read anything, and it can be tested by calling it with numbers whose answer you already know. That is what a module is for. When the next practical asks for compound interest, Program B gets a second function beside the first and main barely changes.

Modularity is also what makes a program more than one person can work on, and what MU's lab is preparing you for when practicals 4, 7 and 8 all say "using a function".

6. Generality

Generality is how much the program can do without being rewritten.

Neither program is very general: both have the deposits typed into them. Program B is closer, because the count is DEPOSITS in one place rather than the literal 3 in four places, so adding a fourth deposit is two edits rather than five. The genuinely general version reads its input, and that is the version the lab asks for in Practical 1.

Generality has a limit. A program that tries to handle every case nobody asked about becomes large and unclear, which costs you two characteristics to buy one. Make it general where the requirement is likely to change, which for these programs means the data, not the formula.

The ordering, and why it is not negotiable

CharacteristicWhat it meansBeaten by
IntegrityThe answer is rightNothing
ClarityA stranger can read itIntegrity
SimplicityNo more machinery than neededIntegrity, clarity
EfficiencyTime and memory usedThe three above, until the program is too slow to use
ModularitySplit into pieces with one job eachNothing much; it usually helps the others
GeneralityHandles more cases unchangedSimplicity, when it is speculative
munotes.in27

Program Characteristics, and Which Ones Are Desirable

The single sentence to remember, because it answers the exam question and it is true: make it right, then make it clear, then make it fast, and only if it is actually too slow.

What this does NOT mean

Fewer lines is not simpler. Program A is shorter and harder to read. Compressing two statements onto one line with a comma operator or a nested assignment makes a file shorter and a program more difficult.

Comments do not create clarity. A badly named variable with a comment explaining it is worse than a well named one with no comment, because the comment can go stale and the name cannot. Rename first, comment second.

Efficiency is not about operators. Writing i++ instead of i = i + 1 does not make a program faster; they compile to the same instructions. Efficiency is decided by how much work the algorithm does, which is chapter 35's territory and the reason recursion has to be understood before it is used.

Modularity is not just "use functions". A function of eighty lines that does four unrelated things is not a module. A module has one job, a name that says what that job is, and inputs and outputs you can describe in a sentence.

Portability is a separate property, and it is the one this book's -std=c17 -Wall -Wextra is for. A program that relies on int being four bytes is not portable even if it is clear, simple and correct on your machine. Chapter 9 gives the sizes the standard actually guarantees.

Quick revision

  • Six characteristics: integrity, clarity, simplicity, efficiency, modularity, generality.
  • Integrity first: a wrong answer produced quickly is still wrong.
  • Clarity comes from names, indentation, comments that say why, and one idea per line.
  • Simplicity is machinery, not line count. A longer program can be simpler.
  • Efficiency is decided by the algorithm, not by the choice of operator.
  • Modularity means one job per function, testable on its own.
  • Generality means the data is input rather than typed in, and constants are named in one place.
  • Make it right, then clear, then fast, and only if it is genuinely too slow.

Test yourself

1. Name the six desirable characteristics of a program.

Integrity, clarity, simplicity, efficiency, modularity, generality.

2. A program gives the right answer and nobody, including its author, can follow it. Which characteristic does it lack, and does it matter?

Clarity. It matters, because every future change to it is a guess, and a guess is how a program loses its integrity.

munotes.in28

Program Characteristics, and Which Ones Are Desirable

3. Which two characteristics is #define DEPOSITS 3 serving?

Clarity, because DEPOSITS says what the 3 counts, and generality, because the count is now in one place and can be changed there.

4. Your program takes two seconds on the data you were given and will never be given more. Should you make it faster?

No. Two seconds is fast enough, and any change you make for speed costs clarity and risks integrity. Efficiency is a requirement, not a virtue in itself.

5. Give one change to Program A that improves clarity and costs nothing else.

Rename a, b and c to principal, rate_percent and years. No line is added and the arithmetic becomes checkable by reading.

What can be asked on this, and how to answer it

"What are the desirable characteristics of a program? Explain." Name all six and give one sentence each, and give an example for at least three. Put integrity first and say why it is first. This is bookwork and the marks are in the explanations, not the list.

"What is meant by program clarity? How is it achieved?" Clarity is how easily a reader who did not write the program can tell what it does. Achieved by meaningful names, consistent indentation, comments that give the reason rather than restate the code, small functions with one job each, and avoiding expressions that do several things at once.

"Distinguish between simplicity and efficiency." Simplicity is about the program as a text: no more machinery than the job needs, so that a reader can follow it. Efficiency is about the program as it runs: time taken and memory used. They usually agree, and where they conflict, correctness and clarity come first and efficiency is improved only when the program is measurably too slow.

"Why is modularity important?" Because a program divided into functions that each do one job can be understood, tested and changed one piece at a time, can be worked on by more than one person, and lets a piece be reused instead of rewritten.

Contents This chapter on its own page

munotes.in29

Chapter Seven

The C Character Set

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

The C character set is the fixed list of characters a C program may be written with, and where a character cannot be typed into a program safely, it is written instead as an escape sequence: a backslash followed by one or more characters.

Why a language fixes its character set

C is meant to compile on machines that have nothing else in common. That was the whole point of it. So the standard names the smallest set of characters every implementation must support, and promises nothing beyond it. Write your program with those characters and it compiles anywhere; reach outside them and you are relying on your own compiler's generosity.

The practical consequence for you is short. English letters, digits and the punctuation on a standard keyboard are safe. A rupee sign, an accented letter, a curly quotation mark pasted in from a word processor, or a character from an Indian script is not part of the basic set, and putting one outside a comment or a string is undefined behaviour.

That last one catches real students. Typing a program in a word processor turns " into a curly quotation mark, and the compiler then reports an error on a line that looks perfectly correct. Write code in a plain text editor.

The four groups

The standard splits the basic set into letters, digits, graphic characters and a few controls.

The 26 uppercase letters.

A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

The 26 lowercase letters.

a b c d e f g h i j k l m n o p q r s t u v w x y z

C is case sensitive, so these are 52 distinct characters, not 26 written two ways. Total and total are different names.

The 10 decimal digits.

0 1 2 3 4 5 6 7 8 9

The standard guarantees that the values of these ten are consecutive, which is what lets c - '0' turn a digit character into the number it stands for. It gives no such guarantee for the letters, which is why well written C never assumes 'z' - 'a' is 25.

The 29 graphic characters, which is the standard's own count and its own list.

! " # % & ' ( ) * + , - . / :
; < = > ? [ \ ] ^ _ { | } ~

Two characters a programmer uses constantly are not in that list of 29: the at sign and the dollar sign. They are not in the basic character set at all. This is why no C keyword or operator uses them, and why $ in an identifier, which some compilers allow, is not portable.

munotes.in30

The C Character Set

The space, and the controls. The space character, plus control characters for horizontal tab, vertical tab and form feed. In the execution set the standard also requires controls for alert, backspace, carriage return and new line.

Escape sequences

Some characters cannot be written directly in a program. A newline cannot go inside a string literal, because the string would run off the end of the line. A double quotation mark cannot go inside a string delimited by double quotation marks. And an alert or a tab has no printable shape to type.

So C writes them with a backslash. The backslash is an escape character: it means "the next character or two is not itself, it is a name".

This program prints the name of each escape sequence and the numeric value it has on the machine that ran it.

#include <stdio.h>

int main(void)
{
    printf("%-16s %-4s %3d\n", "alert",           "\\a",  (int) '\a');
    printf("%-16s %-4s %3d\n", "backspace",       "\\b",  (int) '\b');
    printf("%-16s %-4s %3d\n", "form feed",       "\\f",  (int) '\f');
    printf("%-16s %-4s %3d\n", "new line",        "\\n",  (int) '\n');
    printf("%-16s %-4s %3d\n", "carriage return", "\\r",  (int) '\r');
    printf("%-16s %-4s %3d\n", "horizontal tab",  "\\t",  (int) '\t');
    printf("%-16s %-4s %3d\n", "vertical tab",    "\\v",  (int) '\v');
    printf("%-16s %-4s %3d\n", "backslash",       "\\\\", (int) '\\');
    printf("%-16s %-4s %3d\n", "single quote",    "\\'",  (int) '\'');
    printf("%-16s %-4s %3d\n", "double quote",    "\\\"",  (int) '\"');
    printf("%-16s %-4s %3d\n", "question mark",   "\\?",  (int) '\?');
    printf("%-16s %-4s %3d\n", "null character",  "\\0",  (int) '\0');
    return 0;
}
alert            \a     7
backspace        \b     8
form feed        \f    12
new line         \n    10
carriage return  \r    13
horizontal tab   \t     9
vertical tab     \v    11
backslash        \\    92
single quote     \'    39
double quote     \"    34
question mark    \?    63
null character   \0     0

Read that program carefully, because it contains the joke this whole topic turns on. To print a backslash followed by a, the format string has to contain \\a: the first backslash escapes the second, and the a is then just an a. To get the alert character, the character constant is '\a'. Both appear on the same line, which is the clearest way to see the difference between the two backslashes in your source.

The numbers in the right-hand column are the values in this machine's execution character set, which is ASCII. The standard does not promise those particular numbers, and you should not memorise them. What you should know is that '\0' is zero, always, because the standard requires a null character whose bits are all zero.

munotes.in31

The C Character Set

The table to learn

EscapeNameWhat it does
\nnew lineMoves output to the start of the next line
\thorizontal tabMoves output to the next tab stop
\vvertical tabMoves down, on devices that have the notion
\bbackspaceMoves back one position
\rcarriage returnMoves to the start of the current line, without moving down
\fform feedStarts a new page on a printer
\aalertMakes an audible or visible alert
\\backslashOne backslash
\'single quoteNeeded inside a character constant
\"double quoteNeeded inside a string literal
\?question markHistorical, for trigraphs; harmless
\0null characterThe value zero, which ends every string

Two more forms let you write any character by its value.

  • Octal. \ followed by one, two or three octal digits. '\101' is the character whose value is 101 in base eight, which is 65, which is A.
  • Hexadecimal. \x followed by hexadecimal digits. '\x41' is 65, which is A again.

Both were run to check that claim:

#include <stdio.h>

int main(void)
{
    printf("octal 101 is %d, the character %c\n", (int) '\101', '\101');
    printf("hex 41    is %d, the character %c\n", (int) '\x41', '\x41');
    return 0;
}
octal 101 is 65, the character A
hex 41    is 65, the character A

\0 is not a special case of anything. It is the octal escape with a single octal digit zero, which is why it means the value zero.

\n against \r, which is worth two marks

\n moves to the beginning of the next line. \r moves to the beginning of the current line and stays there, so whatever you print next overwrites what was already on it.

printf("old text\rnew");

On a terminal that prints newtext, because new has overwritten the first three characters of old text and the rest of the old line is still there. This is how a progress counter that seems to update in place is done.

Windows text files end each line with a carriage return followed by a newline, and Unix files with a newline alone. This is why a file written on one and read on the other sometimes shows stray characters. In C you do not usually have to care: \n in a text stream is translated for you by the library.

What this does NOT mean

An escape sequence is one character, not two. '\n' is a single character constant with a single value. sizeof('\n') is not two. The two symbols in the source are how you write it, not what it is.

\ is not a line-continuation character in ordinary text. It does continue a source line when it is the last character before the newline, and that is a separate rule, used for long macros. Inside a string it introduces an escape sequence.

munotes.in32

The C Character Set

A string is not a character. "a" and 'a' are different things: the first is an array of two characters, a and the terminating zero, and the second is a single character. Chapter 12 is this in full, and it is the most common single confusion in the first semester.

The at sign and the dollar sign are not "special characters of C". They are not in C's basic character set at all. A question asking you to list the special characters wants the 29 graphic characters, and those two are not among them.

Not every character in the set is usable everywhere. The set says what you may write the program with. Where each one is allowed is settled by the grammar: a digit may not start an identifier, and a brace may not appear in the middle of an expression.

Quick revision

  • The basic character set: 26 uppercase letters, 26 lowercase letters, 10 digits, 29 graphic characters, the space, and controls for horizontal tab, vertical tab and form feed.
  • C is case sensitive: 52 letters, not 26.
  • The digits are guaranteed to have consecutive values; the letters are not.
  • The at sign and the dollar sign are not in the set.
  • An escape sequence is a backslash plus one or more characters, and denotes one character.
  • Learn \n \t \b \r \f \a \\ \' \" \0.
  • \0 is the null character, value zero, and it terminates every string.
  • \nnn is octal, \xhh is hexadecimal.
  • \n goes to the next line; \r goes to the start of this one.

Test yourself

1. How many characters are in C's basic character set, by group?

26 uppercase letters, 26 lowercase letters, 10 digits, 29 graphic characters, the space character, and control characters for horizontal tab, vertical tab and form feed.

2. What does this print?

printf("a\\b\n");

a\b and then a newline. The \\ is one backslash, so no backspace is involved.

3. Write a printf that prints a double quotation mark, a tab, and a backslash, in that order.

printf("\"\t\\\n");

4. Is '\n' one character or two?

One. It is written with two symbols in the source and is a single character with a single value.

5. Why can you write c - '0' to convert a digit character into its numeric value, but not rely on c - 'a' for letters?

Because the standard guarantees the ten digit characters have consecutive values, and makes no such guarantee about the letters.

6. A program copied out of a word processor will not compile, and the error points at a line that looks right. What is the likely cause?

munotes.in33

The C Character Set

The word processor replaced straight quotation marks with curly ones, or a hyphen with a dash. Those characters are not in C's basic character set. Retype the line in a plain text editor.

What can be asked on this, and how to answer it

"What is the C character set?" Give the four groups with their counts, name the space and the control characters, and say that the set is fixed by the standard so that a program written in it compiles on any conforming implementation. Mention case sensitivity, because it is worth a mark on its own.

"What is an escape sequence? Give any five with their meanings." Define it as a backslash followed by one or more characters, standing for a single character that cannot conveniently be typed, then give five from the table with what each does. \n, \t, \0, \\ and \" are the safest five.

"Explain the difference between \n and \r." \n moves to the start of the next line. \r moves to the start of the current line without advancing, so subsequent output overwrites it.

"What is the significance of \0?" It is the null character, whose value is zero, and it marks the end of a character string in C. Every string literal has one added automatically, and the library's string functions find the end of a string by looking for it.

"Write a program to print a tab-separated table." Use \t between the fields and \n at the end of each row, and say in one line why \t is preferable to counting spaces: the columns stay aligned when the data widths change.

Contents This chapter on its own page

munotes.in34

Chapter Eight

Identifiers and Keywords

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

An identifier is a name you invent for something in your program, and a keyword is a name the language has already taken, which you therefore cannot use for anything of your own.

Why there are rules about names

The compiler reads your program as a stream of characters and has to decide where each name begins and ends, and whether what it just read is a name at all or a piece of the language. Both jobs become impossible if a name may contain anything.

So a name is defined by a rule the compiler can apply mechanically, without understanding anything: it starts with a letter or an underscore, and continues with letters, digits and underscores. The moment something else appears, the name has ended. That is the entire rule, and everything else in this chapter is a consequence of it.

The rules for an identifier

  1. The first character is a letter or an underscore. Not a digit. total1 is legal; 1total is not, because the compiler would already have started reading a number.
  2. After that, letters, digits and underscores, in any order and any mixture.
  3. No other character at all. No space, no hyphen, no dot, no @, no $. first name is two identifiers, and first-name is a subtraction.
  4. It may not be a keyword. The 44 names in the next section are taken.
  5. Case matters. total, Total and TOTAL are three different identifiers.
  6. There is no limit you will reach, but there is a limit you can rely on. The standard requires a compiler to keep at least 63 significant characters for a name inside one file, and at least 31 for a name shared between files. Real compilers keep far more. Never write a name so long that this matters.

Legal: total, _count, sum_of_squares, x1, PERCENT, myVariable2, __internal.

Not legal: 2nd_place (starts with a digit), total marks (space), net-pay (hyphen), float (keyword), roll# (illegal character), rate% (illegal character).

The names you should not use even though they are legal

This is the rule nobody tells first-year students, and it is in the standard. Names beginning with an underscore are reserved for the implementation in two ways:

  • An identifier beginning with two underscores, or with an underscore followed by a capital letter, is reserved for any use. __count and _Total are yours to write and the compiler is entitled to have its own meaning for them.
  • An identifier beginning with one underscore followed by a lower-case letter is reserved at file scope.

The practical rule is one line: do not begin your own names with an underscore. You lose nothing, and you avoid a class of error that produces baffling messages. This also explains why C11's newest keywords look the way they do: _Bool, _Atomic and _Static_assert were put in a space the standard had already reserved, so that adding them could not break any existing program.

munotes.in35

Identifiers and Keywords

The 44 keywords of C17

These are the standard's own eleven rows of four, in the standard's own order.

auto       extern     short      while
break      float      signed     _Alignas
case       for        sizeof     _Alignof
char       goto       static     _Atomic
const      if         struct     _Bool
continue   inline     switch     _Complex
default    int        typedef    _Generic
do         long       union      _Imaginary
double     register   unsigned   _Noreturn
else       restrict   void       _Static_assert
enum       return     volatile   _Thread_local

Eleven rows of four is 44, and that is the number for C17 and for C11.

The 32 of C89, which is the number most textbooks print, is that list with the last column and inline and restrict removed:

auto     break    case     char     const    continue default  do
double   else     enum     extern   float    for      goto     if
int      long     register return   short    signed   sizeof   static
struct   switch   typedef  union    unsigned void     volatile while
  • C99 added 5: inline, restrict, _Bool, _Complex, _Imaginary. That makes 37.
  • C11 added 7: _Alignas, _Alignof, _Atomic, _Generic, _Noreturn, _Static_assert, _Thread_local. That makes 44.
  • C17 added none. It is a correction of C11, which is why the two have the same list.

If the question says "how many keywords does C have", answer 44 for C17 and say in the same breath that the figure is 32 for the original ANSI C of 1989 and that C99 and C11 added the rest. That answer cannot be marked wrong, and it shows you know why two numbers are in circulation.

Of the 44, how many you actually need this semester

Twenty-six, and they are worth separating out, because a list of 44 is intimidating and most of it is for later.

GroupKeywords
Typeschar int float double void short long signed unsigned
User-defined typesstruct union enum typedef
Decisionsif else switch case default
Loopsfor while do
Jumpsbreak continue goto return
Othersizeof const

The remaining keywords are storage classes and qualifiers you will meet in a later semester (auto, extern, static, register, volatile, inline, restrict) and the C11 additions for concurrency and alignment, which a first-year program never needs.

A program that proves case sensitivity

#include <stdio.h>

int main(void)
{
    int total = 10;
    int Total = 20;
    int TOTAL = 30;

    printf("total = %d, Total = %d, TOTAL = %d\n", total, Total, TOTAL);
    printf("their sum is %d\n", total + Total + TOTAL);
    return 0;
}
munotes.in36

Identifiers and Keywords

total = 10, Total = 20, TOTAL = 30
their sum is 60

Three declarations, three separate variables, three different values. The compiler did not complain, because as far as it is concerned these names have nothing to do with one another.

This is legal and it is a terrible idea. The reason to know it is that it explains the error you will actually make: declaring studentName and then writing studentname, and being told the second is undeclared.

What this does NOT mean

A keyword is not a function. sizeof is a keyword and an operator, built into the language. printf is not a keyword; it is an ordinary identifier that happens to name a library function, and you could declare a variable called printf if you were determined to cause trouble.

main is not a keyword. It is an ordinary identifier. The language reserves its meaning as the starting point of a program, not the name itself. This is why it is not in the table of 44.

NULL, TRUE and FALSE are not keywords. NULL is a macro from several standard headers. TRUE and FALSE are not part of C at all; they are things programmers define for themselves. C99 added _Bool and the header stdbool.h, which defines bool, true and false as macros. Macros, not keywords.

An underscore is not a letter, but it is allowed. The grammar lets a name start with one, which is why the reservation rules above exist at all.

A long name is not slow. Identifier names exist only in your source. After compilation the machine code has no idea what anything was called. total_marks_obtained costs exactly what t costs at run time, so choose the readable one.

Quick revision

  • An identifier starts with a letter or underscore, continues with letters, digits and underscores, and contains nothing else.
  • Case sensitive. total and Total are different.
  • It may not be a keyword.
  • C17 has 44 keywords; C89 had 32; C99 added 5 and C11 added 7.
  • Do not begin your own identifiers with an underscore: those names are reserved for the implementation.
  • Guaranteed significance: at least 63 characters within a file, at least 31 for external names.
  • main, printf and NULL are not keywords.

Test yourself

1. Which of these are valid identifiers: _total, 2sum, sum2, net pay, net_pay, double, Double?

Valid: _total, sum2, net_pay, Double. Invalid: 2sum starts with a digit, net pay contains a space, double is a keyword. _total is valid but should be avoided, because leading underscores are reserved.

2. How many keywords are there in C, and why do you see two different answers?

munotes.in37

Identifiers and Keywords

44 in C11 and C17. The 32 that most books print is the C89 list; C99 added five keywords and C11 added seven.

3. Is Main a valid identifier? Is it the entry point of the program?

It is a valid identifier and it is not the entry point. C is case sensitive, so Main is an ordinary name and the linker will still be looking for main.

4. Why are the newest keywords spelled _Bool and _Static_assert instead of bool and static_assert?

Because names beginning with an underscore and a capital letter were already reserved for the implementation, so no existing program could have been using them. Adding bool as a keyword would have broken every program that had a variable of that name.

5. Can you use printf as a variable name?

Yes, because it is not a keyword. You should not: it hides the library function for the rest of that scope, and the next printf call will not compile.

6. int marks-obtained = 50; does not compile. Why?

Because - is not allowed in an identifier. The compiler reads marks, then a subtraction, then obtained, and none of that is a declaration.

What can be asked on this, and how to answer it

"What are identifiers? State the rules for writing an identifier." Define it as a user-defined name for a variable, function, array, structure or label. Then give the six rules, and give two or three invalid examples with the reason each is invalid. The examples are where the marks are; the rules alone read as memorised.

"What are keywords? How do they differ from identifiers?" A keyword is a name reserved by the language with a fixed meaning, and there are 44 in C17. An identifier is a name the programmer invents. The keywords are fixed and few; identifiers are unlimited and chosen. A keyword cannot be used as an identifier.

"List any ten keywords in C." Take them from the twenty-six in the table above, so that every one you write is a keyword you can also explain if asked.

"Is C case sensitive? Give an example." Yes. total and Total are two different identifiers and may be declared in the same program with different values. Give the three-variable program from this chapter.

Contents This chapter on its own page

munotes.in38

Chapter Nine

Data Types and Their Sizes

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

A data type tells the compiler how many bytes an object occupies and how to read the bits in them, and C's basic types are char, int, float and double, widened by the modifiers short, long, signed and unsigned.

Why a type is necessary at all

A byte in memory is eight bits and nothing more. The bits 01000001 are the number 65, and they are also the letter A, and as part of a larger group they are part of some floating-point value. Nothing in the memory says which. The type is where that information lives, and it lives in the compiler, not in the machine.

That single fact explains most of C. The type decides how much space is set aside, how the bits are interpreted, which operations are allowed, and what happens when two different types meet in one expression. It also explains why C is fast and why it will let you do something disastrous without complaint: once the program is compiled, the types are gone, and the machine does what the instructions say.

Ritchie's stated reason for inventing C out of B was precisely this: B was typeless, and a language in which every value is one machine word cannot conveniently talk about a character. Chapter 3 has the history.

The basic types

char holds one character, and is exactly one byte by definition. sizeof(char) is 1, always, on every implementation. A byte need not be eight bits in the standard's terms, but on every machine you will meet it is.

int holds a whole number, and is meant to be the natural size for the machine.

float holds a number with a fractional part, at single precision.

double holds a number with a fractional part, at double precision. It is the one to use.

void holds nothing, and is not a type you can make a variable of. It appears in three places: as the return type of a function that returns nothing, as (void) for a function that takes no parameters, and as void *, the pointer that can point at anything, which is chapter 39.

The four modifiers

Two change the size:

  • short asks for a smaller integer than int.
  • long asks for a larger one. long long, added in C99, asks for larger still.

Two change the meaning of the bits:

  • signed allows negative values. It is the default for int.
  • unsigned does not, and spends the bit that would have carried the sign on magnitude instead, so the largest value roughly doubles.

They combine, and the spelling is flexible: unsigned long int, long unsigned, unsigned long all name the same type. int may be left out whenever a modifier is present, so short means short int and unsigned means unsigned int.

munotes.in39

Data Types and Their Sizes

What the standard actually guarantees

This is the table that is true everywhere. It is the standard's own required minimum magnitudes.

TypeAt least this rangeWhich needs
signed char-127 to +1278 bits
unsigned char0 to 2558 bits
short-32767 to +3276716 bits
unsigned short0 to 6553516 bits
int-32767 to +3276716 bits
unsigned int0 to 6553516 bits
long-2147483647 to +214748364732 bits
unsigned long0 to 429496729532 bits
long long-9223372036854775807 to +922337203685477580764 bits
unsigned long long0 to 1844674407370955161564 bits

Two further guarantees are worth as much as the table:

  1. sizeof(char) is 1. Every other size is measured in units of it.
  2. The sizes are ordered and the ranges nest: char no larger than short, no larger than int, no larger than long, no larger than long long. So a short value always fits in an int, whatever the machine.

Notice what is not guaranteed. int is not four bytes. float is not four bytes. Nothing says int and long differ. On a small embedded processor int really is two bytes and the minimum above really is the range.

What this machine actually has

Ask the machine. That is what sizeof is for: an operator, not a function, that gives the size of a type or of an object in bytes.

#include <stdio.h>

int main(void)
{
    printf("type          bytes\n");
    printf("char          %5zu\n", sizeof(char));
    printf("short         %5zu\n", sizeof(short));
    printf("int           %5zu\n", sizeof(int));
    printf("long          %5zu\n", sizeof(long));
    printf("long long     %5zu\n", sizeof(long long));
    printf("float         %5zu\n", sizeof(float));
    printf("double        %5zu\n", sizeof(double));
    printf("long double   %5zu\n", sizeof(long double));
    printf("void *        %5zu\n", sizeof(void *));
    return 0;
}
type          bytes
char              1
short             2
int               4
long              8
long long         8
float             4
double            8
long double      16
void *            8

Those are the sizes on the 64-bit Linux machine that compiled this book's programs. Four of them differ elsewhere, and these are the four worth knowing about:

  • long is 4 bytes on 64-bit Windows, not 8. Code that assumes a long can hold a large count is not portable between Linux and Windows. long long is 8 everywhere that has it.
  • long double is 16 bytes here, 8 on 64-bit Windows with MSVC, and 12 on 32-bit Linux. It is the least portable type in C.
  • void * is 8 bytes on any 64-bit machine and 4 on a 32-bit one. It is the size of an address, so it tracks the machine, not the standard.
  • int is 4 bytes on every desktop machine and 2 on many microcontrollers.
munotes.in40

Data Types and Their Sizes

%zu is the conversion for a value of type size_t, which is what sizeof produces. Using %d for it is a mistake that -Wall catches and that a great many textbooks make.

Signed against unsigned, and the trap in it

An unsigned type cannot hold a negative value. What it does instead is the thing to learn, because it is defined behaviour and it surprises everybody once.

#include <stdio.h>

int main(void)
{
    unsigned int u = 0;
    int i = 0;

    printf("as unsigned, 0 - 1 is %u\n", u - 1);
    printf("as signed,   0 - 1 is %d\n", i - 1);
    return 0;
}
as unsigned, 0 - 1 is 4294967295
as signed,   0 - 1 is -1

Subtracting one from an unsigned zero does not give minus one and it is not an error. Unsigned arithmetic wraps round: the result is reduced modulo one more than the largest value the type can hold, so it comes out as the largest value. This is exactly why a loop written as

for (unsigned int i = n; i >= 0; i--)

never ends. An unsigned i is always at least zero, so the condition is always true.

Use int for counting unless you have a reason not to. unsigned is for bit patterns, for sizes returned by the library, and for the one case where you genuinely need the extra range.

The floating-point types, and why double

float and double hold approximations. They are stored as a sign, a set of significant digits and an exponent, in the manner of scientific notation, which means a fixed number of significant digits and not a fixed number of decimal places.

The number of decimal digits you can rely on is in <float.h> as FLT_DIG and DBL_DIG. On the machine that ran this book's programs those are 6 and 15, and those two values are typical of every desktop machine, because both types follow the same international format for floating-point arithmetic.

Six digits is not many. Here is what that costs, in a program that is not buggy:

#include <stdio.h>

int main(void)
{
    float f = 123456789.0f;
    double d = 123456789.0;

    printf("as a float  : %.1f\n", (double) f);
    printf("as a double : %.1f\n", d);
    printf("0.1 + 0.2 as a double, to 20 places: %.20f\n", 0.1 + 0.2);
    return 0;
}
as a float  : 123456792.0
as a double : 123456789.0
0.1 + 0.2 as a double, to 20 places: 0.30000000000000004441

The float did not store 123456789. It stored the nearest value it is able to represent, and the last two digits are not the ones that were written. Nothing went wrong: a float has about seven significant decimal digits and was asked for nine.

munotes.in41

Data Types and Their Sizes

Try the same program with 1234567.0f and the float prints it back perfectly. Seven digits happens to be inside what a float can do, which is exactly why this trap is hard to spot: it appears only once the numbers grow, and by then the program is finished and trusted.

And 0.1 + 0.2 is not 0.3, in any language, on any machine that uses binary floating point, because one tenth cannot be written exactly in binary any more than one third can be written exactly in decimal.

The rule that follows is absolute and it is worth marks: never compare two floating-point values with ==. Compare the size of their difference against a small tolerance instead. Chapter 16 gives the form.

Use double, not float. A float saves four bytes, which no program in this course will notice, and costs nine significant digits. float exists for very large arrays and for hardware that is faster at it.

The complete picture

DeclarationTypical size hereWhat it is for
char1One character, or a very small integer
signed char1A small integer, definitely signed
unsigned char1A byte, 0 to 255
short2A small integer
unsigned short2A small non-negative integer
int4The default whole number
unsigned int4Non-negative, or a bit pattern
long8 here, 4 on WindowsA larger integer
long long8A large integer, everywhere
float4Fractions, about 6 digits
double8Fractions, about 15 digits. The default
long double16 hereFractions, more digits, least portable
voidnoneNothing: no value, no parameters, or an untyped pointer

The conversion to print each one

Getting this wrong is undefined behaviour, not a small mistake, because printf reads its arguments according to what the format string says is there. -Wall catches nearly all of it, which is the best argument for compiling with it.

TypeConversion
char as a character%c
char as a number%d
int%d
unsigned int%u
short%hd
long%ld
long long%lld
unsigned long%lu
float%f, and the value is promoted to double
double%f, or %g, or %e
long double%Lf
size_t, from sizeof%zu
a pointer%p

There is no conversion for float in printf, and it is not an omission: a float argument is automatically promoted to double when passed, so %f is correct for both. In scanf, which takes addresses and promotes nothing, %f is for a float and %lf is for a double . Confusing the two in scanf is a real and common bug.

munotes.in42

Data Types and Their Sizes

What this does NOT mean

int is not four bytes. It is four bytes on the machines you will use. The language says at least two. Write sizeof(int) when you need the number and never the literal 4.

sizeof is not a function. It is an operator and it is evaluated by the compiler, not at run time. sizeof x without brackets is legal for an object; brackets are required for a type name.

char is not "the character type" and nothing else. It is a one-byte integer, and whether it is signed is up to the implementation. When you want a small number use signed char or unsigned char and say which. When you want a character, use char.

unsigned does not mean positive. It means non-negative, and it means arithmetic that wraps rather than going negative.

float is not "less accurate for big numbers only". It has about six significant digits at every magnitude. float can hold 3.4 times ten to the thirty-eighth; it just cannot tell that number from its neighbours.

A type is not a promise about the value. Declaring int age does not stop anyone storing -500 in it. Checking the value is your job, and it is part of what integrity meant in chapter 6.

Quick revision

  • Basic types: char, int, float, double, plus void.
  • Modifiers: short, long, signed, unsigned.
  • sizeof(char) is 1 by definition; everything else is measured in those units.
  • The standard fixes minimum ranges, not sizes: int at least -32767 to +32767.
  • Sizes never shrink going char, short, int, long, long long.
  • Here: char 1, short 2, int 4, long 8, long long 8, float 4, double 8, pointer 8.
  • long is 4 bytes on 64-bit Windows. long long is 8 everywhere.
  • Unsigned arithmetic wraps modulo the range; 0u - 1 is the maximum value.
  • float about 6 significant digits, double about 15. Use double.
  • Never compare floating-point values with ==.
  • sizeof gives a size_t; print it with %zu.

Test yourself

1. What is the size of int in C?

The standard does not say. It requires a range of at least -32767 to +32767, so at least two bytes, and it is four bytes on every desktop machine. sizeof(int) is the only reliable answer for a given machine.

2. Which is bigger, int or long?

long is never smaller than int, and on many machines they are the same size. sizeof(long) >= sizeof(int) is guaranteed; > is not.

3. What does this print, and why?

unsigned int u = 5;
printf("%u\n", u - 10);

A very large number, the maximum for unsigned int minus four, which is 4294967291 where unsigned int is 32 bits. Unsigned arithmetic is reduced modulo one more than the maximum, so it cannot produce a negative result.

munotes.in43

Data Types and Their Sizes

4. Why should you not write if (x == 0.1) when x is a double?

Because 0.1 has no exact binary representation, so the value stored is the nearest one that does, and any arithmetic that produced x may have landed on a different nearest value. Test if (fabs(x - 0.1) < 1e-9) instead.

5. How many bytes does char occupy, and how do you know?

One, by definition. sizeof(char) is 1 in every conforming implementation, because the byte is the unit sizeof measures in.

6. Give the right printf conversion for sizeof(int), and say why %d is wrong.

%zu. sizeof produces a size_t, which is an unsigned type that may be wider than int, so %d tells printf to read the wrong number of bytes as the wrong kind of value.

What can be asked on this, and how to answer it

"Explain the data types in C with their sizes and ranges." Give the four basic types and the four modifiers, then the table of typical sizes, and then the sentence that earns the extra marks: the standard fixes minimum ranges rather than sizes, so the correct way to get a size is sizeof. Give int as 2 bytes minimum and 4 in practice, and you have answered both versions of the question.

"What is the difference between signed and unsigned?" A signed type spends one bit on the sign and can hold negative values; an unsigned type of the same width holds only non-negative values and reaches roughly twice as far. Add that unsigned arithmetic wraps modulo the range rather than going negative, and give 0u - 1 as the example.

"Differentiate between float and double." Both hold approximations. float is typically 4 bytes with about 6 significant decimal digits; double is typically 8 with about 15. double is the default for floating-point constants and the type to prefer. float is used to save space in large arrays.

"What is the use of the void data type?" Three uses: the return type of a function that returns no value, the parameter list of a function that takes no arguments, and void *, a pointer to an object of unspecified type.

"What is sizeof? Write a program to find the size of each data type." sizeof is a compile-time operator giving the size in bytes of a type or object, with sizeof(char) equal to 1 by definition. Then give the program in this chapter, and remember %zu.

Contents This chapter on its own page

munotes.in44

Chapter Ten

Constants and Their Types

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

A constant is a value written directly into the program, which cannot change while the program runs, and the way you write it decides its type.

Why the way you write it matters

printf("%d\n", 7 / 2) prints 3. Not 3.5, and it is not a bug. Both operands are integer constants, so C does integer division, which throws the fraction away. Write 7.0 / 2 and you get 3.5, because 7.0 is a double.

The value is the same number in both cases. The type is different, and the type decides what the operators do with it. That is why this chapter comes before the operators.

Integer constants

A whole number with no decimal point and no exponent. It may be written in three bases:

WrittenBaseValue
75decimal75
0113octal, leading zero75
0x4B or 0X4bhexadecimal, leading 0x75

A leading zero means octal, and this catches people. 011 is not eleven, it is nine. Writing a date as 09 or a zero-padded code as 0755 silently changes the number, and if a digit 8 or 9 appears the compiler reports an error on a line that looks fine.

#include <stdio.h>

int main(void)
{
    printf("75 is %d\n", 75);
    printf("0113 is %d\n", 0113);
    printf("0x4B is %d\n", 0x4B);
    printf("011 is %d, not eleven\n", 011);
    return 0;
}
75 is 75
0113 is 75
0x4B is 75
011 is 9, not eleven

Suffixes set the type. Without one, the type is the first of int, long int, long long int that the value fits in.

SuffixMeansExample
noneint, or wider if it must be100
U or uunsigned100U
L or llong100L
UL, LU and so onunsigned long100UL
LLlong long100LL
ULLunsigned long long100ULL

Prefer capital L. A lower-case l is almost indistinguishable from the digit 1 in most fonts, and l against 1 has caused real bugs.

Floating constants

A number with a decimal point, or an exponent, or both. Its type is double by default, which is a fact worth remembering: there is no way to write a float constant without saying so.

WrittenTypeValue
3.14double3.14
3.14f or 3.14Ffloat3.14
3.14Llong double3.14
1e3double1000.0
1.5e-3double0.0015
.5double0.5
5.double5.0

1e3 means 1 times ten to the third. The e may be E, and the exponent may carry a sign. Both .5 and 5. are legal, and both are clearer written as 0.5 and 5.0.

1e3 is a floating constant even though its value is a whole number. printf("%d\n", 1e3) is undefined behaviour, not a rounding question.

munotes.in45

Constants and Their Types

Character constants

A single character in single quotation marks. Its type is int, not char, which surprises everyone the first time.

#include <stdio.h>

int main(void)
{
    printf("'A' as a character is %c\n", 'A');
    printf("'A' as a number is %d\n", 'A');
    printf("sizeof('A') is %zu, and sizeof(char) is %zu\n",
           sizeof('A'), sizeof(char));
    printf("'A' + 1 is %d, and as a character %c\n", 'A' + 1, 'A' + 1);
    return 0;
}
'A' as a character is A
'A' as a number is 65
sizeof('A') is 4, and sizeof(char) is 1
'A' + 1 is 66, and as a character B

The value is the character's code in the execution character set, which is ASCII on any machine you will use. So 'A' is genuinely usable as the number 65, and '0' as 48, and that is what makes c - '0' convert a digit character into its numeric value.

An escape sequence is a character constant: '\n', '\t', '\\', '\0'. Each is one character. Chapter 7 has the list.

'A' and "A" are not the same thing. The first is one character whose value is 65. The second is a string literal: an array of two characters, A and the terminating zero. Chapter 12.

'AB' is not a string and not an error you can rely on. It is a multi-character constant, its value is implementation-defined, and -Wall warns about it. Never write one.

String literals

Text in double quotation marks. Strictly it is a literal rather than a constant, and the distinction is real: a string literal is an array of char, and an array is not a constant expression.

"Hello"

That is six characters in memory, not five: H, e, l, l, o and a terminating \0 that the compiler adds. The count of the terminator is what chapter 12 and the whole string library depend on.

A string literal must not be modified. It may be stored in read-only memory, and a program that writes into one may crash. Declare const char *s = "Hello"; when you only mean to read it.

Two string literals written next to each other are joined into one, which is how a long piece of text is spread over several source lines:

printf("This is one long line of output that would not "
       "fit comfortably in the source, so it is split.\n");

Enumeration constants

An enum declares a set of named integer constants. They are the only constants that are genuinely part of the type system rather than the syntax.

#include <stdio.h>

enum day { MON, TUE, WED, THU, FRI, SAT, SUN };
enum grade { FAIL = 0, PASS = 40, FIRST = 60, DISTINCTION = 70 };

int main(void)
{
    enum day today = WED;

    printf("MON is %d and SUN is %d\n", MON, SUN);
    printf("today is %d\n", today);
    printf("the pass mark is %d and a distinction is %d\n", PASS, DISTINCTION);
    return 0;
}
munotes.in46

Constants and Their Types

MON is 0 and SUN is 6
today is 2
the pass mark is 40 and a distinction is 70

Numbering starts at 0 and increases by one unless you set a value, and after a value that you set, counting continues from it. An enumeration constant has type int.

The three ways to name a constant, and which to use

This is the part examiners like, because the three are genuinely different.

1. #define, a preprocessor macro.

#define PI 3.14159
#define MAX_STUDENTS 60

The preprocessor replaces the name with the text. There is no type, no storage and no scope: the substitution happens everywhere below the directive, and a mistake in it is reported against the line where the name was used, not where it was defined. By convention the name is in capitals.

2. const, a qualified variable.

const double pi = 3.14159;
const int max_students = 60;

This is a real object with a real type, which the compiler will not let you assign to. It obeys scope, it appears in a debugger, and the compiler type-checks every use of it. This is the one to prefer.

3. enum, for a set of related whole numbers.

Best where the values are a group: days, states, menu choices, error codes. It gives them a type as well as values, which #define cannot.

#defineconstenum
Handled byPreprocessorCompilerCompiler
Has a typeNoYesYes, int
Obeys scopeNoYesYes
Type-checkedNoYesYes
Can be a fractionYesYesNo
Usable as an array sizeYesNot portably in CYes
Visible to a debuggerNoYesYes

The array-size row is the one real reason #define survives in C. const int n = 10; int a[n]; is a variable-length array, which C99 allows and which is not the same thing as a fixed-size array; #define N 10 and enum { N = 10 }; both give a genuine constant expression.

What this does NOT mean

A constant is not a variable that cannot change. A constant has no memory of its own; it is a value written in the program. A const variable is a variable that cannot change, which is a different thing with a confusingly similar name.

#define does not create a constant of any type. It creates a text substitution. #define HALF 1/2 used as x HALF becomes x 1/2, which is (x * 1) / 2, and for an integer x of 3 that is 1 rather than 1.5. Chapter 22 is why macro definitions get brackets round them.

munotes.in47

Constants and Their Types

A leading zero is not formatting. 0113 is octal. If you want the decimal number 113, write 113.

'A' is not of type char. In C a character constant has type int. (In C++ it is char, which is one of the few differences that shows up in a first program.)

3.14 is not a float. It is a double. Assigning it to a float narrows it, and 3.14f is how you write a float constant.

A string literal is not modifiable. char *s = "Hello"; s[0] = 'J'; compiles and may crash. Use an array if you mean to change the text: char s[] = "Hello"; copies it into memory you own.

Quick revision

  • Four kinds of constant: integer, floating, character, enumeration. String literals are a fifth thing and are arrays.
  • Integer bases: decimal, octal with a leading 0, hexadecimal with 0x. 011 is 9.
  • Integer suffixes: U, L, LL and combinations. Prefer capital L.
  • A floating constant is a double unless suffixed f or L.
  • A character constant is one character in single quotes and has type int, with the value of the character's code.
  • 'A' is 65 in ASCII; '0' is 48; c - '0' converts a digit character to a number.
  • A string literal in double quotes carries a terminating \0 and must not be modified.
  • enum numbers from 0, continues from any value you set, and has type int.
  • Three ways to name one: #define (text, no type), const (typed, scoped, preferred), enum (a typed set of ints).

Test yourself

1. What is the value of 010 + 10?

  1. 010 is octal for 8, and 10 is decimal ten.

2. What is the type of 3.14, and how do you write the same value as a float?

double. Write 3.14f.

3. What does printf("%d\n", 'a') print, and why is it not an error?

97 on an ASCII machine. A character constant has type int and its value is the character's code, so %d is the correct conversion for it.

4. How many bytes does "Hi" occupy?

Three: H, i and the terminating \0.

5. Give one thing const can do that #define cannot, and one thing #define can do that const cannot.

const has a type, so the compiler checks every use of it, and it obeys scope. #define yields a constant expression usable as a fixed array size, which a const int does not in C.

munotes.in48

Constants and Their Types

6. After enum colour { RED = 5, GREEN, BLUE };, what is BLUE?

  1. Counting continues from the value that was set, so GREEN is 6 and BLUE is 7.

7. Why does 7 / 2 give 3 while 7.0 / 2 gives 3.5?

In the first both operands are integer constants, so integer division discards the fractional part. In the second 7.0 is a double, so the 2 is converted to double and the division is done in floating point.

What can be asked on this, and how to answer it

"What is a constant? Explain the types of constants in C with examples." Define it, then give the four kinds with two examples each, and the rules that fix the type: octal and hexadecimal notation, the integer suffixes, floating constants being double by default, and a character constant having type int. Mention string literals and their terminating \0.

"Distinguish between #define and const." Give the table's rows: preprocessor against compiler, no type against a type, no scope against scope, unchecked against type-checked, and the one advantage #define keeps, which is a constant expression for an array size. Prefer const and say so.

"What is an enumeration? How are its values assigned?" A set of named integer constants declared with enum. Values start at 0, increase by one, and may be set explicitly, after which counting continues from the set value. Give a day or a grade example.

"What is the difference between 'A' and "A"?" 'A' is a character constant of type int with the value 65 in ASCII, occupying the space of an int. "A" is a string literal: an array of two char, A followed by \0. This is a favourite one-mark question.

"Write a program to show that a character constant has a numeric value." Print 'A' with both %c and %d, and 'A' + 1 as well, as this chapter does.

Contents This chapter on its own page

munotes.in49

Chapter Eleven

Variables: Declaration, Definition and Initialisation

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

A variable is a named piece of memory of a known type, and you bring one into existence by writing its type and its name, after which the name stands for whatever value is currently stored there.

Why the compiler has to be told first

The compiler has to know three things before it can compile a single use of a name: how many bytes to set aside, how to read the bits, and therefore what the operators are allowed to do with it. None of that can be worked out from the way the name is used later, because the compiler reads the file once, from the top.

So C requires you to say it in advance. The statement that says it is a declaration, and in the ordinary case it is also the thing that creates the variable.

Declaration and definition

The two words are used loosely everywhere and they are not the same.

A definition creates the variable: it sets aside the memory.

int count;

A declaration describes a variable and promises that a definition exists somewhere else.

extern int count;

extern is the keyword for that promise. It is how two source files share one variable: one file defines it, the others declare it. In a single-file program you will not need extern, and you should read int count; as doing both jobs at once.

The pairing is easier to see with functions, which is where you will actually meet it. double area(double r); with a semicolon is a declaration. The same line with a body in braces is a definition. Chapter 32.

For the exam, the sentence to write is: every definition is also a declaration, but a declaration need not be a definition. A variable may be declared many times and must be defined exactly once.

Declaring a variable

int marks;
double average;
char grade;
int a, b, c;
unsigned long population;

The type comes first, then one or more names separated by commas, then a semicolon. Declaring several on one line is legal and is worth avoiding once the names stop being related: int i, j, k; is fine for loop counters, and int count, total, flag; invites a reader to think those three have something to do with each other.

Where a declaration may go. In C99 and later, anywhere a statement may go. The old rule, that declarations came first in a block, was dropped in 1999 and only the very oldest compilers still enforce it.

Declare a variable where you first need it, not at the top. A variable that comes into existence one line before its first use has a shorter life for a reader to keep track of, and often the declaration and the initial value can then be the same line.

munotes.in50

Variables: Declaration, Definition and Initialisation

Initialisation

Initialisation is giving a variable a value at the moment it is created. Assignment is giving it a value later. They look similar and they are different events.

int marks = 75;          /* initialisation: created and set at once */
int total;               /* created, holding nothing useful       */
total = 0;               /* assignment: set afterwards            */

Several at once, each with its own value:

int a = 1, b = 2, c = 3;

int a = b = c = 0; does not do what it looks like. b and c must already exist; only a is being declared. Write the three separately.

The thing that goes wrong: an uninitialised variable

This is the most important paragraph in the chapter. A variable declared inside a function without an initialiser holds whatever was in that memory already. Not zero. Not a random number in any useful sense. Simply whatever was there.

The program below is wrong on purpose, and the compiler says so.

#include <stdio.h>

int main(void)
{
    int total;
    int count = 5;

    total = total + count;
    printf("total is now %d\n", total);
    return 0;
}

The compiler's own words, which are the point of the listing:

uninitialised.c: In function ‘main’:
uninitialised.c:8:11: warning: ‘total’ is used uninitialized [-Wuninitialized]
    8 |     total = total + count;
      |     ~~~~~~^~~~~~~~~~~~~~~
uninitialised.c:5:9: note: ‘total’ was declared here
    5 |     int total;
      |         ^~~~~

The value that program prints is not a fact about your machine or ours. Reading an uninitialised automatic variable is undefined behaviour: the standard permits the program to print anything, to print a different thing each time, or to do something else entirely. A program that appears to work because the memory happened to contain zero is the worst possible outcome, because it will stop appearing to work on the day it matters.

The habit that prevents it entirely: never declare a variable without a value unless the very next thing you write gives it one. A counter starts at 0. A running total starts at 0. A running product starts at 1. A "largest so far" starts at the first element, not at zero, which is chapter 36's trap.

Where variables live, and what that changes

Three places, and where a variable is declared decides both how long it lives and who can see it.

Automatic, the default: declared inside a function or a block.

void f(void)
{
    int x = 1;       /* created when f starts, gone when f returns */
}

Created on entry, destroyed on exit, and not initialised unless you say so. Every call gets a fresh one. This is what almost every variable in this course is.

munotes.in51

Variables: Declaration, Definition and Initialisation

Static: declared with static inside a function.

void f(void)
{
    static int calls = 0;   /* created once, survives between calls */
    calls++;
}

Created once before the program starts, lives until it ends, keeps its value between calls, and is initialised to zero if you do not initialise it. Visible only inside the function.

Global, also called external: declared outside every function.

int shared = 0;

void f(void) { shared++; }

Lives for the whole program, visible to every function in the file, and is initialised to zero if you do not initialise it.

AutomaticStatic (in a function)Global
DeclaredInside a blockInside a block, with staticOutside all functions
CreatedOn entry to the blockOnce, before mainOnce, before main
DestroyedOn leaving the blockAt program endAt program end
Default valueNone. Undefined00
Visible toThat blockThat functionThe whole file

Use automatic variables. A global is visible to everything, so any function may have changed it, and working out what a program does then means reading all of it. Chapter 21 says the same thing about blocks.

Here are the three side by side, with the counting they make possible:

#include <stdio.h>

int total_calls = 0;              /* global: zero by default, set here for clarity */

int count_up(void)
{
    static int mine = 0;          /* static: survives between calls */
    int fresh = 0;                /* automatic: new every call */

    mine = mine + 1;
    fresh = fresh + 1;
    total_calls = total_calls + 1;
    printf("static %d, automatic %d, global %d\n", mine, fresh, total_calls);
    return mine;
}

int main(void)
{
    count_up();
    count_up();
    count_up();
    return 0;
}
static 1, automatic 1, global 1
static 2, automatic 1, global 2
static 3, automatic 1, global 3

The static variable counts. The automatic one is created afresh each time and cannot. That difference is the whole idea, and it is a standing exam question.

Naming, which is not a side issue

Chapter 6 made clarity the second characteristic after integrity, and a variable's name is where most of a program's clarity comes from.

  • Say what it holds, and in what unit. rate_percent rather than rate. seconds_elapsed rather than time.
  • Length in proportion to scope. i is a perfectly good name for a loop counter that lives three lines. A global needs a name that makes sense far from where it was declared.
  • Be consistent. total_marks and totalMarks in one program means a reader has to remember which you used where. This book uses lower case with underscores, which is the usual style in C.
  • Do not encode the type. int iCount tells a reader what the declaration already told them.
munotes.in52

Variables: Declaration, Definition and Initialisation

What this does NOT mean

Declaring a variable does not give it a value. For an automatic variable it gives you a name for some memory whose contents are whatever they were. This is the single most common first-year bug.

Initialisation is not assignment. Initialisation happens once, when the object is created, and is the only way to set a const object or an array's contents. Assignment happens whenever the statement runs. For const int n = 5; the first is legal and n = 6; is not.

static does not mean constant. A static variable can be changed as freely as any other. What static fixes is how long it lives and who can see it.

A global variable is not "faster because it is not passed". It is a name with a long life and no owner. The cost of passing a parameter is nothing; the cost of not knowing who changed a value is hours.

C does not zero your automatic variables and you must not rely on a compiler that appears to. Many compilers in a debug build fill new memory with a recognisable pattern, which makes the bug visible. The same program in a release build will do something else.

Quick revision

  • A variable is named memory of a known type. int marks; defines one.
  • A definition creates storage. A declaration, with extern, only promises one exists.
  • Every definition is a declaration; not every declaration is a definition.
  • Since C99, a declaration may go anywhere a statement may go. Declare at first use.
  • Initialisation sets a value when the object is created; assignment sets it later.
  • An uninitialised automatic variable holds rubbish, and reading it is undefined behaviour.
  • Static and global variables are zero by default. Automatic ones are not.
  • A static variable in a function is created once and keeps its value between calls.
  • Prefer automatic variables. Prefer names that say what the value is and in what unit.

Test yourself

1. What is the difference between a declaration and a definition?

A definition creates the variable and sets aside memory for it. A declaration states the name and type and promises a definition exists elsewhere, which is what extern int x; does. A variable is defined once and may be declared many times.

2. What is the value of x after int x; inside a function?

Unspecified. The variable holds whatever was in that memory, and reading it before assigning to it is undefined behaviour.

3. And after static int x; inside a function?

Zero. Static and global objects with no initialiser are set to zero before the program starts.

munotes.in53

Variables: Declaration, Definition and Initialisation

4. What does this print, and why?

void f(void) { static int n = 0; n++; printf("%d ", n); }
/* called three times */

1 2 3. The static variable is created once and keeps its value between calls.

5. Why is int a = b = 0; a mistake in a declaration?

Because only a is being declared. b must already exist, and if it does not the line does not compile. Declare and initialise each variable in its own right.

6. Give two reasons to prefer an automatic variable to a global one.

It cannot be changed by any other function, so the reasoning about its value is local. And it exists only while it is needed, so two parts of the program cannot accidentally share it.

What can be asked on this, and how to answer it

"What is a variable? How is it declared and initialised?" Define it as a named location in memory of a given type whose value can change while the program runs. Give the declaration form, the initialisation form, and the distinction between initialisation and assignment. Add that an automatic variable with no initialiser has no defined value, which is the part that earns the extra mark.

"Distinguish between declaration and definition of a variable." Give the two-sentence answer above with int x; against extern int x;, and the rule that one definition and many declarations are allowed.

"Explain the storage classes in C." Name auto, static, extern and register, and for each give where it is written, how long the variable lives, who can see it and what its default value is. The table in this chapter is the answer; register is a request that the value be kept in a processor register, which modern compilers ignore, and its one real consequence is that you cannot take its address.

"What happens if a variable is used without initialising it?" Its value is whatever was in that memory, reading it is undefined behaviour, and the program may print anything or behave differently on different runs or different compilers. Static and global variables are the exception: they are zero.

"Write a program to show the difference between an automatic and a static variable." Give the three-call program in this chapter and point at the two columns of its output.

Contents This chapter on its own page

munotes.in54

Chapter Twelve

Characters and Character Strings

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

A character is a one-byte integer holding a character's code, and a character string is a run of characters in consecutive memory ending with a character whose value is zero, which is how every part of C tells where the string stops.

Why strings work this way in C

C has no string type. This is a deliberate and slightly shocking design decision, and it follows from what the language was for. A string type would need a length stored somewhere, a decision about how long a string may be, and code to manage it, and C was built to be small enough to fit a 1970s machine and close enough to the hardware to write an operating system.

So C provides the minimum that works: a string is just characters in memory, one after another, and the end is marked by a zero byte. Nothing keeps a length. Nothing checks a bound. Every string function in the library works by walking forward until it meets that zero.

Everything good and everything dangerous about strings in C comes from that one sentence.

A character is a small integer

#include <stdio.h>

int main(void)
{
    char grade = 'A';
    char digit = '7';

    printf("grade as a character: %c\n", grade);
    printf("grade as a number   : %d\n", grade);
    printf("digit '7' is the number %d\n", digit);
    printf("'7' turned into 7 is %d\n", digit - '0');
    printf("the next letter after %c is %c\n", grade, grade + 1);
    return 0;
}
grade as a character: A
grade as a number   : 65
digit '7' is the number 55
'7' turned into 7 is 7
the next letter after A is B

A char is one byte and holds an integer. When you print it with %c the library shows you the character that code stands for; with %d it shows the code. Both are the same byte.

digit - '0' is the standard way of turning the character '7' into the number 7, and it works because the standard guarantees the ten digit characters have consecutive values. It has nothing to do with ASCII in particular, and it is the conversion atoi and scanf do internally.

grade + 1 has type int, not char, because C widens a char to an int before doing arithmetic on it. Printing it with %c narrows it back. Chapter 14 is the rule that does this.

Reading and writing one character

Three ways, and they are not interchangeable.

char c;
scanf("%c", &c);        /* takes the very next character, spaces included  */
scanf(" %c", &c);       /* the leading space skips whitespace first        */
c = getchar();          /* takes the next character; returns int, not char */
putchar(c);             /* writes one character                           */
munotes.in55

Characters and Character Strings

The space in " %c" is the fix for the commonest bug in first-semester C. After reading a number with scanf("%d", &n), the newline you pressed is still waiting in the input, so the next "%c" reads the newline instead of what the user types. A leading space in the format tells scanf to skip any amount of whitespace first.

getchar returns an int, not a char, and that is not an oversight. It has to be able to return every possible character and the separate value EOF for "there is no more input", and a char has no spare value to use. Store it in an int.

A string is an array of characters with a zero on the end

#include <stdio.h>

int main(void)
{
    char name[] = "Anita";

    printf("the string is %s\n", name);
    printf("it needs %zu bytes\n", sizeof name);
    printf("byte by byte: ");
    for (int i = 0; i < (int) sizeof name; i++) {
        printf("[%d]=%d ", i, name[i]);
    }
    printf("\n");
    printf("as characters: ");
    for (int i = 0; i < (int) sizeof name - 1; i++) {
        printf("%c", name[i]);
    }
    printf("\n");
    return 0;
}
the string is Anita
it needs 6 bytes
byte by byte: [0]=65 [1]=110 [2]=105 [3]=116 [4]=97 [5]=0
as characters: Anita

Five letters, six bytes. The sixth is the zero the compiler added, and printing the bytes as numbers is the clearest way to see it. name[5] is 0, and it is not the character '0', whose code is 48. The difference between '\0' and '0' is worth a mark on its own.

char name[] = "Anita"; lets the compiler count for you, and it counts the terminator. Writing the size yourself is where mistakes live.

The four ways to declare a string, and which is which

char a[] = "Anita";       /* array of 6, contents copied in, changeable   */
char b[20] = "Anita";     /* array of 20, first 6 used, rest set to zero  */
char c[6] = "Anita";      /* exactly fits: 5 letters and the terminator   */
const char *d = "Anita";  /* a POINTER to the literal, NOT changeable     */

char e[5] = "Anita"; compiles in C and gives you an array of five characters with no terminator. It is not a string, and passing it to printf("%s") walks off the end of the array. Count the terminator.

The difference between a and d is the one that matters and it is not cosmetic:

  • a is an array. The characters are yours, in your memory, and a[0] = 'B'; is legal.
  • d points at the string literal itself, which the program may have put in read-only memory. d[0] = 'B'; compiles without const and may crash.
munotes.in56

Characters and Character Strings

The rule: if you mean to change the text, use an array. If you only mean to read it, use const char *.

What goes wrong, and it is always the terminator

1. The array is one byte too small. char s[5] = "Anita"; has no room for the zero. char s[6] or char s[] is correct.

2. strlen and sizeof are different numbers. strlen counts the characters up to the terminator. sizeof gives the size of the array, terminator and all unused bytes included.

#include <stdio.h>
#include <string.h>

int main(void)
{
    char a[] = "Anita";
    char b[20] = "Anita";

    printf("a: strlen %zu, sizeof %zu\n", strlen(a), sizeof a);
    printf("b: strlen %zu, sizeof %zu\n", strlen(b), sizeof b);
    return 0;
}
a: strlen 5, sizeof 6
b: strlen 5, sizeof 20

3. scanf("%s") stops at the first space. Reading "Anita Desai" with %s gives you "Anita" and leaves the rest waiting. Use fgets when the input has spaces in it.

4. scanf("%s") has no idea how big your array is. It writes as many characters as arrive. scanf("%19s", name) for a char name[20] is the safe form: 19 characters and the terminator.

5. A string cannot be compared with ==. if (a == b) compares two addresses, not two texts, and is almost always false. strcmp is chapter 37.

6. A string cannot be copied with =. a = b; does not compile for arrays. strcpy is chapter 37.

Reading a whole line safely

#include <stdio.h>
#include <string.h>

int main(void)
{
    char line[40];

    if (fgets(line, (int) sizeof line, stdin) != NULL) {
        line[strcspn(line, "\n")] = '\0';     /* cut the newline off */
        printf("you typed [%s], %zu characters\n", line, strlen(line));
    }
    return 0;
}
Anita Desai
you typed [Anita Desai], 11 characters

fgets takes the size of the array, so it cannot overflow it, and it keeps the newline you pressed. The strcspn line finds the position of the first newline and puts a terminator there, which is the ordinary way of trimming it. If there was no newline, strcspn returns the length and the assignment is harmless.

What this does NOT mean

A string is not a data type. There is no string in C. There is an array of char with a convention about its last byte, and a library of functions that all obey that convention.

'\0' is not '0', and neither is "0". '\0' has the value 0. '0' has the value 48. "0" is two bytes: 48 and then 0.

sizeof is not the length. For char b[20] = "Anita";, sizeof b is 20 and strlen(b) is 5.

A char is not guaranteed to be signed. Whether plain char can hold a negative value is left to the implementation. It matters only if you store numbers in a char, in which case write signed char or unsigned char.

munotes.in57

Characters and Character Strings

gets does not exist any more, and never use it. It reads a line with no idea of the size of your array. It was removed from the language in C11, and any book that shows it is describing a language that no longer has it. fgets is the replacement.

Quick revision

  • A char is one byte holding an integer: the character's code.
  • %c prints the character, %d prints the code.
  • c - '0' converts a digit character to its value, guaranteed by the standard.
  • A string is characters in consecutive memory ending in '\0'.
  • "Anita" occupies 6 bytes. n characters need n + 1.
  • char s[] = "..." lets the compiler count, terminator included.
  • char s[] is changeable memory of yours; const char *s points at a literal you must not change.
  • strlen counts to the terminator; sizeof gives the whole array.
  • getchar returns int, so that EOF is distinguishable.
  • scanf(" %c", &c): the leading space skips the leftover newline.
  • scanf("%s") stops at whitespace and does not know your array size. Use %19s or fgets.
  • Never gets. It was removed from the language.

Test yourself

1. How many bytes does char s[] = "Hello"; occupy, and what is in the last one?

Six. The last holds '\0', the value zero, which marks the end of the string.

2. What is the difference between 'A', "A" and '\0'?

'A' is a character constant of type int with the value 65 in ASCII. "A" is a string literal: two bytes, 65 and 0. '\0' is the null character, value 0.

3. char s[5] = "Hello"; compiles. Why is it wrong?

There is no room for the terminator, so s holds five characters and is not a string. Any library function or %s conversion will read past the end of the array.

4. Why does getchar return an int?

So that it can return any character value and also the distinct value EOF, which means there is no more input. A char has no value left over to mean that.

5. After scanf("%d", &n), why does scanf("%c", &c) seem to read nothing?

Because the newline from the Return key is still in the input, and %c reads it. Write scanf(" %c", &c): the leading space skips whitespace first.

6. Your program must read a full name with a space in it. Which input function, and why not the other?

fgets, because scanf("%s") stops at the first whitespace character and would read only the first word.

munotes.in58

Characters and Character Strings

7. For char b[20] = "Anita";, what are strlen(b) and sizeof b?

5 and 20.

What can be asked on this, and how to answer it

"What is a string in C? How is it stored?" A string is an array of characters terminated by the null character '\0'. It is stored as the characters in consecutive bytes with the terminator after the last one, so n characters need n + 1 bytes, and there is no separate length: every library function finds the end by looking for the terminator.

"What is the significance of the null character?" It marks the end of a string. Every string literal gets one automatically, and the library's string functions depend on it. Without one, a function walks past the end of the array, which is undefined behaviour.

"Distinguish between a character and a string." A character is one byte in single quotation marks holding a code; a string is an array of characters in double quotation marks with a terminator. 'A' is one byte's worth of value; "A" is two bytes.

"Distinguish between strlen and sizeof for a character array." strlen is a library function that counts characters up to the terminator, at run time. sizeof is a compile-time operator giving the size of the whole array including the terminator and any unused space.

"Write a program to read a string and display it." Use fgets with sizeof of the array, trim the newline, and print with %s. Say in one line why not gets: it has no bound and was removed from the language in C11.

Contents This chapter on its own page

munotes.in59

Chapter Thirteen

typedef

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

typedef gives an existing type a second name, so that a declaration can say what a value means rather than only how it is stored.

Why you would want to

unsigned long says how a value is stored. It says nothing about what it is. Compare:

unsigned long a;
unsigned long b;

with

typedef unsigned long Bytes;
typedef unsigned long Milliseconds;

Bytes a;
Milliseconds b;

The machine code is identical. What changed is that a reader of the second version knows what the two variables hold, and a reader of the first has to find out.

There is a second reason, and in practice it is the one that makes typedef indispensable: some C type names are genuinely hard to read. The declaration of a pointer to a function returning int and taking two double parameters is int (*f)(double, double), and a typedef turns that into a name you can use like any other.

The syntax, and the trick for remembering it

typedef existing-type new-name;

The trick: write the declaration you would have written, then put typedef in front of it, and the name being declared becomes the name of the type.

unsigned long bytes;            /* declares a VARIABLE called bytes      */
typedef unsigned long bytes;    /* declares a TYPE called bytes          */

That rule is worth more than memorising cases, because it handles the awkward ones automatically:

char line[80];                  /* a variable: an array of 80 char       */
typedef char Line[80];          /* a type: "array of 80 char"            */
Line a, b;                      /* two arrays of 80 char each            */
int *p;                         /* a variable: pointer to int            */
typedef int *IntPtr;            /* a type: "pointer to int"              */
IntPtr p, q;                    /* BOTH are pointers, which is the point */

That last line is the one case where typedef does something #define cannot. With #define IntPtr int , the line IntPtr p, q; expands to int p, q;, which declares one pointer and one plain int. The typedef declares two pointers, because it names a type rather than substituting text.

A working example

#include <stdio.h>

typedef unsigned int Marks;
typedef double Percentage;
typedef char Name[20];

Percentage as_percentage(Marks got, Marks total)
{
    return (Percentage) got * 100.0 / total;
}

int main(void)
{
    Name student = "Anita";
    Marks got = 63;
    Marks total = 75;
    Percentage pc = as_percentage(got, total);

    printf("%s scored %u out of %u, which is %.2f per cent\n",
           student, got, total, pc);
    printf("sizeof(Marks) is %zu and sizeof(unsigned int) is %zu\n",
           sizeof(Marks), sizeof(unsigned int));
    return 0;
}
Anita scored 63 out of 75, which is 84.00 per cent
sizeof(Marks) is 4 and sizeof(unsigned int) is 4

The function signature now reads as what it does: it takes marks and returns a percentage. sizeof(Marks) and sizeof(unsigned int) are the same number, which is the proof that nothing new was created.

munotes.in60

typedef

What typedef does NOT do, with proof

It does not create a new type. The name is an alias, and the two are the same type for every purpose the language has.

#include <stdio.h>

typedef int Metres;
typedef int Feet;

int main(void)
{
    Metres distance = 100;
    Feet height = 6;

    /* Nonsense, and the compiler has no objection: both are int. */
    int nonsense = distance + height;

    printf("adding metres to feet gives %d, and -Wall -Wextra said nothing\n",
           nonsense);
    return 0;
}
adding metres to feet gives 106, and -Wall -Wextra said nothing

That program compiles clean. Adding metres to feet is meaningless and the compiler cannot know it, because Metres and Feet are both spellings of int. So typedef is documentation, not protection. If you need the compiler to keep two quantities apart, you need two different structure types, which is chapter 42.

It does not allocate anything. typedef is a declaration, not a definition of an object. No memory is set aside.

It cannot extend a type. typedef cannot make an integer bigger or add an operation. It only renames.

It is not a macro. The pointer example above is the proof. typedef is handled by the compiler and understands the grammar of declarations; #define is handled by the preprocessor and understands only text.

typedef#define
Handled byCompilerPreprocessor
Knows it is a typeYesNo
Obeys scopeYesNo
X p, q; where X is a pointer typeBoth are pointersFirst a pointer, second is not
Can rename an array or function typeYesNot usefully
Can define a constant or a code fragmentNoYes

Where typedef really earns its place

Three uses account for nearly all of it in real C.

1. Shortening a structure name. In C, a structure's name includes the keyword: struct student. A typedef lets you write Student. This is the commonest use by a wide margin, and chapter 42 gives it.

2. Naming a pointer-to-function type, which is otherwise unreadable.

typedef int (*Comparison)(const void *, const void *);

3. Making a program's own vocabulary portable. The standard library does this everywhere, and you have already used the results: size_t is a typedef for whatever unsigned type is right for a size on this machine, and FILE, time_t, ptrdiff_t and wchar_t are all typedef names. That is why sizeof prints with %zu and not %u: the actual type behind size_t differs between machines, and the typedef is what lets your program not care.

munotes.in61

typedef

That last point is the honest answer to "why does typedef exist": it lets a program name a type by its purpose while the compiler chooses the representation.

Naming style

There is no rule, and there are two common conventions.

  • Capitalised: Marks, Percentage, Student. Used in this book, because it makes a type name visible at a glance.
  • _t suffix: marks_t, student_t. Common in systems code, and it imitates the library's size_t and time_t.

Strictly, names ending in _t are reserved by POSIX for the system's own headers, so the capitalised form is the safer habit. Either way, be consistent within a program.

Quick revision

  • typedef existing-type new-name; gives a type a second name.
  • The trick: write the variable declaration, put typedef in front, and the name declared becomes a type name.
  • It creates no new type: sizeof proves it, and two typedef names for int add together without complaint.
  • It allocates nothing and is handled by the compiler, not the preprocessor.
  • Unlike #define, it knows it is declaring a type, so IntPtr p, q; makes two pointers.
  • It obeys scope, so a typedef inside a function is local to it.
  • Main uses: shortening struct names, naming function-pointer types, and giving a portable name to a machine-dependent type. size_t, FILE and time_t are library typedefs.

Test yourself

1. Write a typedef that names the type "array of 10 double", and declare two such arrays with it.

typedef double Vector[10];
Vector a, b;

2. Is sizeof(Marks) the same as sizeof(unsigned int) after typedef unsigned int Marks;?

Yes. Marks is another name for unsigned int, not a new type.

3. What is wrong with #define IntPtr int * followed by IntPtr p, q;?

It expands to int * p, q;, so p is a pointer and q is a plain int. A typedef would have made both pointers.

4. Will the compiler stop you adding a Metres to a Feet if both are typedef int?

No. They are the same type, so there is nothing for it to object to. Use two distinct structure types if you want that checked.

5. Name three typedef names from the standard library and say why each is a typedef rather than a fixed type.

size_t, time_t and FILE. Each hides a representation that differs between implementations, so a program can name the purpose and let the library choose the type.

6. Can a typedef be written inside a function?

Yes, and it then obeys the same scope rules as any other declaration in that block: it is visible from that point to the end of the block and nowhere else.

munotes.in62

typedef

What can be asked on this, and how to answer it

"What is typedef? Explain with an example." Define it as a way of giving an existing type an additional name, give the syntax, give two examples of which one is an array or pointer type, and say plainly that no new type is created. The last part is where the marks separate.

"Distinguish between typedef and #define." Give the table's rows: compiler against preprocessor, type-aware against text substitution, scoped against not scoped, and the pointer example where the two genuinely differ.

"What are the advantages of typedef?" Readability, because a declaration can name the purpose of a value; brevity, because struct student becomes Student; portability, because a program can use a name whose underlying type the implementation chooses; and maintainability, because a representation can be changed in one line.

"Does typedef create a new data type?" No. It creates a new name for an existing type. The two names are interchangeable everywhere, and the compiler will not distinguish values of them.

Contents This chapter on its own page

munotes.in63

Chapter Fourteen

Type Conversion and Typecasting

Syllabus topic 1, "Introduction: Algorithms, History of C, Structure of C Program. Program Characteristics, Compiler, Linker and preprocessor, pseudo code statements and flowchart symbols, Desirable program characteristics. Program structure. Compilation and Execution of a Program, C Character Set, identifiers and keywords, data types and sizes, constants and its types, variables, Character and character strings, typedef, typecasting"

In one line

Type conversion is C changing the type of a value for you, following fixed rules, and a typecast is you changing it deliberately by writing the type you want in brackets.

Why conversions happen at all

A processor cannot add an integer to a floating-point number. They are stored in completely different ways, and the instruction that adds two integers is not the instruction that adds two floating-point values. So when you write

double average = total / 3;

something has to give, and C has rules for deciding what. The rules are not arbitrary and they are not hard, and they are worth learning properly, because the cost of not knowing them is wrong answers that look plausible.

The two kinds

Implicit conversion, which C does on its own. Also called automatic conversion or coercion. It happens in four places: in an expression with mixed types, on assignment, when passing an argument to a function, and when returning a value.

Explicit conversion, which you write. This is a cast, and the syntax is the type name in brackets before the value:

(double) total
(int) 3.7
(char) 65

Rule 1: integer promotion

Anything narrower than int becomes an int before arithmetic happens to it. char, short, _Bool and the signed and unsigned versions of them are all promoted.

#include <stdio.h>

int main(void)
{
    char a = 100;
    char b = 100;

    printf("as char arithmetic would overflow, but this is int arithmetic: %d\n",
           a + b);
    printf("sizeof a is %zu, but sizeof (a + b) is %zu\n",
           sizeof a, sizeof(a + b));
    return 0;
}
as char arithmetic would overflow, but this is int arithmetic: 200
sizeof a is 1, but sizeof (a + b) is 4

a + b is 200, not an overflow, because both were promoted to int first. And sizeof(a + b) is the size of an int, which is the proof.

This is also why getchar returns an int, why a character constant has type int, and why 'A' + 1 prints as 66 with %d. Chapter 12 met the symptom; this is the cause.

Rule 2: the usual arithmetic conversions

When an operator has two operands of different types, the narrower one is converted to the wider. The order, from widest down, for the types you will use:

long double  >  double  >  float  >  unsigned long long  >  long long
 >  unsigned long  >  long  >  unsigned int  >  int

So in an expression with a double and an int, the int becomes a double and the arithmetic is floating-point. In an expression with two ints, the arithmetic is integer, whatever you are going to do with the answer.

munotes.in64

Type Conversion and Typecasting

That last sentence is where the marks are lost. The type of an expression is decided by the operands, not by what you assign it to.

#include <stdio.h>

int main(void)
{
    int total = 7;
    int count = 2;

    double wrong = total / count;
    double right = (double) total / count;

    printf("7 / 2 assigned to a double  : %.4f\n", wrong);
    printf("(double) 7 / 2              : %.4f\n", right);
    printf("7 %% 2 is %d, so the fraction thrown away was %d/2\n",
           total % count, total % count);
    return 0;
}
7 / 2 assigned to a double  : 3.0000
(double) 7 / 2              : 3.5000
7 % 2 is 1, so the fraction thrown away was 1/2

total / count is integer division, giving 3. Assigning 3 to a double gives 3.0. The fraction was gone before the assignment happened, and no amount of double on the left recovers it.

The fix is to make one operand floating-point. (double) total / count promotes count too, by rule 2. So does total / (double) count, and so does total * 1.0 / count.

Rule 3: assignment converts to the left-hand type

On assignment the value is converted to the type of the variable, and if it will not fit, something is lost. Two cases matter.

Floating-point to integer: the fraction is discarded, not rounded.

#include <stdio.h>
#include <math.h>

int main(void)
{
    printf("(int) 3.7   is %d\n", (int) 3.7);
    printf("(int) 3.2   is %d\n", (int) 3.2);
    printf("(int) -3.7  is %d\n", (int) -3.7);
    printf("round(3.7)  is %.0f\n", round(3.7));
    printf("round(-3.7) is %.0f\n", round(-3.7));
    return 0;
}
(int) 3.7   is 3
(int) 3.2   is 3
(int) -3.7  is -3
round(3.7)  is 4
round(-3.7) is -4

Truncation is towards zero, so (int) -3.7 is -3 and not -4. If you want rounding, use round from <math.h>; if you want the largest integer no greater than the value, use floor.

Integer to a narrower integer: the high bits are discarded. Assigning 300 to a char does not give you 300 and does not give you an error. -Wall -Wextra will usually warn, which is the reason to compile with them.

What a cast is for, and when to use one

Four legitimate uses, and they are worth knowing as a list, because a cast is also a way of telling the compiler to stop objecting, which is a way of hiding a bug.

1. To force floating-point arithmetic. The most common use in this course.

double average = (double) total / count;

2. To discard a fraction deliberately, where truncation is what you mean.

int whole_rupees = (int) amount;

3. To pick the right overload of a library function, or to satisfy a parameter type, most often (double) for a float or (int) for a char in a printf argument.

munotes.in65

Type Conversion and Typecasting

4. To say "I know this is unused", with (void).

(void) unused_parameter;

A cast used to silence a warning you do not understand is a bug you have hidden. -Wall warning about a conversion is usually telling you the truth.

The worked example: the average of three marks

This is the shape of several of the lab's practicals, and it is where the rule bites.

#include <stdio.h>

int main(void)
{
    int m1 = 63, m2 = 58, m3 = 72;
    int total = m1 + m2 + m3;

    printf("total is %d\n", total);
    printf("wrong  : total / 3           = %d\n", total / 3);
    printf("wrong  : assigned to a double = %.4f\n", (double) (total / 3));
    printf("right  : (double) total / 3   = %.4f\n", (double) total / 3);
    printf("right  : total / 3.0          = %.4f\n", total / 3.0);
    return 0;
}
total is 193
wrong  : total / 3           = 64
wrong  : assigned to a double = 64.0000
right  : (double) total / 3   = 64.3333
right  : total / 3.0          = 64.3333

Both correct forms agree, and both differ from the wrong one. Read the third line carefully: casting after the division changes nothing, because the division has already happened. The cast has to be on an operand.

What this does NOT mean

A cast does not change the variable. (int) x produces a new value of type int; x is untouched and still a double. A cast is an operator producing a result, not an instruction to the variable.

Conversion is not rounding. Converting a floating-point value to an integer discards the fractional part, towards zero.

Assigning to a double does not make the arithmetic floating-point. The expression's type was fixed by its operands before the assignment was reached.

An implicit conversion is not an error, and that is the problem. C converts silently in most cases, so a lost fraction or a truncated value does not stop the program. -Wall -Wextra is what makes the risky ones visible.

A cast between unrelated pointer types is not a conversion of the data. It reinterprets the bits. That is a later topic and it is not what this chapter is about.

char to int is not "losing information". It is a widening, and it is exact. The losing direction is int to char.

Quick revision

  • Implicit conversion is C's; explicit conversion is a cast, written (type) value.
  • Rule 1, integer promotion: anything narrower than int becomes int before arithmetic.
  • Rule 2, usual arithmetic conversions: the narrower operand is converted to the wider. long double down to int.
  • Rule 3, assignment: the value is converted to the left-hand type, and may lose data.
  • The type of an expression is decided by its operands, never by what it is assigned to.
  • 7 / 2 is 3 in any context. (double) 7 / 2 is 3.5.
  • Floating-point to integer truncates towards zero: (int) -3.7 is -3. Use round to round.
  • A cast applies to an operand, so it must be written before the division, not around it.
  • Compile with -Wall -Wextra: most damaging conversions warn.
munotes.in66

Type Conversion and Typecasting

Test yourself

1. What does printf("%d\n", 5 / 2) print, and what does printf("%.2f\n", 5 / 2.0) print?

2 and 2.50. The first is integer division; in the second one operand is a double, so the other is converted and the division is floating-point.

2. double avg = sum / n; where both are int gives a whole number. Fix it two ways.

double avg = (double) sum / n; or double avg = sum / (double) n;. Either makes one operand a double, and rule 2 converts the other.

3. What is (int) -2.9?

-2. Truncation is towards zero.

4. What is the type and value of 'a' + 1?

int, and 98 on an ASCII machine. The char is promoted to int before the addition.

5. char c = 300; compiles with a warning. What is in c?

An implementation-defined value: the high bits of 300 do not fit in one byte and are discarded. It is not 300, and the warning is telling you so.

6. Why does (double) (total / 3) not give you the fractional average?

Because the division inside the brackets is integer division and has already thrown the fraction away. The cast then converts the whole number 3 to 3.0.

7. Name the four places an implicit conversion happens.

In a mixed-type expression, on assignment, when passing an argument to a function, and when returning a value.

What can be asked on this, and how to answer it

"What is type conversion? Explain implicit and explicit conversion with examples." Define both, give the three rules with one example each, and give (double) total / count as the explicit case that matters. Say that the expression's type comes from its operands, because that sentence is what the question is really testing.

"What is typecasting? Give its syntax and two uses." A cast converts a value to a named type at a point you choose, written (type) expression. Two uses: forcing floating-point division, and deliberately discarding a fraction. Add that a cast produces a value and does not change the variable.

"Explain the rules of automatic type conversion in an expression." Integer promotion first, then the usual arithmetic conversions towards the wider type, then, on assignment, conversion to the left-hand type with possible loss. Give the ordering from long double down to int.

munotes.in67

Type Conversion and Typecasting

"Write a program to find the average of three integers and explain the conversion involved." Give the program in this chapter, and explain that total / 3 is integer division, so one operand must be made a double before dividing.

"What is the difference between (int) 3.7 and round(3.7)?" The cast truncates towards zero, giving 3. round from <math.h> rounds to the nearest, giving 4.0 as a double. For -3.7 they give -3 and -4.

Contents This chapter on its own page

munotes.in68

Chapter Fifteen

Arithmetic Operators

Syllabus topic 2, "Type of operators: Arithmetic operators, relational and logical operators, Increment and Decrement operators, assignment operators, the conditional operator, Assignment operators and expression, Precedence and order of Evaluation Block Structure, Initialization, C Preprocessor"

In one line

C has five arithmetic operators, +, -, *, / and %, and the two that need care are /, because on two integers it throws the fraction away, and %, because it works only on integers.

The five

OperatorNamea = 17, b = 5
+additiona + b is 22
-subtractiona - b is 12
*multiplicationa * b is 85
/divisiona / b is 3
%remainder, or modulusa % b is 2

All five are binary: they take two operands. - is also available as a unary operator, -x, which gives the negation, and there is a unary +x, which does nothing useful and exists for symmetry.

#include <stdio.h>

int main(void)
{
    int a = 17, b = 5;

    printf("a = %d, b = %d\n", a, b);
    printf("a + b = %d\n", a + b);
    printf("a - b = %d\n", a - b);
    printf("a * b = %d\n", a * b);
    printf("a / b = %d      <- integer division\n", a / b);
    printf("a %% b = %d      <- remainder\n", a % b);
    printf("-a    = %d\n", -a);
    printf("and the identity: (a / b) * b + (a %% b) = %d\n",
           (a / b) * b + (a % b));
    return 0;
}
a = 17, b = 5
a + b = 22
a - b = 12
a * b = 85
a / b = 3      <- integer division
a % b = 2      <- remainder
-a    = -17
and the identity: (a / b) * b + (a % b) = 17

The last line is the identity that defines the pair: (a / b) * b + (a % b) is always a. If you can remember that, you can always work out what % must give.

Integer division, which is the whole difficulty

If both operands are integers, / is integer division: the fractional part is discarded and the answer is an integer. 17 / 5 is 3, not 3.4, and not 3 with something stored elsewhere.

To get a fraction, at least one operand must be a floating-point type. Chapter 14 is the rule; this is how it looks in practice:

#include <stdio.h>

int main(void)
{
    int a = 17, b = 5;

    printf("a / b            = %d\n", a / b);
    printf("(double) a / b   = %.4f\n", (double) a / b);
    printf("a / (double) b   = %.4f\n", a / (double) b);
    printf("a / 5.0          = %.4f\n", a / 5.0);
    printf("1 / 2            = %d     <- both integers\n", 1 / 2);
    printf("1.0 / 2          = %.4f\n", 1.0 / 2);
    return 0;
}
munotes.in69

Arithmetic Operators

a / b            = 3
(double) a / b   = 3.4000
a / (double) b   = 3.4000
a / 5.0          = 3.4000
1 / 2            = 0     <- both integers
1.0 / 2          = 0.5000

1 / 2 is 0. This turns up inside larger expressions and is easy to miss: x 1 / 2 is (x 1) / 2, which for an integer x of 7 is 3, while x (1 / 2) is x 0, which is 0. Neither is 3.5.

The remainder operator

% gives what is left after integer division. It requires both operands to be integers. 7.5 % 2 does not compile; there is fmod in <math.h> for floating-point remainders.

What % is actually used for is worth listing, because these five uses cover nearly every appearance of it in this course:

  1. Is it divisible? n % 2 == 0 tests even. year % 4 == 0 is the first part of the leap-year test in chapter 16.
  2. Extract the last digit. n % 10. With n / 10 to remove it, that is the whole of the reverse-the-digits practical in chapter 27.
  3. Wrap a value into a range. (i + 1) % n steps round a circle of n positions.
  4. Split a quantity into units. Seconds into minutes and seconds, paise into rupees and paise.
  5. Test a multiple. i % 5 == 0 to print every fifth line.
#include <stdio.h>

int main(void)
{
    int seconds = 3725;

    printf("%d seconds is %d h %d m %d s\n",
           seconds, seconds / 3600, (seconds / 60) % 60, seconds % 60);
    printf("the last digit of 3725 is %d\n", 3725 % 10);
    printf("3725 without its last digit is %d\n", 3725 / 10);
    printf("is 3725 even? %d  (1 means yes)\n", 3725 % 2 == 0);
    return 0;
}
3725 seconds is 1 h 2 m 5 s
the last digit of 3725 is 5
3725 without its last digit is 372
is 3725 even? 0  (1 means yes)

Negative operands

Division truncates towards zero, so % takes the sign of the left-hand operand. This is fixed by the standard since C99 and it is worth checking once rather than guessing.

#include <stdio.h>

int main(void)
{
    printf(" 17 /  5 = %3d     17 %% 5 = %3d\n",  17 /  5,  17 %  5);
    printf("-17 /  5 = %3d    -17 %% 5 = %3d\n", -17 /  5, -17 %  5);
    printf(" 17 / -5 = %3d     17 %% -5 = %3d\n",  17 / -5,  17 % -5);
    printf("-17 / -5 = %3d    -17 %% -5 = %3d\n", -17 / -5, -17 % -5);
    return 0;
}
 17 /  5 =   3     17 % 5 =   2
-17 /  5 =  -3    -17 % 5 =  -2
 17 / -5 =  -3     17 % -5 =   2
-17 / -5 =   3    -17 % -5 =  -2
munotes.in70

Arithmetic Operators

So -17 % 5 is -2, not 3. If you need a non-negative remainder for wrapping round a circle, write ((a % n) + n) % n.

Division by zero

Integer division by zero is undefined behaviour. Not zero, not an error you can catch, not infinity. On most machines the program is killed.

There is no way to recover from it afterwards, so the only correct handling is to test first:

if (count != 0) {
    average = (double) total / count;
} else {
    printf("no marks were entered\n");
}

Floating-point division by zero is different: it is defined by the floating-point standard and gives infinity or a "not a number" value rather than killing the program. Do not rely on that. Test the divisor.

The practical: simple interest from input

MU's Practical 1(a) reads the principal, the rate and the number of years from the user and prints the simple interest. Here it is, with the input checked, which is the part that separates a program from a program that works.

#include <stdio.h>

int main(void)
{
    double principal, rate, years;

    printf("Enter principal, rate per year and number of years: ");
    if (scanf("%lf %lf %lf", &principal, &rate, &years) != 3) {
        printf("\nThose were not three numbers.\n");
        return 1;
    }

    double interest = principal * rate * years / 100.0;

    printf("\nPrincipal        : %10.2f\n", principal);
    printf("Rate per year    : %10.2f %%\n", rate);
    printf("Number of years  : %10.2f\n", years);
    printf("Simple interest  : %10.2f\n", interest);
    printf("Amount repayable : %10.2f\n", principal + interest);
    return 0;
}
15000 8.5 2
Enter principal, rate per year and number of years:
Principal        :   15000.00
Rate per year    :       8.50 %
Number of years  :       2.00
Simple interest  :    2550.00
Amount repayable :   17550.00

Four things in that program are the difference between a pass and a good mark.

  • %lf in scanf, not %f. scanf takes an address and does no promotion, so %f means "a float " and %lf means "a double ". Passing a double * to %f writes four bytes into an eight-byte object and is a real bug. In printf the opposite is true: %f is right for both, because a float argument is promoted to double.
  • scanf returns a count, the number of items it successfully read. Checking it against 3 is the whole of the input validation, and it is one line.
  • / 100.0, not / 100. Here it makes no difference because the operands are already double, but writing 100.0 is the habit that saves you when they are not.
  • %% in the format string prints one per cent sign. A single % starts a conversion.
munotes.in71

Arithmetic Operators

For the algorithm and flowchart that MU asks for in the same practical, see chapter 2.

Arithmetic on characters, which is legal and useful

A char is a small integer, so all five operators work on it. Two uses are standard:

#include <stdio.h>

int main(void)
{
    char digit = '7';
    char lower = 'q';

    printf("'7' as a number is %d\n", digit - '0');
    printf("'q' in upper case is %c\n", lower - 'a' + 'A');
    printf("the 5th letter of the alphabet is %c\n", 'A' + 4);
    return 0;
}
'7' as a number is 7
'q' in upper case is Q
the 5th letter of the alphabet is E

lower - 'a' + 'A' works on any machine whose letters are consecutive, which is every machine you will use, but the standard does not guarantee it. toupper from <ctype.h> is the portable form and is chapter 34.

What this does NOT mean

/ is not always integer division. It is integer division only when both operands are integer types. The operator is the same; the operands decide.

% is not "percent". It is the remainder after integer division. To compute a percentage you divide and multiply by 100, in floating point.

% does not work on double. Use fmod from <math.h>.

There is no exponentiation operator in C. 2 ^ 3 is not 8; ^ is bitwise exclusive-or and gives 1. Use pow(2, 3) from <math.h>, or multiply, and for a square prefer x * x to pow(x, 2).

a % b for negative a is not always positive. It takes the sign of a.

Integer overflow is not "wrapping round". For signed integers it is undefined behaviour, unlike unsigned arithmetic which is defined to wrap. INT_MAX + 1 is not reliably INT_MIN.

Quick revision

  • Five operators: + - * / %. Unary - as well.
  • / on two integers discards the fraction. Make one operand a double for a real quotient.
  • % needs integer operands; fmod is the floating-point version.
  • (a / b) * b + (a % b) is always a.
  • Division truncates towards zero, so % takes the sign of the left operand: -17 % 5 is -2.
  • Non-negative wrap: ((a % n) + n) % n.
  • Integer division by zero is undefined behaviour. Test the divisor.
  • No exponentiation operator. ^ is bitwise exclusive-or.
  • Signed overflow is undefined; unsigned arithmetic wraps by definition.
  • In scanf, a double needs %lf. In printf, %f serves both float and double.

Test yourself

1. What are 17 / 5, 17 % 5, 17.0 / 5 and 17 % 5.0?

munotes.in72

Arithmetic Operators

3, 2, 3.4, and the last does not compile: % requires integer operands.

2. Write an expression that gives the tens digit of an integer n.

(n / 10) % 10.

3. What is -7 % 3, and why?

-1. Division truncates towards zero, so -7 / 3 is -2, and the identity (-2) * 3 + (-1) gives -7.

4. A program must print an average to two decimal places from int total and int count. Write the statement.

if (count != 0) printf("%.2f\n", (double) total / count);

5. Convert 3725 seconds into hours, minutes and seconds using only / and %.

Hours 3725 / 3600, minutes (3725 / 60) % 60, seconds 3725 % 60, giving 1 h 2 m 5 s.

6. Why is %lf needed in scanf but not in printf?

scanf receives a pointer and performs no conversion, so the length modifier tells it whether the target is a float or a double. In printf a float argument is automatically promoted to double, so %f handles both.

7. What does 2 ^ 3 give in C?

  1. ^ is bitwise exclusive-or, not exponentiation. Two is 10 and three is 11 in binary, and their exclusive-or is 01.

What can be asked on this, and how to answer it

"Explain the arithmetic operators in C with examples." Give the five with a worked value each, then spend the rest of the answer on the two that need it: integer division and its fix, and the remainder operator with its integer-only restriction and the identity. That is the shape the marks follow.

"What is the difference between / and %?" / gives the quotient and % the remainder of an integer division. / works on any arithmetic type and is integer division only when both operands are integers; % requires integer operands. (a / b) * b + (a % b) is a.

"Write a program to calculate simple interest taking principal, rate and years as input." Give this chapter's program. The examiner's follow-up is nearly always why %lf, or what happens if the user types a letter, so keep the scanf check in.

"What is the output of printf("%d", 5/22)?" 4. / and have equal precedence and group left to right, so it is (5 / 2) 2, which is 2 2. Chapter 20 is precedence in full.

"Write a program to convert a given number of seconds into hours, minutes and seconds." Give the three expressions above and say in one line why % 60 is needed on the minutes: dividing by 60 gives total minutes, and the hours must be taken out of it.

Contents This chapter on its own page

munotes.in73

Chapter Sixteen

Relational and Logical Operators

Syllabus topic 2, "Type of operators: Arithmetic operators, relational and logical operators, Increment and Decrement operators, assignment operators, the conditional operator, Assignment operators and expression, Precedence and order of Evaluation Block Structure, Initialization, C Preprocessor"

In one line

Relational operators compare two values and produce 1 or 0, and logical operators combine such results, with && and || stopping as soon as the answer is certain.

What a condition actually is in C

C has no separate truth value in the way later languages do. A condition is just an integer expression, and the rule is one line:

Zero is false. Every other value is true.

That is all. if (x) runs its body whenever x is not zero, and if (x - 5) runs whenever x is not 5. The relational operators produce 1 for true and 0 for false, so they fit the rule, but nothing requires a condition to come from one.

C99 added _Bool and the header <stdbool.h> with bool, true and false. They are worth using for clarity, and they change nothing underneath: true is 1 and false is 0.

The six relational operators

OperatorMeansa = 5, b = 3
<less thana < b is 0
>greater thana > b is 1
<=less than or equala <= b is 0
>=greater than or equala >= b is 1
==equal toa == b is 0
!=not equal toa != b is 1
#include <stdio.h>

int main(void)
{
    int a = 5, b = 3;

    printf("a < b  is %d\n", a < b);
    printf("a > b  is %d\n", a > b);
    printf("a == b is %d\n", a == b);
    printf("a != b is %d\n", a != b);
    printf("and the result really is an int: sizeof(a > b) is %zu\n",
           sizeof(a > b));
    return 0;
}
a < b  is 0
a > b  is 1
a == b is 0
a != b is 1
and the result really is an int: sizeof(a > b) is 4

= is assignment and == is comparison. if (x = 5) assigns 5 to x and then tests 5, which is not zero, so the body always runs. It compiles, because an assignment is an expression with a value. This is the single most expensive typing error in C, and -Wall warns about it when it looks suspicious.

The habit that prevents it: when comparing a variable with a constant, some programmers write the constant first, if (5 == x), because if (5 = x) will not compile. Use it if it helps you; the more reliable protection is to compile with warnings on and read them.

The three logical operators

&&    logical AND    true when both operands are true
||    logical OR     true when at least one operand is true
!     logical NOT    true when the operand is false
munotes.in74

Relational and Logical Operators

&& and || are binary and produce 1 or 0. ! is unary and also produces 1 or 0, so !0 is 1 and !5 is 0, not -5.

#include <stdio.h>

int main(void)
{
    int age = 20;
    int marks = 65;

    printf("age >= 18 && marks >= 60 is %d\n", age >= 18 && marks >= 60);
    printf("age <  18 || marks >= 60 is %d\n", age <  18 || marks >= 60);
    printf("!(age >= 18)             is %d\n", !(age >= 18));
    printf("!0 is %d and !5 is %d\n", !0, !5);
    return 0;
}
age >= 18 && marks >= 60 is 1
age <  18 || marks >= 60 is 1
!(age >= 18)             is 0
!0 is 1 and !5 is 0

Short circuiting, which is a guarantee and not an optimisation

&& evaluates its left operand first. If that is false, the answer must be false, so the right operand is not evaluated at all. || does the same when the left operand is true. The standard guarantees this, which means you may rely on it, and you will need to.

#include <stdio.h>

int calls = 0;

int noisy(int value)
{
    calls = calls + 1;
    return value;
}

int main(void)
{
    calls = 0;
    printf("0 && noisy(1) is %d, and noisy was called %d time(s)\n",
           0 && noisy(1), calls);

    calls = 0;
    printf("1 || noisy(1) is %d, and noisy was called %d time(s)\n",
           1 || noisy(1), calls);

    calls = 0;
    printf("1 && noisy(1) is %d, and noisy was called %d time(s)\n",
           1 && noisy(1), calls);
    return 0;
}
0 && noisy(1) is 0, and noisy was called 0 time(s)
1 || noisy(1) is 1, and noisy was called 0 time(s)
1 && noisy(1) is 1, and noisy was called 1 time(s)

The counter proves it. And this is what makes the following safe, which is the reason the guarantee matters:

if (count != 0 && total / count > 50)

If count is zero the division never happens. Write the two tests in the other order and the program divides by zero. Chapter 15 said the only correct handling of division by zero is to test first; short circuiting is what lets you do it in one expression.

The leap year, in full

MU's Practical 1(c). The Gregorian rule:

  1. A year divisible by 4 is a leap year,
  2. except that a year divisible by 100 is not,
  3. except that a year divisible by 400 is.

So 2024 is a leap year, 1900 is not, and 2000 is.

Written as one expression:

year % 4 == 0 && (year % 100 != 0 || year % 400 == 0)

The brackets are not optional. && binds tighter than ||, so without them the expression would group as (year % 4 == 0 && year % 100 != 0) || year % 400 == 0, which is a different test. It happens to give the right answer for every year, which is exactly why the mistake survives: it is wrong reasoning that produces right answers, and a viva question about 1900 will expose it.

munotes.in75

Relational and Logical Operators

#include <stdio.h>

int is_leap(int year)
{
    return year % 4 == 0 && (year % 100 != 0 || year % 400 == 0);
}

int main(void)
{
    int years[] = {1900, 1996, 2000, 2023, 2024, 2100, 2400};

    for (int i = 0; i < 7; i++) {
        printf("%d is %sa leap year\n",
               years[i], is_leap(years[i]) ? "" : "not ");
    }
    return 0;
}
1900 is not a leap year
1996 is a leap year
2000 is a leap year
2023 is not a leap year
2024 is a leap year
2100 is not a leap year
2400 is a leap year

And reading the year from the user, which is what the practical asks:

#include <stdio.h>

int main(void)
{
    int year;

    printf("Enter a year: ");
    if (scanf("%d", &year) != 1) {
        printf("\nThat was not a year.\n");
        return 1;
    }
    if (year < 1) {
        printf("\nThere is no year %d in this calendar.\n", year);
        return 1;
    }

    if (year % 4 == 0 && (year % 100 != 0 || year % 400 == 0)) {
        printf("\n%d is a leap year.\n", year);
    } else {
        printf("\n%d is not a leap year.\n", year);
    }
    return 0;
}
1900
Enter a year:
1900 is not a leap year.

For the algorithm and flowchart the same practical asks for, see chapter 2, which draws this test as three decision diamonds.

Two traps worth a mark each

1. a < b < c does not mean what it says in mathematics. It groups as (a < b) < c, and (a < b) is 0 or 1, so the whole thing compares 0 or 1 against c. Write a < b && b < c.

The listing below is wrong on purpose, and gcc says so before it is even run:

#include <stdio.h>

int main(void)
{
    int a = 10, b = 5, c = 1;

    printf("mathematically 10 < 5 < 1 is false\n");
    printf("in C, a < b < c gives %d, because (a < b) is %d and %d < c is %d\n",
           a < b < c, a < b, a < b, (a < b) < c);
    printf("written correctly, a < b && b < c gives %d\n", a < b && b < c);
    return 0;
}
munotes.in76

Relational and Logical Operators

chained.c: In function ‘main’:
chained.c:9:14: warning: comparisons like ‘X<=Y<=Z’ do not have their mathematical meaning [-Wparentheses]
    9 |            a < b < c, a < b, a < b, (a < b) < c);
      |            ~~^~~

It runs anyway, because a warning is not an error, and this is what it prints:

mathematically 10 < 5 < 1 is false
in C, a < b < c gives 1, because (a < b) is 0 and 0 < c is 1
written correctly, a < b && b < c gives 0

2. Never compare floating-point values with ==. Chapter 9 said why: the stored value is the nearest representable one, and arithmetic can land on a different neighbour.

#include <stdio.h>
#include <math.h>

int main(void)
{
    double x = 0.1 + 0.2;

    printf("x == 0.3 is %d\n", x == 0.3);
    printf("fabs(x - 0.3) < 1e-9 is %d\n", fabs(x - 0.3) < 1e-9);
    printf("x is actually %.20f\n", x);
    return 0;
}
x == 0.3 is 0
fabs(x - 0.3) < 1e-9 is 1
x is actually 0.30000000000000004441

fabs(x - y) < tolerance is the form to write, and the tolerance is chosen for the problem.

What this does NOT mean

A relational expression does not produce a special truth type. It produces an int, 1 or 0, which is why it can be printed with %d and added up.

"True" is not 1. Any non-zero value is true as a condition. if (7) runs, and !7 is 0. Only the operators produce 1.

& and | are not && and ||. The single-character forms are bitwise operators: they combine the bits of their operands and they always evaluate both sides. 1 & 2 is 0 while 1 && 2 is 1.

! is not negation. !x is 1 or 0. Arithmetic negation is -x.

Short circuiting is not an optimisation the compiler may skip. It is required by the standard, and programs depend on it for safety.

if (x = 5) is not a syntax error. It is a legal assignment used as a condition, and it is nearly always a typing mistake for ==.

Quick revision

  • Six relational operators: < > <= >= == !=. Each yields int 1 or 0.
  • Zero is false; any non-zero value is true.
  • = assigns, == compares. if (x = 5) is legal and always true.
  • Three logical operators: &&, ||, !.
  • && and || short-circuit, guaranteed by the standard, left operand first.
  • count != 0 && total / count > 50 is safe because of it.
  • && binds tighter than ||, so an OR inside an AND needs brackets.
  • Leap year: y % 4 == 0 && (y % 100 != 0 || y % 400 == 0).
  • a < b < c is a bug. Write a < b && b < c.
  • Never compare floating-point values with ==; use fabs(x - y) < tolerance.
  • & and | are bitwise and evaluate both operands.
munotes.in77

Relational and Logical Operators

Test yourself

1. What is the value of 5 > 3 and what is its type?

1, of type int.

2. What does if (n = 0) do?

It assigns 0 to n and then tests 0, which is false, so the body never runs. It was almost certainly meant to be if (n == 0).

3. Write the leap-year condition and say why the brackets are needed.

year % 4 == 0 && (year % 100 != 0 || year % 400 == 0). Without them, && binding tighter than || would group the expression as (y % 4 == 0 && y % 100 != 0) || y % 400 == 0, which is a different test built on wrong reasoning.

4. Why is if (n != 0 && 100 / n > 5) safe while if (100 / n > 5 && n != 0) is not?

&& evaluates the left operand first and skips the right if the left is false. In the first form the division is reached only when n is non-zero. In the second the division happens first and divides by zero.

5. What is !(-5)?

  1. Minus five is non-zero, so it is true, and the logical NOT of true is 0.

6. Give the difference between && and &.

&& is logical AND: it produces 1 or 0 and does not evaluate its right operand if the left is false. & is bitwise AND: it combines the bits of its operands and always evaluates both.

7. Is 2100 a leap year?

No. It is divisible by 4 and by 100 and not by 400, so the century exception applies.

What can be asked on this, and how to answer it

"Explain the relational and logical operators in C." Give the six relational with a value each, then the three logical, and then the two things the examiner is looking for: that the result is an int of 1 or 0, and that && and || short-circuit. Give the safe-division example for the second.

"What is short-circuit evaluation? Give an example where it matters." Define it, say the standard guarantees it, and give count != 0 && total / count > 50. Say that reversing the order makes the program divide by zero, which is what turns it from a nicety into a rule.

munotes.in78

Relational and Logical Operators

"Distinguish between = and ==." = assigns and yields the value assigned; == compares and yields 1 or 0. if (x = 5) is always true, and if (x == 5) tests. This is a favourite one-mark question.

"Write a program to check whether a year is a leap year." Give this chapter's program with the full three-part rule, and state the three test cases that show you understand it: 2024 yes, 1900 no, 2000 yes.

"Distinguish between logical and bitwise operators." Logical operators treat their operands as true or false, produce 1 or 0, and short-circuit. Bitwise operators work on the individual bits, produce a bit pattern, and always evaluate both operands. 1 && 2 is 1; 1 & 2 is 0.

Contents This chapter on its own page

munotes.in79

Chapter Seventeen

Increment and Decrement Operators

Syllabus topic 2, "Type of operators: Arithmetic operators, relational and logical operators, Increment and Decrement operators, assignment operators, the conditional operator, Assignment operators and expression, Precedence and order of Evaluation Block Structure, Initialization, C Preprocessor"

In one line

++ adds one and -- subtracts one, and where you put the operator decides whether the expression's value is taken before the change or after it.

Why the operators exist

i = i + 1 names i twice. That is fine for a simple variable and it stops being fine the moment the thing being stepped is longer: counts[student][paper] = counts[student][paper] + 1 gives a reader two long expressions to compare character by character, and gives the writer two chances to mistype one.

counts[student][paper]++ names it once. That is the whole argument, and it is a good one. The operators came from B, and B took them from the fact that the PDP-11 had an instruction for exactly this, but they have outlived that reason.

Prefix and postfix

There are four forms and they are two operators each used two ways.

FormNameWhat it doesWhat the expression is worth
++ipre-incrementadds 1 to ithe new value
i++post-incrementadds 1 to ithe old value
--ipre-decrementsubtracts 1 from ithe new value
i--post-decrementsubtracts 1 from ithe old value

The way to remember it: read the operator in the order it is written. ++i is "increment, then use i". i++ is "use i, then increment".

The listing below is the one every textbook prints, and gcc warns about every line of it. It is shown for exactly that reason.

#include <stdio.h>

int main(void)
{
    int i;

    i = 5;
    printf("i = 5; printf with ++i gives %d, and i is now %d\n", ++i, i);

    i = 5;
    printf("i = 5; printf with i++ gives %d, and i is now %d\n", i++, i);

    i = 5;
    printf("i = 5; printf with --i gives %d, and i is now %d\n", --i, i);

    i = 5;
    printf("i = 5; printf with i-- gives %d, and i is now %d\n", i--, i);
    return 0;
}
textbook.c: In function ‘main’:
textbook.c:8:66: warning: operation on ‘i’ may be undefined [-Wsequence-point]
    8 |     printf("i = 5; printf with ++i gives %d, and i is now %d\n", ++i, i);
      |                                                                  ^~~
textbook.c:11:67: warning: operation on ‘i’ may be undefined [-Wsequence-point]
   11 |     printf("i = 5; printf with i++ gives %d, and i is now %d\n", i++, i);
      |                                                                  ~^~
textbook.c:14:66: warning: operation on ‘i’ may be undefined [-Wsequence-point]
   14 |     printf("i = 5; printf with --i gives %d, and i is now %d\n", --i, i);
      |                                                                  ^~~
textbook.c:17:67: warning: operation on ‘i’ may be undefined [-Wsequence-point]
   17 |     printf("i = 5; printf with i-- gives %d, and i is now %d\n", i--, i);
      |                                                                  ~^~

It ran anyway, and on this compiler it printed the numbers the textbooks give:

munotes.in80

Increment and Decrement Operators

i = 5; printf with ++i gives 6, and i is now 6
i = 5; printf with i++ gives 5, and i is now 6
i = 5; printf with --i gives 4, and i is now 4
i = 5; printf with i-- gives 5, and i is now 4

Each of those printf calls passes i twice, once through the operator and once plainly, and the order in which printf's arguments are evaluated is unspecified. The program above is therefore relying on something it should not, and it is shown here only because it is what every textbook shows. The next section is the honest version, and it is the one to learn from.

The honest version: one change per statement

#include <stdio.h>

int main(void)
{
    int i = 5;
    int taken;

    taken = ++i;
    printf("after taken = ++i : taken is %d and i is %d\n", taken, i);

    i = 5;
    taken = i++;
    printf("after taken = i++ : taken is %d and i is %d\n", taken, i);

    i = 5;
    taken = --i;
    printf("after taken = --i : taken is %d and i is %d\n", taken, i);

    i = 5;
    taken = i--;
    printf("after taken = i-- : taken is %d and i is %d\n", taken, i);
    return 0;
}
after taken = ++i : taken is 6 and i is 6
after taken = i++ : taken is 5 and i is 6
after taken = --i : taken is 4 and i is 4
after taken = i-- : taken is 5 and i is 4

In every case i ends up as 6 or 4. The only thing that differs is what was handed to taken, and that is the entire content of the topic.

When the difference does not matter

i++;        /* as a statement on its own */
++i;        /* identical effect          */

As a complete statement, nothing uses the expression's value, so the two forms do exactly the same thing. Both compile to the same instruction. Use whichever you find clearer; this book uses i++ in a for loop because that is what every C program does.

The claim that ++i is faster than i++ is folklore. It was never true for an int on any compiler you will use. It can matter in C++ for a large object with an overloaded operator, and C has no such thing.

When the difference matters: walking an array

This is the real use, and it is worth seeing once now and again in chapter 37.

#include <stdio.h>

int main(void)
{
    int a[5] = {10, 20, 30, 40, 50};
    int i = 0;
    int sum = 0;

    while (i < 5) {
        sum = sum + a[i++];         /* use a[i], THEN step i */
    }
    printf("sum with a[i++] is %d, and i ended at %d\n", sum, i);

    i = 0;
    printf("the first three, taken with a[i++]: ");
    printf("%d ", a[i++]);
    printf("%d ", a[i++]);
    printf("%d\n", a[i++]);
    return 0;
}
munotes.in81

Increment and Decrement Operators

sum with a[i++] is 150, and i ended at 5
the first three, taken with a[i++]: 10 20 30

Each printf above is a separate statement, so each i++ is safely sequenced against the others. That is the discipline: one modification of a variable per statement.

The undefined-behaviour question, and how to answer it

Question banks ask for the value of things like this:

i = 5;
j = i++ + ++i;
printf("%d %d\n", i, j);

There is no answer. The standard says that if two side effects on the same object are unsequenced, the behaviour is undefined, and there are two modifications of i here with nothing sequencing them. Undefined behaviour means the standard places no requirement on the result at all: the program may print anything, may print something different on the next compiler, and is not required to be consistent.

gcc says so itself. This listing is wrong on purpose:

#include <stdio.h>

int main(void)
{
    int i = 5;
    int j;

    j = i++ + ++i;
    printf("i is %d and j is %d, and neither is guaranteed\n", i, j);
    return 0;
}

This book does not print what that program produced, and will not. Its result is undefined, so any number printed here would be a claim about a program the standard makes no promise about, and a reader would reasonably take it for the answer. The compiler's warning is the whole of what can honestly be shown:

unsequenced.c: In function ‘main’:
unsequenced.c:8:10: warning: operation on ‘i’ may be undefined [-Wsequence-point]
    8 |     j = i++ + ++i;
      |         ~^~

The answer to write in an examination, which is correct and complete:

The expression modifies i twice without an intervening sequence point, so by clause 6.5 of the standard the behaviour is undefined. The standard imposes no requirement on the result, so no value can be given, and different compilers will produce different answers. Written correctly as two statements, j = i++; j = j + ++i; the result is defined.

Some examiners want the number their own textbook prints. If you are certain that is what is wanted, give the textbook value and add the sentence above. The sentence cannot be marked wrong and it is the difference between having memorised an answer and understanding one.

Which expressions are safe

ExpressionSafe?Why
i++;YesOne modification, full statement
a[i++] = 0;Yesi modified once
x = i++ + 1;Yesi modified once
printf("%d\n", i++);Yesi modified once
x = i++ + i++;Noi modified twice, unsequenced
a[i] = i++;Noi read and modified, unsequenced
printf("%d %d\n", i++, i);NoArgument order unspecified
i = i++;NoTwo modifications of i
x = i++ && i++;Yes&& sequences its operands
munotes.in82

Increment and Decrement Operators

The last row is the exception that makes the rule clear: &&, | |, the comma operator and the conditional operator ?: all impose an order, so they sequence their operands. Everything else does not.

The one rule to take away

Modify a variable at most once in a statement, and do not also read it elsewhere in that statement. Follow it and you will never meet this problem. Break it and no amount of care about prefix and postfix will save you.

What this does NOT mean

++ does not work on a constant. 5++ does not compile. The operand must be something that can be assigned to.

++i is not "faster". For an int the two forms compile identically.

Postfix does not mean "later". The increment happens as part of evaluating the expression, not at the end of the statement. What is postponed is nothing; what differs is which value the expression yields.

Undefined behaviour is not "whatever your compiler does". It is a licence for the compiler to do anything, including assuming the case cannot arise and optimising on that basis. A program with undefined behaviour cannot be reasoned about at all.

i = i + 1 is not worse than i++. It is longer and perfectly correct. Prefer ++ where the thing being stepped is long enough that repeating it invites a mistake.

Quick revision

  • ++ adds one, -- subtracts one.
  • Prefix ++i yields the new value; postfix i++ yields the old.
  • Read the operator in the order written: ++i is increment-then-use; i++ is use-then-increment.
  • As a whole statement the two forms are identical in effect.
  • Modifying a variable twice in one statement is undefined behaviour, not a puzzle.
  • &&, | |, ?: and the comma operator sequence their operands; nothing else does.
  • The operand must be assignable: 5++ does not compile.
  • -Wall reports an unsequenced modification, and it is right.

Test yourself

1. After int i = 10; int j = i++; what are i and j?

i is 11 and j is 10. Postfix yields the old value.

2. After int i = 10; int j = ++i; what are i and j?

Both 11. Prefix yields the new value.

3. Is there a difference between i++; and ++i; as complete statements?

No. The value of the expression is discarded, so both simply add one.

munotes.in83

Increment and Decrement Operators

4. What is the value of j after int i = 5; int j = i++ + ++i;?

The question has no answer. i is modified twice with nothing sequencing the two, so the behaviour is undefined and the standard permits any result.

5. Is a[i] = i++; safe?

No. i is both read, to index the array, and modified, and the two are unsequenced. Write a[i] = i; i++;.

6. Why is x = i++ && i++; safe when x = i++ + i++; is not?

&& is guaranteed to evaluate its left operand fully, with a sequence point, before the right. Addition imposes no such order.

7. Rewrite sum = sum + a[i++]; without the increment operator.

sum = sum + a[i]; i = i + 1;

What can be asked on this, and how to answer it

"Explain the increment and decrement operators with examples." Give the four forms in a table with what each yields, the read-in-order rule, and a two-line program showing taken = i++ against taken = ++i. Add that as a standalone statement the two are identical, because that is the part most answers leave out.

"Distinguish between prefix and postfix increment." Prefix increments first and yields the new value; postfix yields the old value and then increments. In both cases the variable ends one larger. Give int i = 5; j = ++i; against int i = 5; j = i++; with the values.

"What is the output of i = 5; printf("%d %d", i++, ++i);?" Say that the behaviour is undefined: i is modified twice in one expression with nothing sequencing the two, and the order of evaluation of function arguments is unspecified besides. No value can be given, and different compilers differ. Then give the correct rewrite.

"Can ++ be applied to an expression like (a + b)?" No. The operand must be a modifiable object, and a + b is a value, not a place. (a + b)++ does not compile.

Contents This chapter on its own page

munotes.in84

Chapter Eighteen

Assignment Operators and Expressions

Syllabus topic 2, "Type of operators: Arithmetic operators, relational and logical operators, Increment and Decrement operators, assignment operators, the conditional operator, Assignment operators and expression, Precedence and order of Evaluation Block Structure, Initialization, C Preprocessor"

In one line

= stores a value in a variable and is itself an expression whose value is what was stored, and the ten compound forms such as += do an operation and the store in one step.

Plain assignment

variable = expression;

The right-hand side is evaluated, converted to the type of the left-hand side, and stored. The left-hand side must be something that can be assigned to: a variable, an array element, a structure member, or the thing a pointer points at. The language's word for that is an lvalue, and the name comes from its position on the left of an assignment.

x = 5;          /* fine   */
a[i] = 5;       /* fine   */
5 = x;          /* does not compile: 5 is not a place */
x + 1 = 5;      /* does not compile: a value, not a place */

An assignment is an expression

This is the part of the topic that matters, and it is why MU lists it separately.

The value of x = 5 is 5. An assignment is not a statement that happens to look like one; it is an operator that stores its right operand and then yields that value. Three consequences follow, and all three appear in real C.

1. Chained assignment. a = b = c = 0; works because = groups right to left: c = 0 yields 0, b = that yields 0, a = that yields 0.

#include <stdio.h>

int main(void)
{
    int a, b, c;

    a = b = c = 7;
    printf("a = %d, b = %d, c = %d\n", a, b, c);
    printf("the value of the expression (a = 99) is %d\n", (a = 99));
    printf("and a is now %d\n", a);
    return 0;
}
a = 7, b = 7, c = 7
the value of the expression (a = 99) is 99
and a is now 99

Chaining assigns the same value to each, after conversion at each step. int i; double d; i = d = 3.7; leaves d as 3.7 and i as 3, not both 3.7.

2. Assignment inside a condition. The idiom every C program uses for reading input:

while ((c = getchar()) != EOF) {
    ...
}

Read that from the inside out: getchar() is called, the result is stored in c, the value of the assignment is that same result, and it is compared against EOF. One line does read, store and test.

The inner brackets are not optional. != binds tighter than =, so while (c = getchar() != EOF) would compare first and store the 1 or 0 into c.

3. The if (x = 5) bug. Chapter 16 met it. This is the reason it compiles: the assignment is a legal expression with the value 5, and 5 is not zero, so the condition is true.

munotes.in85

Assignment Operators and Expressions

Where an assignment in a condition is deliberate, put extra brackets round it. That is the convention that tells a reader, and the compiler, that you meant it, and it is what turns off the warning.

The compound assignment operators

Operatorx op= y meansExample with x = 10
+=x = x + (y)x += 3 leaves 13
-=x = x - (y)x -= 3 leaves 7
*=x = x * (y)x *= 3 leaves 30
/=x = x / (y)x /= 3 leaves 3
%=x = x % (y)x %= 3 leaves 1

Five more exist for the bitwise operators, &=, |=, ^=, <<= and >>=. They work the same way. MU names no bitwise operator on this paper, so they are not taught here; know that they exist.

#include <stdio.h>

int main(void)
{
    int x = 10;

    printf("x starts at %d\n", x);
    x += 3;  printf("after x += 3 : %d\n", x);
    x -= 5;  printf("after x -= 5 : %d\n", x);
    x *= 4;  printf("after x *= 4 : %d\n", x);
    x /= 3;  printf("after x /= 3 : %d\n", x);
    x %= 7;  printf("after x %%= 7 : %d\n", x);
    return 0;
}
x starts at 10
after x += 3 : 13
after x -= 5 : 8
after x *= 4 : 32
after x /= 3 : 10
after x %= 7 : 3

Three reasons to prefer the compound form, and the third is the one that is not obvious.

  1. The variable is named once, so a long left-hand side cannot be mistyped on one side only.
  2. It is shorter to read, once you are used to it.
  3. The brackets are implied round the right-hand side. x = a + b is x = x (a + b), not x = x * a + b. This is a guarantee of the language, and it is the opposite of what happens with a careless #define (chapter 22).
#include <stdio.h>

int main(void)
{
    int x = 10, a = 2, b = 3;
    int y = 10;

    x *= a + b;                 /* x = x * (a + b) */
    y = y * a + b;              /* the careless reading */
    printf("x *= a + b   gives %d\n", x);
    printf("y = y * a + b gives %d\n", y);
    return 0;
}
x *= a + b   gives 50
y = y * a + b gives 23
munotes.in86

Assignment Operators and Expressions

The compound operators are not simply text substitution in one other way too: in a[f()] += 1 the expression f() is evaluated once. Written out as a[f()] = a[f()] + 1 it would be evaluated twice.

Assignment converts

The right-hand side is converted to the type of the left-hand side, and chapter 14's third rule applies: something may be lost.

#include <stdio.h>

int main(void)
{
    int i;
    double d;
    char c;

    i = 3.99;      printf("int i = 3.99      gives %d\n", i);
    d = 7 / 2;     printf("double d = 7 / 2  gives %.4f\n", d);
    c = 'A' + 2;   printf("char c = 'A' + 2  gives %c\n", c);
    i = -3.99;     printf("int i = -3.99     gives %d\n", i);
    return 0;
}
int i = 3.99      gives 3
double d = 7 / 2  gives 3.0000
char c = 'A' + 2  gives C
int i = -3.99     gives -3

double d = 7 / 2; is 3.0. The division is integer division and it happened before the assignment was reached. This is the same trap as chapter 14's and it is worth meeting twice.

The comma operator, briefly

, is also an operator: it evaluates its left operand, discards the result, evaluates its right operand, and yields that. It does impose an order, so it is one of the four operators that sequence their operands.

for (i = 0, j = n - 1; i < j; i++, j--)

That is nearly the only place you will see it, and it is the right place: two counters stepped in one for header. Chapter 28.

The commas separating function arguments and separating declarators are not the comma operator. They are punctuation, and they impose no order, which is why printf("%d %d", i++, i) is unsafe.

What this does NOT mean

= is not equality. == is. = stores.

x += 1 is not x++. They have the same effect on x, and x++ as an expression yields the old value while x += 1 yields the new one. As statements they are the same.

Assignment does not return a variable. It yields a value. (a = b) = c does not compile in C, because the result of an assignment is not an lvalue.

Chained assignment does not mean all the variables get the same stored value. Each step converts to its own type. i = d = 3.7 leaves different values in i and d.

x = a + b is not x = x a + b. The right-hand side is bracketed by the language.

An lvalue is not "a variable". An array element, a structure member and *p are all lvalues. A const object is an lvalue that cannot be assigned to.

munotes.in87

Assignment Operators and Expressions

Quick revision

  • = stores the right-hand side, converted to the left-hand type, and yields the stored value.
  • The left-hand side must be an lvalue: a place, not a value.
  • = groups right to left, so a = b = c = 0 works.
  • An assignment is an expression: while ((c = getchar()) != EOF) uses that, and if (x = 5) is the bug that comes from it.
  • Ten compound operators; the five arithmetic ones are += -= *= /= %=.
  • x op= y is x = x op (y): the right-hand side is bracketed, and the left-hand side is evaluated once.
  • Assignment converts, and may lose data. double d = 7 / 2; is 3.0.
  • The comma operator evaluates left then right and yields the right; the commas between function arguments are not it.

Test yourself

1. What is the value of the expression x = 7?

  1. An assignment yields the value it stored.

2. What does a = b = 5; do, and in what order?

b = 5 is evaluated first, because = groups right to left; it yields 5, which is then assigned to a. Both end as 5.

3. Rewrite total = total + marks[i] * weight; with a compound operator, and say what the brackets do.

total += marks[i] weight;. The right-hand side is implicitly bracketed, so this is total = total + (marks[i] weight), which is what was wanted.

4. x = 10; x *= 2 + 3; What is x?

  1. The right-hand side is bracketed, so it is x = x * (2 + 3).

5. Why does while (c = getchar() != EOF) not work?

!= binds tighter than =, so getchar() != EOF is evaluated first and its 1 or 0 is stored in c. The character is lost. Write while ((c = getchar()) != EOF).

6. What is in d after double d = 5 / 2;?

2.0. Both operands are int, so integer division gives 2, which is then converted on assignment.

7. Is (a = b) = c; legal?

No. The result of an assignment in C is a value, not an lvalue, so it cannot appear on the left of another assignment.

What can be asked on this, and how to answer it

"Explain the assignment operators in C." Give plain = with the lvalue requirement and the conversion rule, then the compound operators in a table, then the three properties that earn the extra marks: an assignment is an expression with a value, = groups right to left, and x op= y brackets the right-hand side and evaluates the left-hand side once.

munotes.in88

Assignment Operators and Expressions

"What is the difference between x = x + 5 and x += 5?" They have the same effect. += names the variable once, which matters when the left-hand side is long, and it brackets the right-hand side, so x += a + b differs from x = x + a + b only in that the first cannot be misread. Add that x is evaluated once in the compound form.

"Why does if (a = 10) always execute its body?" Because a = 10 is an assignment expression whose value is 10, and any non-zero value is true. It was almost certainly meant to be if (a == 10).

"What is an lvalue?" An expression that designates an object, so it may appear on the left of an assignment: a variable, an array element, a structure member, or *p. 5 and a + b are not lvalues, and a const object is an lvalue that may not be assigned to.

"Explain the comma operator." It evaluates its left operand, discards the value, evaluates its right operand and yields that, with a guaranteed order between them. Its normal use is stepping two counters in a for header. Note that the commas separating function arguments are punctuation, not this operator.

Contents This chapter on its own page

munotes.in89

Chapter Nineteen

The Conditional Operator

Syllabus topic 2, "Type of operators: Arithmetic operators, relational and logical operators, Increment and Decrement operators, assignment operators, the conditional operator, Assignment operators and expression, Precedence and order of Evaluation Block Structure, Initialization, C Preprocessor"

In one line

condition ? a : b is an expression whose value is a when the condition is true and b when it is false, and it is the only operator in C that takes three operands.

Why an operator and not just if

if is a statement. It chooses which statement to run. An expression cannot contain a statement, so with if alone there is no way to write "this value or that value" in the middle of something else.

printf("you %s\n", marks >= 40 ? "passed" : "failed");

That cannot be done with if without splitting it into several statements or repeating the printf. The conditional operator is what fills the gap, which is why it is sometimes called the ternary operator: it is the only operator in C with three operands, so "the ternary operator" and "the conditional operator" mean the same thing.

The form

condition ? value-if-true : value-if-false

The condition is evaluated. If it is non-zero, the whole expression is the second operand; otherwise it is the third. Exactly one of the two is evaluated, which is the same guarantee && and || give and which the next section proves.

#include <stdio.h>

int main(void)
{
    int marks = 63;
    int a = 12, b = 47;

    printf("marks %d: %s\n", marks, marks >= 40 ? "pass" : "fail");
    printf("the larger of %d and %d is %d\n", a, b, a > b ? a : b);
    printf("the smaller is %d\n", a < b ? a : b);
    printf("%d is %s\n", a, a % 2 == 0 ? "even" : "odd");
    printf("the absolute value of -7 is %d\n", -7 < 0 ? -(-7) : -7);
    return 0;
}
marks 63: pass
the larger of 12 and 47 is 47
the smaller is 12
12 is even
the absolute value of -7 is 7

It sequences, and only one branch runs

#include <stdio.h>

int calls = 0;

int noisy(int value)
{
    calls = calls + 1;
    return value;
}

int main(void)
{
    int x;

    calls = 0;
    x = 1 ? 10 : noisy(20);
    printf("1 ? 10 : noisy(20) gave %d, and noisy ran %d time(s)\n", x, calls);

    calls = 0;
    x = 0 ? noisy(10) : 20;
    printf("0 ? noisy(10) : 20 gave %d, and noisy ran %d time(s)\n", x, calls);
    return 0;
}
1 ? 10 : noisy(20) gave 10, and noisy ran 0 time(s)
0 ? noisy(10) : 20 gave 20, and noisy ran 0 time(s)

The branch not taken is never evaluated. That is what makes this safe:

average = count != 0 ? (double) total / count : 0.0;

The practical: the greatest of three numbers

MU's Practical 1(b), in her own words: find the greatest of three numbers using the conditional operator. There are two correct shapes and both are worth knowing.

munotes.in90

The Conditional Operator

Nested, in one expression. This is the answer the question is asking for.

#include <stdio.h>

int main(void)
{
    int a, b, c;

    printf("Enter three numbers: ");
    if (scanf("%d %d %d", &a, &b, &c) != 3) {
        printf("\nThose were not three numbers.\n");
        return 1;
    }

    int largest = a > b ? (a > c ? a : c) : (b > c ? b : c);

    printf("\nThe largest of %d, %d and %d is %d\n", a, b, c, largest);
    return 0;
}
34 91 57
Enter three numbers:
The largest of 34, 91 and 57 is 91

Read the expression in two halves. a > b decides which of a and b could still win. If a won, the remaining question is a against c; if b won, it is b against c. Three comparisons at most, and the brackets make the structure visible.

The brackets round the inner conditionals are not required, because ?: groups right to left, but write them. Nested conditionals without brackets are the standard example of unreadable C.

Two-step, which is clearer. Worth giving as well, and saying why you did:

int largest = a > b ? a : b;
largest = largest > c ? largest : c;

In the viva, the question after "write it" is usually "which would you prefer and why". The honest answer is the two-step form for readability, and the nested form because the question asked for one expression. Say both.

The same thing with if, for comparison

int largest;
if (a > b) {
    largest = a > c ? a : c;
} else {
    largest = b > c ? b : c;
}

or entirely without the conditional operator:

int largest = a;
if (b > largest) largest = b;
if (c > largest) largest = c;

That last form is the one that scales: for ten numbers it becomes a loop, and the nested conditional does not. It is chapter 36's pattern, and it is the right answer to "find the largest in an array". It is not the answer to this practical, which names the operator.

Where it is genuinely the right tool

1. Choosing a word inside output. The example that started the chapter.

2. Choosing a value in an initialisation, where if would force you to declare without a value first.

const double rate = years > 5 ? 8.5 : 7.25;

3. Guarding a division, shown above.

4. Singular and plural.

#include <stdio.h>

int main(void)
{
    for (int n = 0; n <= 3; n++) {
        printf("%d file%s found\n", n, n == 1 ? "" : "s");
    }
    return 0;
}
munotes.in91

The Conditional Operator

0 files found
1 file found
2 files found
3 files found

Where it is the wrong tool: anywhere the two branches do something rather than produce a value. x > 0 ? printf("pos\n") : printf("neg\n"); works and is bad style, because it uses an expression's value for nothing and hides a decision inside one. Use if.

Nesting, and when to stop

A conditional operator can be nested in either branch, and ?: groups right to left, so:

grade = marks >= 70 ? 'O'
      : marks >= 60 ? 'A'
      : marks >= 55 ? 'B'
      : marks >= 50 ? 'C'
      : marks >= 40 ? 'D'
      :               'F';

Laid out one condition per line like that, it is readable and it is a genuine use: every branch produces a value of the same kind. Written on one line it is unreadable. Chapter 25 gives the same thing as an else if ladder, which is what most programs use.

What this does NOT mean

?: is not a shorter if. It is an expression, so it has a value and must produce one. if is a statement and produces nothing. Where you want two branches that act rather than evaluate, if is correct.

It does not evaluate both branches. Exactly one is evaluated, guaranteed.

The two branches are not free to be any types. They are converted to a common type, which is the type of the whole expression. cond ? 1 : 2.5 has type double and yields 1.0 when the condition is true, not 1.

It is not the only ternary operator, it is the only one C has. "Ternary" describes the number of operands. Since C has exactly one such operator, the two names have come to mean the same thing.

Nesting is not automatically bad. A ladder of conditions each producing a value, laid out one per line, is clear. What is bad is nesting inside the condition rather than inside a branch, and putting it all on one line.

Quick revision

  • condition ? a : b yields a if the condition is non-zero, otherwise b.
  • It is the only operator in C with three operands, so "ternary" and "conditional" name the same thing.
  • Exactly one branch is evaluated, which makes it safe for guarding a division.
  • It is an expression, so it can appear inside a printf argument or an initialiser, where if cannot.
  • It groups right to left, so nesting in the false branch needs no brackets. Write them anyway.
  • The two branches are converted to a common type, which becomes the type of the expression.
  • Greatest of three: a > b ? (a > c ? a : c) : (b > c ? b : c).
  • Use if when the branches act rather than produce a value.
munotes.in92

The Conditional Operator

Test yourself

1. What is the value of 5 > 3 ? 10 : 20?

10.

2. Write one statement that prints "even" or "odd" for an int n.

printf("%s\n", n % 2 == 0 ? "even" : "odd");

3. Write the greatest of three numbers using only the conditional operator.

int largest = a > b ? (a > c ? a : c) : (b > c ? b : c);

4. Is count != 0 ? total / count : 0 safe when count is zero?

Yes. Only the branch selected is evaluated, so the division never happens when count is zero.

5. What is the type and value of 1 ? 1 : 2.5?

Type double, value 1.0. The two branches are converted to a common type before the expression has a value.

6. Why can ?: appear inside a printf argument when if cannot?

Because ?: is an operator and produces a value, and a function argument must be an expression. if is a statement and has no value.

7. Rewrite max = a > b ? a : b; using if.

if (a > b) max = a; else max = b;

What can be asked on this, and how to answer it

"What is the conditional operator? Explain with an example." Give the form, say it is the only ternary operator in C, give one example that produces a value inside a larger expression, and add that exactly one branch is evaluated. That last point is what distinguishes a full answer.

"Write a program to find the greatest of three numbers using the conditional operator." Give this chapter's program with the nested expression, and read the expression out in the two halves. Keep the scanf return check; it is one line and it is the difference between a program and a program that works.

"Distinguish between the conditional operator and the if-else statement." ?: is an expression: it produces a value and can be used wherever a value is wanted, and each branch must be an expression. if-else is a statement: it selects which statement to execute, the branches may contain any number of statements, and it produces no value. Both evaluate only the branch selected.

"Can the conditional operator be nested? Give an example." Yes, in either branch, and it groups right to left. Give the grade ladder laid out one condition per line, and say that beyond three or four branches an else if ladder or a switch is clearer.

Contents This chapter on its own page

munotes.in93

Chapter Twenty

Precedence and Order of Evaluation

Syllabus topic 2, "Type of operators: Arithmetic operators, relational and logical operators, Increment and Decrement operators, assignment operators, the conditional operator, Assignment operators and expression, Precedence and order of Evaluation Block Structure, Initialization, C Preprocessor"

In one line

Precedence decides which operator takes which operands, associativity decides the grouping when two operators of equal precedence meet, and order of evaluation decides when each operand is computed, which for most operators the standard deliberately does not fix.

Precedence: how an expression is grouped

2 + 3 4 is 14, not 20, because has higher precedence than +, so the expression groups as 2 + (3 * 4). Nothing about the order in which 3 and 4 are fetched is involved; the grouping is a fact about the program's structure, decided when it is compiled.

#include <stdio.h>

int main(void)
{
    printf("2 + 3 * 4       = %d\n", 2 + 3 * 4);
    printf("(2 + 3) * 4     = %d\n", (2 + 3) * 4);
    printf("10 - 4 - 3      = %d   <- left to right\n", 10 - 4 - 3);
    printf("2 * 3 %% 4       = %d   <- equal precedence, left to right\n",
           2 * 3 % 4);
    printf("100 / 10 / 2    = %d\n", 100 / 10 / 2);
    return 0;
}
2 + 3 * 4       = 14
(2 + 3) * 4     = 20
10 - 4 - 3      = 3   <- left to right
2 * 3 % 4       = 2   <- equal precedence, left to right
100 / 10 / 2    = 5

Associativity: the tie-breaker

When two operators of the same precedence meet, associativity decides the grouping.

  • Left to right for nearly everything: 10 - 4 - 3 is (10 - 4) - 3, which is 3. Not 10 - (4 - 3), which would be 9.
  • Right to left for the unary operators, the conditional operator and assignment: a = b = c is a = (b = c), which is why chained assignment works.

The table

Highest precedence at the top. Within a row, all the operators have equal precedence.

Level  Operators                                           Associativity
   1   ()  []  .  ->  postfix ++  --                       left to right
   2   prefix ++  --   +  -  !  ~  (type)  *  &  sizeof    right to left
   3   *  /  %                                             left to right
   4   +  -                                                left to right
   5   <<  >>                                              left to right
   6   <  <=  >  >=                                        left to right
   7   ==  !=                                              left to right
   8   &                                                   left to right
   9   ^                                                   left to right
  10   |                                                   left to right
  11   &&                                                  left to right
  12   ||                                                  left to right
  13   ?:                                                  right to left
  14   =  +=  -=  *=  /=  %=  and the rest                  right to left
  15   ,                                                   left to right

Levels 5, 8, 9 and 10 are the shift and bitwise operators. MU names none of them on this paper; they are in the table because leaving them out would put the other levels in the wrong places.

munotes.in94

Precedence and Order of Evaluation

What to actually memorise, because nobody recalls a sixteen-row table under pressure:

  1. Unary before binary.
  2. * / % before + -.
  3. Arithmetic before relational.
  4. Relational before == and !=.
  5. && before ||.
  6. Logical before ?:, and ?: before assignment.
  7. Assignment and the unary operators go right to left; everything else goes left to right.
  8. Comma is last.

Those eight cover every expression on this paper.

The five groupings that are asked about

gcc warns on two of the lines below, and that is the chapter's point made for us: the cases this section calls surprising are the cases the compiler itself thinks you should bracket.

#include <stdio.h>

int main(void)
{
    int a = 10, b = 5, c = 2;

    printf("a + b * c        = %d, grouped a + (b * c)\n", a + b * c);
    printf("a > b == 1       = %d, grouped (a > b) == 1\n", a > b == 1);
    printf("a & b == 0       = %d, grouped a & (b == 0)  <- surprising\n",
           a & b == 0);
    printf("!a + b           = %d, grouped (!a) + b\n", !a + b);
    printf("a - b - c        = %d, grouped (a - b) - c\n", a - b - c);
    printf("c * a %% b        = %d, grouped (c * a) %% b\n", c * a % b);
    return 0;
}
surprising.c: In function ‘main’:
surprising.c:8:63: warning: suggest parentheses around comparison in operand of ‘==’ [-Wparentheses]
    8 |     printf("a > b == 1       = %d, grouped (a > b) == 1\n", a > b == 1);
      |                                                             ~~^~~
surprising.c:10:18: warning: suggest parentheses around comparison in operand of ‘&’ [-Wparentheses]
   10 |            a & b == 0);
      |                ~~^~~~

It compiles and runs, and prints this:

a + b * c        = 20, grouped a + (b * c)
a > b == 1       = 1, grouped (a > b) == 1
a & b == 0       = 0, grouped a & (b == 0)  <- surprising
!a + b           = 5, grouped (!a) + b
a - b - c        = 3, grouped (a - b) - c
c * a % b        = 0, grouped (c * a) % b

The third line is the famous one. & sits below == in the table, so a & b == 0 means a & (b == 0), which is almost never what anyone intends. It is the reason many style guides require brackets round any bitwise operation. The same applies to a & b | c.

munotes.in95

Precedence and Order of Evaluation

Order of evaluation: the part that is not fixed

Now the other half of MU's label, and it is a genuinely different question.

Precedence does not tell you when anything is computed. In f() + g() h(), precedence tells you the expression is f() + (g() h()). It does not tell you whether f, g or h is called first. The standard leaves that unspecified, which means the compiler may choose, may choose differently at a different optimisation level, and need not tell you.

#include <stdio.h>

int order[3];
int next = 0;

int record(int who, int value)
{
    order[next] = who;
    next = next + 1;
    return value;
}

int main(void)
{
    int total = record(1, 2) + record(2, 3) * record(3, 4);

    printf("the expression is %d, which precedence alone settles\n", total);
    printf("but the calls happened in the order %d, %d, %d\n",
           order[0], order[1], order[2]);
    printf("and that order is UNSPECIFIED: another compiler may differ\n");
    return 0;
}
the expression is 14, which precedence alone settles
but the calls happened in the order 1, 2, 3
and that order is UNSPECIFIED: another compiler may differ

The value, 14, is fixed by precedence: 2 + (3 * 4). The order the three functions ran in is the compiler's choice, and the program printing it does not make it a rule.

The same is true of function arguments. printf("%d %d\n", f(), g()) does not say which of f and g runs first.

The four operators that DO fix an order

&&    left operand fully evaluated first, then the right only if needed
||    left operand fully evaluated first, then the right only if needed
?:    condition fully evaluated first, then exactly one branch
,     left operand fully evaluated and discarded, then the right

Those four, and nothing else. Chapters 16, 19 and 18 each proved one of them with a counter.

Undefined behaviour, which is a third thing again

Unspecified order becomes undefined behaviour when two unsequenced operations touch the same object and at least one of them modifies it.

  • f() + g(): unspecified order. Legal, and the value is well defined if neither function touches the other's data.
  • i++ + ++i: undefined behaviour. i is modified twice with nothing sequencing the two.

The difference matters. Unspecified means "one of several possible behaviours, and the standard does not say which". Undefined means "no requirement at all". Chapter 17 has the rule and the compiler's own warning.

What to do about all this

Two habits, and they cost nothing.

1. Bracket for the reader, not for the compiler. a + b * c needs no brackets and reads fine. (a & b) == 0 needs them to be correct and a && (b || c) needs them to be correct. Where a reader would have to consult a table, add brackets.

munotes.in96

Precedence and Order of Evaluation

2. One side effect per statement. Then the unspecified order cannot reach you.

int x = f();
int y = g();
int total = x + y;      /* the order is now yours, and visible */

What this does NOT mean

Precedence is not order of evaluation. Precedence groups; order decides when. This is the whole point of MU's two-part label.

Left-to-right associativity does not mean left-to-right evaluation. f() - g() - h() groups as (f() - g()) - h(), and the three calls may still happen in any order.

Brackets do not force an order of evaluation. (f()) + (g()) is exactly as unspecified as f() + g(). Brackets change grouping, nothing else.

A precedence table does not settle i++ + ++i. That expression is undefined behaviour, and no table can give it a value.

&& is not "higher precedence than ||" in the sense of running first. It is higher precedence, so it groups tighter; separately, it also fixes an order. Two different guarantees that happen to hold for the same operator.

Quick revision

  • Precedence groups an expression; associativity breaks ties at equal precedence; order of evaluation decides when operands are computed.
  • * / % above + -; arithmetic above relational; relational above equality; && above ||; ?: above assignment; comma last.
  • Unary operators, ?: and assignment go right to left. Everything else goes left to right.
  • & and | sit BELOW ==, so a & b == 0 is a & (b == 0). Bracket bitwise operations.
  • Order of evaluation of most operands, and of all function arguments, is unspecified.
  • Only &&, ||, ?: and the comma operator fix an order.
  • Unspecified is not undefined: f() + g() is fine, i++ + ++i is not.
  • Brackets change grouping, never order.

Test yourself

1. What is 2 + 3 * 4 - 6 / 3?

  1. It groups as 2 + (3 * 4) - (6 / 3), which is 2 + 12 - 2.

2. What is 10 - 4 - 3, and why not 9?

  1. - is left-associative, so it groups as (10 - 4) - 3.

3. What does a & b == 0 group as, and is it what a programmer usually wants?

a & (b == 0), because == has higher precedence than &. It is almost never what was wanted; write (a & b) == 0.

4. In f() + g(), which function is called first?

Unspecified. The standard does not fix it, and different compilers or different optimisation levels may differ.

munotes.in97

Precedence and Order of Evaluation

5. Name the only four operators in C that impose an order on their operands.

&&, ||, ?: and the comma operator.

6. Does adding brackets to f() + g() fix the order of the calls?

No. Brackets affect grouping only. To fix the order, put the calls in separate statements.

7. What is 5 > 3 > 1?

  1. It groups as (5 > 3) > 1, which is 1 > 1, which is 0.

What can be asked on this, and how to answer it

"Explain operator precedence and associativity in C with examples." Define both, give the eight rules worth memorising rather than the whole table, and give three worked expressions with their grouping shown in brackets. Include one surprising case, for which a & b == 0 or 5 > 3 > 1 is ideal.

"What is the difference between precedence and order of evaluation?" Precedence and associativity decide how the expression is grouped, and are completely fixed by the language. Order of evaluation decides when each operand is computed, and for most operators the standard leaves it unspecified. Give f() + g() * h(): precedence fixes the value, the order of the three calls is the compiler's choice.

"Which operators guarantee an order of evaluation?" &&, ||, the conditional operator and the comma operator. Say what the guarantee buys: count != 0 && total / count > 5 cannot divide by zero.

"Give the precedence table for C operators." Reproduce the levels from arithmetic downwards, which is what is actually being asked, and say that the bitwise and shift levels sit between the equality operators and &&. Marking the associativity of each level is usually worth a mark on its own.

"Evaluate a = 2, b = 3; printf("%d", a++ * b-- + --a);" Say that the behaviour is undefined, because a is modified twice in one expression with nothing sequencing the two operations, and no precedence table can give it a value. Rewrite it as separate statements to show what a defined version looks like.

Contents This chapter on its own page

munotes.in98

Chapter Twenty-One

Block Structure and Initialization

Syllabus topic 2, "Type of operators: Arithmetic operators, relational and logical operators, Increment and Decrement operators, assignment operators, the conditional operator, Assignment operators and expression, Precedence and order of Evaluation Block Structure, Initialization, C Preprocessor"

In one line

A block is a pair of braces with declarations and statements inside it, a name declared in a block is visible from its declaration to the closing brace and nowhere else, and initialisation is giving an object a value at the moment the block brings it into existence.

What a block is

{
    declarations and statements
}

That is a block, and it is one statement as far as the language is concerned. That last point is what makes if, while and for work: each of them governs exactly one statement, and a block is how you give them more than one.

Blocks appear in four places:

  1. A function body. Always a block.
  2. The body of an if, else, while, for, do or switch.
  3. Anywhere a statement may go, on its own, to limit the life of a variable.
  4. Nested inside another block, to any depth.

Scope: where a name can be seen

A name declared in a block is visible from the point of its declaration to the closing brace of that block. Not before it, and not after.

#include <stdio.h>

int main(void)
{
    int outer = 1;
    printf("in main, outer is %d\n", outer);

    {
        int inner = 2;
        printf("in the inner block, outer is %d and inner is %d\n",
               outer, inner);
    }

    /* inner does not exist here */
    printf("back in main, outer is %d\n", outer);
    return 0;
}
in main, outer is 1
in the inner block, outer is 1 and inner is 2
back in main, outer is 1

The inner block can see outer, because main's block encloses it. main cannot see inner, because inner's scope ended at the brace.

Moving the printf that mentions inner below the closing brace does not compile. That is not a limitation; it is the whole value of the rule. A variable that cannot be seen cannot be accidentally used, and a reader who reaches the closing brace knows they can stop thinking about it.

Lifetime: how long the object exists

Scope is about the name. Lifetime is about the object, and for an automatic variable they coincide: the object is created when control enters the block and destroyed when control leaves it.

That means a fresh object each time. A variable declared inside a loop body is created and destroyed on every pass, so it cannot remember anything between passes.

#include <stdio.h>

int main(void)
{
    for (int i = 0; i < 3; i++) {
        int fresh = 0;          /* new object every pass */
        static int kept = 0;    /* one object, created once */

        fresh = fresh + 1;
        kept = kept + 1;
        printf("pass %d: fresh is %d, kept is %d\n", i, fresh, kept);
    }
    return 0;
}
munotes.in99

Block Structure and Initialization

pass 0: fresh is 1, kept is 1
pass 1: fresh is 1, kept is 2
pass 2: fresh is 1, kept is 3

static changes the lifetime and not the scope. kept still cannot be named outside the loop body; it simply survives out there, waiting.

Shadowing

An inner block may declare a name that already exists outside it. The inner declaration hides the outer one for the rest of that block. This is legal, it is occasionally useful, and it is usually a mistake.

#include <stdio.h>

int value = 100;               /* global */

int main(void)
{
    int value = 10;            /* hides the global */
    printf("in main, value is %d\n", value);

    {
        int value = 1;         /* hides main's */
        printf("in the inner block, value is %d\n", value);
    }

    printf("back in main, value is %d again\n", value);
    return 0;
}
in main, value is 10
in the inner block, value is 1
back in main, value is 10 again

Three objects called value exist at once and each printf reaches the nearest one. Nothing is wrong with that program and it is still a bad idea: a reader has to count braces to know which value a line means, and a typing mistake that was meant for one will silently hit another. Give the inner variable a different name.

The one honest use of shadowing is a loop counter: for (int i = ...) inside a function that also has an i somewhere far away. Even then, a different name is clearer.

for and its own scope

Since C99, a variable declared in a for header belongs to the loop.

for (int i = 0; i < n; i++) {
    ...
}
/* i does not exist here */

This is the form to use. It says the counter is the loop's business and nobody else's, and it means two loops in one function can both use i without interfering. Chapter 28.

Initialisation

The general rule: an object may be given a value at the point it is created, and for some kinds of object that is the only way to give it one.

A scalar: covered in chapter 11.

int marks = 75;
double rate = 8.5;
char grade = 'A';

An array: a braced list. Chapter 36 is arrays proper; this is the initialisation rule.

#include <stdio.h>

int main(void)
{
    int a[5] = {10, 20, 30, 40, 50};   /* all five given            */
    int b[5] = {10, 20};               /* the rest are set to 0     */
    int c[5] = {0};                    /* every element 0           */
    int d[] = {1, 2, 3};               /* size counted for you: 3   */
    int e[5] = {[4] = 99, [0] = 11};   /* C99: by position, any order */

    printf("a: "); for (int i = 0; i < 5; i++) printf("%d ", a[i]);
    printf("\nb: "); for (int i = 0; i < 5; i++) printf("%d ", b[i]);
    printf("\nc: "); for (int i = 0; i < 5; i++) printf("%d ", c[i]);
    printf("\nd has %zu elements: ", sizeof d / sizeof d[0]);
    for (int i = 0; i < 3; i++) printf("%d ", d[i]);
    printf("\ne: "); for (int i = 0; i < 5; i++) printf("%d ", e[i]);
    printf("\n");
    return 0;
}
munotes.in100

Block Structure and Initialization

a: 10 20 30 40 50
b: 10 20 0 0 0
c: 0 0 0 0 0
d has 3 elements: 1 2 3
e: 11 0 0 0 99

Four rules come out of that:

  1. A partly initialised array has the rest set to zero. So int c[5] = {0}; zeroes the whole array, and it is the standard idiom for doing so.
  2. An array with no initialiser at all is not zeroed if it is automatic. int a[5]; inside a function holds rubbish, exactly as a single int does.
  3. int d[] = {1, 2, 3}; lets the compiler count. sizeof d / sizeof d[0] then recovers the count, and that expression is worth memorising.
  4. Designated initialisers, [4] = 99, were added in C99 and let you give elements by position in any order, with everything else zero.

A string: chapter 12's rule, which is the array rule with a terminator.

char name[] = "Anita";      /* 6 bytes, counted for you */
char buf[20] = "Anita";     /* 20 bytes, 14 of them zero */

You may initialise a character array from a string literal and you may not assign one. char s[6]; s = "Anita"; does not compile. Initialisation and assignment are different events, and only the first can fill an array.

const and initialisation

A const object can only ever be given its value at initialisation, because assignment to it is not allowed. That is the point of it.

const double PI = 3.14159;     /* the only chance to set it */

What this does NOT mean

A block is not a function. Both use braces. A function has a name, parameters and a return type, and is called; a block is a piece of a function.

Scope is not lifetime. Scope is where a name is visible; lifetime is how long the object exists. A static variable in a block has block scope and program lifetime.

Entering a block does not clear it. An automatic variable with no initialiser holds whatever was in that memory, every time the block is entered.

munotes.in101

Block Structure and Initialization

Shadowing is not an error. It compiles, and that is the problem.

A partly initialised array is not partly undefined. The elements you did not give are zero. This is the one case where C does zero things for you, and it applies only when there is an initialiser present.

int a[5] = {1}; does not set every element to 1. It sets the first to 1 and the rest to 0. int a[5] = {0}; happens to work as "all zero" only because the fill value is also zero.

Quick revision

  • A block is braces containing declarations and statements, and counts as one statement.
  • A name declared in a block is visible from its declaration to the closing brace.
  • An automatic object is created on entry to its block and destroyed on exit, and is not zeroed.
  • static in a block changes lifetime to the whole program; scope is unchanged.
  • An inner declaration of the same name shadows the outer one. Legal, and to be avoided.
  • Since C99, for (int i = ...) scopes the counter to the loop.
  • An array is initialised with a braced list; elements left out are zero.
  • int a[5] = {0}; zeroes an array; int a[5]; in a function does not.
  • int d[] = {1,2,3}; counts for you, and sizeof d / sizeof d[0] recovers the count.
  • C99 designated initialisers: int e[5] = {[4] = 99};.
  • An array can be initialised from a string literal but never assigned one.
  • A const object must be given its value at initialisation.

Test yourself

1. What is the scope of a variable declared inside an if body?

From its declaration to the closing brace of that body. It cannot be named after the if.

2. Does a variable declared inside a loop body keep its value between passes?

No. It is created and destroyed on each pass. Declare it before the loop, or make it static, if it must persist.

3. What is shadowing? Is it an error?

Declaring a name in an inner block that already exists in an enclosing scope, so that the inner one hides the outer for the rest of the block. It is legal and usually a mistake.

4. After int a[5] = {1, 2}; what is a[4]?

  1. Elements not given in an initialiser are set to zero.

5. After int a[5]; inside a function, what is a[4]?

Unspecified. With no initialiser an automatic array is not zeroed.

6. How do you find the number of elements of int d[] = {4, 8, 15, 16};?

sizeof d / sizeof d[0], which is 4. That works only where the array itself is in scope, not on a parameter, which chapter 41 explains.

munotes.in102

Block Structure and Initialization

7. Why must a const object be initialised rather than assigned?

Because assignment to a const object is not allowed, so initialisation is the only point at which it can be given a value.

What can be asked on this, and how to answer it

"What is a block? Explain block structure in C." Define it as braces enclosing declarations and statements, counting as a single statement. Say where blocks appear, give the scope rule, and give a nested example showing that the inner block sees the outer names and not the reverse. Mention that this is why if and while can govern several statements.

"Explain scope and lifetime with an example." Scope is the region of the program where a name is visible; lifetime is the period for which the object exists. For an automatic variable they coincide with the block. Give the static case as the one where they differ, with the three-pass loop program.

"What is variable shadowing?" An inner declaration hiding an outer name of the same spelling for the extent of the inner block. Give the three-level value program, and say that it is legal and should be avoided because a reader must count braces to know which object a line means.

"How is an array initialised? What happens to elements not given a value?" With a braced list at the point of declaration. Elements omitted are set to zero, provided an initialiser is present; with no initialiser at all an automatic array is not zeroed. Give int a[5] = {0}; as the idiom for zeroing, and the empty-bracket form that lets the compiler count.

Contents This chapter on its own page

munotes.in103

Chapter Twenty-Two

The C Preprocessor

Syllabus topic 2, "Type of operators: Arithmetic operators, relational and logical operators, Increment and Decrement operators, assignment operators, the conditional operator, Assignment operators and expression, Precedence and order of Evaluation Block Structure, Initialization, C Preprocessor"

In one line

The preprocessor is a program that edits your source text before the compiler sees it, driven by lines beginning with #, and it understands no C at all.

The directives

DirectiveWhat it does
#includeInserts another file here
#defineDefines a macro: a name, or a name with parameters
#undefRemoves a definition
#ifdef, #ifndefInclude the following text only if a name is or is not defined
#if, #elif, #else, #endifInclude the following text depending on a constant expression
#lineChanges the line number the compiler reports
#errorStops compilation with a message
#pragmaImplementation-defined instruction to the compiler

A directive occupies its own line, begins with #, and takes no semicolon. Adding one puts a stray semicolon into your program wherever the name is used.

#include

Two forms, and the difference is where the file is looked for.

#include <stdio.h>      /* the system's include directories */
#include "mystuff.h"    /* this file's own directory first, then the system's */

The rule in practice: angle brackets for the standard library and anything installed on the machine, double quotes for headers that are part of your own program.

The headers you will use this semester:

HeaderWhat it declares
<stdio.h>printf, scanf, getchar, putchar, fgets, FILE, EOF
<stdlib.h>abs, atoi, malloc, exit, rand
<string.h>strlen, strcpy, strcmp, strcat, strstr
<math.h>sqrt, pow, fabs, floor, ceil, round, sin
<ctype.h>isdigit, isalpha, toupper, tolower
<limits.h>INT_MAX, INT_MIN, CHAR_BIT
<float.h>DBL_DIG, FLT_MAX
<stdbool.h>bool, true, false

Object-like macros

#define NAME replacement-text

The preprocessor replaces every occurrence of NAME below that line with the replacement text. The name is conventionally in capitals so a reader can see it is a macro.

#include <stdio.h>

#define MAX_STUDENTS 60
#define PI 3.14159
#define GREETING "Welcome to FY B.Sc. IT"
#define PERCENT 100.0

int main(void)
{
    printf("%s\n", GREETING);
    printf("a class holds %d students\n", MAX_STUDENTS);
    printf("a circle of radius 3 has area %.5f\n", PI * 3 * 3);
    printf("45 out of 60 is %.2f per cent\n", 45 * PERCENT / MAX_STUDENTS);
    return 0;
}
Welcome to FY B.Sc. IT
a class holds 60 students
a circle of radius 3 has area 28.27431
45 out of 60 is 75.00 per cent

There is no type and no storage. MAX_STUDENTS is not an int; it is the two characters 6 and 0. Chapter 10's table compares #define with const and says why const is usually better.

Function-like macros, and the brackets

#define NAME(parameters) replacement-text

No space between the name and the opening bracket. #define SQUARE (x) ... defines an object-like macro whose text begins (x).

munotes.in104

The C Preprocessor

Now the part that matters. A macro is text substitution, so the text is dropped into the middle of whatever expression surrounds it, and precedence then applies to the result.

#include <stdio.h>

#define BAD_SQUARE(x)  x * x
#define GOOD_SQUARE(x) ((x) * (x))

#define BAD_HALF(x)    x / 2
#define GOOD_HALF(x)   ((x) / 2.0)

int main(void)
{
    printf("BAD_SQUARE(3)      = %d\n", BAD_SQUARE(3));
    printf("BAD_SQUARE(2 + 1)  = %d   <- wanted 9\n", BAD_SQUARE(2 + 1));
    printf("GOOD_SQUARE(2 + 1) = %d\n", GOOD_SQUARE(2 + 1));
    printf("100 / BAD_SQUARE(5)  = %d   <- wanted 4\n", 100 / BAD_SQUARE(5));
    printf("100 / GOOD_SQUARE(5) = %d\n", 100 / GOOD_SQUARE(5));
    printf("BAD_HALF(7)  = %d   <- integer division\n", BAD_HALF(7));
    printf("GOOD_HALF(7) = %.2f\n", GOOD_HALF(7));
    return 0;
}
BAD_SQUARE(3)      = 9
BAD_SQUARE(2 + 1)  = 5   <- wanted 9
GOOD_SQUARE(2 + 1) = 9
100 / BAD_SQUARE(5)  = 100   <- wanted 4
100 / GOOD_SQUARE(5) = 4
BAD_HALF(7)  = 3   <- integer division
GOOD_HALF(7) = 3.50

Work through the two failures by substitution, which is exactly what the preprocessor did.

  • BAD_SQUARE(2 + 1) becomes 2 + 1 2 + 1. Precedence groups that as 2 + (1 2) + 1, which is 5.
  • 100 / BAD_SQUARE(5) becomes 100 / 5 5. Left to right, that is (100 / 5) 5, which is 100.

Two sets of brackets, always. Round each parameter where it appears, and round the whole body. ((x) * (x)) survives both attacks.

The second macro trap: an argument evaluated twice

#include <stdio.h>

#define SQUARE(x) ((x) * (x))

int calls = 0;

int next(void)
{
    calls = calls + 1;
    return calls;
}

int main(void)
{
    calls = 0;
    int result = SQUARE(next());

    printf("SQUARE(next()) gave %d, and next() ran %d times\n", result, calls);
    printf("a FUNCTION would have called it once and given 1\n");
    return 0;
}
SQUARE(next()) gave 2, and next() ran 2 times
a FUNCTION would have called it once and given 1

SQUARE(next()) expands to ((next()) * (next())), so the function runs twice and multiplies two different numbers. The same happens with SQUARE(i++), which then also modifies i twice in one expression and is undefined behaviour. A real function cannot do this, and that is the strongest argument for writing one.

Macro against function

MacroFunction
Handled byPreprocessorCompiler
Type checkingNoneFull
Argument evaluatedOnce per appearance in the bodyExactly once
Code sizeOne copy per useOne copy in total
Call overheadNoneA call, which inline can remove
Can be debuggedPoorly: the source line is the callYes
Works for any typeYes, which is its one real advantageOne set of types
Can recurseNoYes
munotes.in105

The C Preprocessor

Prefer a function. Use a macro for a constant, for conditional compilation, and for the rare case where the same code must work for several types.

Conditional compilation

#include <stdio.h>

#define DEBUG 1

int main(void)
{
#if DEBUG
    printf("debug: the program has started\n");
#endif

    printf("doing the real work\n");

#ifdef DEBUG
    printf("debug: DEBUG is defined, whatever its value\n");
#endif

#ifndef RELEASE
    printf("RELEASE is not defined\n");
#endif

#if DEBUG > 5
    printf("this line is not compiled at all\n");
#else
    printf("DEBUG is 1, so this branch was compiled\n");
#endif
    return 0;
}
debug: the program has started
doing the real work
debug: DEBUG is defined, whatever its value
RELEASE is not defined
DEBUG is 1, so this branch was compiled

The text in a branch that is not taken is not compiled. It is removed before the compiler sees it, so it need not even be valid C, and an error in it will not be reported. That is what makes conditional compilation useful for code that only works on one operating system, and it is also why a mistake inside a switched-off branch can sit there for years.

#ifdef NAME tests whether the name is defined, not whether it is true. #define DEBUG 0 followed by #ifdef DEBUG takes the branch. Use #if DEBUG when you mean the value.

The standard predefined macros, which are always available and are genuinely useful:

#include <stdio.h>

int main(void)
{
    printf("this is line %d of file %s\n", __LINE__, __FILE__);
    printf("the C standard in use is %ld\n", __STDC_VERSION__);
    return 0;
}
this is line 5 of file lines.c
the C standard in use is 201710

__FILE__ prints whatever name the file was compiled under. That listing was saved as lines.c, so that is what it printed; on your machine it will be the name you chose. __LINE__ is 5 because the printf is the fifth line of the file, counting the #include as line 1.

The include guard

A header included twice would declare everything twice. The standard protection is a macro that records that the file has been seen:

#ifndef STUDENT_H
#define STUDENT_H

/* the contents of the header */

#endif

The first time, STUDENT_H is not defined, so the body is included and the macro is defined. The second time, the body is skipped. Every header in the standard library does this, and you should do it in any header of your own.

What this does NOT mean

The preprocessor does not understand C. It has no idea what a type, an expression or a scope is. It moves text.

A macro is not a constant. It has no type and no address. const gives you one that does.

A macro is not a function. No type checking, no single evaluation of arguments, and no recursion.

munotes.in106

The C Preprocessor

A directive does not take a semicolon. #define N 10; puts the semicolon into the program: int a[N]; becomes int a[10;];.

#ifdef does not test a value. It tests whether a name is defined at all.

Text in an untaken branch is not checked. It is removed, errors and all.

#include does not import a library. It inserts a text file of declarations. The library code is joined on by the linker, which is chapter 5.

Quick revision

  • The preprocessor edits text before compilation, driven by lines starting with #, and takes no semicolon.
  • #include <...> for system headers, #include "..." for your own.
  • #define NAME text substitutes text. No type, no storage.
  • #define NAME(p) text is function-like, and needs no space before the bracket.
  • Bracket every parameter and the whole body: #define SQUARE(x) ((x) * (x)).
  • A macro argument is evaluated once per appearance in the body, so SQUARE(next()) calls next twice.
  • Prefer a function; use a macro for constants, conditional compilation and type-independent code.
  • #ifdef tests definedness; #if tests a value.
  • Untaken conditional branches are removed before compilation and are not checked.
  • Include guard: #ifndef X / #define X / body / #endif.
  • __LINE__, __FILE__ and __STDC_VERSION__ are always defined.

Test yourself

1. What does #define SQ(x) x*x give for SQ(3+1), and why?

3+13+1, which is 7. The text is substituted and then precedence applies. Write ((x) (x)).

2. What is wrong with #define MAX 100;?

The semicolon is part of the replacement text, so int a[MAX]; becomes int a[100;]; and does not compile.

3. Give two differences between a macro and a function.

A macro is expanded as text by the preprocessor with no type checking; a function is compiled once and its arguments are type-checked. And a macro may evaluate an argument more than once, while a function evaluates each argument exactly once.

4. #define DEBUG 0 then #ifdef DEBUG. Is the branch taken?

Yes. #ifdef asks whether the name is defined, not what its value is. Use #if DEBUG to test the value.

5. Write an include guard for a header called student.h.

#ifndef STUDENT_H
#define STUDENT_H
/* contents */
#endif

6. How many times does next() run in SQUARE(next()) where SQUARE(x) is ((x) * (x))?

Twice, because x appears twice in the body and each appearance is replaced by the argument text.

7. Why is #define SQUARE (x) ((x)*(x)) wrong?

The space before the bracket makes it an object-like macro whose replacement text is (x) ((x)*(x)). Remove the space.

munotes.in107

The C Preprocessor

What can be asked on this, and how to answer it

"What is the C preprocessor? List and explain its directives." Define it as a text-processing stage that runs before the compiler, driven by lines beginning with #, understanding no C. Give the table of directives with one line each, and note that no directive takes a semicolon.

"Explain #define with examples of both kinds of macro." Give an object-like macro and a function-like one, then spend the answer on the bracket rule, with SQ(3+1) giving 7 as the worked failure. That example is what the question is really for.

"Distinguish between a macro and a function." Give four rows of the table: preprocessor against compiler, no type checking against full, an argument possibly evaluated several times against exactly once, and code repeated at each use against one copy called. Conclude that a function is preferable unless the code must work for several types.

"What is conditional compilation? Why is it used?" Including or excluding source text at compile time with #if, #ifdef, #ifndef, #else, #elif and #endif. Used for code specific to one machine or operating system, for debugging output that can be switched off, and for include guards. Add that untaken text is removed before the compiler sees it.

"What is an include guard and why is it needed?" A #ifndef / #define / #endif wrapper round a header, so that including the header twice declares its contents once. Without it, a second inclusion would redeclare everything and the compilation would fail.

Contents This chapter on its own page

munotes.in108

Module II

Control Flow, Functions, Pointers and User-defined data types

munotes.in

Chapter Twenty-Three

Statements and Blocks

Syllabus topic 1, "Control Flow: Statements and Blocks, If-Else, Else-If, Switch, Loops- While and For Loops Do-while, Break and Continue, Goto and Labels"

In one line

A statement is one instruction, ended by a semicolon, and a block is braces round several of them so that they count as one.

The semicolon terminates, it does not separate

This is the sentence to hold on to, and it is the difference between C and several other languages.

a = 1;
b = 2;
c = 3;

Every one of those three needs its semicolon, including the last. The semicolon is part of the statement, in the way a full stop is part of a sentence, and not a thing placed between statements.

Two consequences, and the second is a real bug:

  1. The last statement in a block still needs one.
  2. A semicolon on its own is a complete statement, the empty statement, which does nothing. That is legal, and it is the trap.

The kinds of statement

KindExample
Expression statementx = a + b; or printf("hi\n"); or i++;
Empty statement;
Compound statement, a block{ x = 1; y = 2; }
Selectionif, if-else, switch
Iterationwhile, for, do-while
Jumpbreak, continue, return, goto
Labelledcase 1:, default:, again:
Declarationint x = 5;

An expression statement is any expression followed by a semicolon. The expression is evaluated and its value is discarded, which is why x + 1; compiles and does nothing, and why -Wall warns about it.

Where one statement is expected

This is the rule that makes blocks necessary. Each of these governs exactly one statement:

if (condition) statement
if (condition) statement else statement
while (condition) statement
for (a; b; c) statement
do statement while (condition);

So when you want two things done, you need a block, because a block is one statement:

if (marks >= 40) {
    printf("pass\n");
    passed++;
}

Without the braces, only the printf is governed by the if, and passed++ runs whatever the marks are. It is indented as though it belonged, and it does not.

#include <stdio.h>

int main(void)
{
    int marks = 20;
    int passed = 0;

    if (marks >= 40)
        printf("pass\n");
        passed++;              /* NOT part of the if, despite the indentation */

    printf("marks %d, and passed was counted %d time(s)\n", marks, passed);
    return 0;
}

gcc sees the lie in the indentation and says so:

indent.c: In function ‘main’:
indent.c:8:5: warning: this ‘if’ clause does not guard... [-Wmisleading-indentation]
    8 |     if (marks >= 40)
      |     ^~
indent.c:10:9: note: ...this statement, but the latter is misleadingly indented as if it were guarded by the ‘if’
   10 |         passed++;              /* NOT part of the if, despite the indentation */
      |         ^~~~~~

It compiles anyway, and prints:

marks 20, and passed was counted 1 time(s)
munotes.in109

Statements and Blocks

The marks are 20, nothing was printed, and passed was still incremented. The indentation is a lie and the compiler never sees it.

The habit that prevents it: always write the braces, even for one statement. It costs two characters and it removes a whole class of bug, including the one that appears later when you add a second line to the body.

The empty statement, and the semicolon that should not be there

#include <stdio.h>

int main(void)
{
    int count = 0;

    for (int i = 0; i < 5; i++);      /* the stray semicolon */
    {
        count = count + 1;
    }
    printf("count is %d, and the loop body ran %d time(s)\n", count, 0);

    count = 0;
    for (int i = 0; i < 5; i++) {
        count = count + 1;
    }
    printf("written correctly, count is %d\n", count);
    return 0;
}
stray.c: In function ‘main’:
stray.c:7:5: warning: this ‘for’ clause does not guard... [-Wmisleading-indentation]
    7 |     for (int i = 0; i < 5; i++);      /* the stray semicolon */
      |     ^~~
stray.c:8:5: note: ...this statement, but the latter is misleadingly indented as if it were guarded by the ‘for’
    8 |     {
      |     ^
count is 1, and the loop body ran 0 time(s)
written correctly, count is 5

The first for has an empty statement as its body, so it counts to five doing nothing, and the block below it then runs once as an ordinary block. count ends as 1 and not 5.

The same mistake after an if is worse, because the body then always runs:

if (marks >= 40);          /* an if with an empty body */
    printf("pass\n");      /* an ordinary statement: always runs */

When an empty body is what you actually want, write the braces so a reader can see you meant it:

while (getchar() != '\n') {
    /* discard the rest of the line */
}

A declaration is a statement, since C99

Chapter 11 said a declaration may go anywhere a statement may go. Two consequences worth naming here:

for (int i = 0; i < n; i++) { ... }    /* declared in the for header */

if (x > 0) {
    int y = x * 2;                     /* declared where first needed */
    printf("%d\n", y);
}

A declaration may not be the single governed statement of an if or a loop without braces, because there would be no block for it to live in and it would be useless. if (x) int y = 5; does not compile.

Indentation and layout

None of it is required and all of it matters.

  • One statement per line. Two on a line hides one of them.
  • Indent the body of every block by the same amount. Four spaces in this book.
  • The opening brace on the same line as the if or while, the closing brace on a line of its own, lined up under the keyword. This is the layout the first C book used and the one most C uses.
  • Whatever you choose, use it everywhere. Mixed layout is how a missing brace stays hidden.
munotes.in110

Statements and Blocks

A program can legally be written like this, and you should see it once so that you know the compiler does not care:

#include <stdio.h>
int main(void){int i;for(i=0;i<3;i++){printf("%d ",i);}printf("\n");return 0;}
0 1 2

That is the same program as a properly laid out one. It compiles to identical machine code. Layout is entirely for the reader, which is why chapter 6 put clarity second only to integrity.

What this does NOT mean

A semicolon does not separate statements. It terminates one. So the last statement in a block needs one, and a lone semicolon is a statement.

Indentation does not create a block. Only braces do. This is the difference from Python, and it is the source of the passed++ bug above.

A block does not need a semicolon after it. if (x) { ... } has no semicolon after the closing brace. A do-while is the one exception, and the semicolon there belongs to the while, not to the block.

An expression statement is not required to do anything. x + 1; is a legal statement that computes a value and discards it.

A block is not only for if and loops. A block on its own is a legal statement anywhere, and limiting a variable's life is a good reason to write one.

Quick revision

  • A statement is one instruction terminated by a semicolon.
  • The semicolon terminates, it does not separate: the last statement needs one too.
  • A lone ; is the empty statement and does nothing.
  • A block is braces round several statements and counts as one statement.
  • if, else, while, for and do each govern exactly one statement, which is why blocks are needed.
  • Always write the braces, even for a single statement.
  • A stray semicolon after for (...) or if (...) gives it an empty body, and the code below then runs unconditionally.
  • Since C99 a declaration is a statement and may go anywhere one may go.
  • Indentation is for the reader; the compiler ignores it entirely.

Test yourself

1. What does this print?

int x = 5;
if (x > 10)
    printf("big\n");
    printf("done\n");

done. Only the first printf is governed by the if; the second is an ordinary statement that always runs.

2. How many times does the body run?

munotes.in111

Statements and Blocks

for (int i = 0; i < 3; i++);
    printf("hello\n");

The loop body is the empty statement and runs three times doing nothing. printf is not the body and runs once.

3. Is ; a valid C statement?

Yes. It is the empty statement, and it does nothing.

4. Does a block need a semicolon after its closing brace?

No, except after the while clause of a do-while, where the semicolon belongs to the do-while statement.

5. Why is if (x) int y = 5; not legal?

A declaration cannot be the single governed statement of an if. There is no block for the declared name to live in, so the language does not allow it. Write if (x) { int y = 5; ... }.

6. Name the six kinds of statement other than declarations.

Expression, empty, compound (block), selection, iteration, jump. Labelled statements are a seventh.

What can be asked on this, and how to answer it

"What is a statement? What is a compound statement?" A statement is a single instruction terminated by a semicolon. A compound statement, or block, is a sequence of declarations and statements enclosed in braces and treated as one statement, which is how several statements are given to an if or a loop.

"What is the significance of the semicolon in C?" It terminates a statement rather than separating statements, so every statement including the last in a block needs one. A semicolon alone is the empty statement, which is why a misplaced one changes the meaning of a program instead of being an error.

"Find the error: if (a > b); printf("a is bigger");" The semicolon gives the if an empty body, so the printf is an ordinary statement and runs whatever the comparison gives. Remove the semicolon.

"Why should braces be used even for a single statement?" Because the if or loop governs exactly one statement, so a second line added later is silently outside the body, and because the indentation then agrees with what the compiler sees. It costs nothing and removes a whole class of bug.

Contents This chapter on its own page

munotes.in112

Chapter Twenty-Four

if and if-else

Syllabus topic 1, "Control Flow: Statements and Blocks, If-Else, Else-If, Switch, Loops- While and For Loops Do-while, Break and Continue, Goto and Labels"

In one line

if runs a statement when a condition is non-zero, and if-else chooses between two statements, exactly one of which runs.

The forms

if (condition)
    statement

if (condition)
    statement
else
    statement

The condition is any expression. It is evaluated, and if the result is non-zero the first statement runs; otherwise, in the second form, the else statement runs.

There is no then in C. The brackets round the condition are required, and they are not optional decoration: they are what tells the compiler where the condition ends.

With braces, which is how to write it

#include <stdio.h>

int main(void)
{
    int marks = 63;

    if (marks >= 40) {
        printf("Result : pass\n");
        printf("Marks  : %d\n", marks);
    } else {
        printf("Result : fail\n");
        printf("Shortfall by %d marks\n", 40 - marks);
    }
    return 0;
}
Result : pass
Marks  : 63

Chapter 23 gave the reason the braces are there: if governs one statement, and a block is how you give it two.

The condition is any expression

Chapter 16's rule applies in full: zero is false, anything else is true. So all of these are legal conditions, and the first two are worth recognising.

if (n)                  /* true when n is not zero          */
if (!n)                 /* true when n is zero              */
if (n != 0)             /* the same, said plainly           */
if (marks >= 40 && attendance >= 75)
if (strcmp(a, b) == 0)  /* two strings are equal            */

if (n) and if (!n) are idiomatic C and you will read them constantly. When the value is a count or a quantity rather than a truth, if (n != 0) says what you mean and costs four characters.

The dangling else

An else belongs to the nearest unmatched if. Always. Indentation has nothing to do with it, and this is the standard examination trick.

#include <stdio.h>

int main(void)
{
    int a = 5, b = 10;

    /* Indented as though the else belonged to the OUTER if. It does not. */
    if (a > 0)
        if (b > 100)
            printf("a positive, b over a hundred\n");
    else
        printf("this looks like the else of the first if\n");

    printf("a is %d and b is %d\n", a, b);
    return 0;
}

The compiler spots the mismatch between the layout and the meaning:

dangling.c: In function ‘main’:
dangling.c:8:8: warning: suggest explicit braces to avoid ambiguous ‘else’ [-Wdangling-else]
    8 |     if (a > 0)
      |        ^
this looks like the else of the first if
a is 5 and b is 10

Nothing from the first group printed. Read what the compiler read:

if (a > 0) {
    if (b > 100) {
        printf("a positive, b over a hundred\n");
    } else {
        printf("this looks like the else of the first if\n");
    }
}
munotes.in113

if and if-else

a > 0 is true, so the inner if is reached. b > 100 is false, so the else runs, printing the second line. The message that claims to be the outer else is in fact the inner one, and it printed.

To attach an else to the outer if, use braces. There is no other way.

#include <stdio.h>

int main(void)
{
    int a = 5, b = 10;

    if (a > 0) {
        if (b > 100) {
            printf("a positive, b over a hundred\n");
        }
    } else {
        printf("a is not positive\n");
    }
    printf("finished, and nothing above printed\n");
    return 0;
}
finished, and nothing above printed

Now a > 0 is true, the inner if fails, and there is no else for it, so nothing prints from the decision at all. That is what the first program was trying to say.

The one-line rule to write in an exam: an else matches the nearest preceding unmatched if in the same block, regardless of layout; braces are the only way to change that.

Nesting

An if may contain another if in either branch, to any depth. Beyond two levels it becomes hard to read, and the two usual cures are an else if ladder (chapter 25) and pulling the decision into a function that returns early.

#include <stdio.h>

const char *category(int age)
{
    if (age < 0) {
        return "not an age";
    }
    if (age < 13) {
        return "child";
    }
    if (age < 20) {
        return "teenager";
    }
    if (age < 60) {
        return "adult";
    }
    return "senior";
}

int main(void)
{
    int ages[] = {-1, 5, 15, 30, 70};

    for (int i = 0; i < 5; i++) {
        printf("%3d -> %s\n", ages[i], category(ages[i]));
    }
    return 0;
}
 -1 -> not an age
  5 -> child
 15 -> teenager
 30 -> adult
 70 -> senior

That shape, a series of ifs that each return, is called an early return, and it is usually the clearest way to write a decision with several outcomes. Each test reads as "if this, we are done", and nothing is nested.

The three mistakes

1. = for ==. if (x = 5) assigns and is always true. Chapters 16 and 18.

2. A semicolon after the condition. if (x > 0); gives the if an empty body. Chapter 23.

3. Comparing floating-point values with ==. Chapter 9. Use a tolerance.

#include <stdio.h>
#include <math.h>

int main(void)
{
    double x = 0.1 + 0.2;

    if (x == 0.3) {
        printf("exact comparison said equal\n");
    } else {
        printf("exact comparison said NOT equal\n");
    }
    if (fabs(x - 0.3) < 1e-9) {
        printf("comparison with a tolerance said equal\n");
    }
    return 0;
}
munotes.in114

if and if-else

exact comparison said NOT equal
comparison with a tolerance said equal

What this does NOT mean

if does not need else. The else is optional and is often better left out, especially when the if returns.

The braces are not optional in practice. They are optional in the grammar. Leave them out and you meet the passed++ bug of chapter 23 and the dangling else of this one.

Indentation does not bind an else. Only braces do.

The condition does not have to be a comparison. Any expression will do, and zero is false.

if (a > b > c) is not a range test. It groups as (a > b) > c. Chapter 20.

An if is not an expression. It has no value, so it cannot appear inside a printf argument. That is what the conditional operator of chapter 19 is for.

Quick revision

  • if (condition) statement and if (condition) statement else statement.
  • The brackets round the condition are required; there is no then.
  • Zero is false, anything else is true, so if (n) and if (!n) are idiomatic.
  • Each branch governs exactly one statement; use a block for more.
  • An else matches the nearest unmatched if, whatever the indentation.
  • Braces are the only way to attach an else to an outer if.
  • A semicolon after the condition gives an empty body.
  • if (x = 5) assigns and is always true.
  • Never == on floating-point values; compare the size of the difference.
  • A series of ifs that each return is usually clearer than nesting.

Test yourself

1. What does this print, and why?

int a = 5, b = 10;
if (a > 0)
    if (b > 100)
        printf("X\n");
else
    printf("Y\n");

Y. The else belongs to the inner if, not the outer one, whatever the indentation suggests. a > 0 is true, b > 100 is false, so the inner else runs.

2. How do you make that else belong to the outer if?

Put braces round the inner if: if (a > 0) { if (b > 100) printf("X\n"); } else printf("Y\n");

3. What is the difference between if (n) and if (n == 0)?

They are opposites. if (n) is true when n is not zero; if (n == 0) is true when it is.

4. Why is if (x > 0); almost always a bug?

The semicolon is the empty statement, so the if has an empty body and the statement that follows runs regardless of the condition.

5. Rewrite this without nesting:

if (marks >= 40) { if (attendance >= 75) printf("allowed\n"); }
munotes.in115

if and if-else

if (marks >= 40 && attendance >= 75) printf("allowed\n");

6. Can an if statement appear as a function argument?

No. if is a statement and has no value. Use the conditional operator: printf("%s\n", x > 0 ? "pos" : "neg");

What can be asked on this, and how to answer it

"Explain the if and if-else statements with syntax and examples." Give both forms, say the brackets are required and there is no then, give a worked program, and add the rule that each branch governs one statement so a block is needed for more. Mention that zero is false and anything else is true.

"What is the dangling else problem? How is it resolved?" An else is matched to the nearest preceding unmatched if, so in a nested if without braces the else binds to the inner one even when the indentation suggests otherwise. It is resolved with braces, which are the only way to change the match. Give the trace-the-output example.

"Find the output" with nested ifs and no braces. Rewrite the code with the braces the compiler infers, then trace it. Showing that rewrite is what earns the marks, because it demonstrates the rule rather than the answer.

"Distinguish between if-else and the conditional operator." if-else is a statement: it selects which statement runs and has no value, and each branch may hold any number of statements. ?: is an expression: it produces a value and each branch must be a single expression. Both evaluate only the branch selected.

Contents This chapter on its own page

munotes.in116

Chapter Twenty-Five

else-if Ladders

Syllabus topic 1, "Control Flow: Statements and Blocks, If-Else, Else-If, Switch, Loops- While and For Loops Do-while, Break and Continue, Goto and Labels"

In one line

An else-if ladder is a chain of tests where the first one that succeeds runs its block and the rest are skipped, so the order of the tests is part of the logic.

It is not a new construct

if (condition1) {
    ...
} else if (condition2) {
    ...
} else {
    ...
}

Read what that really is. An else governs one statement, and an if is one statement, so the second if is simply the else branch of the first. Written out with all the braces the compiler infers:

if (condition1) {
    ...
} else {
    if (condition2) {
        ...
    } else { ... }
}

A third else if adds a third level inside that innermost else. The ladder layout is the same nesting written flat, and it exists because the nested version marches across the page. Nothing about the language was added; only the layout changed.

Two consequences follow, and both are examinable.

  1. The tests are tried in order and the first true one wins. Everything below it is skipped, including tests that are also true.
  2. The final else is optional and catches everything the tests missed.

The order is the logic

#include <stdio.h>

char grade_wrong(int marks)
{
    if (marks >= 40) return 'D';        /* catches everything above 40 */
    else if (marks >= 50) return 'C';
    else if (marks >= 60) return 'A';
    else if (marks >= 70) return 'O';
    else return 'F';
}

char grade_right(int marks)
{
    if (marks >= 70) return 'O';
    else if (marks >= 60) return 'A';
    else if (marks >= 55) return 'B';
    else if (marks >= 50) return 'C';
    else if (marks >= 40) return 'D';
    else return 'F';
}

int main(void)
{
    int marks[] = {85, 63, 57, 52, 45, 30};

    printf("marks  wrong  right\n");
    for (int i = 0; i < 6; i++) {
        printf("%5d  %5c  %5c\n",
               marks[i], grade_wrong(marks[i]), grade_right(marks[i]));
    }
    return 0;
}
marks  wrong  right
   85      D      O
   63      D      A
   57      D      B
   52      D      C
   45      D      D
   30      F      F

Both ladders have the same five tests and one gives the wrong grade for every mark above 40. A ladder of overlapping tests must be ordered from the most demanding to the least. That is the whole lesson, and it is the commonest logic error in a first-semester program.

Because the earlier tests have already been ruled out, the later ones need no upper bound. marks >= 60 in grade_right means "at least 60 and, since we are here, below 70". Writing marks >= 60 && marks < 70 is correct and redundant, and the redundancy is a second place for a mistake to live.

munotes.in117

else-if Ladders

The practical: the roots of a quadratic equation

MU's Practical 2(a). For ax^2 + bx + c = 0 the discriminant is b^2 - 4ac, and its sign decides the kind of roots. That is three cases, so it is a ladder.

#include <stdio.h>
#include <math.h>

int main(void)
{
    double a, b, c;

    printf("Enter a, b and c for a*x*x + b*x + c = 0: ");
    if (scanf("%lf %lf %lf", &a, &b, &c) != 3) {
        printf("\nThose were not three numbers.\n");
        return 1;
    }
    if (fabs(a) < 1e-12) {
        printf("\nWith a = 0 this is not a quadratic equation.\n");
        return 1;
    }

    double d = b * b - 4 * a * c;

    printf("\nDiscriminant b*b - 4*a*c = %.4f\n", d);
    if (d > 1e-12) {
        double r1 = (-b + sqrt(d)) / (2 * a);
        double r2 = (-b - sqrt(d)) / (2 * a);
        printf("Two distinct real roots: %.4f and %.4f\n", r1, r2);
    } else if (d > -1e-12) {
        double r = -b / (2 * a);
        printf("One repeated real root: %.4f\n", r);
    } else {
        double real = -b / (2 * a);
        double imag = sqrt(-d) / (2 * a);
        printf("Two complex roots: %.4f + %.4fi and %.4f - %.4fi\n",
               real, imag, real, imag);
    }
    return 0;
}
1 -7 12
Enter a, b and c for a*x*x + b*x + c = 0:
Discriminant b*b - 4*a*c = 1.0000
Two distinct real roots: 4.0000 and 3.0000

Four things in that program are the difference between a pass and a good mark.

1. a is checked for zero first. With a zero the formula divides by zero, and the equation is not quadratic anyway. No textbook solution checks this and the viva question is always "what if a is zero".

2. The discriminant is never compared with ==. It is a double computed by arithmetic, so the exactly-zero case will usually miss by a fraction. Chapter 9's rule applies: the tests are d > tolerance for positive, then d > -tolerance for what is left, which means "within the tolerance of zero", and the final else is genuinely negative.

Read those two tests as a ladder and the trick becomes clear: by the time the second test runs, d > 1e-12 has already failed, so d > -1e-12 can only mean -1e-12 < d <= 1e-12. The ladder did the work that an && would otherwise have to.

3. sqrt(-d) and not sqrt(d) in the complex branch. d is negative there, and sqrt of a negative double gives a not-a-number value.

4. 1e-12 is a choice, not a constant of nature. It is small relative to the sizes in this problem. For very large coefficients it would be too small and for very small ones too large; saying so in a viva is worth more than the program.

munotes.in118

else-if Ladders

Run with different coefficients and the other two branches show:

#include <stdio.h>
#include <math.h>

int main(void)
{
    double sets[3][3] = {{1, -7, 12}, {1, -4, 4}, {1, 2, 5}};

    for (int i = 0; i < 3; i++) {
        double a = sets[i][0], b = sets[i][1], c = sets[i][2];
        double d = b * b - 4 * a * c;

        printf("a=%.0f b=%.0f c=%.0f  d=%.0f  ", a, b, c, d);
        if (d > 1e-12) {
            printf("real and distinct: %.4f, %.4f\n",
                   (-b + sqrt(d)) / (2 * a), (-b - sqrt(d)) / (2 * a));
        } else if (d > -1e-12) {
            printf("real and repeated: %.4f\n", -b / (2 * a));
        } else {
            printf("complex: %.4f +/- %.4fi\n",
                   -b / (2 * a), sqrt(-d) / (2 * a));
        }
    }
    return 0;
}
a=1 b=-7 c=12  d=1  real and distinct: 4.0000, 3.0000
a=1 b=-4 c=4  d=0  real and repeated: 2.0000
a=1 b=2 c=5  d=-16  complex: -1.0000 +/- 2.0000i

The leap year as a ladder

Chapter 16 wrote the leap-year rule as one expression. It is also three nested exceptions, and written as a ladder it reads as the rule itself:

#include <stdio.h>

int is_leap(int year)
{
    if (year % 400 == 0) {
        return 1;                 /* a century divisible by 400 IS a leap year */
    } else if (year % 100 == 0) {
        return 0;                 /* any other century is NOT */
    } else if (year % 4 == 0) {
        return 1;                 /* otherwise divisible by 4 IS */
    } else {
        return 0;
    }
}

int main(void)
{
    int years[] = {1900, 1996, 2000, 2023, 2024, 2100, 2400};

    for (int i = 0; i < 7; i++) {
        printf("%d %s\n", years[i], is_leap(years[i]) ? "leap" : "not leap");
    }
    return 0;
}
1900 not leap
1996 leap
2000 leap
2023 not leap
2024 leap
2100 not leap
2400 leap

Notice the order: most specific first. 400 before 100 before 4. Written the other way round, year % 4 would catch 1900 and answer wrongly. It is the same ordering lesson as the grades, in a form MU actually sets.

When to use a ladder, and when not to

SituationUse
Tests on ranges of a valueelse-if ladder
Tests on different variableselse-if ladder
A condition too complex for a tableelse-if ladder
Comparing one integer or char against separate constant valuesswitch, chapter 26
Choosing between two values rather than two actions?:, chapter 19
More than about six branches on one valueswitch, or a table

A ladder on a single variable against single values is what switch is for, and switch says so more clearly:

munotes.in119

else-if Ladders

if (choice == 1) { ... }
else if (choice == 2) { ... }
else if (choice == 3) { ... }

What this does NOT mean

else if is not a keyword. It is else followed by an if statement. C has no elif and no elseif.

A ladder does not test every condition. It stops at the first true one. This is why order matters and why later tests need no lower bound.

The final else is not required. Without one, a value matching no test simply falls through and nothing happens, which is sometimes right and is often a missing case.

A ladder is not a switch. A ladder tests arbitrary conditions; switch compares one integer expression against constants. Neither replaces the other.

Ordering from most to least demanding is not a style preference. Reversed, the ladder gives wrong answers, as grade_wrong shows.

Quick revision

  • An else-if ladder is nested if-else written flat. Nothing was added to the language.
  • The tests run in order; the first true one wins and the rest are skipped.
  • Order overlapping tests from most demanding to least.
  • Later tests need no upper bound, because earlier ones have been ruled out.
  • The final else is optional and catches everything else.
  • Quadratic roots: check a is not zero, then the sign of bb - 4a*c, with a tolerance rather than ==.
  • sqrt(-d) in the complex branch, because d is negative there.
  • Leap year as a ladder: 400, then 100, then 4, in that order.
  • Use switch where one integer is compared against separate constants.

Test yourself

1. Is else if a single keyword?

No. It is an else whose governed statement is another if. The ladder layout is nested if-else written flat.

2. What is wrong with this ladder?

if (m >= 40) g = 'D';
else if (m >= 70) g = 'O';

Everything from 70 upwards satisfies the first test, so the second is never reached. Order from most demanding to least: test 70 first.

3. In the quadratic program, why is the second test d > -1e-12 rather than d == 0?

Because d is a double computed by arithmetic and will rarely be exactly zero even when the roots are repeated. The ladder has already ruled out d > 1e-12, so d > -1e-12 means d is within the tolerance of zero.

4. Why must a be checked before the discriminant?

Because 2 * a is a divisor in every branch, and with a zero the equation is not quadratic at all.

munotes.in120

else-if Ladders

5. Write the leap-year test as a ladder and say why the order cannot be reversed.

Test year % 400 == 0 first, then year % 100 == 0, then year % 4 == 0. Reversed, year % 4 would catch 1900 and report it as a leap year, because the century exception would never be reached.

6. When is a switch better than a ladder?

When a single integer or character expression is being compared against separate constant values, as in a menu. switch states that intent and the compiler can often implement it as a jump table.

What can be asked on this, and how to answer it

"Explain the else-if ladder with syntax and an example." Give the syntax, say plainly that it is nested if-else written flat and that else if is not a keyword, then give the grade program. Include the ordering rule and one example of it going wrong, because that is what the question is really testing.

"Write a program to find the roots of a quadratic equation." Give this chapter's program. Keep the a == 0 check and the tolerance, and be ready to explain both: those are the two follow-up questions.

"Write a program to find the grade of a student from marks." Give grade_right, ordered from the highest boundary down, and say in one line why later tests need no upper bound.

"Distinguish between an else-if ladder and a switch statement." A ladder tests any conditions at all, on any number of variables, in a fixed order. A switch compares one integer or character expression against constant values, in no particular order, and may fall through. Use a ladder for ranges and a switch for discrete values. Add that switch cannot test a double or a string.

Contents This chapter on its own page

munotes.in121

Chapter Twenty-Six

switch

Syllabus topic 1, "Control Flow: Statements and Blocks, If-Else, Else-If, Switch, Loops- While and For Loops Do-while, Break and Continue, Goto and Labels"

In one line

switch compares one integer expression against a list of constant values, jumps to the matching case, and then keeps going until it meets a break.

The form

switch (expression) {
    case constant1:
        statements
        break;
    default:
        statements
        break;
}

Any number of case labels may appear between the braces. The expression is evaluated once. Control jumps to the case whose constant equals it, or to default if none does, and then runs forward from there.

That last clause is the whole of switch. It is a jump into a block, not a set of separate branches. A case is a label, not a box.

The four restrictions

These are the ones students break, and each has a reason.

  1. The expression must be of integer type. An int, a char, a short, an enum. Not a double, because equality on floating-point values is not reliable, and not a string, because a string is an array and comparing arrays is not what == does.
  2. Each case value must be a constant expression, known at compile time. A literal, a #define, an enum constant, or an arithmetic expression of those. Not a variable.
  3. No two case values may be equal. The compiler rejects a duplicate.
  4. There may be at most one default, and it may go anywhere, though the end is where a reader expects it.

A worked example

#include <stdio.h>

int main(void)
{
    for (char grade = 'A'; grade <= 'F'; grade++) {
        printf("%c : ", grade);
        switch (grade) {
            case 'A':
                printf("distinction\n");
                break;
            case 'B':
            case 'C':
                printf("first class\n");
                break;
            case 'D':
                printf("pass\n");
                break;
            default:
                printf("no such grade\n");
                break;
        }
    }
    return 0;
}
A : distinction
B : first class
C : first class
D : pass
E : no such grade
F : no such grade

case 'B': followed immediately by case 'C': is deliberate fall-through, and it is the normal way to give two values the same treatment. There is nothing between the two labels, so a match on B runs straight into C's statements. This is the good use of fall-through, and it is the only one you need this semester.

Fall-through when you did not mean it

Forget a break and control runs into the next case's statements.

#include <stdio.h>

int main(void)
{
    int choice = 2;

    switch (choice) {
        case 1:
            printf("one\n");
        case 2:
            printf("two\n");
        case 3:
            printf("three\n");
        default:
            printf("default\n");
    }
    return 0;
}

The compiler warns about every missing break, because -Wextra turns that check on:

fallthrough.c: In function ‘main’:
fallthrough.c:9:13: warning: this statement may fall through [-Wimplicit-fallthrough=]
    9 |             printf("one\n");
      |             ^~~~~~~~~~~~~~~
fallthrough.c:10:9: note: here
   10 |         case 2:
      |         ^~~~
fallthrough.c:11:13: warning: this statement may fall through [-Wimplicit-fallthrough=]
   11 |             printf("two\n");
      |             ^~~~~~~~~~~~~~~
fallthrough.c:12:9: note: here
   12 |         case 3:
      |         ^~~~
fallthrough.c:13:13: warning: this statement may fall through [-Wimplicit-fallthrough=]
   13 |             printf("three\n");
      |             ^~~~~~~~~~~~~~~~~
fallthrough.c:14:9: note: here
   14 |         default:
      |         ^~~~~~~
munotes.in122

switch

And it prints three lines when the programmer wanted one:

two
three
default

choice was 2, so control jumped to case 2 and then simply carried on downwards through case 3 and default. A case is a label, not a block.

Where you genuinely want fall-through past some statements, say so. gcc, and the standard from C23, accept an attribute:

case 1:
    printf("one, and continuing\n");
    /* fall through */
case 2:
    printf("two\n");
    break;

The comment silences the warning in gcc and, more importantly, tells the next reader that you meant it.

The practical: a menu-driven calculator

MU's Practical 2(b), in her own words: a menu-driven program using switch case to add, subtract, multiply or divide based on the user's choice.

#include <stdio.h>

int main(void)
{
    int choice;
    double a, b;

    printf("1 add\n2 subtract\n3 multiply\n4 divide\n");
    printf("Enter your choice: ");
    if (scanf("%d", &choice) != 1) {
        printf("\nThat was not a choice.\n");
        return 1;
    }
    printf("Enter two numbers: ");
    if (scanf("%lf %lf", &a, &b) != 2) {
        printf("\nThose were not two numbers.\n");
        return 1;
    }

    printf("\n");
    switch (choice) {
        case 1:
            printf("%.4f + %.4f = %.4f\n", a, b, a + b);
            break;
        case 2:
            printf("%.4f - %.4f = %.4f\n", a, b, a - b);
            break;
        case 3:
            printf("%.4f * %.4f = %.4f\n", a, b, a * b);
            break;
        case 4:
            if (b == 0.0) {
                printf("Division by zero is not defined.\n");
            } else {
                printf("%.4f / %.4f = %.4f\n", a, b, a / b);
            }
            break;
        default:
            printf("%d is not one of the four choices.\n", choice);
            break;
    }
    return 0;
}
4
22 7
1 add
2 subtract
3 multiply
4 divide
Enter your choice: Enter two numbers:
22.0000 / 7.0000 = 3.1429

Three things earn the marks here.

1. The division-by-zero check. Chapter 15 said the only correct handling is to test first. Here the operands are double, so a division by zero would not kill the program, and it would print inf, which is not an answer a student should hand in.

b == 0.0 on a double is the one case where comparing with == is right: zero is exactly representable, and what is being asked is whether the user typed a zero, not whether an arithmetic result landed on zero.

2. The default branch. An input the menu does not offer must be reported, not ignored.

3. Both scanf calls are checked. Two lines for the whole of the input validation.

munotes.in123

switch

The practical says "menu driven", and a real menu repeats until the user asks to stop. That needs a loop round the switch, which is chapter 29's do-while, and the full version is in chapter 44.

A case can share a body

Counting vowels is the standard example, and it is worth one program:

#include <stdio.h>

int main(void)
{
    const char *text = "Programming with C";
    int vowels = 0, consonants = 0, spaces = 0, others = 0;

    for (int i = 0; text[i] != '\0'; i++) {
        char c = text[i];
        if (c >= 'A' && c <= 'Z') {
            c = c - 'A' + 'a';
        }
        switch (c) {
            case 'a': case 'e': case 'i': case 'o': case 'u':
                vowels++;
                break;
            case ' ':
                spaces++;
                break;
            default:
                if (c >= 'a' && c <= 'z') {
                    consonants++;
                } else {
                    others++;
                }
                break;
        }
    }
    printf("\"%s\"\n", text);
    printf("vowels %d, consonants %d, spaces %d, others %d\n",
           vowels, consonants, spaces, others);
    return 0;
}
"Programming with C"
vowels 4, consonants 12, spaces 2, others 0

default here does real work rather than reporting an error, which is a legitimate use: the cases pick out the special values and default handles the general one.

switch against an else-if ladder

switchelse-if ladder
TestsOne integer expression against constantsAny conditions at all
Expression typeInteger or char or enum onlyAnything
case valuesConstant, distinct, compile-timeNot applicable
RangesNot directlyNaturally
double or stringsCannotCan
Order of testsIrrelevant; the match is by valuePart of the logic
Falls throughYes, unless you breakNo
Typical useA menu, a command, a stateGrade bands, compound conditions

A case cannot express a range. case 1 ... 5: is a gcc extension, not standard C, and a program that uses it will not compile elsewhere. For ranges, use a ladder.

What this does NOT mean

A case is not a block. It is a label. Control enters at the label and continues until a break or the closing brace.

break is not optional. Leaving it out is legal and is nearly always a bug. -Wextra warns.

switch cannot test a double. Equality on floating-point values is unreliable, so the language does not allow it.

switch cannot test a string. A string is an array, and == on arrays compares addresses. Use strcmp in a ladder.

default is not required and is not necessarily last. It may go anywhere; put it last because that is where a reader looks. Its absence means unmatched values do nothing.

A case value cannot be a variable. It must be a constant expression the compiler can evaluate.

munotes.in124

switch

switch is not always faster than a ladder. A compiler may build a jump table for a dense set of values, which is fast, or a chain of comparisons, which is not. Choose on clarity.

Quick revision

  • switch (expr) evaluates the expression once and jumps to the matching case label.
  • The expression must be integer, char or enum. Never a double or a string.
  • Each case value is a distinct compile-time constant.
  • Control falls through from one case to the next unless stopped by break.
  • Empty consecutive labels, case 'B': case 'C':, are the deliberate fall-through you want.
  • default handles everything unmatched; it is optional and conventionally last.
  • -Wextra warns about an implicit fall-through. A / fall through / comment silences it and documents intent.
  • switch cannot express a range; use an else-if ladder.
  • A break inside a switch leaves the switch, not any enclosing loop. Chapter 30.

Test yourself

1. What does this print when n is 2?

switch (n) {
    case 1: printf("one ");
    case 2: printf("two ");
    case 3: printf("three ");
}

two three . Control enters at case 2 and falls through case 3, because there is no break.

2. Can switch be used on a double?

No. The controlling expression must have integer type, because equality comparison of floating-point values is not reliable.

3. Can two case labels have the same value?

No. The compiler rejects a duplicate, because there would be no way to decide which to jump to.

4. Is default compulsory? Must it come last?

Neither. Without it, an unmatched value does nothing. It may appear anywhere, and it is put last by convention.

5. How do you make case 'B' and case 'C' do the same thing?

Write the labels one after the other with nothing between them: case 'B': case 'C': statements break;

6. Why can a case value not be a variable?

Because the compiler must know every case value while compiling, so that it can decide where each value jumps to. A variable's value is not known until the program runs.

7. How would you handle marks bands 40 to 49, 50 to 59 and 60 to 100 with a switch?

Not directly, because a case cannot express a range. Either use an else-if ladder, or switch on marks / 10 so that each band becomes one or more distinct integers.

What can be asked on this, and how to answer it

"Explain the switch statement with syntax and an example." Give the syntax, say the expression is evaluated once and control jumps to the matching label, and stress that execution continues from there until a break. Give the four restrictions and a worked example with a deliberate shared case.

munotes.in125

switch

"What is fall-through in a switch? Give one useful and one harmful example." Control continuing from one case into the next when no break stops it. Useful: two labels sharing a body, case 'B': case 'C':. Harmful: a missing break running the next case's statements as well, which -Wextra warns about.

"Write a menu-driven program using switch to perform addition, subtraction, multiplication and division." Give this chapter's program. Keep the division-by-zero check and the default; they are the two things the examiner asks about.

"Distinguish between switch and if-else." Give the table's rows: one integer expression against constants versus arbitrary conditions, constant distinct labels versus any test, no ranges versus ranges, fall-through versus none, and the order of tests being irrelevant in one and part of the logic in the other.

"Can a switch be nested?" Yes, and a break in the inner one leaves only the inner switch. Nesting more than one level is hard to read; a function for the inner decision is usually better.

Contents This chapter on its own page

munotes.in126

Chapter Twenty-Seven

while Loops

Syllabus topic 1, "Control Flow: Statements and Blocks, If-Else, Else-If, Switch, Loops- While and For Loops Do-while, Break and Continue, Goto and Labels"

In one line

while (condition) statement tests the condition first and runs the statement again and again for as long as the condition stays non-zero, so a while loop may run zero times.

The form, and the three parts every loop needs

while (condition)
    statement

The condition is tested. If non-zero, the statement runs, and then the condition is tested again. When the test fails, control passes to whatever follows the loop.

Every correct loop has three things, and a loop that fails to end is always missing the third:

  1. Initialisation. Something set up before the loop.
  2. A condition that can become false.
  3. Progress, inside the body, towards making it false.
#include <stdio.h>

int main(void)
{
    int i = 1;                    /* 1. initialisation */

    while (i <= 5) {              /* 2. condition      */
        printf("%d squared is %d\n", i, i * i);
        i++;                      /* 3. progress       */
    }
    printf("the loop finished with i at %d\n", i);
    return 0;
}
1 squared is 1
2 squared is 4
3 squared is 9
4 squared is 16
5 squared is 25
the loop finished with i at 6

The loop ended with i at 6, not 5. It had to: the condition is tested with i at 6, fails, and only then does the loop stop. That off-by-one is the first thing to check when a loop gives an answer one too big or one too small.

It may run zero times

This is the property that distinguishes while from do-while, and it is usually what you want.

#include <stdio.h>

int main(void)
{
    int n = 0;
    int count = 0;

    while (count < n) {
        printf("this never prints\n");
        count++;
    }
    printf("with n = %d the body ran %d time(s)\n", n, count);
    return 0;
}
with n = 0 the body ran 0 time(s)

The condition was false before the first pass, so the body never ran. A loop that must run at least once is chapter 29.

The practical: reversing the digits of a number

MU's Practical 3(a), and her own words name the loop: "using while loop to reverse the digits of a number".

The method is two operators from chapter 15. n % 10 gives the last digit; n / 10 removes it. Do both until nothing is left, building the answer up as you go.

#include <stdio.h>

int main(void)
{
    int n;

    printf("Enter a whole number: ");
    if (scanf("%d", &n) != 1) {
        printf("\nThat was not a whole number.\n");
        return 1;
    }

    int original = n;
    int negative = n < 0;
    if (negative) {
        n = -n;
    }

    int reversed = 0;
    while (n > 0) {
        int digit = n % 10;
        reversed = reversed * 10 + digit;
        n = n / 10;
    }
    if (negative) {
        reversed = -reversed;
    }

    printf("\n%d reversed is %d\n", original, reversed);
    return 0;
}
munotes.in127

while Loops

5024
Enter a whole number:
5024 reversed is 4205

Trace it by hand, which is what an examiner will ask you to do at the board:

Passn at the startdigitreversed becomesn becomes
1502440 × 10 + 4 = 4502
250224 × 10 + 2 = 4250
350042 × 10 + 0 = 4205
455420 × 10 + 5 = 42050

n is 0, the condition fails, and reversed is 4205.

Three things in that program that a bare textbook version leaves out, and each is a viva question.

1. Zero. while (n > 0) with n of 0 runs zero times, and reversed stays 0, which is right.

2. A negative number. -5024 % 10 is -4 in C, because the remainder takes the sign of the left operand (chapter 15). Without the sign handling, the loop would never end, since n > 0 is false at once and the answer would be 0. The program takes the sign off, reverses, and puts it back.

3. A trailing zero is lost, and that is arithmetic, not a bug. 5024 reversed is 4205; 50240 reversed is 4205 as well, because 04205 is 4205. If the leading zero must be kept, the answer is a string and not a number.

The practical: the factorial

MU's Practical 3(b). The factorial of n is the product of every whole number from 1 to n, and 0 factorial is 1 by definition.

#include <stdio.h>

int main(void)
{
    int n;

    printf("Enter a whole number: ");
    if (scanf("%d", &n) != 1) {
        printf("\nThat was not a whole number.\n");
        return 1;
    }
    if (n < 0) {
        printf("\nThe factorial of a negative number is not defined.\n");
        return 1;
    }
    if (n > 20) {
        printf("\n%d! is too large for a 64-bit integer. The largest that "
               "fits is 20!.\n", n);
        return 1;
    }

    unsigned long long factorial = 1;
    int i = 1;
    while (i <= n) {
        factorial = factorial * i;
        i++;
    }

    printf("\n%d! = %llu\n", n, factorial);
    return 0;
}
12
Enter a whole number:
12! = 479001600

The n > 20 check is the part worth having. Factorials grow faster than anything else in a first-semester program, and a textbook solution using int gives silently wrong answers from 13 onwards. Here is the limit, measured:

#include <stdio.h>
#include <limits.h>

int main(void)
{
    printf("an int holds up to        %d\n", INT_MAX);
    printf("13! is 6227020800, which does NOT fit in an int\n");
    printf("an unsigned long long holds up to %llu\n", ULLONG_MAX);

    unsigned long long f = 1;
    for (int i = 1; i <= 21; i++) {
        f = f * (unsigned long long) i;
        if (i >= 19) {
            printf("%2d! = %llu%s\n", i, f, i == 21 ? "   <- WRONG, it wrapped" : "");
        }
    }
    return 0;
}
munotes.in128

while Loops

an int holds up to        2147483647
13! is 6227020800, which does NOT fit in an int
an unsigned long long holds up to 18446744073709551615
19! = 121645100408832000
20! = 2432902008176640000
21! = 14197454024290336768   <- WRONG, it wrapped

21 factorial does not fit in 64 bits, and unsigned arithmetic wraps rather than reporting anything (chapter 9). The answer printed for 21 is not the factorial of 21 and the program did not notice. That is why the check comes before the loop.

The infinite loop, and the four ways to write one by accident

while (1) { ... }        /* deliberate, and legitimate with a break inside */

The accidental ones:

int i = 1;
while (i <= 5) {
    printf("%d\n", i);      /* 1. no progress: i never changes */
}

while (i != 5) { i += 2; }  /* 2. steps past the target: 1, 3, 5 works, 2, 4, 6 does not */

unsigned int j = 5;
while (j >= 0) { j--; }     /* 3. an unsigned value is never negative */

while (x != 0.3) { ... }    /* 4. floating-point equality may never hold */

The third is the one to remember. Chapter 9 showed that 0u - 1 is the maximum value, so an unsigned condition of >= 0 is always true. Use int for a counter that may go below zero.

Reading input until it ends

The standard shape, and one you will use constantly:

#include <stdio.h>

int main(void)
{
    int c;
    int lines = 0, chars = 0;

    while ((c = getchar()) != EOF) {
        chars++;
        if (c == '\n') {
            lines++;
        }
    }
    printf("%d character(s) in %d line(s)\n", chars, lines);
    return 0;
}
one
two
three
14 character(s) in 3 line(s)

Three things: c is an int so that EOF is distinguishable (chapter 12), the brackets round the assignment are required (chapter 18), and the loop ends by itself when the input runs out, which is why a while and not a for is right here.

What this does NOT mean

A while loop does not always run. If the condition is false at the start, the body runs zero times.

The condition is not rechecked during the body. It is tested between passes. Changing the variable halfway through the body does not stop the loop early; break does.

munotes.in129

while Loops

The counter does not end at the last value used. It ends at the first value that failed the test.

while (n) is not the same as while (n > 0) for a signed n. while (n) is true for negative values too, which is what makes the reverse-digits loop fail to end on a negative input.

An infinite loop is not always a bug. while (1) with a break is a normal shape, and it is how a menu repeats.

Quick revision

  • while (condition) statement tests first, so it may run zero times.
  • A correct loop has initialisation, a condition that can fail, and progress towards failing it.
  • The counter ends at the first value that failed the test, not the last one used.
  • Reverse digits: d = n % 10; r = r * 10 + d; n = n / 10; until n is 0.
  • A negative input needs the sign taken off first, because % keeps the sign of the left operand.
  • Factorial: f = f * i with i from 1 to n, and 0! is 1. 20! is the largest that fits in 64 bits.
  • Unsigned arithmetic wraps, so while (j >= 0) on an unsigned j never ends.
  • while ((c = getchar()) != EOF) reads to the end of input; c must be an int.
  • while (1) with a break is a legitimate deliberate infinite loop.

Test yourself

1. How many times does this body run?

int i = 5;
while (i < 5) { i++; }

Zero. The condition is false before the first pass.

2. What is i after int i = 1; while (i <= 5) i++;?

  1. The loop stops when the test fails, which is the first time i is 6.

3. Trace n = 407 through the reverse-digits loop.

Pass 1: digit 7, reversed 7, n 40. Pass 2: digit 0, reversed 70, n 4. Pass 3: digit 4, reversed 704, n 0. Answer 704.

4. Why does the reverse-digits program handle the sign separately?

Because in C the remainder takes the sign of the left operand, so -5024 % 10 is -4, and while (n > 0) is false at once for a negative n. Taking the sign off first makes the loop work, and it is put back at the end.

5. What is wrong with unsigned int i = 10; while (i >= 0) i--;?

An unsigned value is never negative, so the condition is always true and the loop never ends. Once i reaches 0, decrementing wraps it to the maximum value.

6. Why must c be an int in while ((c = getchar()) != EOF)?

getchar must be able to return every character value and also EOF, which is a distinct negative value. A char has no spare value for it, so a char would either never compare equal to EOF or would compare equal to a real character.

munotes.in130

while Loops

7. What is the largest factorial that fits in an unsigned long long, and what happens beyond it?

20 factorial. Beyond that the value wraps modulo 2 to the 64, so the program prints a number that is not the factorial and reports no error.

What can be asked on this, and how to answer it

"Explain the while loop with syntax and an example." Give the syntax, say the condition is tested before each pass so the body may run zero times, and name the three parts a loop needs. Give a counting example and point out that the counter ends one past the last value used.

"Write a program using a while loop to reverse the digits of a number." Give this chapter's program and the trace table. Be ready for the three follow-ups: zero, a negative number, and a trailing zero.

"Write a program to find the factorial of a number." Give the while version with unsigned long long, the check for a negative input and the check for n > 20. Saying why 20 is the limit is what separates this from a copied answer.

"Distinguish between while and do-while." while tests before the body, so it may run zero times. do-while tests after, so it always runs at least once, and its while clause ends with a semicolon. Give input validation as the case for do-while.

"What is an infinite loop? Give two ways one happens by accident." A loop whose condition never becomes false. Accidentally: no progress in the body, and a condition that cannot fail, such as >= 0 on an unsigned variable. Add that while (1) with a break is a deliberate and useful form.

Contents This chapter on its own page

munotes.in131

Chapter Twenty-Eight

for Loops

Syllabus topic 1, "Control Flow: Statements and Blocks, If-Else, Else-If, Switch, Loops- While and For Loops Do-while, Break and Continue, Goto and Labels"

In one line

for (initialisation; condition; step) statement is a while loop with its three parts collected into one line, which is why it is the loop to use whenever you know how many passes there will be.

The form, and what it is equivalent to

for (initialisation; condition; step)
    statement

means exactly this:

initialisation;
while (condition) {
    statement
    step;
}

That equivalence is the whole chapter. Everything a for loop does follows from it, including the order the three clauses run in:

  1. Initialisation, once, before anything else.
  2. Condition, tested before every pass. If false, the loop ends.
  3. The body.
  4. The step, after the body, then back to 2.
#include <stdio.h>

int main(void)
{
    for (int i = 1; i <= 5; i++) {
        printf("%d cubed is %d\n", i, i * i * i);
    }

    printf("counting down: ");
    for (int i = 5; i >= 1; i--) {
        printf("%d ", i);
    }
    printf("\n");

    printf("in threes: ");
    for (int i = 0; i <= 20; i += 3) {
        printf("%d ", i);
    }
    printf("\n");
    return 0;
}
1 cubed is 1
2 cubed is 8
3 cubed is 27
4 cubed is 64
5 cubed is 125
counting down: 5 4 3 2 1
in threes: 0 3 6 9 12 15 18

The counter declared in the header belongs to the loop and does not exist after it (chapter 21). That is the C99 form and the one to use. If your college compiler rejects it, it is set to C89: declare int i; above the loop and write for (i = 1; ...).

Any clause may be empty

All three are optional, and the semicolons are not.

#include <stdio.h>

int main(void)
{
    int i;

    i = 1;
    for (; i <= 3; i++) {                 /* no initialisation */
        printf("a%d ", i);
    }
    printf("\n");

    for (i = 1; i <= 3; ) {               /* no step: the body does it */
        printf("b%d ", i);
        i++;
    }
    printf("\n");

    for (i = 1; ; i++) {                  /* no condition: always true */
        if (i > 3) {
            break;
        }
        printf("c%d ", i);
    }
    printf("\n");

    i = 0;
    for (;;) {                            /* all three empty */
        i++;
        if (i > 3) {
            break;
        }
        printf("d%d ", i);
    }
    printf("\n");
    return 0;
}
a1 a2 a3
b1 b2 b3
c1 c2 c3
d1 d2 d3

for (;;) is the idiomatic deliberate infinite loop in C, and it means the same as while (1). Either is fine; for (;;) is what you will read in other people's code.

Two counters, with the comma operator

Chapter 18 said the comma operator's real home is a for header. Here it is.

#include <stdio.h>
#include <string.h>

int main(void)
{
    char word[] = "abcdefg";
    int n = (int) strlen(word);

    printf("before: %s\n", word);
    for (int i = 0, j = n - 1; i < j; i++, j--) {
        char t = word[i];
        word[i] = word[j];
        word[j] = t;
    }
    printf("after : %s\n", word);
    return 0;
}
munotes.in132

for Loops

before: abcdefg
after : gfedcba

Two counters walk towards each other and the loop stops when they meet. That is the standard way to reverse an array in place, and chapter 37 uses it on the palindrome test.

The practical: the Fibonacci series

MU's Practical 3(c). Each term is the sum of the two before it, starting 0 and 1.

#include <stdio.h>

int main(void)
{
    int n;

    printf("How many terms? ");
    if (scanf("%d", &n) != 1 || n < 1) {
        printf("\nThat was not a count of terms.\n");
        return 1;
    }

    printf("\n");
    long long a = 0, b = 1;
    for (int i = 1; i <= n; i++) {
        printf("%lld ", a);
        long long next = a + b;
        a = b;
        b = next;
    }
    printf("\n");
    return 0;
}
12
How many terms?
0 1 1 2 3 5 8 13 21 34 55 89

Read the three lines in the body: print the current term, work out the next one, then shift the pair along. The temporary next is necessary. Without it, a = b; b = a + b; would use the new a in the second line and give the wrong series.

long long rather than int, because the series passes two thousand million at the 47th term. Chapter 35 gives the same series recursively and shows why that version is unusable past about the 40th term.

The practical: patterns of asterisks

MU's Practical 2(c). This is what nested loops are for: the outer loop counts the rows and the inner loop draws one row.

#include <stdio.h>

int main(void)
{
    int n = 5;

    printf("1. a right triangle\n");
    for (int row = 1; row <= n; row++) {
        for (int col = 1; col <= row; col++) {
            printf("*");
        }
        printf("\n");
    }

    printf("\n2. an inverted triangle\n");
    for (int row = n; row >= 1; row--) {
        for (int col = 1; col <= row; col++) {
            printf("*");
        }
        printf("\n");
    }

    printf("\n3. a pyramid\n");
    for (int row = 1; row <= n; row++) {
        for (int space = 1; space <= n - row; space++) {
            printf(" ");
        }
        for (int col = 1; col <= 2 * row - 1; col++) {
            printf("*");
        }
        printf("\n");
    }
    return 0;
}
1. a right triangle
*
**
***
****
*****

2. an inverted triangle
*****
****
***
**
*

3. a pyramid
    *
   ***
  *****
 *******
*********
munotes.in133

for Loops

The pyramid is the one worth understanding rather than memorising. Row row needs n - row spaces and then 2 * row - 1 asterisks. Check it: row 1 gets 4 spaces and 1 asterisk, row 5 gets 0 spaces and 9 asterisks. Every pattern question is that same arithmetic: work out, for row row, how many of each thing, and the loops write themselves.

And the number patterns MU also sets:

#include <stdio.h>

int main(void)
{
    printf("Floyd's triangle\n");
    int value = 1;
    for (int row = 1; row <= 4; row++) {
        for (int col = 1; col <= row; col++) {
            printf("%d ", value);
            value++;
        }
        printf("\n");
    }

    printf("\nthe multiplication table of 1 to 5\n");
    for (int i = 1; i <= 5; i++) {
        for (int j = 1; j <= 5; j++) {
            printf("%4d", i * j);
        }
        printf("\n");
    }
    return 0;
}
Floyd's triangle
1
2 3
4 5 6
7 8 9 10

the multiplication table of 1 to 5
   1   2   3   4   5
   2   4   6   8  10
   3   6   9  12  15
   4   8  12  16  20
   5  10  15  20  25

for or while

Both can do anything the other can. Which to use is a question about what the loop is.

Use for whenUse while when
The number of passes is known before the loop startsIt is not
A counter walks a rangeThe loop ends on an event, such as input running out
Walking an arrayWaiting for a condition to change
The three parts are short enough to read on one lineThe condition is complex

The honest rule: a for loop says "this many times", a while loop says "until this happens". Choose the one that tells the truth about your loop.

What this does NOT mean

The step does not run before the first pass. Order is initialisation, condition, body, step.

The step does not run after the last body. It does: the step runs, then the condition fails. That is why the counter ends one past the last value used.

A for loop is not required to count. Any expression may be the step, and any condition may be the test.

The semicolons are not optional even when the clauses are. for (;;) has two.

for (int i = ...) is not universally available. It is C99. Some college machines compile as C89 and will reject it.

A nested loop's counters must not share a name. They may, and the inner one then shadows the outer (chapter 21), which breaks the outer loop's counting. Use row and col, not i and i.

Quick revision

  • for (init; condition; step) statement.
  • Order: init once, then condition, body, step, condition, body, step, and so on.
  • It is exactly init; while (condition) { statement step; }.
  • Any clause may be empty; the semicolons stay. for (;;) is a deliberate infinite loop.
  • The counter ends at the first value that failed the condition.
  • for (int i = ...) scopes the counter to the loop, and is C99.
  • The comma operator steps two counters: for (int i = 0, j = n - 1; i < j; i++, j--).
  • Nested loops: the outer counts rows, the inner draws one row.
  • Pyramid row row of n: n - row spaces then 2 * row - 1 asterisks.
  • Fibonacci needs a temporary for the next term, and long long past the 46th.
  • Use for for a known number of passes, while for an event.
munotes.in134

for Loops

Test yourself

1. In what order do the three clauses of a for loop run?

Initialisation once, then for each pass the condition, then the body, then the step.

2. Rewrite for (i = 0; i < n; i++) sum += a[i]; as a while loop.

i = 0;
while (i < n) { sum += a[i]; i++; }

3. What does for (;;) do?

Loops for ever: with no condition the test is treated as true. It is ended with a break or a return.

4. How many times does the inner printf run?

for (int i = 1; i <= 3; i++)
    for (int j = 1; j <= 4; j++)
        printf("*");

Twelve: three outer passes times four inner passes.

5. For a pyramid of n rows, how many spaces and asterisks does row row need?

n - row spaces and 2 * row - 1 asterisks.

6. Why does the Fibonacci loop need a temporary variable?

Because both a and b must change together. Writing a = b; b = a + b; uses the already-updated a in the second statement and produces the wrong series.

7. What is i after for (int i = 0; i < 5; i++) { }, and can you print it afterwards?

Inside the loop it would end at 5, but the name does not exist after the loop, so it cannot be printed. Declare i before the loop if you need its final value.

What can be asked on this, and how to answer it

"Explain the for loop with syntax and an example." Give the syntax, the order of the three clauses, and the while equivalent, which is the part that shows understanding. Then a counting example. Add that any clause may be empty.

"Write a program to print the Fibonacci series." Give this chapter's for version, name the temporary and say why it is needed, and mention the type limit.

munotes.in135

for Loops

"Write a program to print a pyramid of stars." Give the program and, more importantly, the arithmetic: n - row spaces and 2 * row - 1 asterisks. An examiner who asks for a different pattern is asking for the same method.

"Distinguish between for and while." Both are pre-tested loops and either can do the other's job. A for collects initialisation, condition and step into one header and suits a known number of passes; a while suits a loop that ends on an event. Give the mechanical equivalence as the closing line.

"What is a nested loop? Give an example." A loop inside another loop's body. The outer loop runs once per row, the inner once per item in that row, so the inner body runs the product of the two counts. Give the multiplication table or the right triangle.

Contents This chapter on its own page

munotes.in136

Chapter Twenty-Nine

do-while

Syllabus topic 1, "Control Flow: Statements and Blocks, If-Else, Else-If, Switch, Loops- While and For Loops Do-while, Break and Continue, Goto and Labels"

In one line

do statement while (condition); runs the statement first and tests afterwards, so the body always runs at least once.

The form

do
    statement
while (condition);

The body runs. Then the condition is tested. If it is non-zero, the body runs again. The semicolon at the end is part of the statement and is not optional.

#include <stdio.h>

int main(void)
{
    int i = 1;

    do {
        printf("%d ", i);
        i++;
    } while (i <= 5);
    printf("\n");

    int n = 0;
    int count = 0;
    do {
        count++;
    } while (count < n);
    printf("with n = %d the body still ran %d time(s)\n", n, count);
    return 0;
}
1 2 3 4 5
with n = 0 the body still ran 1 time(s)

The second loop is the point. count < n was false before the loop began, and the body ran anyway, because a do-while tests afterwards. A while loop with the same condition runs zero times, which chapter 27 showed.

The three loops compared

whilefordo-while
Condition testedBefore the bodyBefore the bodyAfter the body
Minimum passes001
Parts collected in the headerCondition onlyAll threeCondition only
Ends with a semicolonNoNoYes
SuitsAn eventA known countSomething that must happen once

All three are interchangeable in the sense that any loop can be written with any of them. Which one you choose tells the reader what kind of loop it is, and that is the only reason to prefer one.

The missing semicolon

#include <stdio.h>

int main(void)
{
    int i = 1;

    do {
        printf("%d ", i);
        i++;
    } while (i <= 5)
    printf("\n");
    return 0;
}
nosemi.c: In function ‘main’:
nosemi.c:10:21: error: expected ‘;’ before ‘printf’
   10 |     } while (i <= 5)
      |                     ^
      |                     ;
   11 |     printf("\n");
      |     ~~~~~~

The compiler is looking for the semicolon that ends the do-while statement and finds a printf instead. Modern gcc is helpful about it: it points at the end of line 10, which is exactly where the semicolon belongs, prints the ; it wants on a line of its own, and then shows the printf that made it notice. An older compiler will simply report a syntax error on line 11, which is why chapter 5 said an error is reported where the parser got stuck rather than where the mistake is.

What a do-while is actually for: asking until the answer is usable

A while loop would need the reading code written twice, once before the loop and once inside it. A do-while needs it once.

#include <stdio.h>

int main(void)
{
    int marks;
    int attempts = 0;

    do {
        printf("Enter marks out of 100: ");
        if (scanf("%d", &marks) != 1) {
            printf("\nInput ended.\n");
            return 1;
        }
        attempts++;
        if (marks < 0 || marks > 100) {
            printf("\n%d is not between 0 and 100. Try again.\n", marks);
        }
    } while (marks < 0 || marks > 100);

    printf("\nAccepted %d after %d attempt(s).\n", marks, attempts);
    return 0;
}
munotes.in137

do-while

150
-8
63
Enter marks out of 100:
150 is not between 0 and 100. Try again.
Enter marks out of 100:
-8 is not between 0 and 100. Try again.
Enter marks out of 100:
Accepted 63 after 3 attempt(s).

Read the shape, because it is the one to reuse: ask, check, and repeat while the answer is unusable. The scanf return is checked separately, so that a loop cannot spin for ever on input that has ended.

That last point matters. A validation loop with no escape when the input runs out is an infinite loop, and it is the commonest way a student's program hangs in a practical examination.

The menu shape

MU's Practical 2(b) asked for a menu-driven calculator, and chapter 26 gave the switch. A real menu repeats, and a menu must be shown at least once, so it is a do-while round a switch.

#include <stdio.h>

int main(void)
{
    int choice;

    do {
        printf("1 greet\n2 count to three\n0 quit\n");
        printf("Choice: ");
        if (scanf("%d", &choice) != 1) {
            printf("\nInput ended.\n");
            return 1;
        }
        printf("\n");
        switch (choice) {
            case 1:
                printf("Hello.\n");
                break;
            case 2:
                for (int i = 1; i <= 3; i++) {
                    printf("%d ", i);
                }
                printf("\n");
                break;
            case 0:
                printf("Goodbye.\n");
                break;
            default:
                printf("%d is not on the menu.\n", choice);
                break;
        }
        printf("\n");
    } while (choice != 0);
    return 0;
}
1
2
9
0
1 greet
2 count to three
0 quit
Choice:
Hello.

1 greet
2 count to three
0 quit
Choice:
1 2 3

1 greet
2 count to three
0 quit
Choice:
9 is not on the menu.

1 greet
2 count to three
0 quit
Choice:
Goodbye.

case 0 prints its goodbye and the loop condition then ends the loop. Using break in case 0 to try to leave the loop would not work: a break inside a switch leaves the switch, not the loop around it. Chapter 30 is exactly that trap.

The do { } while (0) idiom

You will meet this and it is worth recognising once:

#define SWAP(a, b) do { int t = (a); (a) = (b); (b) = t; } while (0)

A macro that expands to several statements has to behave as one statement, so that if (x) SWAP(p, q); else ... works. A bare block would leave the semicolon at the call site dangling; do { ... } while (0) runs once and swallows the semicolon. It is not a loop. It is the only way to write a multi-statement macro safely, and chapter 22's advice still holds: prefer a function.

munotes.in138

do-while

What this does NOT mean

A do-while is not a while written backwards. The difference is real: the body runs before the first test, so it always runs at least once.

The semicolon is not optional. It is the only control structure in C that ends with one.

while (condition); on its own is not a do-while. It is a while loop with an empty body, which is chapter 23's stray-semicolon bug.

A do-while does not guarantee valid input. It guarantees one pass. Making the input valid is the condition's job, and handling input that has ended is yours.

break inside a switch inside a do-while does not leave the loop. It leaves the switch.

Quick revision

  • do statement while (condition); tests after the body, so the body always runs at least once.
  • The trailing semicolon is required. It is the only loop that needs one.
  • while and for may run zero times; do-while may not.
  • Use it when something must happen once: asking for input, showing a menu, a retry.
  • Validation shape: ask, check, repeat while the answer is unusable, with an escape for input that ends.
  • A menu is a do-while round a switch, with 0 meaning quit.
  • do { ... } while (0) in a macro is not a loop; it makes several statements behave as one.

Test yourself

1. How many times does this body run?

int i = 10;
do { printf("%d ", i); i++; } while (i < 5);

Once. The body runs before the condition is tested, and the condition is false the first time.

2. What is the one syntactic feature that distinguishes do-while from the other loops?

The semicolon after the while clause.

3. Write a loop that asks for a positive number and keeps asking until it gets one.

do {
    printf("Enter a positive number: ");
    if (scanf("%d", &n) != 1) return 1;
} while (n <= 0);

4. Why is a do-while the right loop for a menu?

Because the menu must be shown at least once before the user can choose anything, and then repeated for as long as they do not choose to quit.

5. In a do-while containing a switch, how do you leave the loop from inside a case?

Set the loop's condition variable so the test fails, or use a flag, or return from the function. A break leaves only the switch.

6. What is do { ... } while (0) used for?

munotes.in139

do-while

To make a multi-statement macro behave as a single statement, so that it can be used as the body of an if without the semicolon at the call site causing a syntax error. It is not a loop.

What can be asked on this, and how to answer it

"Explain the do-while loop with syntax and an example." Give the syntax including the semicolon, say the condition is tested after the body so it runs at least once, and give a program where the condition is false from the start to prove it. Contrast with while in one line.

"Distinguish between while and do-while." while tests before the body and may run zero times; do-while tests after and always runs at least once. do-while ends with a semicolon. while suits a loop that may not need to run; do-while suits input validation and menus.

"Write a program to accept marks and validate them using a do-while loop." Give this chapter's validation program. Keep the scanf return check and say why: without it, input that has ended makes the loop spin for ever.

"Which loop is a post-test loop?" do-while. while and for are pre-test loops.

Contents This chapter on its own page

munotes.in140

Chapter Thirty

break and continue

Syllabus topic 1, "Control Flow: Statements and Blocks, If-Else, Else-If, Switch, Loops- While and For Loops Do-while, Break and Continue, Goto and Labels"

In one line

break leaves the innermost loop or switch at once, and continue abandons the rest of the current pass and goes on to the next one.

break

#include <stdio.h>

int main(void)
{
    int a[] = {4, 17, 8, 23, 42, 9};
    int n = 6;
    int target = 23;
    int found_at = -1;

    for (int i = 0; i < n; i++) {
        printf("looking at a[%d] = %d\n", i, a[i]);
        if (a[i] == target) {
            found_at = i;
            break;                      /* stop: there is nothing left to learn */
        }
    }

    if (found_at >= 0) {
        printf("%d is at index %d\n", target, found_at);
    } else {
        printf("%d is not in the array\n", target);
    }
    return 0;
}
looking at a[0] = 4
looking at a[1] = 17
looking at a[2] = 8
looking at a[3] = 23
23 is at index 3

The loop examined four elements and stopped. It did not look at 42 or 9, because the question was answered. That is the honest use of break: a search, where continuing would be work with no purpose.

found_at starting at -1 is the idiom that goes with it. After the loop, -1 means "not found" and anything else is the position. Chapter 36 uses the same pattern.

continue

#include <stdio.h>

int main(void)
{
    int a[] = {4, -17, 8, -23, 42, 0, 9};
    int n = 7;
    int sum = 0, counted = 0;

    for (int i = 0; i < n; i++) {
        if (a[i] <= 0) {
            continue;                   /* skip this one, go on to the next */
        }
        sum = sum + a[i];
        counted++;
    }
    printf("the %d positive value(s) sum to %d\n", counted, sum);
    return 0;
}
the 4 positive value(s) sum to 63

continue did not end the loop; it ended the pass. It is a way of saying "this one does not interest us" at the top of a body, which keeps the rest of the body from being wrapped in an if.

The same loop written without continue:

for (int i = 0; i < n; i++) {
    if (a[i] > 0) {
        sum = sum + a[i];
        counted++;
    }
}

For two lines, that version is clearer. continue earns its place when the body is long, or when there are several separate reasons to skip.

The difference between continue in a for and in a while

This is the one to get right, and it is where continue causes a hung program.

In a for loop, continue jumps to the step clause, so the counter still advances.

In a while loop, continue jumps straight to the condition, so anything in the body that was going to advance the counter is skipped.

munotes.in141

break and continue

#include <stdio.h>

int main(void)
{
    printf("for loop, continue skips to the step: ");
    for (int i = 1; i <= 6; i++) {
        if (i % 2 == 0) {
            continue;
        }
        printf("%d ", i);
    }
    printf("\n");

    printf("while loop, written correctly: ");
    int i = 0;
    while (i < 6) {
        i++;                            /* advance FIRST */
        if (i % 2 == 0) {
            continue;
        }
        printf("%d ", i);
    }
    printf("\n");
    return 0;
}
for loop, continue skips to the step: 1 3 5
while loop, written correctly: 1 3 5

The for loop is safe because i++ is in the header and continue runs it. The while loop was made safe by advancing i before the continue can be reached.

Written the other way round it never ends:

int i = 0;
while (i < 6) {
    if (i % 2 == 0) {
        continue;          /* jumps to the condition; i never changes */
    }
    printf("%d ", i);
    i++;
}

i starts at 0, which is even, so continue runs, the condition is tested, i is still 0, and so on for ever. The rule: in a while loop, make the progress happen before any continue can skip it.

break in a switch inside a loop

Chapter 29 said a break in a case does not leave the loop. Here is the proof.

#include <stdio.h>

int main(void)
{
    printf("with break in the case:\n");
    for (int i = 1; i <= 4; i++) {
        switch (i) {
            case 3:
                printf("  i is 3, and break here leaves the SWITCH\n");
                break;
            default:
                printf("  i is %d\n", i);
                break;
        }
    }

    printf("\nto leave the loop, use a flag:\n");
    int done = 0;
    for (int i = 1; i <= 4 && !done; i++) {
        switch (i) {
            case 3:
                printf("  i is 3, setting done\n");
                done = 1;
                break;
            default:
                printf("  i is %d\n", i);
                break;
        }
    }
    printf("\n");
    return 0;
}
with break in the case:
  i is 1
  i is 2
  i is 3, and break here leaves the SWITCH
  i is 4

to leave the loop, use a flag:
  i is 1
  i is 2
  i is 3, setting done

The first loop ran all four passes. The break in case 3 ended the switch for that pass and the loop carried on, which is almost never what a student writing that code intends.

Three ways out of a loop from inside a switch:

  1. A flag, tested in the loop's condition, as above.
  2. return, if the loop is the last thing the function does.
  3. goto, which chapter 31 covers and which is the honest choice for leaving several nested levels at once.
munotes.in142

break and continue

break in nested loops

break leaves one level. In a nested loop it leaves the inner loop only.

#include <stdio.h>

int main(void)
{
    printf("break leaves only the inner loop:\n");
    for (int i = 1; i <= 3; i++) {
        for (int j = 1; j <= 3; j++) {
            if (j == 2) {
                break;
            }
            printf("  i=%d j=%d\n", i, j);
        }
    }

    printf("\nleaving both, with a flag:\n");
    int stop = 0;
    for (int i = 1; i <= 3 && !stop; i++) {
        for (int j = 1; j <= 3; j++) {
            if (i == 2 && j == 2) {
                stop = 1;
                break;
            }
            printf("  i=%d j=%d\n", i, j);
        }
    }
    return 0;
}
break leaves only the inner loop:
  i=1 j=1
  i=2 j=1
  i=3 j=1

leaving both, with a flag:
  i=1 j=1
  i=1 j=2
  i=1 j=3
  i=2 j=1

The first nest printed j=1 three times: the inner break fired on every outer pass and the outer loop kept going.

Where break and continue are not allowed

  • continue in a switch that is not inside a loop is an error. There is no pass to continue.
  • break outside any loop or switch is an error.
  • break in a switch inside a loop applies to the switch. continue in the same place applies to the loop, because a switch has no passes. That asymmetry is worth a mark: the two keywords do not both stop at the switch.
#include <stdio.h>

int main(void)
{
    for (int i = 1; i <= 4; i++) {
        switch (i) {
            case 2:
                continue;               /* skips to the loop's next pass */
            default:
                break;                  /* leaves the switch only */
        }
        printf("reached the end of pass %d\n", i);
    }
    return 0;
}
reached the end of pass 1
reached the end of pass 3
reached the end of pass 4

Pass 2 printed nothing: continue skipped the rest of the loop body, printf included. The other three passes reached the printf, because their break left only the switch.

What this does NOT mean

break does not leave a function. return does.

break does not leave more than one level. Use a flag, a return or a goto.

continue does not restart the loop from the beginning. It moves to the next pass: the step clause in a for, or the condition in a while.

continue in a while does not run anything at the bottom of the body. That is how a while loop with a continue hangs.

break in a case does not leave the enclosing loop. It leaves the switch.

Neither is bad practice. break in a search and continue as a guard at the top of a body are both clearer than the alternatives. What is unclear is several of them scattered through a long body.

munotes.in143

break and continue

Quick revision

  • break leaves the innermost enclosing loop or switch.
  • continue ends the current pass and goes on to the next.
  • In a for, continue runs the step clause. In a while, it goes straight to the condition.
  • In a while, advance the counter before any continue can skip it, or the loop never ends.
  • break in a case leaves the switch, not the loop. continue in a case does affect the loop.
  • break leaves one level of nesting. Use a flag, a return or a goto for more.
  • break never leaves a function; return does.
  • The search idiom: found_at = -1 before the loop, set and break on a match.

Test yourself

1. What does break do inside a switch that is inside a for loop?

It ends the switch for that pass. The loop continues with its next pass.

2. Why does this loop never end?

int i = 0;
while (i < 5) {
    if (i == 0) continue;
    printf("%d ", i);
    i++;
}

continue jumps to the condition, so i++ at the bottom of the body is never reached while i is 0. i stays 0 for ever. Advance i before the continue.

3. How many lines does this print?

for (int i = 1; i <= 3; i++)
    for (int j = 1; j <= 3; j++) {
        if (j == 2) break;
        printf("%d %d\n", i, j);
    }

Three. The inner break fires on each outer pass after printing j=1, and the outer loop runs three times.

4. How do you leave both loops of a nest at once?

Set a flag tested by the outer loop's condition, return from the function, or goto a label after the outer loop.

5. What is the difference between break and continue in one line?

break ends the loop; continue ends only the current pass.

6. Is continue allowed in a switch?

Only if the switch is inside a loop, and it then applies to the loop, not the switch. Outside a loop it is an error.

What can be asked on this, and how to answer it

"Explain break and continue with examples." Define both in one line each, give a search using break and a skip using continue, and then the two things that earn the extra marks: continue runs the step clause in a for but not the bottom of a while body, and break in a case leaves the switch rather than the loop.

"What is the difference between break and continue?" break ends the innermost loop or switch immediately. continue abandons the rest of the current pass and proceeds to the next. Add that break also works in a switch while continue in a switch refers to an enclosing loop.

munotes.in144

break and continue

"How do you exit from a nested loop?" break leaves only one level. Use a flag in the outer condition, a return if the loop is the last thing in the function, or a goto to a label after the loops. Say which you would choose and why.

"Trace the output" of a loop with break or continue in it. Write the pass number and the variable's value in a small table and mark where control jumped. That table is the answer; the printed lines fall out of it.

Contents This chapter on its own page

munotes.in145

Chapter Thirty-One

goto and Labels

Syllabus topic 1, "Control Flow: Statements and Blocks, If-Else, Else-If, Switch, Loops- While and For Loops Do-while, Break and Continue, Goto and Labels"

In one line

goto label; jumps to a labelled statement in the same function, and it is the only unrestricted jump C has.

The form

A label is an identifier followed by a colon, attached to a statement:

again:
    printf("here\n");

and the jump is:

goto again;

Three rules:

  1. The label must be in the same function. There is no jumping between functions.
  2. A label may appear before or after the goto, so the jump may go forwards or backwards.
  3. A label has function scope: it is visible everywhere in the function, including inside blocks it is not in.

The practical: a program using goto

MU asks for one, so here is one that is not artificial: a loop written with goto, which is what goto was for before while existed.

#include <stdio.h>

int main(void)
{
    int i = 1;

start:
    if (i > 5) {
        goto finished;
    }
    printf("%d ", i);
    i++;
    goto start;

finished:
    printf("\nthe loop ended with i at %d\n", i);
    return 0;
}
1 2 3 4 5
the loop ended with i at 6

That is exactly the while loop of chapter 27, written out in the jumps a while loop compiles into. Seeing it once is worth something: it shows that a loop is not a primitive but a tidy way of writing two jumps.

And the multiplication table, which is the form the practical is usually set in:

#include <stdio.h>

int main(void)
{
    int n = 7;
    int i = 1;

loop:
    printf("%d x %d = %d\n", n, i, n * i);
    i++;
    if (i <= 5) {
        goto loop;
    }
    return 0;
}
7 x 1 = 7
7 x 2 = 14
7 x 3 = 21
7 x 4 = 28
7 x 5 = 35

Why it is avoided

The argument is not that goto is mysterious. It is that a program using it freely cannot be read a piece at a time.

With if, loops and functions, control enters a block at the top and leaves at the bottom. That means you can read a block, understand it, and then forget how it works while you read the next one. With goto, control may arrive at any labelled statement from anywhere in the function, so understanding any one part requires knowing every place that jumps into it.

Edsger Dijkstra put this in a letter in 1968 titled "Go To Statement Considered Harmful", which is where the phrase comes from. The substance of it is the paragraph above.

Two concrete consequences for a student:

  • A loop written with goto has its three parts, initialisation, condition and progress, scattered. A while loop has them where a reader expects them, so a missing one is visible.
  • A goto that jumps into a block past a declaration with an initialiser skips the initialisation. The variable then exists and holds nothing, which is chapter 11's undefined behaviour arriving by a new route. Compilers warn about the worst cases; they do not catch all of them.
munotes.in146

goto and Labels

The one use that is genuinely defensible

Leaving several nested levels at once. Chapter 30 showed that break leaves one level and that a flag is the usual alternative. When there are three levels, the flags multiply and the condition of every loop grows a clause. A goto to a label after the loops says the thing directly.

#include <stdio.h>

int main(void)
{
    int target = 23;
    int grid[3][4] = {{4, 17, 8, 2},
                      {11, 23, 5, 9},
                      {1, 6, 14, 30}};

    for (int i = 0; i < 3; i++) {
        for (int j = 0; j < 4; j++) {
            if (grid[i][j] == target) {
                printf("found %d at row %d, column %d\n", target, i, j);
                goto found;
            }
        }
    }
    printf("%d is not in the grid\n", target);
    goto done;

found:
    printf("the search stopped as soon as it had the answer\n");
done:
    printf("finished\n");
    return 0;
}
found 23 at row 1, column 1
the search stopped as soon as it had the answer
finished

The same search with flags needs a stop variable, a clause in the outer condition and a test after the loops. The goto version is shorter and, unusually for goto, clearer. This is the shape the Linux kernel uses for error handling, and it is the one case where experienced C programmers reach for goto without apology.

Note the discipline that makes it readable: the jump goes forwards, and only out of the nest. A backward goto, or one that jumps into a loop, is where the trouble lives.

What to write instead

Instead ofUse
A backward jump to repeat somethingwhile, for or do-while
A forward jump to skip the rest of a passcontinue
A forward jump out of one loopbreak
A forward jump out of a functionreturn
A forward jump out of several nested loopsgoto to a label after them, or a function with return
A jump to shared cleanup on an errorgoto, or a function

The last two rows are why goto is still in the language, and everything above them is why you will hardly ever write one.

What this does NOT mean

goto cannot leave a function. The label must be in the same function. setjmp and longjmp are the library's answer to that and are not first-semester material.

A label is not a variable. It shares the identifier namespace with nothing else, so a label called count and a variable called count can coexist. That is legal and confusing.

munotes.in147

goto and Labels

A label needs a statement. done: immediately before a closing brace was not allowed before C23. Write done: ; with an empty statement, or put something after it.

goto is not faster. A while loop compiles to the same jumps.

goto is not forbidden. MU sets a practical on it, and the multi-level exit is a real use. What is discouraged is using it in place of the structured statements.

A goto does not re-run declarations it jumps over. Jumping past int x = 5; leaves x existing and uninitialised.

Quick revision

  • goto label; jumps to a labelled statement in the same function.
  • A label is identifier: attached to a statement, and has function scope.
  • The jump may go forwards or backwards; it may not leave the function.
  • Any loop can be written with goto, which is what a loop compiles into.
  • It is avoided because control can arrive at a label from anywhere, so a program cannot be read a piece at a time.
  • Jumping over a declaration with an initialiser skips the initialisation.
  • The defensible use is leaving several nested loops at once, jumping forwards to a label after them.
  • Prefer break, continue, return and the loops for everything else.

Test yourself

1. Write the syntax of a label and of a goto.

A label is name: before a statement. The jump is goto name;.

2. Can a goto jump from one function into another?

No. The label must be in the same function as the goto.

3. Write a loop that prints 1 to 5 using goto and no loop keyword.

int i = 1;
start:
    if (i > 5) goto done;
    printf("%d ", i);
    i++;
    goto start;
done: ;

4. Why is goto discouraged?

Because control can reach a labelled statement from anywhere in the function, so no part of the program can be understood on its own. Structured statements enter at the top and leave at the bottom, which is what makes them readable.

5. Give one case where goto is the best available choice.

Leaving two or three nested loops at once, by jumping forwards to a label after them. The alternative is a flag per level and an extra clause in every loop condition.

6. What is the danger in goto skip; int x = 5; skip: printf("%d", x);?

The jump skips the initialisation, so x exists and holds nothing. Reading it is undefined behaviour.

What can be asked on this, and how to answer it

"Explain the goto statement with syntax and an example." Give the label and jump syntax, the three rules, and a program. The loop written with goto is the best example because it shows what the statement is for. Say that the label must be in the same function.

munotes.in148

goto and Labels

"Write a program using the goto statement." Give the multiplication table or the 1-to-5 loop from this chapter. Add one sentence saying you would normally write it as a while loop, which shows you know why the practical exists.

"Why should goto be avoided?" Because it destroys the property that a block is entered at the top and left at the bottom, so a reader must know every jump into a label to understand the code at it. Add the concrete case: a loop's initialisation, condition and progress become scattered, and a jump over a declaration skips its initialisation.

"Compare goto with break and continue." All three transfer control. break leaves the innermost loop or switch, continue starts the next pass of the innermost loop, and both are restricted to one level and one direction. goto may jump anywhere in the function, in either direction, which is why the first two are preferred where they fit.

Contents This chapter on its own page

munotes.in149

Chapter Thirty-Two

What a Function Is: Definition, Call, Arguments, Return

Syllabus topic 2, "Basics of functions. User defined and Library functions"

In one line

A function is a named block with a list of parameters and a return type, which you call by name to get its work done and, if it returns one, its value.

Why programs are made of functions

Three reasons, and the third is the one that matters most in an examination.

To stop repeating yourself. Code written once is corrected once.

To give a piece of work a name. simple_interest(p, r, t) says what the line does; the same arithmetic written out does not.

To make a program you can reason about a piece at a time. A function has a small, stated interface: these values in, this value out. Once you trust it you can stop thinking about how it works, which is the only way anyone manages a program larger than a page. Chapter 6 called that modularity and put it fifth in the list of characteristics; in practice it is what makes the other four achievable.

The three parts

Every function you use involves three separate things, and confusing them is the commonest source of a compiler error in this topic.

1. The declaration, or prototype. Tells the compiler the name, the return type and the parameter types. Ends in a semicolon; has no body.

double area_of_circle(double radius);

2. The definition. The same header, with the body. This is the function.

double area_of_circle(double radius)
{
    return 3.14159 * radius * radius;
}

3. The call. Uses it.

double a = area_of_circle(5.0);

A function is declared as many times as you like and defined exactly once. Chapter 11 made the same distinction for variables.

A worked example, with the practical's program

MU's Practical 4(a) is the area of a square using a function. Here it is with the three parts labelled.

#include <stdio.h>

double area_of_square(double side);        /* 1. declaration */

int main(void)
{
    double side = 6.5;
    double area = area_of_square(side);    /* 3. call */

    printf("a square of side %.2f has area %.2f\n", side, area);
    printf("and of side 10, area %.2f\n", area_of_square(10.0));
    return 0;
}

double area_of_square(double side)         /* 2. definition */
{
    return side * side;
}
a square of side 6.50 has area 42.25
and of side 10, area 100.00

Move the definition above main and the declaration becomes unnecessary, because the compiler will have read the real thing before it meets the call. Both arrangements are correct. The declaration is the one that keeps working when the program grows past one file, and it is the habit to form.

What happens when a function is called

Worth knowing in order, because it explains both the return value and chapter 40's whole subject.

  1. The arguments are evaluated. In an unspecified order, if there is more than one (chapter 20).
  2. Each argument's value is copied into the matching parameter.
  3. Control transfers to the function's body, which runs.
  4. A return hands a value back and control returns to the caller.
  5. The call expression takes that value.
munotes.in150

What a Function Is: Definition, Call, Arguments, Return

Step 2 is the sentence to hold on to: C passes arguments by value. The function gets a copy. Changing a parameter inside the function cannot change the caller's variable. Chapter 40 is the proof and the way round it.

return

return expression;      /* in a function with a return type */
return;                 /* in a void function */

return does two things at once: it ends the function and it supplies the value. A function may have several return statements, and the first one reached wins.

#include <stdio.h>

int larger(int a, int b)
{
    if (a > b) {
        return a;               /* the function ends here */
    }
    return b;
}

void report(int n)              /* returns nothing */
{
    if (n < 0) {
        printf("%d is negative, nothing to report\n", n);
        return;                 /* an early exit, no value */
    }
    printf("%d has %d digit(s)\n", n, n < 10 ? 1 : n < 100 ? 2 : 3);
}

int main(void)
{
    printf("larger of 12 and 47 is %d\n", larger(12, 47));
    report(-5);
    report(7);
    report(842);
    return 0;
}
larger of 12 and 47 is 47
-5 is negative, nothing to report
7 has 1 digit(s)
842 has 3 digit(s)

The value returned is converted to the function's return type. A function declared int that does return 3.7; returns 3.

Falling off the end of a non-void function without returning anything is undefined behaviour if the caller uses the value. main is the exception: reaching its closing brace returns 0.

void, in both of its places

void greet(void);

The first void is the return type: this function gives nothing back, so its call cannot be used as a value. The second is the parameter list: it takes no arguments.

void greet(); is not the same thing in C17. An empty parameter list means the function takes an unspecified number of arguments, so the compiler checks nothing at the call and greet(1, 2, 3) compiles.

#include <stdio.h>

void takes_nothing(void);
void unspecified();

int main(void)
{
    takes_nothing();
    unspecified();
    unspecified(1, 2, 3);        /* compiles: the empty list checks nothing */
    return 0;
}

void takes_nothing(void)
{
    printf("declared (void): the compiler rejects any argument\n");
}

void unspecified()
{
    printf("declared (): the compiler accepted three arguments it will ignore\n");
}
declared (void): the compiler rejects any argument
declared (): the compiler accepted three arguments it will ignore
declared (): the compiler accepted three arguments it will ignore

Always write (void). The empty list is a survival from the language before the first standard, and C23 has finally redefined it to mean (void). Until every compiler you meet is a C23 one, (void) is the form that gets your calls checked.

munotes.in151

What a Function Is: Definition, Call, Arguments, Return

Parameters and arguments

Two words for two different things, and MU may ask for the distinction.

  • A parameter is the name in the function's definition. It is a variable local to the function.
  • An argument is the value in the call.
double area_of_square(double side)   /* side is a PARAMETER */
...
area_of_square(6.5)                  /* 6.5 is an ARGUMENT */

"Formal parameter" and "actual parameter" are the older names for the same pair, and some examiners use them. Parameter and argument are the standard's own words.

A function used inside an expression

Because a call is an expression, it can go anywhere a value can.

#include <stdio.h>

int square(int n)
{
    return n * n;
}

int main(void)
{
    printf("square(4) + square(3) = %d\n", square(4) + square(3));
    printf("square(square(2))     = %d\n", square(square(2)));
    printf("inside a condition   : %s\n", square(5) > 20 ? "yes" : "no");

    int a[3];
    for (int i = 0; i < 3; i++) {
        a[i] = square(i + 1);
    }
    printf("filled an array      : %d %d %d\n", a[0], a[1], a[2]);
    return 0;
}
square(4) + square(3) = 25
square(square(2))     = 16
inside a condition   : yes
filled an array      : 1 4 9

A void function cannot be used this way, because it has no value. printf("%d", greet()); does not compile.

What this does NOT mean

A declaration is not a definition. One promises the function exists, the other is the function.

A function does not have to return a value. A void function returns nothing, and its call cannot be used as a value.

Arguments are not shared with the caller. They are copied. A function cannot change the caller's variable through an ordinary parameter.

int f() is not int f(void) in C17. The first checks nothing at the call.

A function cannot be defined inside another function. C has no nested function definitions. Declarations may be nested, and definitions may not.

The order of the arguments' evaluation is not left to right. It is unspecified.

return is not only for the end. Several returns are normal and often clearer than one exit with a flag.

Quick revision

  • A function has a name, a return type, a parameter list and a body.
  • Declaration (prototype) ends in a semicolon; definition has the body; the call uses it.
  • Declared many times, defined once.
  • A call is an expression and may appear anywhere a value may.
  • Arguments are evaluated in an unspecified order and copied into the parameters: C passes by value.
  • return expression; ends the function and supplies the value, converted to the return type.
  • return; alone is for a void function.
  • void as a return type means no value; (void) as a parameter list means no arguments.
  • f() in C17 means an unspecified argument list and turns off checking. Write f(void).
  • Parameter is the name in the definition; argument is the value in the call.
  • Functions cannot be nested in C.
munotes.in152

What a Function Is: Definition, Call, Arguments, Return

Test yourself

1. What is the difference between a function declaration and a function definition?

A declaration gives the name, return type and parameter types and ends in a semicolon; it promises the function exists. A definition gives the same and the body; it is the function. One definition, any number of declarations.

2. Why does main usually need a prototype for a function defined below it?

Because the compiler reads the file once from the top, so at the call it must already know the return type and parameter types in order to compile and check the call.

3. What is the difference between void f() and void f(void) in C17?

f(void) takes no arguments and the compiler rejects any. f() says the argument list is unspecified, so the compiler checks nothing and f(1, 2) compiles. C23 makes the two the same.

4. Distinguish parameter from argument.

A parameter is the variable named in the function's definition. An argument is the value supplied in the call and copied into the parameter.

5. Can a function change the caller's variable through an ordinary parameter?

No. The argument's value is copied, so the function works on its own copy. A pointer parameter is the way round it, which is chapter 40.

6. How many return statements may a function have?

Any number. The first one reached ends the function.

7. What does a function declared int return if its body says return 2.9;?

  1. The value is converted to the return type, and conversion to an integer truncates.

What can be asked on this, and how to answer it

"What is a function? Explain its parts with an example." Define it, name the four parts of a definition, and give the three-part example: declaration, definition, call. Say that a function is declared many times and defined once, and that a call is an expression.

"Write a program to find the area of a square using a function." Give this chapter's program with the prototype above main and the definition below it. That layout is what the examiner expects, and it lets you explain why the prototype is there.

"What is a function prototype? Why is it needed?" A declaration giving the name, return type and parameter types, ending in a semicolon. It is needed because the compiler reads the file once and must know the function's interface at the point of the call in order to check and compile it.

munotes.in153

What a Function Is: Definition, Call, Arguments, Return

"Distinguish between actual and formal parameters." The formal parameters are the names in the function's definition, which are local variables of the function. The actual parameters, or arguments, are the values in the call, which are copied into them. Add that C copies, so the function cannot change the caller's variables through them.

"Explain the return statement." It ends the function and supplies the value of the call, converted to the declared return type. return; with no value is for a void function. A function may have several, and the first reached takes effect.

Contents This chapter on its own page

munotes.in154

Chapter Thirty-Three

User-Defined Functions

Syllabus topic 2, "Basics of functions. User defined and Library functions"

In one line

A user-defined function is one you write, and writing a good one is a matter of choosing what goes in, what comes out, and nothing else.

The shape of a definition

return-type name(parameter-list)
{
    declarations and statements
    return expression;
}

Every part is a decision:

  • The return type is what the caller gets. void if nothing.
  • The name says what the function does. A verb for an action, a noun for a value: print_report, area_of_square, is_leap.
  • The parameter list is everything the function needs, each with its type. (void) if nothing.
  • The body is a block, so everything chapter 21 said about blocks applies.

MU's Practical 4(a), and then the same program made general

#include <stdio.h>

double area_of_square(double side)
{
    return side * side;
}

double perimeter_of_square(double side)
{
    return 4.0 * side;
}

int main(void)
{
    double sides[] = {2.5, 6.0, 10.0};

    printf("%8s %10s %10s\n", "side", "area", "perimeter");
    for (int i = 0; i < 3; i++) {
        printf("%8.2f %10.2f %10.2f\n",
               sides[i], area_of_square(sides[i]), perimeter_of_square(sides[i]));
    }
    return 0;
}
    side       area  perimeter
    2.50       6.25      10.00
    6.00      36.00      24.00
   10.00     100.00      40.00

Two functions, each with one job, each returning a value and printing nothing. That last point is the single most useful habit in this chapter: a function that computes should not print. Keep the calculation and the output apart and the function becomes testable, reusable and describable in one line.

Reading input, the way the practical wants it

#include <stdio.h>

double area_of_square(double side)
{
    return side * side;
}

int main(void)
{
    double side;

    printf("Enter the side of the square: ");
    if (scanf("%lf", &side) != 1) {
        printf("\nThat was not a number.\n");
        return 1;
    }
    if (side <= 0) {
        printf("\nA side must be greater than zero.\n");
        return 1;
    }

    printf("\nArea = %.4f\n", area_of_square(side));
    return 0;
}
7.5
Enter the side of the square:
Area = 56.2500

Notice where the validation is: in main, not in area_of_square. The function's job is arithmetic and it should work for any number it is given. Deciding what the user is allowed to type is the caller's job. Mixing the two is how a function becomes untestable.

Call by value, proved

#include <stdio.h>

void try_to_change(int n)
{
    printf("   inside, n starts at %d\n", n);
    n = 999;
    printf("   inside, n is now   %d\n", n);
}

void try_to_swap(int a, int b)
{
    int t = a;
    a = b;
    b = t;
    printf("   inside, a is %d and b is %d\n", a, b);
}

int main(void)
{
    int value = 5;

    printf("before try_to_change, value is %d\n", value);
    try_to_change(value);
    printf("after  try_to_change, value is %d\n", value);

    int x = 10, y = 20;
    printf("\nbefore try_to_swap, x is %d and y is %d\n", x, y);
    try_to_swap(x, y);
    printf("after  try_to_swap, x is %d and y is %d\n", x, y);
    return 0;
}
munotes.in155

User-Defined Functions

before try_to_change, value is 5
   inside, n starts at 5
   inside, n is now   999
after  try_to_change, value is 5

before try_to_swap, x is 10 and y is 20
   inside, a is 20 and b is 10
after  try_to_swap, x is 10 and y is 20

Nothing changed in main. The parameters n, a and b are local variables of their functions, initialised from copies of the arguments. The swap worked perfectly on the copies and the copies were then thrown away.

This is not a defect. It is what makes a function safe to call: you know it cannot alter your variables behind your back. When you do want a function to change the caller's variable, you pass its address, which is chapter 40 and MU's Practical 7.

One exception, and it is the one that confuses everybody. An array argument is not copied. What is passed is the address of its first element, so a function can change the caller's array. Chapters 36 and 41 are that in full.

Returning more than one value, when you cannot

A function returns one value. The three ways round it, in the order you will meet them:

  1. A pointer parameter, so the function writes into the caller's variable. Chapter 40.
  2. A structure, returned by value, holding several members. Chapter 42.
  3. Two functions, each returning one thing, which is often the right answer.
#include <stdio.h>

int quotient(int a, int b)
{
    return a / b;
}

int remainder_of(int a, int b)
{
    return a % b;
}

int main(void)
{
    int a = 47, b = 5;

    printf("%d / %d is %d remainder %d\n",
           a, b, quotient(a, b), remainder_of(a, b));
    return 0;
}
47 / 5 is 9 remainder 2

remainder is the name of a function in <math.h>, so this one is called remainder_of. Shadowing a library name is legal and confusing; -Wall does not warn.

Rules of thumb that hold up

One job per function. If the name needs "and" in it, it is two functions.

Short. A function that does not fit on a screen is hard to check. There is no magic number; the test is whether you can say what it does in one sentence.

Few parameters. Beyond four, a reader cannot keep the order straight and neither can the writer. That is usually a sign the parameters belong together in a structure.

Do not print from a function that computes. Return the value and let the caller decide what to do with it.

Validate at the edge. The function that reads from the user checks the input; the functions that calculate assume it is already sensible.

munotes.in156

User-Defined Functions

Name the return, not the mechanism. is_leap(year) reads better than check_year(year), and a function returning a truth value is best named as a question.

#include <stdio.h>

int is_leap(int year)
{
    return year % 4 == 0 && (year % 100 != 0 || year % 400 == 0);
}

int days_in_month(int month, int year)
{
    int days[] = {0, 31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31};

    if (month < 1 || month > 12) {
        return 0;
    }
    if (month == 2 && is_leap(year)) {
        return 29;
    }
    return days[month];
}

int main(void)
{
    printf("February 1900 has %d days\n", days_in_month(2, 1900));
    printf("February 2024 has %d days\n", days_in_month(2, 2024));
    printf("April    2024 has %d days\n", days_in_month(4, 2024));
    printf("month 13 gives %d, which the caller must treat as an error\n",
           days_in_month(13, 2024));
    return 0;
}
February 1900 has 28 days
February 2024 has 29 days
April    2024 has 30 days
month 13 gives 0, which the caller must treat as an error

days_in_month calls is_leap. That is the point of functions: each one is written once and the next one is built on it. is_leap was chapter 16's expression, now with a name.

What this does NOT mean

A function does not need parameters. (void) is a perfectly good parameter list.

A function does not need a return value. void is a perfectly good return type, and a function that only prints should have it.

A parameter is not the caller's variable. It is a local variable initialised from a copy.

You cannot define a function inside another. C has no nested definitions.

Two functions cannot have the same name. C has no overloading. area(double) and area(int) in one program is a duplicate definition.

A function is not slow. The call costs a few instructions, and a compiler often removes even that. Choosing one long function over four short ones for speed is a bad trade.

Quick revision

  • return-type name(parameters) { body }.
  • (void) for no parameters; void return type for no value.
  • One job per function; name it for what it gives, not how.
  • Keep calculation separate from output: a computing function should not print.
  • Validate input at the edge, in the function that reads it.
  • C passes by value: parameters are local copies, so a swap by value does not swap.
  • An array argument is the exception: the address is passed, so the function can change it.
  • One return value; use a pointer parameter, a structure, or two functions.
  • No nested function definitions, and no two functions with the same name.
munotes.in157

User-Defined Functions

Test yourself

1. Write a function that returns 1 if its argument is even and 0 otherwise.

int is_even(int n) { return n % 2 == 0; }

2. Why does a swap function taking two int parameters not swap the caller's variables?

Because the arguments are copied into the parameters, which are local variables. The function swaps its own copies and they are discarded when it returns.

3. Should a function that calculates an average also print it?

No. Return the average and let the caller print. Separating them makes the function testable and reusable.

4. Can C have two functions called area?

No. C has no overloading; two definitions of one name is an error. Give them different names.

5. Where should input validation go: in the calculating function or in main?

In whichever function reads the input, normally main. The calculating function should work for any value of its declared type.

6. How can a function give the caller two results?

By taking pointers and writing through them, by returning a structure, or by being split into two functions.

7. In days_in_month, why does month 13 return 0 rather than printing an error?

Because the function's job is to give the number of days; reporting to the user is the caller's job. Returning 0, a day count that cannot be real, lets the caller detect it.

What can be asked on this, and how to answer it

"What is a user-defined function? Write a program using one." Define it as a function written by the programmer rather than supplied by the library, give the definition syntax, and give this chapter's area-of-a-square program with a prototype, a definition and a call. Say what each part is.

"Explain call by value with an example." Say the argument's value is copied into the parameter, so the parameter is a local variable and changing it cannot affect the caller. Give the swap that fails, with its output. Then say that passing an address is the way to allow a change, and name it as call by reference.

"What are the advantages of using functions?" Avoiding repetition, giving a piece of work a name, letting a program be understood and tested a piece at a time, allowing reuse in other programs, and making the program shorter to change because a correction is made in one place.

"Write a program using a function to find the number of days in a month." Give days_in_month with is_leap beside it, and point out that one function calls the other, which is the answer to "why functions" in a single example.

"Can a function return more than one value?" Not directly: a return supplies one value. Use pointer parameters, return a structure, or split the work into two functions.

Contents This chapter on its own page

munotes.in158

Chapter Thirty-Four

Library Functions

Syllabus topic 2, "Basics of functions. User defined and Library functions"

In one line

The standard library is a set of functions every C implementation must provide, declared in headers you #include and joined to your program by the linker.

Why they are not part of the language

C the language is small: types, operators, statements, functions. It has no printf, no sqrt and no strlen. Those are functions, written in C, compiled into a library and linked in, and you could write them yourself.

That is not a technicality. It is why the language is small enough to learn in a semester, and it is why printf needs a declaration from a header before you can call it: to the compiler it is an ordinary function that happens to be somebody else's.

The headers you need this semester

HeaderWhat it gives you
<stdio.h>printf, scanf, getchar, putchar, puts, fgets, FILE, EOF, fopen
<stdlib.h>abs, labs, atoi, atof, rand, srand, malloc, free, exit
<string.h>strlen, strcpy, strncpy, strcmp, strcat, strstr, strchr, memset
<math.h>sqrt, pow, fabs, floor, ceil, round, sin, cos, tan, log, exp, fmod
<ctype.h>isdigit, isalpha, isalnum, isspace, isupper, islower, toupper, tolower
<limits.h>INT_MAX, INT_MIN, CHAR_BIT, LONG_MAX
<float.h>DBL_DIG, DBL_MAX, FLT_DIG
<stdbool.h>bool, true, false
<time.h>time, clock, localtime, strftime

The practical: square root and absolute value

MU's Practical 4(c).

#include <stdio.h>
#include <stdlib.h>
#include <math.h>

int main(void)
{
    double x;
    int n;

    printf("Enter a number and a whole number: ");
    if (scanf("%lf %d", &x, &n) != 2) {
        printf("\nThose were not two numbers.\n");
        return 1;
    }

    printf("\n");
    if (x < 0) {
        printf("sqrt(%.4f) is not a real number\n", x);
    } else {
        printf("sqrt(%.4f)  = %.6f\n", x, sqrt(x));
    }
    printf("fabs(%.4f)  = %.4f      <- math.h, takes a double\n", x, fabs(x));
    printf("abs(%d)         = %d           <- stdlib.h, takes an int\n", n, abs(n));
    printf("pow(%.4f, 3) = %.4f\n", x, pow(x, 3.0));
    printf("floor(%.4f) = %.1f, ceil = %.1f, round = %.1f\n",
           x, floor(x), ceil(x), round(x));
    return 0;
}
7.6
-42
Enter a number and a whole number:
sqrt(7.6000)  = 2.756810
fabs(7.6000)  = 7.6000      <- math.h, takes a double
abs(-42)         = 42           <- stdlib.h, takes an int
pow(7.6000, 3) = 438.9760
floor(7.6000) = 7.0, ceil = 8.0, round = 8.0

How to compile that on Linux:

cc -std=c17 -Wall -Wextra prog.c -o prog -lm

The -lm links the maths library, and it must come after the source file. cc -lm prog.c -o prog fails, because the linker works left to right and drops a library nothing has asked for yet. The error is undefined reference to 'sqrt', and it comes from the linker, not the compiler, exactly as chapter 5 said. On macOS the maths functions are in the standard library and no flag is needed, which is why a program that builds on a Mac may not build in the lab.

munotes.in159

Library Functions

abs against fabs, which is a real trap

#include <stdio.h>
#include <stdlib.h>
#include <math.h>

int main(void)
{
    double d = -7.85;

    printf("fabs(-7.85)      = %.4f   <- correct\n", fabs(d));
    printf("abs((int) -7.85) = %d        <- the fraction is gone\n", abs((int) d));
    printf("labs(-3000000000L) = %ld\n", labs(-3000000000L));
    return 0;
}
fabs(-7.85)      = 7.8500   <- correct
abs((int) -7.85) = 7        <- the fraction is gone
labs(-3000000000L) = 3000000000

abs takes an int and returns an int. Passing it a double converts the argument, discarding the fraction before abs ever sees it. Use fabs for a double and labs for a long.

Character classification, which saves writing your own

#include <stdio.h>
#include <ctype.h>

int main(void)
{
    const char *text = "MU B.Sc. IT 2026!";
    int letters = 0, digits = 0, spaces = 0, punct = 0, upper = 0;

    for (int i = 0; text[i] != '\0'; i++) {
        unsigned char c = (unsigned char) text[i];

        if (isalpha(c)) { letters++; }
        if (isdigit(c)) { digits++; }
        if (isspace(c)) { spaces++; }
        if (isupper(c)) { upper++; }
        if (ispunct(c)) { punct++; }
    }
    printf("\"%s\"\n", text);
    printf("letters %d (of which upper case %d), digits %d, spaces %d, punctuation %d\n",
           letters, upper, digits, spaces, punct);

    printf("toupper: ");
    for (int i = 0; text[i] != '\0'; i++) {
        putchar(toupper((unsigned char) text[i]));
    }
    printf("\n");
    return 0;
}
"MU B.Sc. IT 2026!"
letters 7 (of which upper case 6), digits 4, spaces 3, punctuation 3
toupper: MU B.SC. IT 2026!

The (unsigned char) cast is not decoration. Every <ctype.h> function takes an int that must be either a value representable as an unsigned char or EOF. On a machine where plain char is signed, a character above 127 is a negative int, and passing it is undefined behaviour. Casting to unsigned char first is the correct idiom, and it is the reason this book's listings are compiled under both char conventions.

toupper returns an int, and on a character that is not a lower-case letter it returns it unchanged. That is what makes the loop above safe on spaces and punctuation.

Reading the documentation

The most useful skill in this chapter is not a list of functions. It is knowing how to look one up, because nobody remembers argument orders.

On a Linux machine, man 3 sqrt gives the header, the prototype, what it returns and what -lm is needed. The 3 is the section for library functions; without it you may get a shell command of the same name.

munotes.in160

Library Functions

For each function, four things are worth reading before you use it:

  1. Which header declares it.
  2. The prototype: parameter types and the return type.
  3. What it returns on failure. sqrt of a negative gives a not-a-number value; atoi of text that is not a number gives 0, which is indistinguishable from the text "0".
  4. Whether it changes anything you passed it. strcpy writes into its first argument; strlen writes nothing.

Why not to write your own

#include <stdio.h>
#include <math.h>

double my_sqrt(double x)
{
    if (x < 0) {
        return -1.0;
    }
    double guess = x / 2.0;
    for (int i = 0; i < 40; i++) {
        if (guess == 0.0) {
            break;
        }
        guess = (guess + x / guess) / 2.0;
    }
    return guess;
}

int main(void)
{
    double values[] = {2.0, 9.0, 1e10};

    printf("%12s %18s %18s\n", "x", "my_sqrt", "sqrt");
    for (int i = 0; i < 3; i++) {
        printf("%12g %18.10f %18.10f\n",
               values[i], my_sqrt(values[i]), sqrt(values[i]));
    }
    return 0;
}
           x            my_sqrt               sqrt
           2       1.4142135624       1.4142135624
           9       3.0000000000       3.0000000000
       1e+10  100000.0000000000  100000.0000000000

Forty rounds of Newton's method agree with the library to ten decimal places, which is a decent result for twelve lines. The library version is still the one to use: it is correct for every input including the awkward ones, it is as accurate as the hardware allows, and on most machines it is a single processor instruction. Writing your own is an excellent exercise and a poor habit.

What this does NOT mean

Library functions are not part of the language. They are ordinary functions in a compiled library, which is why they need a header and a link step.

A header is not the library. The header declares; the library defines. Chapter 5.

#include <math.h> is not enough on Linux. You also need -lm, after the source file.

abs is not for floating-point values. fabs is.

sqrt of a negative number is not an error you can see. It returns a not-a-number value and sets errno. Test the argument first.

atoi does not report failure. It returns 0 for text that is not a number. strtol reports properly, and is the one to use when the input matters.

A <ctype.h> function does not take a char. It takes an int, and a plain char should be cast to unsigned char first.

Quick revision

  • The library is functions, not language. Header declares, library defines, linker joins.
  • <stdio.h> input and output, <stdlib.h> general utilities, <string.h> strings, <math.h> maths, <ctype.h> characters, <limits.h> and <float.h> limits.
  • Linux: cc prog.c -o prog -lm for <math.h>, and -lm after the source.
  • abs is <stdlib.h> and int; fabs is <math.h> and double; labs for long.
  • floor rounds down, ceil rounds up, round to nearest, a cast to int truncates towards zero.
  • <ctype.h> takes an int: cast a char to unsigned char first.
  • toupper returns non-letters unchanged.
  • man 3 name gives the header, the prototype and the link flag.
  • Check what a function returns on failure before trusting it.
munotes.in161

Library Functions

Test yourself

1. What is the difference between a library function and a user-defined function?

A library function is supplied with the implementation, already compiled, and is declared in a standard header. A user-defined function is written by the programmer in their own source. Both are called the same way.

2. Which header declares sqrt, and what else does a Linux build need?

<math.h>, and the link flag -lm placed after the source file.

3. What is wrong with abs(-7.85)?

abs takes an int, so the argument is converted and the fraction is discarded before abs runs, giving 7 rather than 7.85. Use fabs.

4. Give the four functions that turn 7.6 into a whole number, and their answers.

floor gives 7.0, ceil gives 8.0, round gives 8.0, and (int) gives 7. For -7.6 they give -8.0, -7.0, -8.0 and -7.

5. Why cast to unsigned char before calling isdigit?

Because the function takes an int that must be representable as an unsigned char or be EOF, and on a machine where plain char is signed a character above 127 would be a negative value, which is undefined behaviour.

6. atoi("hello") returns 0. Why is that a problem?

Because atoi("0") also returns 0, so a failure cannot be told from a valid zero. Use strtol, which reports where it stopped parsing.

7. Which header gives INT_MAX?

<limits.h>.

What can be asked on this, and how to answer it

"What are library functions? Name any five with their headers." Define them as functions supplied by the implementation, declared in standard headers and linked in. Then five with headers and one line each: printf from <stdio.h>, strlen from <string.h>, sqrt from <math.h>, abs from <stdlib.h>, toupper from <ctype.h>.

"Write a program to find the square root and the absolute value of a number using library functions." Give this chapter's program, and mention -lm. The abs against fabs distinction is the most likely follow-up.

"Distinguish between library and user-defined functions." Library functions come with the compiler, are already compiled, are declared in standard headers and cannot be changed. User-defined functions are written by the programmer, compiled with the program, and declared by the programmer. Both are called identically, which is the point worth adding.

munotes.in162

Library Functions

"What is the difference between floor, ceil, round and a cast to int?" floor gives the largest integer value not greater than the argument, ceil the smallest not less, round the nearest with halves away from zero, and a cast truncates towards zero. Give 7.6 and -7.6 through all four, because the negative case is where they separate.

"Why does a program using sqrt fail to link with undefined reference to sqrt?" Because <math.h> only declares the function; its compiled code is in the maths library, which must be named to the linker with -lm after the source file. Add that the error comes from the linker rather than the compiler.

Contents This chapter on its own page

munotes.in163

Chapter Thirty-Five

Recursion

Syllabus topic 2, "Basics of functions. User defined and Library functions"

In one line

A recursive function is one that calls itself, and it works only if every call moves towards a base case that does not call again.

The two parts, neither optional

  1. A base case. A value the function answers directly, with no further call.
  2. A recursive case that calls the function with an argument closer to the base case.

Leave out the base case, or fail to move towards it, and the calls go on until the memory set aside for them runs out. That is a stack overflow, and the program is killed.

The practical: factorial

MU's Practical 3(b) again, and 4(b). Chapter 27 did it with a loop; this is the definition written directly.

The mathematical definition is already recursive: 0! = 1, and n! = n * (n-1)!.

#include <stdio.h>

unsigned long long factorial(int n)
{
    if (n <= 1) {
        return 1;                       /* base case: 0! and 1! are both 1 */
    }
    return (unsigned long long) n * factorial(n - 1);   /* recursive case */
}

int main(void)
{
    for (int n = 0; n <= 10; n++) {
        printf("%2d! = %llu\n", n, factorial(n));
    }
    printf("20! = %llu   <- the largest that fits in 64 bits\n", factorial(20));
    return 0;
}
 0! = 1
 1! = 1
 2! = 2
 3! = 6
 4! = 24
 5! = 120
 6! = 720
 7! = 5040
 8! = 40320
 9! = 362880
10! = 3628800
20! = 2432902008176640000   <- the largest that fits in 64 bits

Trace factorial(4) by hand, because this is the standard board question:

factorial(4) = 4 * factorial(3)
                   factorial(3) = 3 * factorial(2)
                                      factorial(2) = 2 * factorial(1)
                                                         factorial(1) = 1
                                      factorial(2) = 2 * 1  = 2
                   factorial(3) = 3 * 2  = 6
factorial(4) = 4 * 6  = 24

Read it in two directions. On the way down, each call suspends itself and waits. On the way back up, each waiting call takes the answer it was waiting for and multiplies. Nothing is computed until the base case is reached.

What is actually happening in memory

Each call needs its own copy of the parameters and local variables, because each call is a separate piece of work. Those copies live on the call stack: a region of memory that grows as calls are made and shrinks as they return.

#include <stdio.h>

int depth = 0;

unsigned long long factorial(int n)
{
    depth++;
    printf("%*scall factorial(%d), depth %d\n", depth * 2, "", n, depth);
    if (n <= 1) {
        printf("%*sbase case, returning 1\n", depth * 2, "");
        depth--;
        return 1;
    }
    unsigned long long result = (unsigned long long) n * factorial(n - 1);
    printf("%*sfactorial(%d) returns %llu\n", depth * 2, "", n, result);
    depth--;
    return result;
}

int main(void)
{
    printf("factorial(4) is %llu\n", factorial(4));
    return 0;
}
munotes.in164

Recursion

  call factorial(4), depth 1
    call factorial(3), depth 2
      call factorial(2), depth 3
        call factorial(1), depth 4
        base case, returning 1
      factorial(2) returns 2
    factorial(3) returns 6
  factorial(4) returns 24
factorial(4) is 24

Five calls are alive at the deepest point, each with its own n. That is the answer to "how does recursion work": the same code, several sets of variables, one per call.

%*s prints an empty string padded to a width given as an argument, which is a tidy way to indent by depth. It is not part of the lesson; it is how the trace was produced.

Recursion against iteration

#include <stdio.h>

unsigned long long fact_loop(int n)
{
    unsigned long long f = 1;
    for (int i = 2; i <= n; i++) {
        f = f * (unsigned long long) i;
    }
    return f;
}

unsigned long long fact_rec(int n)
{
    return n <= 1 ? 1 : (unsigned long long) n * fact_rec(n - 1);
}

int main(void)
{
    for (int n = 5; n <= 20; n += 5) {
        printf("%2d! loop %20llu   recursive %20llu   same? %s\n",
               n, fact_loop(n), fact_rec(n),
               fact_loop(n) == fact_rec(n) ? "yes" : "no");
    }
    return 0;
}
 5! loop                  120   recursive                  120   same? yes
10! loop              3628800   recursive              3628800   same? yes
15! loop        1307674368000   recursive        1307674368000   same? yes
20! loop  2432902008176640000   recursive  2432902008176640000   same? yes
RecursionIteration
Reads likeThe mathematical definitionA procedure
MemoryOne stack frame per callFixed
SpeedA call per stepNo call
RiskStack overflow if too deepAn endless loop
Best forTrees, and definitions that are recursiveEverything else

Anything recursive can be written iteratively and anything iterative can be written recursively. The choice is about which one says what the problem is. Factorial is genuinely recursive in its definition and is still better written as a loop, because a loop expresses "multiply these numbers together" and costs nothing.

The one where recursion is a disaster: Fibonacci

MU sets Fibonacci as Practical 3(c), and chapter 28 wrote it as a loop. The recursive version is the standard example of recursion done badly, and the cost is worth measuring rather than being told.

#include <stdio.h>

long long calls = 0;

long long fib_rec(int n)
{
    calls++;
    if (n < 2) {
        return n;
    }
    return fib_rec(n - 1) + fib_rec(n - 2);
}

long long fib_loop(int n)
{
    long long a = 0, b = 1;
    for (int i = 0; i < n; i++) {
        long long next = a + b;
        a = b;
        b = next;
    }
    return a;
}

int main(void)
{
    printf("%4s %12s %14s %12s\n", "n", "fib", "calls made", "loop passes");
    for (int n = 5; n <= 35; n += 5) {
        calls = 0;
        long long r = fib_rec(n);
        printf("%4d %12lld %14lld %12d\n", n, r, calls, n);
    }
    printf("\nthe loop gives the same answers:\n");
    for (int n = 5; n <= 35; n += 5) {
        printf("fib(%d) = %lld\n", n, fib_loop(n));
    }
    return 0;
}
munotes.in165

Recursion

   n          fib     calls made  loop passes
   5            5             15            5
  10           55            177           10
  15          610           1973           15
  20         6765          21891           20
  25        75025         242785           25
  30       832040        2692537           30
  35      9227465       29860703           35

the loop gives the same answers:
fib(5) = 5
fib(10) = 55
fib(15) = 610
fib(20) = 6765
fib(25) = 75025
fib(30) = 832040
fib(35) = 9227465

Look at the third column. fib_rec(35) made nearly thirty million calls to compute a number the loop reaches in thirty-five additions. The reason is visible in the tree: fib(5) calls fib(4) and fib(3), and fib(4) calls fib(3) again. The same sub-problems are recomputed over and over, and the table shows the cost exactly: the call count is multiplied by about eleven for every five added to n, a factor of roughly 1.6 per step. The loop's cost goes up by one.

This is the honest answer to "is recursion good". Recursion is a tool for problems whose sub-problems do not overlap. Where they do overlap, either use a loop or remember the answers, and remembering them is a later semester's topic.

Where recursion is the right answer

Two examples worth having, both short.

#include <stdio.h>
#include <string.h>

/* Reverse a string in place, one pair of characters per call. */
void reverse(char *s, int left, int right)
{
    if (left >= right) {
        return;                         /* base case: met in the middle */
    }
    char t = s[left];
    s[left] = s[right];
    s[right] = t;
    reverse(s, left + 1, right - 1);
}

/* The greatest common divisor, which IS defined recursively. */
int gcd(int a, int b)
{
    if (b == 0) {
        return a;
    }
    return gcd(b, a % b);
}

/* The sum of the digits of a number. */
int digit_sum(int n)
{
    if (n < 10) {
        return n;
    }
    return n % 10 + digit_sum(n / 10);
}

int main(void)
{
    char word[] = "recursion";
    reverse(word, 0, (int) strlen(word) - 1);
    printf("reversed : %s\n", word);
    printf("gcd(48, 18) = %d\n", gcd(48, 18));
    printf("gcd(270, 192) = %d\n", gcd(270, 192));
    printf("digit_sum(9875) = %d\n", digit_sum(9875));
    return 0;
}
reversed : noisrucer
gcd(48, 18) = 6
gcd(270, 192) = 6
digit_sum(9875) = 29

gcd is the best example in the chapter. Euclid's algorithm is a recursive statement: the gcd of a and b is the gcd of b and a % b, and the gcd of a and 0 is a. The recursive code is the definition, character for character.

munotes.in166

Recursion

What goes wrong

No base case. The calls never stop.

A base case that is never reached. factorial(-1) with a base case of n == 0 recurses towards minus infinity. n <= 1 is the safer test, and checking for a negative argument is better still.

Too deep. Each call uses stack space, and the stack is finite. Recursion depth of a few thousand is usually safe; a few million is not. A loop has no such limit.

Recomputing the same thing, as Fibonacci does.

What this does NOT mean

Recursion is not a loop. It is repeated function calls, each with its own variables.

Recursion is not more powerful than iteration. Each can express what the other can.

Recursion is not slower because of the arithmetic. It is slower because of the calls and, in the Fibonacci case, because of the repeated work.

A recursive function does not need to return a value. reverse above returns nothing.

The base case is not always n == 0. It is whatever value the function can answer without calling itself.

A function calling another function which calls the first is still recursion. That is indirect recursion, and it needs a base case just as much.

Quick revision

  • A recursive function calls itself and needs a base case plus progress towards it.
  • Each call has its own parameters and locals, held on the call stack.
  • Compute nothing on the way down; the answers are built on the way back up.
  • factorial(n): base n <= 1 returns 1, else n * factorial(n - 1).
  • gcd(a, b): base b == 0 returns a, else gcd(b, a % b).
  • Recursion and iteration can each do the other's job; choose the one that states the problem.
  • Naive recursive Fibonacci recomputes sub-problems and the call count is multiplied by about 1.6 per step, so by about eleven every five.
  • Missing or unreachable base case, or too great a depth, means a stack overflow.
  • Recursion suits trees and genuinely recursive definitions; loops suit the rest.

Test yourself

1. What two things must every recursive function have?

A base case that returns without calling itself, and a recursive case whose argument is closer to the base case.

2. Trace factorial(4).

4 factorial(3), which is 4 (3 factorial(2)), which is 4 (3 (2 factorial(1))). factorial(1) is 1, so the value unwinds as 2, then 6, then 24.

3. What happens if the base case is left out?

munotes.in167

Recursion

The function calls itself for ever, the stack runs out of space, and the program is killed. That is a stack overflow.

4. Write gcd recursively.

int gcd(int a, int b) { return b == 0 ? a : gcd(b, a % b); }

5. Why is recursive Fibonacci a bad idea?

Because fib(n-1) and fib(n-2) both recompute the same smaller values, so the number of calls grows roughly by a factor of 1.6 per step. The loop version does n additions.

6. How many calls to factorial are alive at the deepest point of factorial(5)?

Six: the calls for 5, 4, 3, 2, 1 and the one that hits the base case, depending on how the base is written. With a base case of n <= 1, five calls are made and the fifth returns without calling again.

7. Can a recursive function be void?

Yes. The reverse function in this chapter returns nothing; the base case simply returns.

What can be asked on this, and how to answer it

"What is recursion? Write a program to find the factorial of a number using recursion." Define it, name the base case and the recursive case, give the program, and then give the trace of factorial(4). The trace is what shows you understand it rather than remember it.

"Write a program using a recursive function." Any of the four in this chapter answers it. gcd is the best choice, because you can say in one line that Euclid's algorithm is itself recursive, so the code is the definition.

"Distinguish between recursion and iteration." Recursion repeats by calling itself, needs a base case, and uses stack space proportional to the depth. Iteration repeats with a loop, needs a terminating condition, and uses fixed memory. Either can do the other's job; recursion suits recursive definitions and trees, iteration suits everything else.

"What is a stack overflow? How does it arise in recursion?" The call stack is the memory holding one frame per live call, and it is finite. A recursion with no base case, an unreachable base case, or too great a depth exhausts it and the program is killed.

"Explain how recursion works internally." Each call gets its own frame on the call stack holding its parameters and local variables. Calls on the way down are suspended and waiting; when the base case returns, each waiting call resumes, combines the returned value and returns in turn. Give the indented trace.

Contents This chapter on its own page

munotes.in168

Chapter Thirty-Six

Arrays

Syllabus topic 3, "Pointer and Addresses, Pointer and Function Arguments, Pointer and Arrays."

In one line

An array is a fixed number of objects of one type in consecutive memory, reached by an index counting from zero.

Declaring and using one

int marks[5];                       /* five ints, uninitialised */
int marks[5] = {63, 58, 72, 41, 89};
int marks[] = {63, 58, 72, 41, 89}; /* the compiler counts: 5 */

The index runs from 0 to size - 1. marks[5] in an array of five is past the end, and this is where most first-semester bugs live.

#include <stdio.h>

int main(void)
{
    int marks[5] = {63, 58, 72, 41, 89};
    int n = (int) (sizeof marks / sizeof marks[0]);

    printf("the array holds %d element(s)\n", n);
    for (int i = 0; i < n; i++) {
        printf("marks[%d] = %d\n", i, marks[i]);
    }
    printf("the first is marks[0] = %d and the last is marks[%d] = %d\n",
           marks[0], n - 1, marks[n - 1]);
    return 0;
}
the array holds 5 element(s)
marks[0] = 63
marks[1] = 58
marks[2] = 72
marks[3] = 41
marks[4] = 89
the first is marks[0] = 63 and the last is marks[4] = 89

sizeof marks / sizeof marks[0] is the size of the whole array divided by the size of one element, which is the number of elements. Memorise it. It works only where the array itself is in scope; inside a function that received the array as a parameter it does not, which is chapter 41's subject.

Why counting starts at zero

Not a convention chosen to be awkward. The index is an offset from the start: marks[0] is at the start, marks[3] is three elements along. Chapter 41 shows that marks[i] is defined as *(marks + i), and with that definition zero is the only sensible first index.

The practical consequence is the pair of facts to hold together:

  • The first element is a[0].
  • The last element of an array of n is a[n - 1].
  • A loop over it is for (int i = 0; i < n; i++), with < and not <=.

There is no bounds checking

C does not check that an index is inside the array. Reading or writing outside it is undefined behaviour: it may print rubbish, may corrupt another variable, may crash, and may appear to work.

#include <stdio.h>

int main(void)
{
    int a[3] = {10, 20, 30};

    /* Every one of these is undefined behaviour. */
    printf("a[3] is %d\n", a[3]);
    printf("a[100] is %d\n", a[100]);
    a[3] = 99;
    return 0;
}

That program compiles, under -Wall -Wextra, with no warning at all. That is worth sitting with for a moment: the compiler can see that a has three elements and can see the constant 100, and it still says nothing, because checking an index is not its job and the language does not ask it to. Some compilers warn about the obvious cases at higher optimisation levels, and none of them catch the general case, where the index is a variable.

munotes.in169

Arrays

This book does not print what that program produced and will not. The values are not facts about any machine: the standard places no requirement on them, and a different compiler, a different optimisation level or a different day may give something else. The compiler's warning is the whole of what can honestly be shown.

What to do instead: check the index yourself where it could be wrong.

#include <stdio.h>

int at(const int *a, int n, int i, int fallback)
{
    if (i < 0 || i >= n) {
        printf("  index %d is outside 0 to %d\n", i, n - 1);
        return fallback;
    }
    return a[i];
}

int main(void)
{
    int a[3] = {10, 20, 30};

    printf("at(a, 3, 1)  = %d\n", at(a, 3, 1, -1));
    printf("at(a, 3, 5)  = %d\n", at(a, 3, 5, -1));
    printf("at(a, 3, -2) = %d\n", at(a, 3, -2, -1));
    return 0;
}
at(a, 3, 1)  = 20
  index 5 is outside 0 to 2
at(a, 3, 5)  = -1
  index -2 is outside 0 to 2
at(a, 3, -2) = -1

The practical: ten students

MU's Practical 5(a), roll numbers and names of ten students using an array. Names need an array of character arrays, which is the natural place to meet a two-dimensional array of char.

#include <stdio.h>

#define STUDENTS 5
#define NAME_LEN 20

int main(void)
{
    int roll[STUDENTS] = {101, 102, 103, 104, 105};
    char name[STUDENTS][NAME_LEN] = {"Anita Desai", "Rahul Mehta",
                                     "Fatima Shaikh", "Joseph D'Souza",
                                     "Priya Nair"};
    int marks[STUDENTS] = {63, 58, 72, 41, 89};

    printf("%-6s %-16s %6s\n", "Roll", "Name", "Marks");
    for (int i = 0; i < STUDENTS; i++) {
        printf("%-6d %-16s %6d\n", roll[i], name[i], marks[i]);
    }
    return 0;
}
Roll   Name              Marks
101    Anita Desai          63
102    Rahul Mehta          58
103    Fatima Shaikh        72
104    Joseph D'Souza       41
105    Priya Nair           89

MU asks for ten; five are shown so the output fits a page, and STUDENTS is the only thing to change. That is chapter 6's generality: the count is named in one place.

The four things you do to an array

These four patterns cover almost every array question on this paper, and each is worth knowing as a shape rather than as a program.

1. Total and average.

2. Largest and smallest. Start from the first element, not from zero. Starting from zero gives the wrong answer for an array of all-negative values, which is the classic trap.

munotes.in170

Arrays

3. Search. Chapter 30's found_at = -1 idiom.

4. Count how many satisfy something.

#include <stdio.h>

int main(void)
{
    int a[] = {-5, -17, -3, -42, -8};
    int n = (int) (sizeof a / sizeof a[0]);

    int total = 0;
    for (int i = 0; i < n; i++) {
        total += a[i];
    }

    int largest_wrong = 0;              /* the classic mistake */
    int largest = a[0];                 /* correct */
    int smallest = a[0];
    for (int i = 1; i < n; i++) {
        if (a[i] > largest) { largest = a[i]; }
        if (a[i] < smallest) { smallest = a[i]; }
    }
    for (int i = 0; i < n; i++) {
        if (a[i] > largest_wrong) { largest_wrong = a[i]; }
    }

    int target = -3, found_at = -1;
    for (int i = 0; i < n; i++) {
        if (a[i] == target) { found_at = i; break; }
    }

    int below = 0;
    for (int i = 0; i < n; i++) {
        if (a[i] < -10) { below++; }
    }

    printf("total %d, average %.2f\n", total, (double) total / n);
    printf("largest %d (starting from 0 would have said %d)\n",
           largest, largest_wrong);
    printf("smallest %d\n", smallest);
    printf("%d found at index %d\n", target, found_at);
    printf("%d value(s) below -10\n", below);
    return 0;
}
total -75, average -15.00
largest -3 (starting from 0 would have said 0)
smallest -42
-3 found at index 2
2 value(s) below -10

largest_wrong is 0, and 0 is not in the array. That is what starting a maximum at zero does. Start at a[0] and loop from 1.

The practical: sorting

MU's Practical 5(b), sorting into ascending or descending order. Bubble sort is the one to know: compare each adjacent pair and swap them if they are the wrong way round, and repeat until a pass makes no swap.

#include <stdio.h>

void print_array(const char *label, const int *a, int n)
{
    printf("%-12s", label);
    for (int i = 0; i < n; i++) {
        printf("%5d", a[i]);
    }
    printf("\n");
}

void bubble_sort(int *a, int n, int ascending)
{
    for (int pass = 0; pass < n - 1; pass++) {
        int swapped = 0;
        for (int i = 0; i < n - 1 - pass; i++) {
            int wrong_way = ascending ? a[i] > a[i + 1] : a[i] < a[i + 1];
            if (wrong_way) {
                int t = a[i];
                a[i] = a[i + 1];
                a[i + 1] = t;
                swapped = 1;
            }
        }
        printf("  after pass %d: ", pass + 1);
        for (int i = 0; i < n; i++) {
            printf("%5d", a[i]);
        }
        printf("\n");
        if (!swapped) {
            printf("  no swap in that pass, so it is sorted\n");
            break;
        }
    }
}

int main(void)
{
    int a[] = {42, 8, 17, 4, 23, 15};
    int n = (int) (sizeof a / sizeof a[0]);
    int b[6];

    for (int i = 0; i < n; i++) { b[i] = a[i]; }

    print_array("original", a, n);
    printf("ascending:\n");
    bubble_sort(a, n, 1);
    print_array("sorted", a, n);

    printf("descending:\n");
    bubble_sort(b, n, 0);
    print_array("sorted", b, n);
    return 0;
}
munotes.in171

Arrays

original       42    8   17    4   23   15
ascending:
  after pass 1:     8   17    4   23   15   42
  after pass 2:     8    4   17   15   23   42
  after pass 3:     4    8   15   17   23   42
  after pass 4:     4    8   15   17   23   42
  no swap in that pass, so it is sorted
sorted          4    8   15   17   23   42
descending:
  after pass 1:    42   17    8   23   15    4
  after pass 2:    42   17   23   15    8    4
  after pass 3:    42   23   17   15    8    4
  after pass 4:    42   23   17   15    8    4
  no swap in that pass, so it is sorted
sorted         42   23   17   15    8    4

Three details that earn marks.

1. n - 1 - pass. After pass 1 the largest value is at the end, after pass 2 the two largest are, and so on. There is no point comparing them again.

2. The swapped flag. A pass with no swap means the array is in order, so the rest of the passes can be skipped. Without it, the loop always does n - 1 passes.

3. One function, both orders. The ascending parameter chooses the comparison. Writing two nearly identical functions is what chapter 6 called a failure of simplicity.

An array parameter is written int *a here, and int a[] means exactly the same thing in a parameter list. Chapter 41 explains why, and it is also why bubble_sort can change the caller's array while chapter 33's swap could not change two ints.

What this does NOT mean

An array is not a variable that holds many values. It is many objects with one name, and the name is not itself a value you can assign.

You cannot copy an array with =. b = a; does not compile. Copy element by element, or with memcpy.

You cannot compare arrays with ==. That compares addresses.

The size is not part of what a function receives. Pass it as a separate parameter. sizeof inside the function gives the size of a pointer.

An array's size must be known when it is declared, and cannot change afterwards. C99 allows a size from a variable, a variable-length array, which is a different thing and not for a first program.

a[n] is not the last element. a[n - 1] is.

An uninitialised array does not hold zeros. Only one with a partial initialiser does, and a static or global one.

munotes.in172

Arrays

Quick revision

  • An array is a fixed number of same-type objects in consecutive memory.
  • Indexes run 0 to n - 1. Loop with i < n, never i <= n.
  • sizeof a / sizeof a[0] gives the count, where the array itself is in scope.
  • No bounds checking. Out of range is undefined behaviour, not an error.
  • Start a maximum or minimum at a[0] and loop from 1, never at 0.
  • Search idiom: found_at = -1, set and break on a match.
  • Cannot be copied with = or compared with ==.
  • A function receives the address, so it can change the caller's array, and must be told the size.
  • Bubble sort: adjacent compares, n - 1 - pass inner limit, and a swapped flag to stop early.

Test yourself

1. For int a[10];, what are the valid indexes?

0 to 9. a[10] is past the end.

2. How do you find the number of elements of an array?

sizeof a / sizeof a[0], provided the array itself is in scope rather than a parameter that received it.

3. What is wrong with int max = 0; before a loop looking for the largest element?

If every element is negative, the answer comes out as 0, which is not in the array. Start with max = a[0] and loop from index 1.

4. What happens when you read a[100] of a ten-element array?

Undefined behaviour. It may print a meaningless value, corrupt another variable or crash, and it is not an error the language reports.

5. Why can you not write b = a; to copy an array?

An array name is not a modifiable value. Copy element by element, or use memcpy.

6. In bubble sort, why is the inner loop bounded by n - 1 - pass?

Because each pass moves the largest remaining value to its final place at the end, so those positions need not be compared again.

7. How many passes does bubble sort need on an already sorted array with the swapped flag?

One. The first pass makes no swap, so the flag ends the loop.

What can be asked on this, and how to answer it

"What is an array? Explain its declaration and initialisation with examples." Define it, give the three declaration forms, state that indexes run from 0 to n - 1, and give the initialisation rules from chapter 21, including that omitted elements are zero. Add that there is no bounds checking, because that is the fact with consequences.

"Write a program to find the largest and smallest element of an array." Give the program with largest = a[0] and the loop from 1, and say in one line why starting at 0 is wrong. That sentence is worth a mark on its own.

munotes.in173

Arrays

"Write a program to sort an array in ascending order." Give bubble sort with the n - 1 - pass bound and the swapped flag, and show the array after each pass, which is what an examiner asks you to trace.

"Write a program to store and display the roll numbers and names of ten students." Give this chapter's program: an int array for the roll numbers and a two-dimensional char array for the names, printed in a loop with a formatted header.

"What is meant by array out of bounds? What does C do about it?" Using an index outside 0 to n - 1. C does nothing: there is no check, and the behaviour is undefined. The program may print rubbish, damage other data or crash, and checking is the programmer's responsibility.

Contents This chapter on its own page

munotes.in174

Chapter Thirty-Seven

Strings as Arrays, and the String Library

Syllabus topic 3, "Pointer and Addresses, Pointer and Function Arguments, Pointer and Arrays."

In one line

A string is an array of char ending in '\0', and <string.h> is a set of functions that all find the end of a string by looking for that terminator.

What chapter 12 established, in one paragraph

A string is characters in consecutive memory with a zero byte after the last one. There is no stored length. "Anita" occupies six bytes. char s[] = "Anita"; copies it into memory you own and may change; const char *s = "Anita"; points at the literal and must not be changed. strlen counts to the terminator, sizeof gives the whole array.

Walking a string yourself

Every function in <string.h> is a loop you could write. Writing two of them once is the best way to understand the rest.

#include <stdio.h>
#include <string.h>

size_t my_strlen(const char *s)
{
    size_t n = 0;
    while (s[n] != '\0') {
        n++;
    }
    return n;
}

int my_strcmp(const char *a, const char *b)
{
    size_t i = 0;
    while (a[i] != '\0' && a[i] == b[i]) {
        i++;
    }
    return (unsigned char) a[i] - (unsigned char) b[i];
}

int main(void)
{
    const char *words[] = {"apple", "apply", "app", "Apple"};

    for (int i = 0; i < 4; i++) {
        printf("%-8s my_strlen %zu, strlen %zu\n",
               words[i], my_strlen(words[i]), strlen(words[i]));
    }
    printf("\n");
    printf("my_strcmp(apple, apply) = %d, strcmp = %d\n",
           my_strcmp("apple", "apply"), strcmp("apple", "apply"));
    printf("my_strcmp(apply, apple) = %d, strcmp = %d\n",
           my_strcmp("apply", "apple"), strcmp("apply", "apple"));
    printf("my_strcmp(apple, apple) = %d, strcmp = %d\n",
           my_strcmp("apple", "apple"), strcmp("apple", "apple"));
    return 0;
}
apple    my_strlen 5, strlen 5
apply    my_strlen 5, strlen 5
app      my_strlen 3, strlen 3
Apple    my_strlen 5, strlen 5

my_strcmp(apple, apply) = -20, strcmp = -1
my_strcmp(apply, apple) = 20, strcmp = 1
my_strcmp(apple, apple) = 0, strcmp = 0

Look at the last three lines of that output before reading on. my_strcmp returned -20 and 20 where the library returned -1 and 1, and both are correct. The standard fixes only the sign of strcmp's result and whether it is zero, never the magnitude. -20 is the actual difference between 'e' and 'y'; the library's -1 is what its own faster implementation happens to produce. Further down this chapter the same library returns -160 for another pair of strings, for no reason you could predict. That is why the only correct test for equality is strcmp(a, b) == 0 and the only correct test for order is the sign.

Both are three-line loops. That is the whole of the string library's design: the terminator is the contract, and every function walks forward until it meets one.

The cast to unsigned char in the comparison is the standard's own requirement, and it is the same reason chapter 34 casts before calling a <ctype.h> function: on a machine where plain char is signed, a character above 127 would compare as negative.

munotes.in175

Strings as Arrays, and the String Library

The functions you need

FunctionWhat it doesReturns
strlen(s)Counts characters up to '\0'size_t, the length
strcpy(d, s)Copies s into d, terminator includedd
strncpy(d, s, n)Copies at most n charactersd
strcat(d, s)Appends s to the end of dd
strncat(d, s, n)Appends at most n charactersd
strcmp(a, b)Comparesnegative, 0, positive
strncmp(a, b, n)Compares the first n charactersnegative, 0, positive
strchr(s, c)Finds the first ca pointer to it, or NULL
strstr(s, t)Finds the first t inside sa pointer to it, or NULL
strcspn(s, set)How many characters before any of setsize_t

strcmp does not return 1 for "different". It returns a negative number if a sorts before b, zero if they are equal, and a positive number if a sorts after b. The test for equality is strcmp(a, b) == 0, and writing if (strcmp(a, b)) means "if they differ", which reads backwards and is a real source of bugs.

#include <stdio.h>
#include <string.h>

int main(void)
{
    char dest[30] = "";

    strcpy(dest, "Programming");
    printf("after strcpy : \"%s\", length %zu\n", dest, strlen(dest));

    strcat(dest, " with C");
    printf("after strcat : \"%s\", length %zu\n", dest, strlen(dest));

    char *found = strstr(dest, "with");
    printf("strstr found \"with\" at offset %ld\n", (long) (found - dest));

    char *ch = strchr(dest, 'g');
    printf("strchr found the first 'g' at offset %ld\n", (long) (ch - dest));

    printf("strcmp(\"abc\", \"abd\") = %d  (negative: abc sorts first)\n",
           strcmp("abc", "abd"));
    printf("equality is tested as strcmp(a, b) == 0, which gives %d\n",
           strcmp(dest, "Programming with C") == 0);
    return 0;
}
after strcpy : "Programming", length 11
after strcat : "Programming with C", length 18
strstr found "with" at offset 12
strchr found the first 'g' at offset 3
strcmp("abc", "abd") = -1  (negative: abc sorts first)
equality is tested as strcmp(a, b) == 0, which gives 1

The danger, which is the whole reason these functions are taught carefully

strcpy and strcat do not know how big the destination is. They write until they have copied the terminator, and if the destination is too small they write past the end of it. That is undefined behaviour and it is the classic security defect in C programs.

#include <stdio.h>
#include <string.h>

int main(void)
{
    char small[6];

    strcpy(small, "Programming with C");   /* 18 characters into 6 bytes */
    printf("%s\n", small);
    return 0;
}

gcc does catch this one, and says so precisely:

overflow.c: In function ‘main’:
overflow.c:8:5: warning: ‘__builtin_memcpy’ writing 19 bytes into a region of size 6 overflows the destination [-Wstringop-overflow=]
    8 |     strcpy(small, "Programming with C");   /* 18 characters into 6 bytes */
      |     ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
overflow.c:6:10: note: destination object ‘small’ of size 6
    6 |     char small[6];
      |          ^~~~~
munotes.in176

Strings as Arrays, and the String Library

It caught it because both numbers are in the source: it can see that small is six bytes and that the literal is nineteen with its terminator. That is the limit of what a compiler can do here. Make the source string something it cannot work out, a name read from the user for instance, and the same strcpy compiles in silence and corrupts memory when it runs. Compare chapter 36, where a[100] on a three-element array drew no warning at all.

So the warning is worth having and is not a safety net. It is not run here, because what it does is undefined and any result printed would be a claim about a program the standard makes no promise about.

The safe forms:

#include <stdio.h>
#include <string.h>

int main(void)
{
    char source[40];
    char dest[10];

    if (fgets(source, (int) sizeof source, stdin) == NULL) { return 1; }
    source[strcspn(source, "\n")] = '\0';

    /* strncpy: at most 9 characters, then terminate BY HAND */
    strncpy(dest, source, sizeof dest - 1);
    dest[sizeof dest - 1] = '\0';
    printf("strncpy into char[10] : \"%s\" (length %zu)\n", dest, strlen(dest));

    /* snprintf: the one that always terminates, and reports what it wanted */
    char other[10];
    int wanted = snprintf(other, sizeof other, "%s", source);
    printf("snprintf into char[10]: \"%s\" (wanted %d characters)\n",
           other, wanted);
    return 0;
}
Programming with C
strncpy into char[10] : "Programmi" (length 9)
snprintf into char[10]: "Programmi" (wanted 18 characters)

strncpy does not add a terminator if it ran out of room. That is the trap, and most textbooks present strncpy as the safe version without saying so. The dest[sizeof dest - 1] = '\0'; line is not optional. snprintf always terminates and returns the length it would have needed, which is why it is the better choice.

The practical: extracting a substring

MU's Practical 6(a). C has no substring function, so it is a loop or a careful strncpy.

#include <stdio.h>
#include <string.h>

/* Copy `count` characters starting at `start` of `s` into `out`.
   `out` must have room for count + 1. Returns how many were copied. */
int substring(const char *s, int start, int count, char *out)
{
    int len = (int) strlen(s);
    int copied = 0;

    if (start < 0 || start >= len || count <= 0) {
        out[0] = '\0';
        return 0;
    }
    while (copied < count && s[start + copied] != '\0') {
        out[copied] = s[start + copied];
        copied++;
    }
    out[copied] = '\0';
    return copied;
}

int main(void)
{
    const char *text = "Programming with C";
    char part[40];

    int n = substring(text, 0, 11, part);
    printf("from 0, 11 characters : \"%s\" (%d copied)\n", part, n);

    n = substring(text, 12, 4, part);
    printf("from 12, 4 characters : \"%s\" (%d copied)\n", part, n);

    n = substring(text, 12, 100, part);
    printf("from 12, 100 asked for: \"%s\" (%d copied, the string ran out)\n",
           part, n);

    n = substring(text, 50, 4, part);
    printf("from 50, out of range : \"%s\" (%d copied)\n", part, n);
    return 0;
}
munotes.in177

Strings as Arrays, and the String Library

from 0, 11 characters : "Programming" (11 copied)
from 12, 4 characters : "with" (4 copied)
from 12, 100 asked for: "with C" (6 copied, the string ran out)
from 50, out of range : "" (0 copied)

Three things that make this an answer rather than a sketch: the start and the count are both checked, the loop stops at the terminator as well as at the count, and the result is always terminated even when nothing was copied.

The practical: palindrome

MU's Practical 6(b). A palindrome reads the same backwards. Two pointers walk towards each other, which is chapter 28's two-counter for loop.

#include <stdio.h>
#include <string.h>
#include <ctype.h>

int is_palindrome(const char *s)
{
    int left = 0;
    int right = (int) strlen(s) - 1;

    while (left < right) {
        if (s[left] != s[right]) {
            return 0;
        }
        left++;
        right--;
    }
    return 1;
}

/* The version an examiner usually wants next: ignore case and anything
   that is not a letter or a digit. */
int is_palindrome_loose(const char *s)
{
    int left = 0;
    int right = (int) strlen(s) - 1;

    while (left < right) {
        while (left < right && !isalnum((unsigned char) s[left])) { left++; }
        while (left < right && !isalnum((unsigned char) s[right])) { right--; }
        if (tolower((unsigned char) s[left]) != tolower((unsigned char) s[right])) {
            return 0;
        }
        left++;
        right--;
    }
    return 1;
}

int main(void)
{
    const char *tests[] = {"madam", "level", "hello", "a", "",
                           "Madam", "Never odd or even"};

    printf("%-20s %-8s %s\n", "string", "strict", "loose");
    for (int i = 0; i < 7; i++) {
        printf("%-20s %-8s %s\n", tests[i][0] ? tests[i] : "(empty)",
               is_palindrome(tests[i]) ? "yes" : "no",
               is_palindrome_loose(tests[i]) ? "yes" : "no");
    }
    return 0;
}
string               strict   loose
madam                yes      yes
level                yes      yes
hello                no       no
a                    yes      yes
(empty)              yes      yes
Madam                no       yes
Never odd or even    no       yes

"Madam" is not a palindrome strictly, because M and m are different characters. Whether that is the wanted answer depends on the question, and saying so is worth a mark. The empty string and a single character are palindromes under both, because the loop never runs.

The practical: strlen and strcmp

MU's Practical 6(c), with the input read and the results shown.

#include <stdio.h>
#include <string.h>

int main(void)
{
    char a[40], b[40];

    printf("Enter the first string : ");
    if (fgets(a, (int) sizeof a, stdin) == NULL) { return 1; }
    a[strcspn(a, "\n")] = '\0';

    printf("Enter the second string: ");
    if (fgets(b, (int) sizeof b, stdin) == NULL) { return 1; }
    b[strcspn(b, "\n")] = '\0';

    printf("\nstrlen(\"%s\") = %zu\n", a, strlen(a));
    printf("strlen(\"%s\") = %zu\n", b, strlen(b));

    int r = strcmp(a, b);
    printf("strcmp = %d, so ", r);
    if (r == 0) {
        printf("the two strings are equal\n");
    } else if (r < 0) {
        printf("\"%s\" sorts before \"%s\"\n", a, b);
    } else {
        printf("\"%s\" sorts after \"%s\"\n", a, b);
    }
    return 0;
}
munotes.in178

Strings as Arrays, and the String Library

apple
apply
Enter the first string : Enter the second string:
strlen("apple") = 5
strlen("apply") = 5
strcmp = -160, so "apple" sorts before "apply"

Counting things in a string

The other standard question, and it is one loop.

#include <stdio.h>
#include <ctype.h>
#include <string.h>

int main(void)
{
    const char *s = "MU B.Sc. IT Semester 1";
    int vowels = 0, consonants = 0, digits = 0, spaces = 0, others = 0, words = 0;
    int in_word = 0;

    for (int i = 0; s[i] != '\0'; i++) {
        unsigned char c = (unsigned char) s[i];

        if (isspace(c)) {
            spaces++;
            in_word = 0;
        } else {
            if (!in_word) { words++; }
            in_word = 1;
            if (isdigit(c)) {
                digits++;
            } else if (isalpha(c)) {
                char l = (char) tolower(c);
                if (l == 'a' || l == 'e' || l == 'i' || l == 'o' || l == 'u') {
                    vowels++;
                } else {
                    consonants++;
                }
            } else {
                others++;
            }
        }
    }
    printf("\"%s\"\n", s);
    printf("length %zu, words %d\n", strlen(s), words);
    printf("vowels %d, consonants %d, digits %d, spaces %d, other %d\n",
           vowels, consonants, digits, spaces, others);
    return 0;
}
"MU B.Sc. IT Semester 1"
length 22, words 5
vowels 5, consonants 10, digits 1, spaces 4, other 2

The in_word flag is the standard way to count words: a word begins at a non-space character that follows a space or the start of the string.

What this does NOT mean

strcmp does not return 1 or 0. It returns a negative number, zero or a positive number. Test == 0 for equality.

strlen is not sizeof. strlen walks to the terminator at run time; sizeof is the array's size at compile time.

strncpy is not simply "the safe strcpy". It does not terminate if it filled the destination. Terminate by hand, or use snprintf.

strcpy(d, s) does not check anything. The destination must already be big enough, and that is your responsibility.

A string cannot be compared with == or copied with =. Those work on addresses and on nothing respectively.

munotes.in179

Strings as Arrays, and the String Library

strlen does not count the terminator. strlen("Anita") is 5 while the array needs 6 bytes.

gets does not exist. It was removed in C11. Use fgets.

Quick revision

  • A string is a char array ending in '\0'; there is no stored length.
  • strlen counts to the terminator; sizeof gives the array.
  • strcmp(a,b) == 0 means equal. Negative means a sorts first.
  • strcpy and strcat do not know the destination's size; they can write past its end.
  • strncpy may leave the destination unterminated. Terminate it yourself, or use snprintf.
  • Every string function is a loop that stops at the terminator, and my_strlen is three lines.
  • Substring: check the start and the count, copy while both the count and the terminator allow, and always terminate.
  • Palindrome: two indexes walking towards each other; decide whether case and punctuation count.
  • Count words with an in_word flag.
  • Cast a char to unsigned char before any <ctype.h> call.

Test yourself

1. What does strcmp("abc", "abd") return, and what does the sign mean?

A negative number, because 'c' is less than 'd', so "abc" sorts before "abd". The exact magnitude is not specified in a useful way; only the sign and zero matter.

2. For char s[20] = "hello";, what are strlen(s) and sizeof s?

5 and 20.

3. Why is strncpy(d, s, sizeof d) still unsafe?

If s is at least as long as d, strncpy fills d with no terminator, so d is not a string. Copy at most sizeof d - 1 and set the last byte to '\0', or use snprintf.

4. Write a loop that finds the length of a string without strlen.

size_t n = 0;
while (s[n] != '\0') n++;

5. Is "Madam" a palindrome?

Not if case matters, because 'M' and 'm' are different characters. It is if you compare case-insensitively. Say which convention you are using.

6. How do you test two strings for equality?

if (strcmp(a, b) == 0). if (a == b) compares addresses, and if (strcmp(a, b)) is true when they differ.

7. Why must a char be cast to unsigned char before tolower?

Because the function takes an int that must be representable as an unsigned char or be EOF, and on a machine where plain char is signed a byte above 127 would be negative, which is undefined behaviour.

What can be asked on this, and how to answer it

"Explain any five string handling functions with examples." Take strlen, strcpy, strcat, strcmp and strstr, give the prototype, one line of what it does and a worked call with its result. For strcmp give the three possible signs, because that is the one examiners probe.

munotes.in180

Strings as Arrays, and the String Library

"Write a program to find whether a string is a palindrome." Give the two-index loop version. Mention the case and punctuation question and say which convention your program uses; that one sentence is often the difference between full and partial marks.

"Write a program to extract a portion of a string." Give the substring function from this chapter with its bounds checks, and say that C has no substring function so this is what a programmer writes.

"Write a program using strlen and strcmp." Give this chapter's program with fgets, the newline trimmed, and the three-way report on the sign of strcmp.

"What is the difference between strcpy and strncpy?" strcpy copies until the terminator with no regard for the destination's size. strncpy copies at most n characters, and if it reaches the limit it does not add a terminator. Neither is safe without care; snprintf always terminates.

"How would you count the vowels, consonants and words in a string?" One pass with <ctype.h> tests, and an in_word flag for the words. Give the program and say why the flag is needed: a word begins at a non-space that follows a space.

Contents This chapter on its own page

munotes.in181

Chapter Thirty-Eight

Two-Dimensional Arrays and Matrices

Syllabus topic 3, "Pointer and Addresses, Pointer and Function Arguments, Pointer and Arrays."

In one line

A two-dimensional array is an array whose elements are themselves arrays, written a[rows][columns] and stored one whole row after another.

Declaring one

int grid[3][4];                     /* 3 rows of 4 columns: 12 ints */
int grid[3][4] = {{1, 2, 3, 4},
                  {5, 6, 7, 8},
                  {9, 10, 11, 12}};
int grid[][4] = {{1, 2}, {3, 4}};   /* rows counted: 2. Columns MUST be given */
int zeros[3][4] = {0};              /* every element 0 */

The number of columns may never be left out. The compiler needs it to work out where row i begins, and chapter 41 gives the arithmetic. The number of rows can be counted from the initialiser.

The index order is [row][column], and both count from zero.

#include <stdio.h>

#define ROWS 3
#define COLS 4

int main(void)
{
    int grid[ROWS][COLS] = {{1, 2, 3, 4},
                            {5, 6, 7, 8},
                            {9, 10, 11, 12}};

    printf("the whole array is %zu bytes, one row is %zu, one element %zu\n",
           sizeof grid, sizeof grid[0], sizeof grid[0][0]);
    printf("so it holds %zu rows of %zu columns\n",
           sizeof grid / sizeof grid[0],
           sizeof grid[0] / sizeof grid[0][0]);

    for (int r = 0; r < ROWS; r++) {
        for (int c = 0; c < COLS; c++) {
            printf("%4d", grid[r][c]);
        }
        printf("\n");
    }
    printf("grid[1][2] is %d\n", grid[1][2]);
    return 0;
}
the whole array is 48 bytes, one row is 16, one element 4
so it holds 3 rows of 4 columns
   1   2   3   4
   5   6   7   8
   9  10  11  12
grid[1][2] is 7

sizeof grid / sizeof grid[0] gives the rows and sizeof grid[0] / sizeof grid[0][0] gives the columns. Same idea as chapter 36, one level down.

How it is stored

Memory is one-dimensional. A two-dimensional array is stored in row-major order: the whole of row 0, then the whole of row 1, and so on.

#include <stdio.h>

int main(void)
{
    int grid[2][3] = {{10, 20, 30}, {40, 50, 60}};
    int *flat = &grid[0][0];

    printf("reading it as one run of 6 ints: ");
    for (int i = 0; i < 6; i++) {
        printf("%d ", flat[i]);
    }
    printf("\n");
    printf("grid[1][0] is %d, and flat[3] is %d: the same object\n",
           grid[1][0], flat[3]);
    printf("so grid[r][c] is at offset r * 3 + c\n");
    return 0;
}
reading it as one run of 6 ints: 10 20 30 40 50 60
grid[1][0] is 40, and flat[3] is 40: the same object
so grid[r][c] is at offset r * 3 + c

That is the fact behind everything else in the chapter: grid[r][c] lives at offset r * COLS + c from the start. It is why the column count must be given to a function, and why a loop over rows in the outer position and columns in the inner is the one that walks memory in order.

munotes.in182

Two-Dimensional Arrays and Matrices

The practical: reading an m by n matrix

MU's Practical 8(a).

#include <stdio.h>

#define MAX 10

int main(void)
{
    int a[MAX][MAX];
    int m, n;

    printf("Enter the number of rows and columns: ");
    if (scanf("%d %d", &m, &n) != 2) {
        printf("\nThose were not two numbers.\n");
        return 1;
    }
    if (m < 1 || m > MAX || n < 1 || n > MAX) {
        printf("\nRows and columns must be between 1 and %d.\n", MAX);
        return 1;
    }

    printf("Enter %d value(s), row by row: ", m * n);
    for (int r = 0; r < m; r++) {
        for (int c = 0; c < n; c++) {
            if (scanf("%d", &a[r][c]) != 1) {
                printf("\nRan out of numbers at row %d, column %d.\n", r, c);
                return 1;
            }
        }
    }

    printf("\nthe %d by %d matrix is\n", m, n);
    for (int r = 0; r < m; r++) {
        for (int c = 0; c < n; c++) {
            printf("%5d", a[r][c]);
        }
        printf("\n");
    }
    return 0;
}
2 3
1 2 3 4 5 6
Enter the number of rows and columns: Enter 6 value(s), row by row:
the 2 by 3 matrix is
    1    2    3
    4    5    6

int a[MAX][MAX] with m and n used only up to MAX is the standard way to handle a size the user chooses. The array is as big as it could need to be and only part of it is used. The bounds check is what makes that safe, and it is the line an examiner looks for.

Passing a matrix to a function

The parameter must give the column count:

void print_matrix(int a[][COLS], int rows);      /* correct     */
void print_matrix(int a[ROWS][COLS], int rows);  /* also fine   */
void print_matrix(int a[][], int rows);          /* will not compile */

The reason is the offset formula: to find a[r][c] the function must compute r * COLS + c, so it must know COLS. The row count is not needed for the arithmetic, which is why it may be left out, and is passed separately so the function knows when to stop.

The practical: multiplying two matrices using a function

MU's Practical 8(b). Two things have to be right: the shape rule, and the triple loop.

The shape rule. An m by n matrix times a p by q matrix is defined only when n == p, and the result is m by q. Each element of the result is a sum of n products:

c[i][j] = a[i][0]*b[0][j] + a[i][1]*b[1][j] + ... + a[i][n-1]*b[n-1][j]
#include <stdio.h>

#define MAX 10

void print_matrix(const char *label, int a[][MAX], int rows, int cols)
{
    printf("%s (%d by %d)\n", label, rows, cols);
    for (int r = 0; r < rows; r++) {
        for (int c = 0; c < cols; c++) {
            printf("%6d", a[r][c]);
        }
        printf("\n");
    }
}

/* Returns 1 on success, 0 if the shapes do not allow multiplication. */
int multiply(int a[][MAX], int m, int n,
             int b[][MAX], int p, int q,
             int c[][MAX])
{
    if (n != p) {
        return 0;
    }
    for (int i = 0; i < m; i++) {
        for (int j = 0; j < q; j++) {
            int sum = 0;
            for (int k = 0; k < n; k++) {
                sum += a[i][k] * b[k][j];
            }
            c[i][j] = sum;
        }
    }
    return 1;
}

int main(void)
{
    int a[MAX][MAX] = {{1, 2, 3},
                       {4, 5, 6}};
    int b[MAX][MAX] = {{7, 8},
                       {9, 10},
                       {11, 12}};
    int c[MAX][MAX];
    int m = 2, n = 3, p = 3, q = 2;

    print_matrix("A", a, m, n);
    print_matrix("B", b, p, q);

    if (multiply(a, m, n, b, p, q, c)) {
        print_matrix("A times B", c, m, q);
    } else {
        printf("cannot multiply a %d by %d by a %d by %d\n", m, n, p, q);
    }

    /* the same two matrices the other way round: 3 by 2 times 2 by 3 */
    int d[MAX][MAX];
    if (multiply(b, p, q, a, m, n, d)) {
        print_matrix("B times A", d, p, n);
    }

    /* and a pair whose shapes do not allow it */
    if (!multiply(a, m, n, a, m, n, c)) {
        printf("\nA times A is refused: %d columns cannot meet %d rows\n", n, m);
    }
    return 0;
}
munotes.in183

Two-Dimensional Arrays and Matrices

A (2 by 3)
     1     2     3
     4     5     6
B (3 by 2)
     7     8
     9    10
    11    12
A times B (2 by 2)
    58    64
   139   154
B times A (3 by 3)
    39    54    69
    49    68    87
    59    82   105

A times A is refused: 3 columns cannot meet 2 rows

Check one element by hand, because that is what a viva asks. c[0][0] is row 0 of A against column 0 of B:

1*7 + 2*9 + 3*11 = 7 + 18 + 33 = 58

And c[1][1] is row 1 of A against column 1 of B:

4*8 + 5*10 + 6*12 = 32 + 50 + 72 = 154

Three things that make this an answer rather than a sketch.

1. The shape check comes first. Without it the loops read b[k][j] for k beyond p, which is outside the data and is chapter 36's undefined behaviour.

2. sum is a local of the inner pair of loops, set to 0 for each element. A sum declared once outside and not reset is the commonest bug in this program.

munotes.in184

Two-Dimensional Arrays and Matrices

3. A times B and B times A are different matrices, and here they are different shapes as well. Matrix multiplication is not commutative, and printing both makes the point without a paragraph.

The other matrix programs

Addition, transpose and the diagonal sum are the rest of what gets asked, and each is a few lines once the index order is clear.

#include <stdio.h>

#define MAX 10

void print_matrix(const char *label, int a[][MAX], int rows, int cols)
{
    printf("%s\n", label);
    for (int r = 0; r < rows; r++) {
        for (int c = 0; c < cols; c++) {
            printf("%5d", a[r][c]);
        }
        printf("\n");
    }
}

int main(void)
{
    int a[MAX][MAX] = {{1, 2, 3}, {4, 5, 6}, {7, 8, 9}};
    int b[MAX][MAX] = {{9, 8, 7}, {6, 5, 4}, {3, 2, 1}};
    int sum[MAX][MAX], t[MAX][MAX];
    int n = 3;

    for (int r = 0; r < n; r++) {
        for (int c = 0; c < n; c++) {
            sum[r][c] = a[r][c] + b[r][c];
            t[c][r] = a[r][c];              /* transpose: swap the indexes */
        }
    }
    print_matrix("A + B", sum, n, n);
    print_matrix("transpose of A", t, n, n);

    int main_diag = 0, other_diag = 0;
    for (int i = 0; i < n; i++) {
        main_diag += a[i][i];
        other_diag += a[i][n - 1 - i];
    }
    printf("main diagonal of A sums to %d\n", main_diag);
    printf("other diagonal of A sums to %d\n", other_diag);
    return 0;
}
A + B
   10   10   10
   10   10   10
   10   10   10
transpose of A
    1    4    7
    2    5    8
    3    6    9
main diagonal of A sums to 15
other diagonal of A sums to 15

The transpose is one line: t[c][r] = a[r][c]. Read it as "what was at row r, column c goes to row c, column r", and every transpose question is answered.

Addition needs both matrices to be the same shape; multiplication needs the columns of the first to equal the rows of the second. Those are different rules and confusing them is a standing mistake.

Three and more dimensions

int cube[2][3][4];      /* 2 layers of 3 rows of 4 columns: 24 ints */

The same rules apply, with one more index and one more loop. All dimensions but the first must be given in a parameter. You will not need three dimensions this semester.

What this does NOT mean

A two-dimensional array is not an array of pointers. It is one block of memory, rows columns sizeof(element) bytes, laid out row after row. An array of pointers is a different thing.

a[2][3] is not a[2, 3]. C has no comma subscript. a[2, 3] uses the comma operator and means a[3].

munotes.in185

Two-Dimensional Arrays and Matrices

The column count is not optional in a parameter. The offset arithmetic needs it.

The row count is not needed for the arithmetic, which is why int a[][COLS] compiles. It is still needed by the function, so pass it.

Addition and multiplication do not have the same shape rule. Addition needs identical shapes; multiplication needs the inner dimensions to agree.

A times B is not B times A. Matrix multiplication is not commutative, and the two may not even have the same shape.

Quick revision

  • int a[rows][cols] is an array of arrays, stored row after row in row-major order.
  • a[r][c] is at offset r * cols + c.
  • Indexes count from zero; the last row is rows - 1.
  • In a parameter, every dimension except the first must be given: int a[][COLS].
  • Rows: sizeof a / sizeof a[0]. Columns: sizeof a[0] / sizeof a[0][0].
  • Declare a[MAX][MAX] and use only m by n of it when the user chooses the size, with a bounds check.
  • Multiplication: m by n times p by q needs n == p and gives m by q.
  • c[i][j] is the sum over k of a[i][k] * b[k][j], with sum reset for every element.
  • Transpose: t[c][r] = a[r][c].
  • Main diagonal a[i][i]; other diagonal a[i][n - 1 - i].

Test yourself

1. How many ints does int a[4][5]; hold, and how many bytes where int is 4 bytes?

20 ints, 80 bytes.

2. What is the valid range of each index of int a[3][4];?

The first index 0 to 2, the second 0 to 3.

3. Why must the column count appear in a function parameter?

Because the function computes the address of a[r][c] as the start plus r * columns + c, so it cannot find a row without knowing how long a row is.

4. Can a 2 by 3 matrix be multiplied by another 2 by 3 matrix?

No. Multiplication needs the columns of the first to equal the rows of the second, and 3 is not 2. Addition of two 2 by 3 matrices is fine.

5. What shape is the product of a 4 by 2 and a 2 by 7 matrix?

4 by 7.

6. Write the line that transposes a into t.

t[c][r] = a[r][c]; inside loops over r and c.

7. What is the commonest bug in a matrix multiplication program?

Not resetting the running sum to zero for each element of the result, so every element after the first is too large.

munotes.in186

Two-Dimensional Arrays and Matrices

What can be asked on this, and how to answer it

"What is a two-dimensional array? Explain its declaration, initialisation and storage." Define it as an array of arrays, give the declaration and the braced initialiser, say that indexes count from zero, and state row-major storage with the offset formula r * cols + c. The offset formula is the part that shows understanding.

"Write a program to read an m by n matrix and display it." Give this chapter's program with MAX, the bounds check and the nested scanf loop with its return value checked.

"Write a program to multiply two matrices using a function." Give the shape rule first, then the function with the triple loop and the n != p check, then verify one element by hand. The hand check is what an examiner asks for next.

"Write a program to find the transpose of a matrix." One line inside two loops, t[c][r] = a[r][c], and note that an m by n matrix transposes to n by m.

"Distinguish between matrix addition and matrix multiplication in terms of shapes." Addition needs both matrices to have the same number of rows and the same number of columns, and works element by element. Multiplication needs the columns of the first to equal the rows of the second, gives a result with the rows of the first and the columns of the second, and each element is a sum of products.

Contents This chapter on its own page

munotes.in187

Chapter Thirty-Nine

Pointers and Addresses

Syllabus topic 3, "Pointer and Addresses, Pointer and Function Arguments, Pointer and Arrays."

In one line

Every object is somewhere in memory, that somewhere is its address, and a pointer is a variable that holds an address.

Why C has them at the front of the language

Chapter 33 showed a swap function that could not swap, because arguments are copied. That is the first reason: to let a function change something belonging to its caller, you give it the address rather than the value.

There are three more, and together they are most of what C is for.

  • Arrays. a[i] is defined in terms of pointer arithmetic, so an array and a pointer are closely related. Chapter 41.
  • Strings. A string is handled as a pointer to its first character. Chapter 37 used that without naming it.
  • Memory you ask for while the program runs. A later semester's subject, and impossible without pointers.

The two operators

OperatorNameGivenGives
&address-ofan objectits address
*indirection, or dereferencean addressthe object at it

They are opposites. *&x is x.

int x = 42;
int *p = &x;       /* p holds the address of x        */
printf("%d", *p);  /* prints 42: the object p points at */

The in int p is part of the declaration and says "p is a pointer to int". The in p = 7 is the operator and means "the object p points at". Same symbol, two jobs, and telling them apart is most of the difficulty of this chapter.

Read a declaration as a claim about the dereference: int p; says that p is an int. That reading scales to every declaration C can write.

Declaring one

int *p;             /* pointer to int */
double *q;          /* pointer to double */
char *s;            /* pointer to char */
int *a, b;          /* a is a pointer, b is an int */
int *a, *b;         /* both pointers */

The last two lines are why this book writes the next to the name rather than next to the type: int a, b; looks as though both are pointers and neither reading changes what the compiler does.

Seeing it work

#include <stdio.h>

int main(void)
{
    int x = 42;
    int *p = &x;

    printf("x is %d\n", x);
    printf("*p is %d, which is the same object\n", *p);

    *p = 99;                       /* change x through p */
    printf("after *p = 99, x is %d\n", x);

    x = 7;                         /* change x directly */
    printf("after x = 7, *p is %d\n", *p);

    printf("&x and p hold the same address: %d\n", p == &x);
    printf("*&x is %d, and &*p == &x is %d\n", *&x, &*p == &x);
    printf("sizeof x is %zu, sizeof p is %zu\n", sizeof x, sizeof p);
    return 0;
}
munotes.in188

Pointers and Addresses

x is 42
*p is 42, which is the same object
after *p = 99, x is 99
after x = 7, *p is 7
&x and p hold the same address: 1
*&x is 7, and &*p == &x is 1
sizeof x is 4, sizeof p is 8

The two middle blocks are the point: x and *p are two names for one object. Changing either changes both, because there is only one thing there.

Addresses themselves

An address is a number, and on the machine this book is built on it is 8 bytes, which is sizeof p above. Printing one is done with %p and a cast to void *.

A printed address is not reproducible. Modern operating systems place a program at a different base address on every run, deliberately, so the same program prints different numbers each time. What is stable is the difference between the addresses of two objects in the same array, and that is what the next listing shows.

#include <stdio.h>

int main(void)
{
    int a[5] = {10, 20, 30, 40, 50};

    printf("sizeof(int) is %zu\n", sizeof(int));
    printf("the gap between consecutive elements, in bytes:\n");
    for (int i = 1; i < 5; i++) {
        printf("  &a[%d] - &a[%d] = %ld byte(s)\n", i, i - 1,
               (long) ((char *) &a[i] - (char *) &a[i - 1]));
    }
    printf("and in elements: &a[4] - &a[0] = %ld\n", (long) (&a[4] - &a[0]));
    return 0;
}
sizeof(int) is 4
the gap between consecutive elements, in bytes:
  &a[1] - &a[0] = 4 byte(s)
  &a[2] - &a[1] = 4 byte(s)
  &a[3] - &a[2] = 4 byte(s)
  &a[4] - &a[3] = 4 byte(s)
and in elements: &a[4] - &a[0] = 4

Two subtractions, two different answers, and both are right. Subtracting char gives bytes; subtracting int gives elements. That is pointer arithmetic, and chapter 41 is built on it.

For completeness, here is a program that prints an address, and what happened when it was run three times in a row on Ubuntu 24.04 with gcc 13.3:

#include <stdio.h>

int main(void)
{
    int x = 42;
    int *p = &x;

    printf("x is at %p\n", (void *) p);
    return 0;
}
$ cc -std=c17 -Wall -Wextra addresses.c -o addresses
$ ./addresses
x is at 0xffffdb4ba9ac
$ ./addresses
x is at 0xfffff0b91e9c
$ ./addresses
x is at 0xfffff37ed9fc

Three runs of one unchanged program, three different addresses. Nothing is wrong: the operating system places the program somewhere different every time, on purpose, as a security measure. That is exactly why those lines are in a transcript rather than in an output block: this book proves what its programs print, and an address is the one thing it cannot. It is shown so that an address stops being an abstraction, and so that you are not puzzled when your own number differs from your neighbour's.

munotes.in189

Pointers and Addresses

NULL

A pointer must point at something, and sometimes there is nothing to point at. NULL, from <stdio.h> and several other headers, is the value that means "points at nothing".

#include <stdio.h>
#include <string.h>

int main(void)
{
    int *p = NULL;

    printf("p == NULL is %d\n", p == NULL);
    printf("as a condition, a null pointer is false: %d\n", p ? 1 : 0);

    const char *text = "Programming with C";
    char *found = strchr(text, 'z');
    if (found == NULL) {
        printf("strchr found no 'z', and said so by returning NULL\n");
    }
    found = strchr(text, 'w');
    if (found != NULL) {
        printf("strchr found 'w' at offset %ld\n", (long) (found - text));
    }
    return 0;
}
p == NULL is 1
as a condition, a null pointer is false: 0
strchr found no 'z', and said so by returning NULL
strchr found 'w' at offset 12

Dereferencing a null pointer is undefined behaviour and in practice kills the program. Every library function that returns a pointer and can fail returns NULL on failure, so the test is not optional:

char *found = strstr(text, key);
if (found != NULL) {
    ...
}

A null pointer is false as a condition and any other pointer is true, so if (p) and if (p != NULL) are the same test. Write whichever you find clearer; this book writes the second in new code because it says what is being tested.

const and pointers

Two different things can be constant, and the difference matters as soon as you write a function.

const int *p;       /* p may be changed; *p may not */
int *const p;       /* p may not be changed; *p may */
const int *const p; /* neither */

Read it right to left: const int p is "p is a pointer to a const int". int const p means the same as const int *p; both spellings are in use.

#include <stdio.h>

void print_all(const int *a, int n)     /* promises not to change the array */
{
    for (int i = 0; i < n; i++) {
        printf("%d ", a[i]);
    }
    printf("\n");
}

void double_all(int *a, int n)          /* is allowed to change it */
{
    for (int i = 0; i < n; i++) {
        a[i] = a[i] * 2;
    }
}

int main(void)
{
    int a[4] = {1, 2, 3, 4};

    print_all(a, 4);
    double_all(a, 4);
    print_all(a, 4);
    return 0;
}
1 2 3 4
2 4 6 8

const int * in a parameter is a promise to the caller, checked by the compiler. Use it on every pointer parameter a function does not write through. It costs nothing and it tells a reader which arguments can come back changed.

munotes.in190

Pointers and Addresses

What goes wrong

An uninitialised pointer. int p; p = 5; writes to whatever address happened to be in p. Chapter 11's undefined behaviour, now able to damage anything. Set a pointer when you declare it, to a real address or to NULL.

Dereferencing NULL. Test before use.

A pointer to something that has gone. Returning the address of a local variable gives a pointer to memory that no longer belongs to anybody.

#include <stdio.h>

int *broken(void)
{
    int local = 42;
    return &local;              /* the object dies when the function returns */
}

int main(void)
{
    int *p = broken();
    printf("p is not NULL: %d, and reading *p is undefined\n", p != NULL);
    return 0;
}
dangling.c: In function ‘broken’:
dangling.c:6:12: warning: function returns address of local variable [-Wreturn-local-addr]
    6 |     return &local;              /* the object dies when the function returns */
      |            ^~~~~~

The compiler catches this one, and the program is not run here: *p refers to an object whose lifetime ended. That is a dangling pointer, and the compiler's warning is the whole of what can honestly be shown.

What this does NOT mean

A pointer is not an integer. It holds an address, its arithmetic counts in elements rather than in bytes, and converting between the two is not portable.

in a declaration is not the dereference operator. In int p it declares a pointer; in *p = 5 it dereferences one.

& is not the bitwise AND here. As a unary operator it takes an address; as a binary operator it is bitwise AND.

NULL is not zero the number, exactly. It is a null pointer constant. Writing p = 0; is legal and means the same; NULL says what you meant.

A pointer does not have to be initialised, and must be. An uninitialised pointer holds an arbitrary address.

const int *p does not make p constant. It makes what p points at read-only through p.

An address printed once is not a fact about your machine. It changes between runs.

Quick revision

  • An address is where an object is; a pointer is a variable holding an address.
  • &x gives the address of x; *p gives the object p points at. They are opposites.
  • int p; declares p such that p is an int. Read the declaration as a claim about the dereference.
  • int *a, b; declares one pointer and one int.
  • x and *p are two names for one object when p is &x.
  • A pointer is 8 bytes on a 64-bit machine, and %p needs a cast to void *.
  • Subtracting two int gives a count of elements; casting to char first gives bytes.
  • NULL means "points at nothing". Dereferencing it is undefined behaviour.
  • Every library function returning a pointer that can fail returns NULL; test it.
  • const int *p promises not to write through p. Use it on read-only parameters.
  • Never return the address of a local variable.
munotes.in191

Pointers and Addresses

Test yourself

1. What do & and * do?

&x yields the address of x. *p yields the object at the address in p. Each undoes the other.

2. In int p = &x;, what is the type of p and what is the type of p?

p is int , a pointer to int. p is an int.

3. What does int *a, b; declare?

a as a pointer to int and b as an int. The * applies to one declarator only.

4. After int x = 5; int p = &x; p = 9; what is x?

  1. x and *p are the same object.

5. Why must the result of strstr be tested against NULL?

Because strstr returns NULL when it does not find the text, and dereferencing a null pointer is undefined behaviour and normally kills the program.

6. What is the difference between const int p and int const p?

The first says the object may not be written through p, while p itself may be pointed elsewhere. The second says p may not be pointed elsewhere, while the object may be changed.

7. Why is returning &local from a function wrong?

The local object's lifetime ends when the function returns, so the returned address refers to memory that no longer holds it. That is a dangling pointer and reading through it is undefined behaviour.

What can be asked on this, and how to answer it

"What is a pointer? Explain with an example." Define it as a variable holding the address of another object. Give the declaration, & and , and a short program showing that x and p are the same object. Say that the in the declaration and the in the dereference are different uses of the symbol, because that is what confuses most answers.

"Explain the address-of and indirection operators." & applied to an object gives its address; applied to a pointer gives the object. They are inverse, so &x is x. Note that & cannot be applied to a constant or an expression, only to an object.

"What is a null pointer? What is its use?" A pointer holding NULL, which points at no object. It is used as a "nothing here" value, particularly as the failure return of library functions that give back pointers. Dereferencing it is undefined behaviour, so it must be tested first.

munotes.in192

Pointers and Addresses

"What is a dangling pointer?" A pointer to an object whose lifetime has ended: the classic case is the address of a local variable returned from a function. Reading or writing through it is undefined behaviour.

"Distinguish between const int p and int const p." Give the right-to-left reading and one sentence each, as in Test yourself question 6, and add why the first is used on function parameters.

Contents This chapter on its own page

munotes.in193

Chapter Forty

Pointers as Function Arguments: Call by Value and Call by Reference

Syllabus topic 3, "Pointer and Addresses, Pointer and Function Arguments, Pointer and Arrays."

In one line

To let a function change one of your variables, pass its address, and the function changes the object through the pointer.

The problem, restated

Chapter 33 proved that this cannot work:

void try_to_swap(int a, int b)
{
    int t = a; a = b; b = t;
}

a and b are local variables of try_to_swap, initialised from copies. The swap succeeds on the copies, which are then discarded.

The fix is not a different kind of parameter. It is a different kind of argument: instead of the value, pass where the value lives.

The practical: swap both ways

MU's Practical 7, with both methods in one program so the difference is visible in the output.

#include <stdio.h>

/* Call by value: the function gets copies. */
void swap_by_value(int a, int b)
{
    int t = a;
    a = b;
    b = t;
    printf("   inside swap_by_value : a is %d, b is %d\n", a, b);
}

/* Call by reference, as MU names it: the function gets addresses. */
void swap_by_reference(int *a, int *b)
{
    int t = *a;
    *a = *b;
    *b = t;
    printf("   inside swap_by_reference: *a is %d, *b is %d\n", *a, *b);
}

int main(void)
{
    int x = 10, y = 20;

    printf("before swap_by_value     : x is %d, y is %d\n", x, y);
    swap_by_value(x, y);
    printf("after  swap_by_value     : x is %d, y is %d   <- unchanged\n", x, y);

    printf("\nbefore swap_by_reference : x is %d, y is %d\n", x, y);
    swap_by_reference(&x, &y);
    printf("after  swap_by_reference : x is %d, y is %d   <- swapped\n", x, y);
    return 0;
}
before swap_by_value     : x is 10, y is 20
   inside swap_by_value : a is 20, b is 10
after  swap_by_value     : x is 10, y is 20   <- unchanged

before swap_by_reference : x is 10, y is 20
   inside swap_by_reference: *a is 20, *b is 10
after  swap_by_reference : x is 20, y is 10   <- swapped

Four lines of that program are the whole topic, and each is worth naming.

In the programWhat it is
void swap_by_reference(int a, int b)The parameters are pointers to int
swap_by_reference(&x, &y)The arguments are the addresses of x and y
int t = *a;Read the object a points at
a = b;Write the object a points at

Forget the & at the call and the compiler stops you, because an int is not an int . Forget a inside the function and you swap the two pointers instead of the two objects, which compiles and does nothing useful. That second mistake is the one to watch for.

#include <stdio.h>

void broken_swap(int *a, int *b)
{
    int *t = a;
    a = b;                  /* swaps the POINTERS, which are local copies */
    b = t;
    printf("   inside, *a is %d and *b is %d\n", *a, *b);
}

int main(void)
{
    int x = 10, y = 20;

    printf("before: x is %d, y is %d\n", x, y);
    broken_swap(&x, &y);
    printf("after : x is %d, y is %d   <- still unchanged\n", x, y);
    return 0;
}
munotes.in194

Pointers as Function Arguments: Call by Value and Call by Reference

before: x is 10, y is 20
   inside, *a is 20 and *b is 10
after : x is 10, y is 20   <- still unchanged

Note what did not happen: -Wall -Wextra said nothing at all about that function. It is a perfectly well formed program that assigns to two local variables, and there is nothing for a compiler to object to. Compare the listings in chapters 23, 24 and 26, where gcc does catch the mistake. The only way to catch this one is to read it.

The pointers themselves are parameters, and parameters are copies. Swapping them swaps two local variables, exactly as chapter 33's version swapped two local ints. The * is what reaches through to the caller's object.

Why "call by reference" is a borrowed name

C has one parameter-passing mechanism: the argument's value is copied into the parameter. When the argument is &x, the value copied is an address, and the function can then reach x through it.

So:

ArgumentParameter holdsCan change the caller's variable
Call by valuexa copy of x's valueNo
Call by reference, so called&xa copy of x's addressYes, through *

Both rows are call by value. The second passes a value that happens to be an address. Languages with genuine call by reference, such as C++ with int &a, let you write a = b in the function and change the caller's variable with no & at the call and no inside. C does not, and the & and you have to write are the visible evidence.

MU's paper uses "call by value" and "call by reference", so use her terms in an answer. Add the sentence "in C this is done by passing a pointer by value", which is correct and shows you know what the mechanism is.

The other reason to pass a pointer: more than one result

Chapter 33 said a function returns one value. A pointer parameter is the usual way round it.

#include <stdio.h>

/* Returns 1 on success, and writes the results through the pointers. */
int divide(int a, int b, int *quotient, int *remainder)
{
    if (b == 0) {
        return 0;
    }
    *quotient = a / b;
    *remainder = a % b;
    return 1;
}

void min_max(const int *a, int n, int *smallest, int *largest)
{
    *smallest = a[0];
    *largest = a[0];
    for (int i = 1; i < n; i++) {
        if (a[i] < *smallest) { *smallest = a[i]; }
        if (a[i] > *largest)  { *largest = a[i]; }
    }
}

int main(void)
{
    int q, r;

    if (divide(47, 5, &q, &r)) {
        printf("47 / 5 is %d remainder %d\n", q, r);
    }
    if (!divide(47, 0, &q, &r)) {
        printf("47 / 0 was refused, and q and r were left alone\n");
    }

    int a[] = {42, 8, 17, 4, 23};
    int lo, hi;
    min_max(a, 5, &lo, &hi);
    printf("smallest %d, largest %d\n", lo, hi);
    return 0;
}
munotes.in195

Pointers as Function Arguments: Call by Value and Call by Reference

47 / 5 is 9 remainder 2
47 / 0 was refused, and q and r were left alone
smallest 4, largest 42

That is the standard C shape for a function that can fail and also produce results: the return value says whether it worked, and the results come back through pointers. You have already used it: scanf returns how many items it read and puts the values where its pointer arguments say.

Which is why scanf takes &

scanf("%d", &n);

Now it is obvious. scanf has to put a value into your variable, and C copies arguments, so it cannot be given n: it would receive a copy and fill that in. It is given &n, the address, and writes through it.

And the exception you have already met: scanf("%s", name) for a character array takes no &, because an array argument is already an address. Chapter 41 is why.

Not writing through a pointer

Chapter 39 introduced const on a pointer parameter. It matters most here, where a reader has to know which arguments can come back changed.

void print_all(const int *a, int n);   /* will not change your array */
void double_all(int *a, int n);        /* will */
int  min_max(const int *a, int n, int *lo, int *hi);  /* reads a, writes lo and hi */

Read a function's parameter list and you can tell what it may do. That is worth more than any comment, because the compiler enforces it.

What goes wrong

A missing & at the call. The compiler catches it.

A missing * in the function. It may not catch it, as broken_swap showed.

Passing the address of something that has gone. Chapter 39's dangling pointer.

Not checking a pointer parameter for NULL. A function that dereferences a pointer it was given should either document that the pointer must not be null or test it.

#include <stdio.h>

int safe_double(int *n)
{
    if (n == NULL) {
        return 0;
    }
    *n = *n * 2;
    return 1;
}

int main(void)
{
    int x = 21;

    printf("safe_double(&x) returned %d and x is %d\n", safe_double(&x), x);
    printf("safe_double(NULL) returned %d, and nothing was touched\n",
           safe_double(NULL));
    return 0;
}
munotes.in196

Pointers as Function Arguments: Call by Value and Call by Reference

safe_double(&x) returned 1 and x is 42
safe_double(NULL) returned 0, and nothing was touched

What this does NOT mean

C does not have two parameter-passing modes. It has one, by value. Passing an address is a use of it.

A pointer parameter is not the caller's variable. It is a copy of the address, which is why swapping two pointer parameters changes nothing outside.

& at the call is not optional for a scalar. Without it the types do not match.

An array argument does not need &. It already yields an address.

Passing a pointer is not "faster". For an int it is the same cost or slightly more. For a large structure it is genuinely cheaper, which is the other reason to do it.

const on a pointer parameter is not a comment. The compiler refuses a write through it.

Quick revision

  • C passes every argument by value. Always.
  • To let a function change your variable, pass &x and use *p inside.
  • Parameters are int a; the call is f(&x); the body reads and writes a.
  • Swapping the pointers instead of the objects compiles and does nothing: the * is what matters.
  • "Call by reference" in C means passing a pointer by value. Say so, and use MU's term.
  • A function that must give back more than one result takes pointers for the extras, and returns whether it worked.
  • scanf takes &n for the same reason, and no & for an array because an array is already an address.
  • const int *a in a parameter list promises not to write through it, and the compiler enforces it.
  • Test a pointer parameter against NULL if the caller could reasonably pass one.

Test yourself

1. Why can a function taking two int parameters not swap the caller's variables?

Because the arguments are copied into the parameters, which are local variables. The function swaps its own copies and they are discarded when it returns.

2. Write a swap function that works, and its call.

void swap(int *a, int *b) { int t = *a; *a = *b; *b = t; }
swap(&x, &y);

3. Does C have call by reference?

No. C has call by value only. What is called call by reference in C is passing a pointer, which is itself passed by value; the & at the call and the * in the function are the evidence.

4. What does this function do to the caller's variables?

void f(int *a, int *b) { int *t = a; a = b; b = t; }
munotes.in197

Pointers as Function Arguments: Call by Value and Call by Reference

Nothing. It swaps two local pointer variables. The objects they point at are untouched.

5. Why does scanf need &n but not &name for a character array?

scanf must write into your object, and C copies arguments, so it needs the address. An array used as an argument already yields the address of its first element, so no & is needed or wanted.

6. How does a function return two results?

By taking a pointer for each extra result and writing through them, usually with the return value reporting success. Returning a structure is the other way.

7. What does const add to void print_all(const int *a, int n)?

A compiler-checked promise that the function will not write through a, so the caller knows the array comes back unchanged.

What can be asked on this, and how to answer it

"Write a program to swap two numbers using call by value and call by reference." Give this chapter's program with both functions and the output showing that one worked and the other did not. Then add the sentence that earns the extra mark: C has only call by value, and the second method passes a pointer by value.

"Distinguish between call by value and call by reference." Give the table: the argument is the value against the address, the parameter holds a copy of the value against a copy of the address, and the function cannot against can change the caller's variable. Note that in C both are call by value, and that a genuine reference parameter exists in C++ and not in C.

"Why does scanf require the address of a variable?" Because arguments are copied, so scanf given the value could only fill in its own copy. Given the address it writes into the caller's object. Add that an array argument already yields an address, so %s needs no &.

"How can a function return more than one value?" By taking pointers to the caller's variables and writing the extra results through them, or by returning a structure. Give the divide function with its quotient, remainder and success return.

"What is the output of this program?" with a swap that swaps pointers rather than objects. Say that the parameters are local copies of the addresses, so swapping them changes nothing in the caller, and give the unchanged values.

Contents This chapter on its own page

munotes.in198

Chapter Forty-One

Pointers and Arrays

Syllabus topic 3, "Pointer and Addresses, Pointer and Function Arguments, Pointer and Arrays."

In one line

In almost every expression an array's name yields the address of its first element, and a[i] is defined as *(a + i), which is why an array parameter is a pointer and why a function cannot find out how long an array is.

Pointer arithmetic

Adding an integer to a pointer moves it by that many elements, not bytes. The compiler knows the size of the type and does the multiplication.

#include <stdio.h>

int main(void)
{
    int a[5] = {10, 20, 30, 40, 50};
    int *p = &a[0];

    printf("p points at %d\n", *p);
    printf("*(p + 1) is %d\n", *(p + 1));
    printf("*(p + 3) is %d\n", *(p + 3));

    p++;                            /* moves forward by one int, 4 bytes */
    printf("after p++, *p is %d\n", *p);

    p += 2;
    printf("after p += 2, *p is %d\n", *p);

    int *q = &a[4];
    printf("q - p is %ld element(s)\n", (long) (q - p));
    printf("in bytes that is %ld\n", (long) ((char *) q - (char *) p));
    return 0;
}
p points at 10
*(p + 1) is 20
*(p + 3) is 40
after p++, *p is 20
after p += 2, *p is 40
q - p is 1 element(s)
in bytes that is 4

p++ moved four bytes because p is an int . On a double it would move eight. That is the whole of pointer arithmetic: the unit is the type.

The identity

The language defines the subscript operator in terms of pointer arithmetic:

a[i]  is  *(a + i)

and addition is commutative, so a[i], (a + i), (i + a) and even i[a] all mean the same thing.

#include <stdio.h>

int main(void)
{
    int a[5] = {10, 20, 30, 40, 50};

    printf("a[2]      = %d\n", a[2]);
    printf("*(a + 2)  = %d\n", *(a + 2));
    printf("*(2 + a)  = %d\n", *(2 + a));
    printf("2[a]      = %d   <- legal, and never write it\n", 2[a]);
    printf("\nwalking with an index and with a pointer:\n");
    for (int i = 0; i < 5; i++) {
        printf("  a[%d] = %2d   *(a + %d) = %2d\n", i, a[i], i, *(a + i));
    }
    return 0;
}
a[2]      = 30
*(a + 2)  = 30
*(2 + a)  = 30
2[a]      = 30   <- legal, and never write it

walking with an index and with a pointer:
  a[0] = 10   *(a + 0) = 10
  a[1] = 20   *(a + 1) = 20
  a[2] = 30   *(a + 2) = 30
  a[3] = 40   *(a + 3) = 40
  a[4] = 50   *(a + 4) = 50

2[a] is a curiosity and a favourite trick question. It is legal because it means *(2 + a), and you should never write it.

munotes.in199

Pointers and Arrays

The identity also explains why indexes start at zero: a[0] is *(a + 0), which is the first element with no offset at all.

Array decay, and the three exceptions

In almost every expression, an array's name is converted to a pointer to its first element. That conversion is called decay, and it is why a can be passed to a function expecting int *.

The three places it does not happen:

  1. sizeof a, which gives the size of the whole array.
  2. &a, which gives the address of the whole array, of type "pointer to array of 5 int".
  3. A string literal initialising a character array, char s[] = "Anita";, which copies the characters.
#include <stdio.h>

void by_parameter(int a[5])            /* looks like an array: it is a pointer */
{
    printf("   inside, sizeof a is %zu  <- the size of a POINTER\n", sizeof a);
    printf("   inside, so sizeof a / sizeof a[0] is %zu, which is WRONG\n",
           sizeof a / sizeof a[0]);
}

int main(void)
{
    int a[5] = {10, 20, 30, 40, 50};

    printf("in main, sizeof a is %zu  <- the whole array\n", sizeof a);
    printf("in main, sizeof a / sizeof a[0] is %zu, which is right\n",
           sizeof a / sizeof a[0]);
    printf("sizeof &a[0] is %zu, sizeof &a is %zu: both pointers\n",
           sizeof &a[0], sizeof &a);
    by_parameter(a);
    return 0;
}

gcc names this trap explicitly, which is worth reading twice:

decay.c: In function ‘by_parameter’:
decay.c:5:77: warning: ‘sizeof’ on array function parameter ‘a’ will return size of ‘int *’ [-Wsizeof-array-argument]
    5 |     printf("   inside, sizeof a is %zu  <- the size of a POINTER\n", sizeof a);
      |                                                                             ^
decay.c:3:23: note: declared here
    3 | void by_parameter(int a[5])            /* looks like an array: it is a pointer */
      |                   ~~~~^~~~
decay.c:7:19: warning: ‘sizeof’ on array function parameter ‘a’ will return size of ‘int *’ [-Wsizeof-array-argument]
    7 |            sizeof a / sizeof a[0]);
      |                   ^
decay.c:3:23: note: declared here
    3 | void by_parameter(int a[5])            /* looks like an array: it is a pointer */
      |                   ~~~~^~~~
in main, sizeof a is 20  <- the whole array
in main, sizeof a / sizeof a[0] is 5, which is right
sizeof &a[0] is 8, sizeof &a is 8: both pointers
   inside, sizeof a is 8  <- the size of a POINTER
   inside, so sizeof a / sizeof a[0] is 2, which is WRONG

That is the answer to "why must I pass the length". Inside the function the parameter is a pointer, so sizeof gives 8 and the division gives 2, which is not the number of elements and is not an error the compiler reports. Chapter 36 said the sizeof trick works only where the array itself is in scope; this is the proof.

munotes.in200

Pointers and Arrays

int a[5], int a[] and int a are the same parameter declaration. The 5 is documentation and the compiler ignores it. Write int a when you want a reader to know what it really is, and int a[] when you want them to know it is used as an array.

Where an array and a pointer genuinely differ

#include <stdio.h>

int main(void)
{
    int a[5] = {10, 20, 30, 40, 50};
    int *p = a;                     /* p points at a[0]; no & needed */

    printf("a[1] is %d and p[1] is %d: subscripting works on both\n", a[1], p[1]);

    p++;                            /* a pointer can be moved */
    printf("after p++, p[0] is %d\n", p[0]);
    p = a;                          /* and pointed somewhere else */

    /* a++ would not compile: an array name is not a modifiable value */
    printf("sizeof a is %zu, sizeof p is %zu\n", sizeof a, sizeof p);
    printf("&a[0] == a is %d, and p == a is %d\n", &a[0] == a, p == a);
    return 0;
}
a[1] is 20 and p[1] is 20: subscripting works on both
after p++, p[0] is 20
sizeof a is 20, sizeof p is 8
&a[0] == a is 1, and p == a is 1
Array int a[5]Pointer int *p
What it is5 objects1 object holding an address
sizeof20 here8 here
Can be assignedNoYes
Can be incrementedNoYes
SubscriptingYesYes
Decays to a pointerYes, in most expressionsAlready one
Memory it namesIts ownSomebody else's

The last row is the one to remember. An array owns its memory; a pointer borrows somebody's.

Strings, finally explained

Chapter 37's whole library now makes sense. A string function takes a char *, and every call you made passed an array that decayed to one.

#include <stdio.h>
#include <string.h>

int main(void)
{
    char word[] = "Anita";
    char *p = word;

    printf("the string is %s\n", word);
    printf("through a pointer: %s\n", p);
    printf("word[0] is %c and *p is %c\n", word[0], *p);
    printf("the 4th character: word[3] is %c, *(p + 3) is %c\n",
           word[3], *(p + 3));

    printf("walking to the terminator with a pointer: ");
    for (char *q = word; *q != '\0'; q++) {
        printf("%c", *q);
    }
    printf("\n");

    printf("strlen with pointer arithmetic: ");
    char *end = word;
    while (*end != '\0') { end++; }
    printf("%ld, and strlen says %zu\n", (long) (end - word), strlen(word));
    return 0;
}
the string is Anita
through a pointer: Anita
word[0] is A and *p is A
the 4th character: word[3] is t, *(p + 3) is t
walking to the terminator with a pointer: Anita
strlen with pointer arithmetic: 5, and strlen says 5
munotes.in201

Pointers and Arrays

for (char q = word; q != '\0'; q++) is the idiomatic C loop over a string, and it is worth being able to read even if you write the indexed form. *q is the character, and the loop ends when it is the terminator.

An array of pointers, and a pointer to an array

Two declarations that look alike and are not.

int *a[5];      /* an array of 5 pointers to int */
int (*p)[5];    /* a pointer to an array of 5 int */

[] binds tighter than , so int a[5] is an array first. The brackets in the second force the other reading. You will meet the first as an array of strings:

#include <stdio.h>

int main(void)
{
    const char *names[] = {"Anita", "Rahul", "Fatima", "Joseph"};
    int n = (int) (sizeof names / sizeof names[0]);

    printf("%d name(s), each a pointer to a string literal:\n", n);
    for (int i = 0; i < n; i++) {
        printf("  names[%d] = %-8s (%zu characters)\n",
               i, names[i], strlen(names[i]));
    }
    return 0;
}
noheader.c: In function ‘main’:
noheader.c:11:29: warning: implicit declaration of function ‘strlen’ [-Wimplicit-function-declaration]
   11 |                i, names[i], strlen(names[i]));
      |                             ^~~~~~
noheader.c:2:1: note: include ‘<string.h>’ or provide a declaration of ‘strlen’
    1 | #include <stdio.h>
  +++ |+#include <string.h>
    2 |
noheader.c:11:29: warning: incompatible implicit declaration of built-in function ‘strlen’ [-Wbuiltin-declaration-mismatch]
   11 |                i, names[i], strlen(names[i]));
      |                             ^~~~~~
noheader.c:11:29: note: include ‘<string.h>’ or provide a declaration of ‘strlen’
4 name(s), each a pointer to a string literal:
  names[0] = Anita    (5 characters)
  names[1] = Rahul    (5 characters)
  names[2] = Fatima   (6 characters)
  names[3] = Joseph   (6 characters)

That listing is missing #include <string.h>, and the warning is chapter 5's implicit declaration arriving again: gcc has never been told what strlen looks like, so it guesses, and the guess is wrong for a function returning size_t. It happened to print the right numbers here and that is luck, not correctness. The corrected version follows, and the only difference is the second #include.

#include <stdio.h>
#include <string.h>

int main(void)
{
    const char *names[] = {"Anita", "Rahul", "Fatima", "Joseph"};
    int n = (int) (sizeof names / sizeof names[0]);

    printf("%d name(s), each a pointer to a string literal:\n", n);
    for (int i = 0; i < n; i++) {
        printf("  names[%d] = %-8s (%zu characters)\n",
               i, names[i], strlen(names[i]));
    }
    return 0;
}
4 name(s), each a pointer to a string literal:
  names[0] = Anita    (5 characters)
  names[1] = Rahul    (5 characters)
  names[2] = Fatima   (6 characters)
  names[3] = Joseph   (6 characters)

const char *names[] is four pointers, each pointing at a literal somewhere in the program's read-only data. It is not four arrays of characters, which is why sizeof names is 32 on this machine and not the total length of the words.

munotes.in202

Pointers and Arrays

What this does NOT mean

An array is not a pointer. It decays to one in most expressions, and sizeof and & show the difference.

A pointer is not an array. It has no elements of its own.

a[i] is not "a special array syntax". It is defined as *(a + i).

p + 1 is not "one byte further". It is one element further.

int a[5] as a parameter does not mean five elements. The compiler ignores the 5; the parameter is a pointer.

sizeof inside a function does not give the array's size. It gives the pointer's size, and no warning is issued.

char s and char s[] are not interchangeable as declarations of objects. As parameters they are the same; as local variables char s[] = "x" makes an array you may change and char s = "x" makes a pointer to a literal you may not.

Quick revision

  • Pointer arithmetic counts in elements: p + 1 moves by sizeof(*p) bytes.
  • a[i] is defined as (a + i), so (a + i), *(i + a) and i[a] all work.
  • Subtracting two pointers into the same array gives a count of elements.
  • An array's name decays to a pointer to its first element in almost every expression.
  • The three exceptions: sizeof a, &a, and a string literal initialising a char array.
  • Inside a function, an array parameter is a pointer, so sizeof gives 8 and the length must be passed.
  • int a[5], int a[] and int *a are the same parameter declaration.
  • An array cannot be assigned or incremented; a pointer can.
  • for (char q = s; q; q++) is the idiomatic walk over a string.
  • int a[5] is an array of pointers; int (p)[5] is a pointer to an array.

Test yourself

1. What is a[3] defined as?

*(a + 3).

2. If p is an int * and int is 4 bytes, how many bytes does p + 2 differ from p by?

Eight. Pointer arithmetic counts in elements.

3. Inside void f(int a[10]), what is sizeof a?

The size of a pointer, 8 on a 64-bit machine. The parameter is a pointer, whatever the brackets say.

4. Why must an array's length be passed to a function separately?

Because the array decays to a pointer, and a pointer carries no length. sizeof inside the function measures the pointer.

5. Give two things you can do with a pointer that you cannot do with an array name.

Assign to it, and increment it.

6. What is the difference between char s[] = "hi"; and char *s = "hi";?

munotes.in203

Pointers and Arrays

The first is an array of three characters copied into memory you own and may change. The second is a pointer to a string literal, which you must not change.

7. What does int *a[5] declare?

An array of five pointers to int. A pointer to an array of five int would be int (*a)[5].

What can be asked on this, and how to answer it

"Explain the relationship between pointers and arrays." State the identity a[i] is *(a + i), say that an array's name decays to a pointer to its first element in most expressions, and give the three exceptions. Then the consequence: an array parameter is a pointer, so the length must be passed.

"What is pointer arithmetic?" Adding an integer to a pointer moves it by that many elements, using the size of the pointed-to type; subtracting two pointers into the same array gives the number of elements between them. Give a worked example with sizeof(int) and the byte difference.

"Distinguish between an array and a pointer." Give the table: an array is several objects and owns its memory, a pointer is one object holding an address; sizeof differs; an array cannot be assigned or incremented; both can be subscripted; an array decays to a pointer.

"Write a program to print the elements of an array using a pointer." Give a loop with *(a + i) or with a moving pointer, and say in one line that a[i] is the same thing written differently.

"Why does sizeof give the wrong answer inside a function?" Because the parameter is a pointer, not an array: the array decayed at the call. sizeof therefore measures the pointer, and it does so silently, which is why the length is a separate parameter in every C library function that takes an array.

Contents This chapter on its own page

munotes.in204

Chapter Forty-Two

Structures

Syllabus topic 4, "User-defined data types- structure and union"

In one line

A structure groups values of different types into one object with named members, so that things that belong together can be handled as one thing.

Why it exists

An array holds many values of one type. A book has a title, an author, a subject and an identifier: four values of three types that belong to one book. Without structures you would keep four parallel arrays and hope their indexes stayed in step.

char title[50][100];      /* the title of book i   */
char author[50][100];     /* the author of book i  */
int  id[50];              /* the id of book i      */

Nothing in that code says the three are connected. Sort one and you have destroyed the other two. A structure says it in the type.

Declaring one

struct book {
    char title[60];
    char author[40];
    char subject[40];
    int id;
};

That declares a type, struct book, and creates no object. The semicolon after the closing brace is required and is forgotten constantly.

Then objects of it:

struct book b1;
struct book b2 = {"Programming with C", "Kernighan and Ritchie",
                  "Computer Science", 101};
struct book b3 = {.id = 102, .title = "Let Us C"};   /* C99, by name */

struct book is the type's full name, including the keyword. That is why typedef is so common with structures, and chapter 13 promised this:

typedef struct book Book;
Book b1;                    /* now one word */

or in one declaration:

typedef struct {
    char title[60];
    int id;
} Book;

Reaching the members

Two operators, and the second waits for a pointer.

OperatorUsed onExample
.a structureb1.id
->a pointer to a structurep->id

p->id means (p).id. The brackets in that form are required, because . binds tighter than , so p.id would mean (p.id).

The practical: books

MU's Practical 9. She asks for Title, Author, Subject and Book ID, and the details of two printed.

#include <stdio.h>
#include <string.h>

struct book {
    char title[60];
    char author[40];
    char subject[40];
    int id;
};

void print_book(const struct book *b)
{
    printf("  Book ID : %d\n", b->id);
    printf("  Title   : %s\n", b->title);
    printf("  Author  : %s\n", b->author);
    printf("  Subject : %s\n", b->subject);
}

int main(void)
{
    struct book first = {"The C Programming Language",
                         "Kernighan and Ritchie",
                         "Computer Science", 101};
    struct book second;

    /* filling one member at a time, which is what scanf into a struct looks like */
    second.id = 102;
    strcpy(second.title, "Programming with C");
    strcpy(second.author, "E. Balagurusamy");
    strcpy(second.subject, "Computer Science");

    printf("First book:\n");
    print_book(&first);
    printf("\nSecond book:\n");
    print_book(&second);

    printf("\nreaching a member directly: first.title is \"%s\"\n", first.title);
    printf("through a pointer: (&first)->id is %d\n", (&first)->id);
    return 0;
}
First book:
  Book ID : 101
  Title   : The C Programming Language
  Author  : Kernighan and Ritchie
  Subject : Computer Science

Second book:
  Book ID : 102
  Title   : Programming with C
  Author  : E. Balagurusamy
  Subject : Computer Science

reaching a member directly: first.title is "The C Programming Language"
through a pointer: (&first)->id is 101
munotes.in205

Structures

print_book takes const struct book *. A pointer rather than the structure itself, because a structure is copied when passed by value and this one is 144 bytes; const because the function only reads. Chapter 40's advice, applied.

A structure is copied

This is where a structure differs from an array, and it is worth a program of its own.

#include <stdio.h>

struct point { int x; int y; };

void by_value(struct point p)
{
    p.x = 999;
    printf("   inside by_value, p.x is %d\n", p.x);
}

void by_pointer(struct point *p)
{
    p->x = 999;
    printf("   inside by_pointer, p->x is %d\n", p->x);
}

struct point moved(struct point p, int dx, int dy)
{
    p.x += dx;                       /* changes the copy */
    p.y += dy;
    return p;                        /* and returns it */
}

int main(void)
{
    struct point a = {1, 2};
    struct point b;

    b = a;                           /* a whole-structure copy, in one line */
    b.y = 50;
    printf("after b = a and b.y = 50: a is (%d, %d), b is (%d, %d)\n",
           a.x, a.y, b.x, b.y);

    printf("before by_value  : a.x is %d\n", a.x);
    by_value(a);
    printf("after  by_value  : a.x is %d   <- unchanged\n", a.x);

    printf("before by_pointer: a.x is %d\n", a.x);
    by_pointer(&a);
    printf("after  by_pointer: a.x is %d   <- changed\n", a.x);

    struct point c = moved(a, 10, 20);
    printf("moved(a, 10, 20) gave (%d, %d), and a is still (%d, %d)\n",
           c.x, c.y, a.x, a.y);
    return 0;
}
after b = a and b.y = 50: a is (1, 2), b is (1, 50)
before by_value  : a.x is 1
   inside by_value, p.x is 999
after  by_value  : a.x is 1   <- unchanged
before by_pointer: a.x is 1
   inside by_pointer, p->x is 999
after  by_pointer: a.x is 999   <- changed
moved(a, 10, 20) gave (1009, 22), and a is still (999, 2)

Three facts out of that one program:

  1. b = a copies every member, including arrays inside the structure. An array cannot be assigned; a structure containing one can.
  2. Passing by value copies, so by_value could not change the caller's structure.
  3. A structure can be returned by value, which is the second answer to chapter 33's "how do I return two values".

Two structures still cannot be compared with ==. Compare member by member, or use memcmp and understand the padding problem below.

An array of structures

This is the shape of almost every real program, and it is what chapter 44 is built on.

munotes.in206

Structures

#include <stdio.h>

#define STUDENTS 4

struct student {
    int roll;
    char name[20];
    int marks[3];
    double average;
};

int main(void)
{
    struct student class[STUDENTS] = {
        {101, "Anita Desai",   {63, 71, 58}, 0.0},
        {102, "Rahul Mehta",   {48, 52, 61}, 0.0},
        {103, "Fatima Shaikh", {88, 79, 91}, 0.0},
        {104, "Priya Nair",    {35, 42, 40}, 0.0}
    };

    for (int i = 0; i < STUDENTS; i++) {
        int total = 0;
        for (int j = 0; j < 3; j++) {
            total += class[i].marks[j];
        }
        class[i].average = (double) total / 3;
    }

    printf("%-6s %-16s %5s %5s %5s %9s %s\n",
           "Roll", "Name", "P1", "P2", "P3", "Average", "Result");
    for (int i = 0; i < STUDENTS; i++) {
        printf("%-6d %-16s %5d %5d %5d %9.2f %s\n",
               class[i].roll, class[i].name,
               class[i].marks[0], class[i].marks[1], class[i].marks[2],
               class[i].average, class[i].average >= 40 ? "pass" : "fail");
    }

    int best = 0;
    for (int i = 1; i < STUDENTS; i++) {
        if (class[i].average > class[best].average) {
            best = i;
        }
    }
    printf("\nhighest average: %s with %.2f\n",
           class[best].name, class[best].average);
    return 0;
}
Roll   Name                P1    P2    P3   Average Result
101    Anita Desai         63    71    58     64.00 pass
102    Rahul Mehta         48    52    61     53.67 pass
103    Fatima Shaikh       88    79    91     86.00 pass
104    Priya Nair          35    42    40     39.00 fail

highest average: Fatima Shaikh with 86.00

class[i].marks[j] is a member of a structure in an array, and that member is itself an array. Read it left to right: element i of class, its marks member, element j of that. Nesting like this is normal and needs no new rules.

The "find the best" loop is chapter 36's largest-element pattern with class[i].average in place of a[i], and it starts at index 0 for the same reason.

Nesting structures

A member may itself be a structure.

#include <stdio.h>

struct date {
    int day;
    int month;
    int year;
};

struct employee {
    int id;
    char name[30];
    struct date joined;
    double salary;
};

int main(void)
{
    struct employee e = {7001, "Joseph D'Souza", {15, 6, 2021}, 48500.0};

    printf("%s (id %d)\n", e.name, e.id);
    printf("joined on %02d/%02d/%d\n",
           e.joined.day, e.joined.month, e.joined.year);
    printf("salary %.2f\n", e.salary);

    e.joined.year = 2022;
    printf("after correction, the year is %d\n", e.joined.year);
    return 0;
}
Joseph D'Souza (id 7001)
joined on 15/06/2021
salary 48500.00
after correction, the year is 2022

e.joined.year is read left to right and the nesting can go as deep as you like.

sizeof a structure, and padding

sizeof a structure is at least the sum of its members and is usually more. The compiler inserts unused bytes, called padding, so that each member starts at an address the processor likes.

#include <stdio.h>

struct badly_ordered {
    char  a;        /* 1 byte  */
    int   b;        /* 4 bytes */
    char  c;        /* 1 byte  */
};

struct well_ordered {
    int   b;        /* 4 bytes */
    char  a;        /* 1 byte  */
    char  c;        /* 1 byte  */
};

int main(void)
{
    printf("sizeof(char) %zu + sizeof(int) %zu + sizeof(char) %zu = %zu\n",
           sizeof(char), sizeof(int), sizeof(char),
           sizeof(char) + sizeof(int) + sizeof(char));
    printf("sizeof(struct badly_ordered) is %zu\n", sizeof(struct badly_ordered));
    printf("sizeof(struct well_ordered)  is %zu\n", sizeof(struct well_ordered));
    printf("same members, different order, different size\n");
    return 0;
}
munotes.in207

Structures

sizeof(char) 1 + sizeof(int) 4 + sizeof(char) 1 = 6
sizeof(struct badly_ordered) is 12
sizeof(struct well_ordered)  is 8
same members, different order, different size

The exact numbers depend on the machine and the compiler; what is guaranteed is that the total may exceed the sum. Two consequences:

  1. Never assume a structure's size. Use sizeof.
  2. Never compare two structures with memcmp. The padding bytes hold whatever they held, so two structures with identical members can compare unequal. Compare member by member.

What this does NOT mean

A structure declaration does not create an object. struct book { ... }; declares a type. struct book b; creates one.

The semicolon after the closing brace is not optional.

struct book is not shortened to book automatically. In C the keyword is part of the name. A typedef is what shortens it. (C++ differs, which is why many books get this wrong.)

p->x is not different from (*p).x. It is the same thing with better syntax.

Two structures cannot be compared with ==. Compare members.

sizeof a structure is not the sum of its members. Padding.

A structure is not passed by reference. It is copied, unlike an array. Pass a pointer when it is large or when the function must change it.

Quick revision

  • struct name { members }; declares a type; the trailing semicolon is required.
  • struct name object; creates one. typedef shortens the type name.
  • . on a structure, -> on a pointer to one. p->m is (*p).m.
  • Initialise with a braced list, or by member name with C99 designators.
  • A structure is copied on assignment, on being passed and on being returned.
  • That copying is why b = a works for a structure containing an array, though an array cannot be assigned.
  • Pass a pointer, and const if read-only, when the structure is large or must be changed.
  • Cannot be compared with ==; compare member by member, never with memcmp.
  • sizeof a structure is at least the sum of its members, because of padding.
  • Structures nest, and an array of structures is the ordinary way to hold records.

Test yourself

1. What is the difference between a structure declaration and a structure variable?

The declaration introduces a type and allocates nothing. A variable of that type is an object and occupies memory.

munotes.in208

Structures

2. How do you reach member id through a pointer p to a structure?

p->id, which is the same as (*p).id.

3. Can you assign one structure to another?

Yes. b = a copies every member, including any arrays inside it.

4. Can you compare two structures with ==?

No. Compare member by member. memcmp is not a substitute, because padding bytes may differ.

5. Why is sizeof(struct { char a; int b; char c; }) usually 12 rather than 6?

Because the compiler inserts padding so that each member is suitably aligned, and pads the whole structure so that an array of them stays aligned.

6. Write a structure for a book with a title, an author and an id, and create one.

struct book { char title[60]; char author[40]; int id; };
struct book b = {"Programming with C", "Balagurusamy", 101};

7. Why pass a large structure by pointer rather than by value?

Because passing by value copies every byte of it at every call. A pointer copies eight bytes, and const on it keeps the promise that the function will not change the original.

What can be asked on this, and how to answer it

"What is a structure? Explain its declaration, initialisation and member access with an example." Define it as a user-defined type grouping members of possibly different types, give the declaration with its trailing semicolon, the braced initialiser, and both . and ->. Mention the typedef habit and say why it exists.

"Write a program to store and display the details of a book using a structure." Give this chapter's Practical 9 program, with the four members MU names and two books printed through a function taking const struct book *.

"Distinguish between a structure and an array." An array holds many elements of one type reached by an index; a structure holds members of possibly different types reached by name. An array cannot be assigned or returned, and a structure can. An array argument decays to a pointer, while a structure argument is copied.

"Explain nested structures with an example." A member of a structure may itself be a structure. Give struct employee containing struct date joined, and show e.joined.year.

"What is structure padding? Why does it happen?" Unused bytes the compiler inserts between members, and after the last one, so that each member begins at an address its type requires and so that an array of the structure stays aligned. It means sizeof a structure may exceed the sum of its members, so a size must never be assumed and two structures must not be compared with memcmp.

munotes.in209

Structures

"How is a structure passed to a function?" By value, which copies it, or by pointer, which does not. Give both, and say that a pointer with const is the usual choice for anything larger than a few members.

Contents This chapter on its own page

munotes.in210

Chapter Forty-Three

Unions

Syllabus topic 4, "User-defined data types- structure and union"

In one line

A union looks like a structure but all its members occupy the same memory, so it holds exactly one of them at a time and is as large as its largest member.

The difference in one picture

A structure with an int, a double and a 20-character array holds all three at once and is at least 32 bytes, their sum. A union with the same three members holds one of them, so it needs only as much room as its largest member, 20 bytes, rounded up to 24 so that the double inside it stays aligned.

#include <stdio.h>

struct all_three {
    int i;
    double d;
    char s[20];
};

union one_of_three {
    int i;
    double d;
    char s[20];
};

int main(void)
{
    printf("sizeof(int) %zu, sizeof(double) %zu, sizeof(char[20]) %zu\n",
           sizeof(int), sizeof(double), sizeof(char[20]));
    printf("their sum is %zu\n",
           sizeof(int) + sizeof(double) + sizeof(char[20]));
    printf("sizeof(struct all_three)   is %zu  <- room for all of them\n",
           sizeof(struct all_three));
    printf("sizeof(union one_of_three) is %zu  <- room for the largest\n",
           sizeof(union one_of_three));
    return 0;
}
sizeof(int) 4, sizeof(double) 8, sizeof(char[20]) 20
their sum is 32
sizeof(struct all_three)   is 40  <- room for all of them
sizeof(union one_of_three) is 24  <- room for the largest

The union is the size of its largest member, rounded up for alignment: 20 became 24 here. The structure is at least the sum of its members, and more with padding, which is why 32 became 40.

Declaring and using one

The syntax is the structure's with union in place of struct.

union value {
    int i;
    double d;
    char s[20];
};

union value v;
v.i = 42;              /* now the int member is the live one */
v.d = 3.14;            /* now the double member is: the int is gone */

Only one member is meaningful at a time. Writing one overwrites whatever was there, because there is only one piece of memory.

#include <stdio.h>
#include <string.h>

union value {
    int i;
    double d;
    char s[20];
};

int main(void)
{
    union value v;

    v.i = 42;
    printf("after v.i = 42      : v.i is %d\n", v.i);

    v.d = 3.14159;
    printf("after v.d = 3.14159 : v.d is %.5f\n", v.d);

    strcpy(v.s, "Programming");
    printf("after strcpy to v.s : v.s is \"%s\"\n", v.s);

    v.i = 7;
    printf("after v.i = 7       : v.i is %d, and v.s is now \"%s\"\n", v.i, v.s);
    return 0;
}
after v.i = 42      : v.i is 42
after v.d = 3.14159 : v.d is 3.14159
after strcpy to v.s : v.s is "Programming"
after v.i = 7       : v.i is 7, and v.s is now ""

Look at the last line. Writing v.i changed the first bytes of v.s, because they are the same bytes. The string is no longer what was put there, and no warning was given. A union has no memory of which member you wrote.

munotes.in211

Unions

Reading a member other than the one last written is, for most combinations of type, not something the standard defines. The last line of that output is therefore not a fact to rely on; it is a demonstration that the storage is shared.

Why a union exists

Three real reasons.

1. To save space when only one of several things is ever present. A record that holds either an integer amount or a text note, never both, is half the size as a union.

2. To describe a value that genuinely has several forms. A configuration setting is a number, or a string, or a truth value.

3. To look at the bytes of one type as another. This is the use that gets written about, and it is the one to be careful with.

The tagged union, which is the correct pattern

A union on its own is unusable in a program of any size, because nothing records which member is live. The fix is to put the union in a structure beside an enum that says.

#include <stdio.h>
#include <string.h>

enum kind { IS_INT, IS_DOUBLE, IS_TEXT };

struct setting {
    char name[20];
    enum kind kind;
    union {
        int i;
        double d;
        char text[24];
    } value;
};

void print_setting(const struct setting *s)
{
    printf("%-14s = ", s->name);
    switch (s->kind) {
        case IS_INT:
            printf("%d (a whole number)\n", s->value.i);
            break;
        case IS_DOUBLE:
            printf("%.4f (a fraction)\n", s->value.d);
            break;
        case IS_TEXT:
            printf("\"%s\" (text)\n", s->value.text);
            break;
        default:
            printf("(unknown kind)\n");
            break;
    }
}

int main(void)
{
    struct setting settings[3];

    strcpy(settings[0].name, "max_students");
    settings[0].kind = IS_INT;
    settings[0].value.i = 60;

    strcpy(settings[1].name, "pass_percent");
    settings[1].kind = IS_DOUBLE;
    settings[1].value.d = 40.0;

    strcpy(settings[2].name, "college");
    settings[2].kind = IS_TEXT;
    strcpy(settings[2].value.text, "Mumbai University");

    for (int i = 0; i < 3; i++) {
        print_setting(&settings[i]);
    }
    printf("\none setting occupies %zu bytes\n", sizeof(struct setting));
    return 0;
}
max_students   = 60 (a whole number)
pass_percent   = 40.0000 (a fraction)
college        = "Mumbai University" (text)

one setting occupies 48 bytes

That is the pattern to remember, and it is the honest answer to "what is a union for". The enum is the tag; the switch on it is the only safe way to read the union. An anonymous union as a member, as above, is C11 and lets you write s->value.i rather than naming the union type separately.

Looking at the bytes of a value

The third use. This one is machine-dependent by nature and the chapter says so.

#include <stdio.h>

union bytes {
    unsigned int n;
    unsigned char b[4];
};

int main(void)
{
    union bytes v;

    v.n = 0x12345678u;
    printf("the value is %u, in hexadecimal %x\n", v.n, v.n);
    printf("its 4 bytes, lowest address first: ");
    for (int i = 0; i < 4; i++) {
        printf("%02x ", v.b[i]);
    }
    printf("\n");
    if (v.b[0] == 0x78) {
        printf("the least significant byte came first: this machine is little-endian\n");
    } else if (v.b[0] == 0x12) {
        printf("the most significant byte came first: this machine is big-endian\n");
    } else {
        printf("neither: an unusual byte order\n");
    }
    return 0;
}
munotes.in212

Unions

the value is 305419896, in hexadecimal 12345678
its 4 bytes, lowest address first: 78 56 34 12
the least significant byte came first: this machine is little-endian

Byte order is a property of the machine. Almost every machine you will use, including every desktop and phone, is little-endian, and some network equipment is big-endian. The union told us which without any special library, and that is a legitimate and common use.

Reading a unsigned char array member is the one case the standard does permit: examining an object's representation as bytes through unsigned char is explicitly allowed. That is why this listing uses unsigned char and not char or int.

Structure against union

StructureUnion
MembersAll exist at onceOne at a time
MemorySum of members, plus paddingSize of the largest member
Writing one memberLeaves the others aloneOverwrites them all
Members live atDifferent offsetsThe same offset
InitialiserValues for several membersOnly the first member, unless named
Which member is validAll of themWhichever you wrote last, and nothing records it
Normal useA record with several fieldsOne of several alternatives, with a tag
#include <stdio.h>

struct s { int a; int b; };
union u { int a; int b; };

int main(void)
{
    struct s st = {1, 2};
    union u un = {7};                /* only the FIRST member may be given */
    union u un2 = {.b = 9};          /* C99: by name */

    printf("struct: a is %d, b is %d, and they are independent\n", st.a, st.b);
    st.a = 100;
    printf("after st.a = 100: a is %d, b is still %d\n", st.a, st.b);

    printf("union : a is %d\n", un.a);
    un.b = 500;
    printf("after un.b = 500: b is %d and a is %d: the same bytes\n",
           un.b, un.a);
    printf("un2 was initialised by name: b is %d\n", un2.b);
    printf("sizeof struct s is %zu, sizeof union u is %zu\n",
           sizeof(struct s), sizeof(union u));
    return 0;
}
struct: a is 1, b is 2, and they are independent
after st.a = 100: a is 100, b is still 2
union : a is 7
after un.b = 500: b is 500 and a is 500: the same bytes
un2 was initialised by name: b is 9
sizeof struct s is 8, sizeof union u is 4
munotes.in213

Unions

un.a and un.b are both int here, so reading either after writing either is well defined and gives the same value. That is the one case where reading a member you did not write is safe: the two have the same type.

What this does NOT mean

A union is not a structure that saves space. It holds one member, not all of them compressed.

A union does not record which member is live. You must, with a tag.

Reading a member you did not write is not reliable. For most type combinations the standard does not define it. The exception is examining the bytes through unsigned char.

A union is not the size of the sum of its members. It is the size of the largest, possibly rounded up.

A union initialiser does not set several members. It sets the first, or one named with a C99 designator.

A union is not a variant type with checking. Languages that have one do the checking for you; C does not.

Quick revision

  • A union's members share one piece of memory. Size is that of the largest member.
  • Syntax is the structure's with union, including the trailing semicolon.
  • Members are reached with . and -> exactly as for a structure.
  • Writing one member overwrites the others; only the last written is meaningful.
  • Reading another member is not defined for most type pairs; examining the bytes through unsigned char is allowed.
  • A braced initialiser sets the first member only, or a named one in C99.
  • Nothing records which member is live: put the union in a structure with an enum tag and switch on it.
  • A union can be copied and passed like a structure.
  • The endianness test is a real and legitimate use.

Test yourself

1. What is the size of union u { char c; int i; double d; }; where char is 1, int is 4 and double is 8?

8, the size of the largest member, possibly rounded up for alignment.

2. What happens when you write to one member of a union and then read another?

You read the bytes the first member left, reinterpreted as the second member's type. For most type combinations the standard does not define the result, so it must not be relied on.

3. How is a union different from a structure?

A structure's members exist at once at different offsets and its size is at least their sum. A union's members share one offset, only one is meaningful at a time, and its size is that of the largest.

4. Why is a union usually put inside a structure?

Because nothing in a union records which member is live. The structure adds an enum tag, and code then switches on the tag before reading.

munotes.in214

Unions

5. Can a union be initialised with a list of values for all its members?

No. A braced initialiser gives a value to the first member only, or to one named member with a C99 designated initialiser.

6. Give one legitimate use of reading a different member than you wrote.

Examining an object's bytes through an unsigned char array member, which the standard permits and which is how the endianness of a machine can be tested.

7. Is sizeof(union u) ever larger than its largest member?

Yes. It may be rounded up so that an array of the union stays properly aligned.

What can be asked on this, and how to answer it

"What is a union? Explain with an example." Define it as a user-defined type whose members share the same memory, so that one is held at a time and the size is that of the largest. Give the declaration, show that writing one member changes the others, and say that nothing records which member is live.

"Distinguish between a structure and a union." Give the table's rows: all members against one, sum against largest, independent against overwritten, different offsets against the same offset, and the initialiser difference. Finish with the practical consequence: a structure is a record, a union is a choice.

"Write a program to show that the members of a union share memory." Give the four-step program: write the int, write the double, write the string, then write the int again and print the string. Add the sentence that the last read is a demonstration and not a guaranteed value.

"What are the applications of a union?" Saving memory where only one of several values is ever present, representing a value that has several possible forms, and examining the representation of a value as bytes. Give the tagged union as the pattern used in practice.

"What is the size of a union and why?" The size of its largest member, possibly rounded up for alignment, because all members occupy the same storage and that storage must be big enough for any of them.

Contents This chapter on its own page

munotes.in215

Chapter Forty-Four

Putting It Together: A Menu-Driven Program

Syllabus topic 4, "User-defined data types- structure and union"

In one line

A menu-driven program is a do-while loop round a switch, with the data in an array of structures and each menu item in a function of its own.

What this chapter is for

Every technique in this book has been shown on its own. A program is not a collection of techniques; it is a set of decisions about how they fit together. This chapter makes those decisions explicitly, because that is what a mini project is being marked on.

The five decisions, and the chapter that each one comes from:

DecisionChapter
The data is an array of structures with a count42, 36
Each menu item is one function with one job33
Functions that change the data take a pointer40
The menu is a do-while round a switch29, 26
Every input is read and checked in one place15, 29

The design, before the code

The record. An account has a number, a holder's name and a balance. That is a structure.

The store. A fixed array of them and a count of how many are in use. MAX_ACCOUNTS is the capacity; count is the population.

Finding an account. One function, returning the index or -1. Chapter 30's search idiom. Everything else uses it, so the search is written once.

The operations. Open, deposit, withdraw, display one, display all. Five functions, each doing one thing and printing its own report.

Who owns the data. main owns the array. Every function that changes it takes a pointer to it; every function that only reads takes a const pointer. That is chapter 40, and it means you can tell from the parameter list which operations can alter the bank.

Writing those five paragraphs before any code is the single most useful habit this book can leave you with. Every one of them is a sentence a viva examiner will ask you to say out loud.

The program

#include <stdio.h>
#include <string.h>

#define MAX_ACCOUNTS 50
#define NAME_LEN     30

struct account {
    int number;
    char name[NAME_LEN];
    double balance;
};

struct bank {
    struct account accounts[MAX_ACCOUNTS];
    int count;
    int next_number;
};

/* ---- input, read and checked in ONE place ---- */

/* Reads one whole line and discards it, so a bad entry cannot be read twice. */
static void discard_line(void)
{
    int c;
    while ((c = getchar()) != '\n' && c != EOF) {
        /* nothing: the line is being thrown away */
    }
}

static int read_int(const char *prompt, int *out)
{
    printf("%s", prompt);
    if (scanf("%d", out) != 1) {
        discard_line();
        printf("  that was not a whole number\n");
        return 0;
    }
    discard_line();
    return 1;
}

static int read_double(const char *prompt, double *out)
{
    printf("%s", prompt);
    if (scanf("%lf", out) != 1) {
        discard_line();
        printf("  that was not an amount\n");
        return 0;
    }
    discard_line();
    return 1;
}

static int read_name(const char *prompt, char *out, int size)
{
    printf("%s", prompt);
    if (fgets(out, size, stdin) == NULL) {
        return 0;
    }
    out[strcspn(out, "\n")] = '\0';
    return out[0] != '\0';
}

/* ---- the bank ---- */

static void bank_init(struct bank *b)
{
    b->count = 0;
    b->next_number = 1001;
}

/* Chapter 30's search idiom: the index, or -1. */
static int find_account(const struct bank *b, int number)
{
    for (int i = 0; i < b->count; i++) {
        if (b->accounts[i].number == number) {
            return i;
        }
    }
    return -1;
}

static void print_one(const struct account *a)
{
    printf("  %-6d %-30s %12.2f\n", a->number, a->name, a->balance);
}

static void open_account(struct bank *b)
{
    char name[NAME_LEN];
    double opening;

    if (b->count >= MAX_ACCOUNTS) {
        printf("  the bank is full: %d accounts is the limit\n", MAX_ACCOUNTS);
        return;
    }
    if (!read_name("  Holder's name: ", name, (int) sizeof name)) {
        printf("  a name is needed\n");
        return;
    }
    if (!read_double("  Opening balance: ", &opening)) {
        return;
    }
    if (opening < 0) {
        printf("  an opening balance cannot be negative\n");
        return;
    }

    struct account *a = &b->accounts[b->count];
    a->number = b->next_number;
    strncpy(a->name, name, sizeof a->name - 1);
    a->name[sizeof a->name - 1] = '\0';
    a->balance = opening;
    b->count++;
    b->next_number++;

    printf("  opened account %d for %s with %.2f\n",
           a->number, a->name, a->balance);
}

static void deposit(struct bank *b)
{
    int number;
    double amount;

    if (!read_int("  Account number: ", &number)) { return; }
    int at = find_account(b, number);
    if (at < 0) {
        printf("  there is no account %d\n", number);
        return;
    }
    if (!read_double("  Amount to deposit: ", &amount)) { return; }
    if (amount <= 0) {
        printf("  a deposit must be more than zero\n");
        return;
    }
    b->accounts[at].balance += amount;
    printf("  deposited %.2f; account %d now holds %.2f\n",
           amount, number, b->accounts[at].balance);
}

static void withdraw(struct bank *b)
{
    int number;
    double amount;

    if (!read_int("  Account number: ", &number)) { return; }
    int at = find_account(b, number);
    if (at < 0) {
        printf("  there is no account %d\n", number);
        return;
    }
    if (!read_double("  Amount to withdraw: ", &amount)) { return; }
    if (amount <= 0) {
        printf("  a withdrawal must be more than zero\n");
        return;
    }
    if (amount > b->accounts[at].balance) {
        printf("  account %d holds only %.2f, so %.2f cannot be withdrawn\n",
               number, b->accounts[at].balance, amount);
        return;
    }
    b->accounts[at].balance -= amount;
    printf("  withdrew %.2f; account %d now holds %.2f\n",
           amount, number, b->accounts[at].balance);
}

static void show_one(const struct bank *b)
{
    int number;

    if (!read_int("  Account number: ", &number)) { return; }
    int at = find_account(b, number);
    if (at < 0) {
        printf("  there is no account %d\n", number);
        return;
    }
    printf("  %-6s %-30s %12s\n", "Number", "Holder", "Balance");
    print_one(&b->accounts[at]);
}

static void show_all(const struct bank *b)
{
    if (b->count == 0) {
        printf("  no accounts have been opened yet\n");
        return;
    }
    double total = 0.0;
    printf("  %-6s %-30s %12s\n", "Number", "Holder", "Balance");
    for (int i = 0; i < b->count; i++) {
        print_one(&b->accounts[i]);
        total += b->accounts[i].balance;
    }
    printf("  %d account(s), holding %.2f in total\n", b->count, total);
}

static void menu(void)
{
    printf("\n--- Bank management ---\n");
    printf("1 open an account\n");
    printf("2 deposit\n");
    printf("3 withdraw\n");
    printf("4 show one account\n");
    printf("5 show all accounts\n");
    printf("0 quit\n");
}

int main(void)
{
    struct bank b;
    int choice;

    bank_init(&b);

    do {
        menu();
        if (!read_int("Choice: ", &choice)) {
            choice = -1;
            continue;
        }
        switch (choice) {
            case 1: open_account(&b); break;
            case 2: deposit(&b);      break;
            case 3: withdraw(&b);     break;
            case 4: show_one(&b);     break;
            case 5: show_all(&b);     break;
            case 0: printf("  closing. Goodbye.\n"); break;
            default: printf("  %d is not on the menu\n", choice); break;
        }
    } while (choice != 0);
    return 0;
}
munotes.in216

Putting It Together: A Menu-Driven Program

And a session that exercises every branch, including the ones that must refuse:

munotes.in217

Putting It Together: A Menu-Driven Program

1
Anita Desai
5000
1
Rahul Mehta
2500
5
2
1001
1500
3
1002
9000
3
1002
500
4
1001
2
9999
6
0

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   Holder's name:   Opening balance:   opened account 1001 for Anita Desai with 5000.00

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   Holder's name:   Opening balance:   opened account 1002 for Rahul Mehta with 2500.00

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   Number Holder                              Balance
  1001   Anita Desai                         5000.00
  1002   Rahul Mehta                         2500.00
  2 account(s), holding 7500.00 in total

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   Account number:   Amount to deposit:   deposited 1500.00; account 1001 now holds 6500.00

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   Account number:   Amount to withdraw:   account 1002 holds only 2500.00, so 9000.00 cannot be withdrawn

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   Account number:   Amount to withdraw:   withdrew 500.00; account 1002 now holds 2000.00

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   Account number:   Number Holder                              Balance
  1001   Anita Desai                         6500.00

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   Account number:   there is no account 9999

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   6 is not on the menu

--- Bank management ---
1 open an account
2 deposit
3 withdraw
4 show one account
5 show all accounts
0 quit
Choice:   closing. Goodbye.
munotes.in218

Putting It Together: A Menu-Driven Program

Read the output against the program

Walk the session and notice what the program refused to do, because that is what it is being marked on.

  • Two accounts opened, numbered automatically from 1001. The user never chooses a number, so two accounts cannot collide.
  • Show all printed both and totalled them.
  • A deposit of 1500 into 1001 succeeded and reported the new balance.
  • A withdrawal of 9000 from 1002 was refused, because the account held 2500, and the balance was left alone.
  • A withdrawal of 500 from 1002 succeeded.
  • A deposit into account 9999 was refused, because there is no such account, and it was refused before the amount was asked for. The session therefore never supplies an amount for it, which is exactly what happens at a real keyboard: the user is returned to the menu without having typed one.
  • Menu choice 6 was reported and ignored.
  • 0 ended the loop.

That last-but-one point is a design decision worth stating in a viva: deposit finds the account first and asks for the amount second. Asking for the amount and then discovering the account does not exist wastes the user's time and leaves a number in the input that the next read would have to deal with.

The parts of the program that are not about banking

Three pieces of this program are general and you should take them to every menu-driven program you write.

1. discard_line. After a failed scanf, the text that failed is still in the input. Without throwing the rest of the line away, the next read fails on the same text, and the menu spins. This is the single commonest reason a student's menu program hangs in a practical examination.

static void discard_line(void)
{
    int c;
    while ((c = getchar()) != '\n' && c != EOF) { }
}

The c != EOF half is what stops it spinning for ever when the input has ended.

2. One reader per type, returning success. read_int, read_double and read_name each prompt, read, check and report. Every caller is then one line: if (!read_int(...)) { return; }. Without them, the same six lines of checking appear in five functions.

3. const on the reading operations. show_one and show_all take const struct bank *. The compiler will refuse a write through them, so a reader of the code knows those two cannot alter the bank. Chapter 40.

munotes.in219

Putting It Together: A Menu-Driven Program

What this program does not do, and what you would add next

Being honest about the limits is part of presenting a mini project.

  • Nothing is saved. Close the program and the accounts are gone. Files are the next chapter of your degree, not of this book, and fopen, fprintf and fscanf are what you would add.
  • The capacity is fixed at MAX_ACCOUNTS. Memory obtained while the program runs, with malloc, is the general answer and is a later semester's topic.
  • An account cannot be closed. Adding it means either moving the later elements down or marking the record as unused, and the second is usually better.
  • Money is a double. Real banking systems hold amounts in the smallest unit as an integer, in paise, because a double cannot represent 0.10 exactly and repeated arithmetic accumulates error. Chapter 9 gave the reason. For a first-semester mini project a double is expected; knowing why it is wrong is what a viva rewards.

An examiner who asks "what would you improve" is inviting exactly that list. Having it ready is worth more than any extra feature.

Where each technique in this book appears

TechniqueChapterWhere in the program
Structure42struct account, struct bank
Nested structure42struct bank contains an array of struct account
Array of structures36, 42b->accounts
Pointer to a structure39, 42every function's struct bank *b
const pointer parameter40show_one, show_all, find_account
->42b->count, a->balance
Function returning an index or -130, 33find_account
do-while29the menu loop
switch with default26dispatching the choice
for loop over an array28, 36show_all, find_account
while loop to end of line27discard_line
scanf return value checked15read_int, read_double
fgets and strcspn12, 37read_name
strncpy with a hand-written terminator37open_account
Compound assignment18balance += amount
Named constants10, 22MAX_ACCOUNTS, NAME_LEN
Formatted output widths15every printf with %-6d and %12.2f

That table is the answer to "explain your project". Forty-four chapters, one program.

What this does NOT mean

A mini project is not a long main. It is a set of small functions with one main that dispatches.

A menu is not a while loop. It must be shown at least once, so it is a do-while.

Validation is not optional. Most of the marks in a practical examination are lost to programs that work on correct input and misbehave on anything else.

munotes.in220

Putting It Together: A Menu-Driven Program

break in a case does not end the menu. It ends the switch. The loop ends because choice is 0, which its condition tests. Chapter 30.

A fixed-size array is not a defect in a first project. It is a documented limit. Saying what the limit is and what you would do instead is better than pretending otherwise.

Quick revision

  • A menu-driven program is a do-while round a switch, with 0 to quit.
  • Data in a structure; many records in an array of them with a separate count.
  • One function per menu item, each with one job.
  • A function that changes the data takes a pointer; one that only reads takes a const pointer.
  • Search once, in one function, returning the index or -1.
  • Read and check every input in one place, with one small reader per type returning success.
  • After a failed scanf, discard the rest of the line, or the menu spins for ever.
  • Find the record before asking for the amount.
  • Refuse an overdraft, a non-positive amount and an unknown account, and say so.
  • Know the program's limits: nothing is saved, the capacity is fixed, and money in a double is not exact.

Test yourself

1. Why is the menu loop a do-while rather than a while?

Because the menu must be shown at least once before the user can make any choice.

2. Why does a failed scanf need the rest of the line thrown away?

Because the characters that failed to convert are still in the input, so the next read fails on the same text and the menu repeats for ever.

3. Which functions in the program cannot change the bank, and how can you tell?

find_account, show_one and show_all, because their parameter is const struct bank * and the compiler will refuse any write through it.

4. Why does deposit look up the account before asking for the amount?

So that an unknown account number is reported at once, without asking for an amount that will not be used.

5. What is find_account returning when it returns -1, and why -1?

That there is no account with that number. -1 is used because it cannot be a valid array index, so the caller can test for it unambiguously.

6. Why should a real banking program not hold money in a double?

Because values such as 0.10 have no exact binary representation, so repeated arithmetic accumulates error. Amounts are held as whole numbers of the smallest unit, in paise, in an integer type.

7. How would you add the ability to close an account?

Either move every later element of the array down one and decrement the count, or add a member marking the record as unused and skip it everywhere. The second is simpler and does not disturb the other records.

munotes.in221

Putting It Together: A Menu-Driven Program

What can be asked on this, and how to answer it

"Create a mini project on a bank management system. The program should be menu driven." Give this chapter's program, or a program of the same shape. Before writing code, say the five design decisions out loud: array of structures with a count, one function per operation, pointers for the operations that change the data, do-while round a switch, and input checked in one place. That is what the marks are for.

"Explain the structure of your project." Use the last table: name each technique and say where in the program it appears. An answer that walks the program top to bottom without naming the techniques is worth much less.

"What happens if the user types a letter where a number is expected?" scanf returns 0, the reader function reports it, the offending line is discarded and the menu is shown again. Say that without discarding the line the program would loop for ever, because that is the part examiners test.

"What are the limitations of your project and how would you improve it?" Nothing is saved between runs, so add file handling; the capacity is fixed, so allocate memory as needed; accounts cannot be closed; and money should be held in an integer number of paise rather than a double.

"Why did you use a structure rather than separate arrays?" Because the account number, the holder's name and the balance belong to one account, and a structure says so in the type. With separate arrays nothing keeps the three in step, and one sort or one deletion would separate them.

Contents This chapter on its own page

munotes.in222

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Report or request
Done!