munotes®

Major Practical 1 Notes | B.Sc. (Information Technology) Semester 1 | Mumbai University | munotes

Official Notes munotes.in

Major Practical 1

B.SC. (INFORMATION TECHNOLOGY) · SEMESTER 1

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Information Technology)

For B.Sc. (Information Technology) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in First Year

Major Practical 1

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I Programming with C

  1. How This Paper Is Examined, and What Your Journal Must Contain 1
  2. Your First Program: Writing It, Compiling It and Running It 5
  3. Writing the Algorithm, and Drawing the Flowchart 9
  4. Practical 1(a): Simple Interest 13
  5. Practical 1(b): The Greatest of Three Numbers, with the Conditional Operator 17
  6. Practical 1(c): Is This a Leap Year? 20
  7. Practical 2(a): The Roots of a Quadratic Equation 23
  8. Practical 2(b): A Menu Driven Calculator, with switch 27
  9. Practical 2(c): Patterns of Asterisks 31
  10. Practical 3(a): Reversing the Digits of a Number 34
  11. Practical 3(b): The Factorial of a Number 38
  12. Practical 3(c): The Fibonacci Series 41
  13. Practical 4(a): The Area of a Square, Using a Function 44
  14. Practical 4(b): A Recursive Function 47
  15. Practical 4(c): sqrt and abs, and the Headers They Need 51
  16. Practical 4(d): The goto Statement 54
  17. Practical 5(a): Ten Students' Roll Numbers and Names 57
  18. Practical 5(b): Sorting an Array 61
  19. Practical 6(a): Extracting Part of a String 65
  20. Practical 6(b): Is This String a Palindrome? 68
  21. Practical 6(c): strlen and strcmp 72
  22. Practical 7: Swapping Two Numbers, by Value and by Reference 76
  23. Practical 8(a): Reading a Matrix of m Rows and n Columns 80
  24. Practical 8(b): Multiplying Two Matrices in a Function 83
  25. Practical 9: A Structure, and Two Records of It 87
  26. Practical 10: Designing the Bank Management System 91
  27. Practical 10: The Bank Management System, Written and Run 95

Module II

  1. Getting Into MySQL, and What a Database Is 100
  2. Practical 1: The ER Diagram: Entities, Attributes and Keys 104
  3. Practical 1: Relationships and Cardinality 109
  4. Practical 1: Generalization and Specialization 112
  5. Practical 2: Viewing Databases, Creating One and Listing Its Tables 115
  6. Practical 2: Creating a Table, and Choosing Its Data Types 118
  7. Practical 2: The Constraints, and What Each One Refuses 122
  8. Practical 2: Inserting, Updating and Deleting Rows 127
  9. Practical 3: Altering a Table That Already Holds Data 131
  10. Practical 3: Dropping, Truncating and Renaming 136
  11. Practical 3: Backing Up a Database, and Restoring It 140
  12. Practical 4: Simple Queries 144
  13. Practical 4: Aggregate Functions, GROUP BY and HAVING 149
  14. Practical 5: Date Functions 153
  15. Practical 5: String Functions 157
  16. Practical 5: Math Functions 161
  17. Practical 6: The Inner Join 165
  18. Practical 6: The Outer Join 170
  19. Practical 7: Subqueries with IN 175
  20. Practical 7: Subqueries with EXISTS 180
  21. Practical 8: Turning the ER Model Into Tables 184
  22. Practical 8: Normalizing to Third Normal Form 189
  23. Practical 9: Views 194
  24. Practical 10: Granting and Revoking Permissions 199
  25. Practical 10: COMMIT and ROLLBACK 203
munotes.in

Module I

Programming with C

munotes.in

Chapter One

How This Paper Is Examined, and What Your Journal Must Contain

Syllabus topic Major Practical 1, particulars rows 5, 6, 13 and 14: "2 credits (60 Hours of Practical work in a semester)", "30 Hours (C Programming Practical) + 30 Hours(DBMS - Practical)", "50 Marks", and the format of the practical slip

In one line

This is a practical paper. There is no written theory examination for it at all: you are marked on what you can do at a machine, on the notebook you have been keeping all semester, and on the questions you are asked while the examiner stands beside you.

In the wording you can use in a viva: Major Practical 1 is a Major Practical course of 2 credits, 60 hours of practical work, assessed out of 50 marks, of which 40 per cent is internal continuous assessment and 60 per cent is the semester end practical examination.

The two halves

The 60 hours are split down the middle, and the two halves have nothing to do with each other on the timetable even though this book joins them up.

HalfHoursWhat it is
Module 130C Programming Practical: ten practicals, twenty-eight separate programs
Module 230DBMS Practical: ten practicals, worked at a database prompt

Module 1 is the laboratory half of your theory paper Programming with C, so most of what it needs you will also hear in a lecture. Module 2 is different, and this is worth knowing early: there is no database theory paper in your first semester. The twelve Course Objectives MU prints for this paper run from drawing an entity relationship diagram to granting and revoking permissions, and the practical is the only place you are taught any of it. If you skip Module 2 practicals you have not fallen behind in a lab, you have missed a subject.

The 50 marks, and where each one comes from

ComponentMarksHow it is earned
Internal continuous assessment202.5 marks per practical for performing it and submitting the write-up
Semester end practical examination30one sitting, two hours, at a machine
Total50

The internal 20 is not a single test. MU's own words are that you are expected to attend each practical and submit the written practical of the previous session, and that 2.5 marks are awarded for each practical performance and write-up submission, totalling 50 marks, which are then converted to 20.

Read that arithmetic once more, because it tells you exactly how many practicals are being counted: 50 divided by 2.5 is twenty. Ten in Module 1 and ten in Module 2. Every single one of them carries marks, and they are marked in two parts, the doing and the writing.

Notice also which write-up is due when. You submit the previous session's work at the start of the next session. You are therefore always one practical behind in ink, and the commonest way to lose internal marks in this paper is not failing to write a program, it is turning up without last week's write-up.

munotes.in1

How This Paper Is Examined, and What Your Journal Must Contain

The examination slip

Two hours. Thirty marks. MU prints the slip itself:

QuestionFromMarks
Q1Module 113
Q2Module 212
Q3Journal and viva5
Total30

Three things follow from that table and all three change how you should prepare.

Both halves are examined on the same day, every sitting. Thirteen and twelve is as near equal as thirty divides. You cannot pass this paper on C alone, and you cannot pass it on SQL alone.

One question per half. You will be given one C program to write and one database task to perform. Which one is luck; that is exactly why all twenty practicals matter, and why this book works every one of them rather than the popular ones.

Five of the thirty marks are not a program at all. They are your journal and your answers to questions about it. That is a sixth of the paper available to anybody who has kept a proper notebook and can explain their own work out loud.

The sentence that stops a student at the door

MU prints it in row 14, above the slip: a certified copy of the journal is compulsory to appear for the practical examination.

Certified means signed by your subject teacher. It is signed practical by practical, through the semester, as you submit them. It cannot be certified the week before the examination, because the signatures are dated against sessions you either attended or did not.

So the journal is not a formality that carries five marks. It is the ticket. Without it you are not marked out of 25 instead of 30; you do not sit the examination.

What one journal entry contains

MU does not print a format, and different colleges ask for small variations, so ask your own teacher in the first week and follow what they say. What follows is the entry every college expects some version of, and it is the shape this book uses for all twenty practicals.

The heading. Practical number, the date you performed it, and the aim in one line, copied from the syllabus rather than invented. For Practical 1(a) the aim is "To calculate simple interest taking principal, rate of interest and number of years as input from user."

The algorithm. Numbered steps, in plain English, with a Start and a Stop. MU asks for one by name in three of Practical 1's programs, which means the examiner may ask for one in any of them.

The flowchart. Drawn, not described. The symbols are settled and there are only six you need; they are in [Writing the Algorithm, and Drawing the Flowchart].

munotes.in2

How This Paper Is Examined, and What Your Journal Must Contain

The program. Written out in full, in ink, including the header lines. Indentation copied faithfully, because the person signing it reads the indentation to see whether you understood the structure.

The output. What actually appeared on the screen, including whatever you typed in. Not what the program should have printed. If the output in the journal is impossible for the program above it, that is the first thing a viva will ask about.

The conclusion. One or two lines saying what the practical demonstrated. Not "the program was executed successfully", which says nothing. Something like: "Interest is computed as P into R into N divided by 100, and the division must be written in floating point or the fraction is lost."

Worked example: one entry, end to end

Suppose you performed Practical 3(b), the factorial, on 12 August. The entry reads:

Practical 3(b) 12 August 2025

Aim: Write a program to calculate the factorial of a given number.

Algorithm: 1. Start. 2. Read n. 3. Set fact to 1 and i to 1. 4. While i is less than or equal to n, set fact to fact into i and increase i by 1. 5. Print fact. 6. Stop.

Flowchart: drawn, with the loop arrow returning above the comparison.

Program: the listing from [Practical 3(b): The Factorial of a Number], copied in full.

Output: Enter a number: 6 then 6! = 720

Conclusion: A factorial is built by repeated multiplication, so a loop that multiplies a running product by each value in turn is enough. The product grows very fast, and an int stops being able to hold it at 13.

That entry takes twenty minutes to write and is worth 2.5 marks of the internal plus its share of the five in the examination. There are twenty of them in the semester.

What beginners get wrong

Thinking the practical exam is about typing fast. It is two hours for one C program and one database task. Time is not the constraint; knowing which program to write is.

Leaving the journal to the end. The signatures are dated. A journal written in one weekend cannot be certified for sessions that have already passed.

Writing "output: program executed successfully". The output is a specific thing that specific run printed. An examiner reading a journal full of that sentence has been handed a reason to ask harder viva questions.

Preparing only the popular practicals. Q1 and Q2 are drawn from the whole module. Twenty practicals, two questions, and no way to know which two.

Quick revision

  • Practical paper, 2 credits, 60 hours, 50 marks. No theory examination.
  • 30 hours C plus 30 hours DBMS. Ten practicals in each.
  • Internal 20 marks: 2.5 per practical for performance plus the write-up, 50 converted to 20.
  • Semester end 30 marks, two hours: Q1 Module 1 for 13, Q2 Module 2 for 12, Q3 journal and viva for 5.
  • A certified copy of the journal is compulsory to appear.
  • The previous session's write-up is submitted at the next session.
  • A journal entry: heading and aim, algorithm, flowchart, program, real output, conclusion.
munotes.in3

How This Paper Is Examined, and What Your Journal Must Contain

Test yourself

1. How many marks is the semester end practical examination, and how is it divided? Thirty marks in two hours: 13 for a question from Module 1, 12 for a question from Module 2, and 5 for the journal and viva.

2. You have performed every practical but your journal is unsigned. What happens? You cannot appear for the practical examination. A certified copy of the journal is a condition of entry, not a component of the marks.

3. Your internal is out of 20 but marks are awarded at 2.5 each. How many practicals are counted? Twenty, because 20 practicals at 2.5 marks is 50, and MU converts that 50 to 20.

4. Which half of this paper has no theory paper behind it in Semester 1? Module 2, the DBMS half. Its Course Objectives are taught in the practical alone.

5. What is due at the beginning of each practical session? The written write-up of the previous session's practical.

Contents This chapter on its own page

munotes.in4

Chapter Two

Your First Program: Writing It, Compiling It and Running It

Syllabus topic Major Practical 1, Module 1, row 5: "30 Hours (C Programming Practical)". The mechanics every one of her ten practicals is performed with

In one line

A C program is a plain text file. A program called a compiler turns that file into a second file the machine can run, and then you run it. Three steps, three commands, and they never change for the rest of this book.

In the wording you can use in a viva: C is a compiled language. The source code is written in a text file with a .c extension, translated by a compiler into machine code stored in an executable file, and the executable is then run.

The three steps

StepWhat you doWhat comes out
Writetype the program into a file called interest.ca source file
Compilecc interest.c -o interestan executable called interest
Run./interestyour program's output on screen

The -o interest part names the executable. Leave it out and the compiler still works, but it names the result a.out, which tells you nothing when you have twenty of them in one folder. In a college laboratory you may be using Turbo C, Code::Blocks or Dev-C++ instead, where the three steps are menu items or keys rather than commands; they are still the same three steps, and the compiler underneath is doing the same work.

The shortest program that works

#include <stdio.h>

int main(void)
{
    printf("Practical 1\n");
    return 0;
}
Practical 1

Five lines, and every one of them is worth a viva question.

#include <stdio.h> brings in the declarations for the standard input and output library. printf is in that library, not in the language itself, and without this line the compiler does not know what printf is. The # marks it as a line for the preprocessor, which runs before the compiler proper and does nothing but text substitution.

int main(void) is where the program starts. The name main is fixed. int says it hands back a whole number when it finishes, and void in the brackets says it takes nothing. The C standard allows exactly two forms for it, int main(void) and the one that takes the command line arguments, and this book uses the first.

The braces hold the body. Everything between { and } is what main does, in order, top to bottom.

printf("Practical 1\n"); prints. The text inside the double quotes goes to the screen exactly as written, except for \n, which is not a backslash and an n: it is one character, the newline, and it moves the cursor to the start of the next line. Leave it out and your next output continues on the same line, which looks like a bug in the marking.

return 0; hands 0 back to the operating system, meaning the program finished normally. Any other number conventionally means something went wrong. Every program in this book ends with it.

munotes.in5

Your First Program: Writing It, Compiling It and Running It

The semicolon ends a statement. It is not optional, it is not a line separator, and a missing one is the error you will make most often this semester.

The three things a compiler can say

This is the part nobody teaches and every student needs on the first day.

It says nothing. The compilation succeeded. There is now an executable, and you run it. Silence is success.

It reports an error. Compilation failed, no executable was produced, and the old one from your last attempt is still sitting there. This is how a student ends up running yesterday's program and concluding that their edit did nothing.

It reports a warning. Compilation succeeded and an executable exists, but the compiler thinks you have written something you did not mean. It is almost always right.

#include <stdio.h>

int main(void)
{
    int marks = 30;
    printf("Marks: %f\n", marks);
    return 0;
}

That compiles. cc accepts it, produces an executable, and warns:

Marks: 0.000000

The warning says the format %f expects a double and the argument is an int. The program then prints nonsense, because printf read the bytes of an integer as though they were the bytes of a floating point number. What nonsense it prints is not fixed: the standard calls a mismatch between a conversion specification and its argument undefined behaviour, so another compiler, or the same compiler on another machine, may print a different number, or the same program may print a different number twice. Nothing crashed. Nothing was reported at run time. A student who ignores warnings has thrown away the only help they were going to get.

So compile with the warnings switched on, always:

cc -std=c17 -Wall -Wextra interest.c -o interest

-Wall and -Wextra turn on the two standard groups of warnings. -std=c17 fixes which version of the language the compiler judges you against, which matters because C has had five and they disagree in small ways.

Worked example: an error, read and fixed

You type this, missing one semicolon:

#include <stdio.h>

int main(void)
{
    printf("Total marks: ")
    printf("50\n");
    return 0;
}

The compiler refuses it, and this is exactly what it says:

interest.c:5:28: error: expected ';' after expression
    5 |     printf("Total marks: ")
      |                            ^
      |                            ;
1 error generated.

Read it in three parts and it stops being frightening.

interest.c:5:28 is the file, the line and the column. Line 5, column 28. Go there, and notice that the compiler has also printed the line and drawn an arrow at the exact character position, with the character it wanted underneath.

error: means nothing was built. Compare this with warning:, where something was.

munotes.in6

Your First Program: Writing It, Compiling It and Running It

expected ';' after expression is the complaint. The ^ underneath points at the position, and the ; under that is the compiler telling you what to type there.

One caution that will save you an hour. The line number is where the compiler noticed, which is not always where you went wrong. A missing semicolon on line 5 is reported on line 5 here, but a missing closing brace at the end of a function is often reported at the end of the file, many lines below the mistake. When a reported line looks perfectly correct, look at the line above it.

How to read the output blocks in this book

Every program in this book was compiled and run, and the block printed under it is what the program sent to the screen. From the next chapter onward many of the programs ask you for a number first, and there is one thing to know about those blocks before you meet one.

What you type is not in them. When you type 12000 and press Enter, the characters appear on screen because the terminal echoes them, not because the program printed them. So a run that looks like this on your machine:

Enter the principal: 12000
Enter the rate of interest: 8.5

is printed in this book as the program's own output, with the prompts running together and the typed numbers absent. The values that were fed in are shown separately, above the output, so you can reproduce the run exactly.

In your journal, write what appeared on the screen, with the numbers you typed in their places. That is the honest record of the run, and it is what the person signing your journal expects to see.

What beginners get wrong

Believing an unchanged output means an unchanged program. It usually means the compile failed and you ran the previous executable. Look for error: before you look at your logic.

Treating warnings as noise. Every warning in this book is either fixed or shown on purpose as a mistake to avoid. A listing that warns teaches a habit that will fail on somebody else's compiler.

Forgetting \n. Output runs together and looks wrong when the program is right.

Writing Main or MAIN. C distinguishes capitals from small letters everywhere: in main, in printf, and in every name you invent. Total and total are two different variables.

Saving the file as interest.txt. Some editors add the extension silently. The compiler will tell you it does not know what to do with it.

What goes in your journal

Nothing on its own. This chapter is the method you use for all twenty practicals, not a practical. Write the three steps on the inside cover of the notebook where you can find them, and start the journal at Practical 1.

munotes.in7

Your First Program: Writing It, Compiling It and Running It

Quick revision

  • C is compiled. Write the .c file, compile it, run the executable.
  • cc -std=c17 -Wall -Wextra prog.c -o prog then ./prog.
  • #include <stdio.h> is needed for printf and scanf.
  • Execution starts at main, whose body is inside braces, and which ends with return 0;.
  • Every statement ends with a semicolon.
  • \n is one character, the newline.
  • Silence means success. error: means nothing was built. warning: means something was built and is probably wrong.
  • The reported line number is where the compiler noticed, not always where the mistake is.

Test yourself

1. What is the difference between a source file and an executable? The source file is the text you write, in C. The executable is what the compiler produces from it, in the machine's own instructions, and it is the file you run.

2. Why is #include <stdio.h> needed? It brings in the declarations of the standard input and output library. printf and scanf live there, not in the language, so without it the compiler does not know what they are.

3. What does return 0; at the end of main mean? That the program finished normally. The value goes back to the operating system, where a non-zero value conventionally reports a failure.

4. A program compiled with a warning. Can you run it? Yes, an executable was produced. But the compiler has told you the program does not mean what you wrote, so fix it before you write it into your journal.

5. Your program prints the same wrong answer after you edited it. What is the first thing to check? Whether the compile succeeded. If it reported an error, no new executable was made and you have just run the old one.

Contents This chapter on its own page

munotes.in8

Chapter Three

Writing the Algorithm, and Drawing the Flowchart

Syllabus topic Module 1, Practical 1(a), (b) and (c): "Write algorithm & draw flowchart for the same"

In one line

An algorithm is the solution written out as numbered steps in plain English. A flowchart is the same solution drawn as boxes joined by arrows. You write both before you write any C, and both go in your journal.

In the wording you can use in a viva: an algorithm is a finite sequence of unambiguous steps which, carried out in order, solves a given problem in a finite time. A flowchart is its diagrammatic representation, drawn with standard symbols connected by flow lines.

Why both, when the program says the same thing

Because they say it to different people, and because writing them first is what stops you writing the wrong program.

The algorithm is language free. It does not know what C is. A student who can write the algorithm for sorting an array can write that program in C this year and in Python next year, and the thinking is done once. The flowchart shows the shape of the solution at a glance: you can see a loop in a flowchart from across the room, and you cannot see it in fifty lines of code.

There is also the plain examination reason. MU asks for them. Three of Practical 1's programs say "Write algorithm and draw flowchart for the same" in her own words, which means the examiner may ask for either in any practical, and your journal is marked on them all semester.

The five properties an algorithm must have

PropertyWhat it meansWhat breaks it
Finitenessit stopsa loop with no way out
Definitenessevery step means exactly one thing"calculate the interest somehow"
Inputzero or more values are given to ita step that uses a value nobody supplied
Outputat least one value comes outa program that computes and prints nothing
Effectivenessevery step can actually be carried out"guess the answer and check it"

Definiteness is the one students lose marks on. "Add the numbers" is not a step if there are three numbers and you have not said in what order or into what. "Set sum to a plus b plus c" is a step.

How an algorithm is written

Numbered steps. One instruction each. It begins with Start and ends with Stop, and everything in between is either a value read, a value computed, a value printed, or a decision.

For the simple interest program the whole thing is six steps:

1. Start

2. Read P, R and N

3. Set SI to (P into R into N) divided by 100

4. Print SI

5. Stop

Notice what is not there. No float, no scanf, no semicolons. Those are C's problem, and C comes after.

munotes.in9

Writing the Algorithm, and Drawing the Flowchart

Where a step depends on a condition, say so with if and otherwise, and number the branches:

3. If a is greater than b, go to step 4, otherwise go to step 6

Where steps repeat, say what makes them repeat and what makes them stop:

4. While n is greater than 0, repeat steps 5 and 6

5. Set rev to rev into 10 plus the remainder of n divided by 10

6. Set n to n divided by 10

The symbols, and what each one is for

There are six. You will not need a seventh in this paper.

The six flowchart symbols used in this book

Figure 3.1 The six symbols, and what each one is for

The terminator is the rounded box. A chart has exactly one Start and, in this paper, one Stop. Two Stops is not wrong in general, but a chart with three of them usually means the solution has not been thought through.

Input and output is the parallelogram, leaning right. Anything read from the user, anything printed to the screen. Some colleges teach a separate symbol for a printed document; you will not need it.

The process is the plain rectangle. A calculation, or a value put into a variable. One rectangle can hold two or three closely related assignments, as the loop body in this book's charts does.

The decision is the diamond. One arrow in, exactly two out, and the two must be labelled. Label them Yes and No, and write the question in the diamond with a question mark, so that reading the chart out loud makes sense: "a greater than b? Yes, go left."

The connector is the small circle with a letter in it. Use it when the chart will not fit in one column: put a circle marked A where the flow leaves, and another circle marked A where it arrives. It is a label, not a step.

The flow line is the arrow. It carries the order. Arrows go down and to the right by default, and an arrow that goes up is a loop returning, which is exactly what you want a reader to notice.

Worked example: from problem to chart, in four steps

The problem: read a number and print it with its digits reversed.

Step 1: write down what is given and what is wanted. Given: a whole number n. Wanted: the same digits, printed in the opposite order.

Step 2: solve it by hand once, and watch what you do. Take 4271. You peel the last digit off, 1, and start building the answer. Then 7, and the answer so far becomes 17. Then 2, giving 172. Then 4, giving 1724. You stop when there is nothing left to peel. Two operations keep appearing: the last digit of n, and n with its last digit removed.

munotes.in10

Writing the Algorithm, and Drawing the Flowchart

Step 3: write the steps.

1. Start

2. Read n

3. Set rev to 0

4. While n is greater than 0, repeat steps 5 and 6

5. Set rev to rev into 10 plus the remainder of n divided by 10

6. Set n to n divided by 10

7. Print rev

8. Stop

Step 4: draw it. Every read and print becomes a parallelogram, every assignment a rectangle, and the "while" becomes a diamond with the loop body hanging under its Yes arm and an arrow returning from the bottom of the body to just above the diamond.

A while loop drawn as a flowchart

Figure 3.2 The shape every while loop has: test at the top, body below, arrow returning above the test

That returning arrow is the whole point of the picture. It goes back to a point above the diamond, not into the diamond's side, because the test is performed again before the body runs again. A chart whose arrow returns into the body instead of above the test is drawing a different program, and it is the commonest drawing mistake in a first-semester journal.

The two loop shapes, and how to tell them apart in a chart

Test at the topTest at the bottom
In Cwhiledo while
In the chartdiamond above the body, arrow returns above the diamondbody first, diamond below it, arrow returns above the body
Body runszero or more timesone or more times, always at least once
Use it whenthe loop may not need to run at allthe loop must run once before the question can be asked, as a menu must

What beginners get wrong

Writing C in the algorithm. scanf("%d", &n); is not a step. "Read n" is. An algorithm with semicolons in it has not done the job the algorithm exists to do.

A decision with one arrow out. A diamond always has two, and both are labelled. A single unlabelled arrow out of a diamond is a decision that decides nothing.

Unlabelled branches. Yes and No must be written on the arrows. A reader cannot guess, and neither can an examiner.

Forgetting Stop. A chart that runs off the bottom of the page is not finished.

Drawing the chart after the program. Then it is a picture of your code, which is not what it is for. Draw it first and the code writes itself.

What goes in your journal

For every practical where MU asks for them, and for any other where you have room: the numbered algorithm in ink, then the flowchart drawn with a ruler. Use a stencil if your college gives you one. The symbols must be the right shapes: a rectangle where a diamond belongs is marked wrong even when the logic is right.

munotes.in11

Writing the Algorithm, and Drawing the Flowchart

Quick revision

  • Algorithm: numbered steps, plain English, Start and Stop, no C.
  • Five properties: finiteness, definiteness, input, output, effectiveness.
  • Six symbols: terminator, input and output, process, decision, connector, flow line.
  • A decision has exactly two exits and both are labelled Yes and No.
  • A loop's returning arrow goes back to above the test, never into the body.
  • while tests at the top and may run zero times; do while tests at the bottom and runs at least once.
  • Both go in the journal, before the program.

Test yourself

1. Define an algorithm. A finite sequence of unambiguous steps which, carried out in order, solves a given problem in a finite time.

2. Which symbol is used for reading a value from the user, and which for a calculation? A parallelogram for input and output, a rectangle for a process such as a calculation.

3. How many arrows leave a decision box, and what must they carry? Exactly two, and each carries a label, Yes and No.

4. Your flowchart's loop arrow returns into the middle of the loop body. What is wrong? The test is then never performed again, so the chart shows a loop that never ends. The arrow must return above the decision.

5. Give the algorithm for finding the largest of two numbers.

  1. Start. 2. Read a and b. 3. If a is greater than b, print a, otherwise print b. 4. Stop.

6. Why is an algorithm written before the program and not after? Because it settles the solution in language anybody can check, and because the program is then a translation rather than an invention. It is also what MU asks for in the journal.

Contents This chapter on its own page

munotes.in12

Chapter Four

Practical 1(a): Simple Interest

Syllabus topic Module 1, Practical 1(a): "To calculate simple interest taking principal, rate of interest and number of years as input from user. Write algorithm & draw flowchart for the same."

Aim

To calculate simple interest, taking the principal, the rate of interest and the number of years as input from the user, and to write the algorithm and draw the flowchart for the same.

The formula, and where it comes from

Simple interest is interest paid on the original sum only, never on interest already earned.

SI = (P R N) / 100

P is the principal, the money borrowed or deposited. R is the rate per cent per year. N is the number of years. The division by 100 is there because R is given as a percentage: 8.5 per cent means 8.5 hundredths, and dividing by 100 turns the percentage into the fraction the arithmetic actually needs.

One worked case by hand, so you know what the program should print. Principal 12000, rate 8.5, three years. 12000 multiplied by 8.5 is 102000. Multiplied by 3 is 306000. Divided by 100 is 3060.

Algorithm

1. Start

2. Read P, R and N

3. Set SI to (P into R into N) divided by 100

4. Print SI

5. Stop

Flowchart

Flowchart for the simple interest program

Figure 4.1 Practical 1(a): five steps, no decisions and no loops

Five boxes and four arrows. There is no diamond in this chart, because nothing in the problem depends on a condition, and no arrow returns upwards, because nothing repeats. That is what a straight line program looks like.

The program

#include <stdio.h>

int main(void)
{
    float p, r, n, si;

    printf("Enter the principal: ");
    scanf("%f", &p);
    printf("Enter the rate of interest: ");
    scanf("%f", &r);
    printf("Enter the number of years: ");
    scanf("%f", &n);

    si = (p * r * n) / 100;

    printf("Principal = %.2f\n", p);
    printf("Rate = %.2f percent\n", r);
    printf("Years = %.2f\n", n);
    printf("Simple interest = %.2f\n", si);

    return 0;
}
12000
8.5
3
Enter the principal: Enter the rate of interest: Enter the number of years: Principal = 12000.00
Rate = 8.50 percent
Years = 3.00
Simple interest = 3060.00

The run above was fed 12000, 8.5 and 3. Your screen puts each of those numbers after its own prompt as you type it; the block here is only what the program itself printed, for the reason given in [Your First Program: Writing It, Compiling It and Running It].

The program, line by line

float p, r, n, si; declares four variables. A variable is a named place in memory that holds a value you can change. float is the type: a number that may have a fractional part. A rate of 8.5 per cent is not a whole number, and neither is most interest, so float and not int.

Every variable in C is declared before it is used, and the declaration says what type it is. That is not bureaucracy: the type is how the machine knows how many bytes to set aside and how to interpret them.

munotes.in13

Practical 1(a): Simple Interest

printf("Enter the principal: "); prints the prompt. There is no \n at the end on purpose, so that the cursor stays on the same line and the number you type appears next to the words.

scanf("%f", &p); reads one number from the keyboard and stores it in p.

Two parts of that line deserve a paragraph each, because they are where first-semester marks are lost.

"%f" is the format specifier, and it tells scanf what kind of value to expect. %f is a float, %d a whole number, %c a single character, %s a string. Give scanf the wrong one and it reads the bytes wrongly; there is no error message.

&p is the address of p, and the & is not optional. scanf has to change the variable, and to change something it needs to be told where it lives, not what is currently in it. Leaving the & out compiles on many compilers with a warning and then behaves unpredictably at run time. There is one exception you will meet later: a string read with %s needs no &, because the name of an array already is an address.

si = (p r n) / 100; is the calculation. The is multiplication, / is division, and the brackets say do the multiplying first. The brackets are not needed here, because and / have the same precedence and run left to right, but they are worth writing: they make the formula look like the formula.

printf("Simple interest = %.2f\n", si); prints. Inside a printf the % markers are filled in from the values listed after the string, in order. %.2f means a float rounded to two decimal places, which is what money wants.

The trap: integer division

Declare everything int instead and the program goes wrong in a way nothing warns you about. Take a principal of 1250 at 7 per cent for 3 years. By hand: 1250 by 7 is 8750, by 3 is 26250, divided by 100 is 262.5.

#include <stdio.h>

int main(void)
{
    int p = 1250, r = 7, n = 3;

    printf("Integer division: %d\n", (p * r * n) / 100);
    printf("Floating point: %.2f\n", (float) p * r * n / 100);

    return 0;
}
Integer division: 262
Floating point: 262.50

Half a rupee has gone, and no compiler said a word. When both sides of a / are whole numbers C performs integer division: it throws the fractional part away rather than rounding it. 26250 divided by 100 is 262 with a remainder of 50, and the remainder is simply dropped.

munotes.in14

Practical 1(a): Simple Interest

The second line shows the fix when you cannot change the declarations. (float) p is a cast, which converts that one value to a float for this one expression. As soon as one side of an operator is a float, C converts the other side too and the whole calculation is done in floating point.

The worse version of the same trap is at the other end, when the value is read rather than computed:

#include <stdio.h>

int main(void)
{
    int rate = 8.5;

    printf("You meant 8.5, and the int is holding %d\n", rate);
    return 0;
}
You meant 8.5, and the int is holding 8

The fraction is gone before any arithmetic happens at all, so every figure the program prints afterwards is wrong by six per cent. Here the compiler does warn, because the constant 8.5 is visible in the source; when the 8.5 arrives through scanf into an int there is nothing in the source to warn about and the loss is silent. That is why the working program above declares all four variables float.

What beginners get wrong

Leaving out the & in scanf. The variable is not filled in and the value you see is whatever was in memory.

Using %d for a float, or %f for an int. The compiler warns with -Wall, and the printed value is nonsense. Match the specifier to the type every time.

Declaring the rate as int. A rate of 8.5 becomes 8, silently.

Writing si = p r n / 100 with everything int and believing the answer. Integer division truncates; it does not round.

Forgetting that printf prompts have no newline. That is deliberate here, but if you copy a prompt with \n into a program that reads three values, the layout looks wrong next to the journal.

Quick revision

  • SI = (P R N) / 100, and the 100 is there because R is a percentage.
  • float for anything with a fraction; int only for whole numbers.
  • scanf("%f", &p): format specifier, and the & is the address.
  • %d int, %f float, %c character, %s string.
  • %.2f prints two decimal places.
  • Integer divided by integer is integer division: the remainder is discarded, not rounded.
  • A cast (float) p forces floating point arithmetic.
  • 12000 at 8.5 per cent for 3 years is 3060.

What goes in your journal

Aim, the five-step algorithm, the flowchart with its five boxes, the program, the output with your own values written in, and a conclusion naming the one thing the practical demonstrated: that the rate and the result must be held in a floating point type or the division throws the fraction away.

munotes.in15

Practical 1(a): Simple Interest

Test yourself

1. Why is & needed in scanf("%f", &p) but not in printf("%f", p)? scanf must change p, so it needs the address of p. printf only reads the value, so the value itself is enough.

2. int a = 7, b = 2; What does a / b give, and why?

  1. Both operands are integers, so C performs integer division and discards the fractional part rather than rounding it.

3. What does %.2f do? Prints a floating point value rounded to two places after the decimal point.

4. You typed 8.5 for the rate but the interest came out as though the rate were 8. What is the likeliest cause? The rate was declared int, so the fractional part was lost when the value was stored.

5. Rewrite step 3 of the algorithm for compound interest. Set CI to P into ((1 plus R divided by 100) raised to N) minus P. The reading and printing steps do not change.

Contents This chapter on its own page

munotes.in16

Chapter Five

Practical 1(b): The Greatest of Three Numbers, with the Conditional Operator

Syllabus topic Module 1, Practical 1(b): "Write a program to find greatest of three numbers using conditional operator. Write algorithm & draw flowchart for the same."

Aim

To find the greatest of three numbers using the conditional operator, and to write the algorithm and draw the flowchart for the same.

The conditional operator

C has exactly one operator that takes three operands, and this is it. It is written with a question mark and a colon, and it is often called the ternary operator for that reason.

condition ? value_if_true : value_if_false

Read it as a question. "Is the condition true? Then this, otherwise that." The whole thing is an expression: it has a value, so it can be assigned, printed or passed to a function, which is exactly what an if statement cannot do.

#include <stdio.h>

int main(void)
{
    int a = 34, b = 78;

    int big = (a > b) ? a : b;

    printf("The larger of %d and %d is %d\n", a, b, big);
    printf("Directly inside printf: %d\n", (a > b) ? a : b);

    return 0;
}
The larger of 34 and 78 is 78
Directly inside printf: 78

The second line is the point. You cannot put an if statement inside printf, because a statement has no value. You can put a conditional expression there, because it has one.

Three numbers, not two

With three values there are three candidates, so the question has to be asked twice. Ask first which of a and b is bigger, and then compare the winner with c.

big = (a > b) ? ((a > c) ? a : c) : ((b > c) ? b : c);

That is one conditional expression whose two arms are themselves conditional expressions, which is called nesting. Read it outward in:

  • a > b decides which branch is taken.
  • If yes, a is ahead of b, so the answer is whichever of a and c is bigger.
  • If no, b is at least as big as a, so the answer is whichever of b and c is bigger.

The brackets around the inner expressions are not required by the language, because ?: groups from the right, but write them. A nested conditional without brackets is readable to a compiler and not to a person, and the person reading it is marking your journal.

Algorithm

1. Start

2. Read a, b and c

3. If a is greater than b, go to step 4, otherwise go to step 7

4. If a is greater than c, set big to a, otherwise set big to c

5. Print big

6. Stop

7. If b is greater than c, set big to b, otherwise set big to c

8. Go to step 5

Flowchart

Flowchart for finding the greatest of three numbers

Figure 5.1 Two decisions, four possible assignments, and one printed result

munotes.in17

Practical 1(b): The Greatest of Three Numbers, with the Conditional Operator

Two diamonds, four rectangles, and every arm ends in a box. Notice that big = c appears twice. That is not a mistake: c wins both when a beat b and c beat a, and when b beat a and c beat b, and a flowchart shows each path separately rather than merging them early.

The program

#include <stdio.h>

int main(void)
{
    int a, b, c, big;

    printf("Enter three numbers: ");
    scanf("%d %d %d", &a, &b, &c);

    big = (a > b) ? ((a > c) ? a : c) : ((b > c) ? b : c);

    printf("Numbers: %d, %d, %d\n", a, b, c);
    printf("The greatest is %d\n", big);

    return 0;
}
34 78 56
Enter three numbers: Numbers: 34, 78, 56
The greatest is 78

One scanf reads all three, because "%d %d %d" asks for three integers and the spaces in the format string match any amount of whitespace, including the Enter you press between them. You may type them on one line or on three; scanf cannot tell the difference and does not care.

The same answer with if-else

#include <stdio.h>

int main(void)
{
    int a = 34, b = 78, c = 56, big;

    if (a > b) {
        if (a > c)
            big = a;
        else
            big = c;
    } else {
        if (b > c)
            big = b;
        else
            big = c;
    }

    printf("The greatest is %d\n", big);
    return 0;
}
The greatest is 78

Fourteen lines against one. That is why MU asks for the operator: it is the compact form, and a student who can write the nested conditional can certainly write the if-else.

Conditional operator against if-else

Conditional operatorif-else
It isan expression, and has a valuea statement, and has none
Can be assigned, printed, passedyesno
Can hold several statements per branchno, one expression eachyes, a whole block
Operandsthree, which is why it is called ternarynot an operator at all
Good forchoosing between two valueschoosing between two courses of action

The rule of thumb: if the two branches produce a value, use the operator. If they do things, use if-else.

What it does NOT mean

It is not a shorter if. It is a different thing that happens to look similar. (x > 0) ? printf("yes") : printf("no"); compiles, but writing statements inside it is abuse, and a nested chain used for control flow is unreadable within three levels.

The colon is not an else keyword. ? and : are two halves of one operator and neither works without the other.

It does not need brackets round the condition. a > b ? a : b is legal, because > binds tighter than ?:. Brackets are still worth writing.

munotes.in18

Practical 1(b): The Greatest of Three Numbers, with the Conditional Operator

What beginners get wrong

Missing the second colon in a nested expression, which produces a long and unhelpful error because the compiler keeps looking for it.

Comparing with = instead of ==. (a = b) ? a : b assigns and then tests, which is legal C and always wrong here.

Writing a > b > c. That is legal and means something else: a > b gives 1 or 0, and 1 or 0 is then compared with c. There is no three-way comparison in C.

Forgetting the case where two numbers are equal. With 78, 78 and 56 this program prints 78, which is right; "greatest" does not mean "strictly greater than both others".

Quick revision

  • condition ? value_if_true : value_if_false, the only ternary operator in C.
  • It is an expression with a value, so it can be assigned or printed.
  • Greatest of three: (a > b) ? ((a > c) ? a : c) : ((b > c) ? b : c).
  • ?: groups from the right; bracket the inner expressions anyway.
  • scanf("%d %d %d", &a, &b, &c) reads three integers separated by any whitespace.
  • a > b > c does not mean what it looks like.

What goes in your journal

Aim, the eight-step algorithm, the flowchart with both diamonds, the program using the conditional operator, the output for your own three numbers, and a conclusion that states what the operator is: a three-operand expression that yields one of two values, which is why it can be assigned where an if-else cannot.

Test yourself

1. Why is the conditional operator called ternary? Because it takes three operands: the condition, the value when true and the value when false. It is the only operator in C that does.

2. Write an expression that gives the smaller of x and y. (x < y) ? x : y

3. Can you write int m = if (a > b) a; else b;? No. if is a statement and has no value, so it cannot appear on the right of an assignment. The conditional operator can.

4. What does a > b > c evaluate to when a is 5, b is 3 and c is 1? a > b is 1, and 1 > 1 is 0. So the whole expression is 0, which is false, even though 5 is greater than 3 which is greater than 1.

5. Rewrite the three-number expression to find the smallest. small = (a < b) ? ((a < c) ? a : c) : ((b < c) ? b : c);

Contents This chapter on its own page

munotes.in19

Chapter Six

Practical 1(c): Is This a Leap Year?

Syllabus topic Module 1, Practical 1(c): "Write a program to check if the year entered is leap year or not. Write algorithm & draw flowchart for the same."

Aim

To check whether a year entered by the user is a leap year, and to write the algorithm and draw the flowchart for the same.

The rule, and why it is not just "divisible by 4"

The Earth takes about 365.2422 days to go round the Sun. A calendar of 365 days therefore runs fast by roughly a quarter of a day a year, which is why an extra day is added every fourth year. But a quarter of a day is not exactly right either: four years of 365.2422 days is 1460.97 days, and adding a whole day every four years overshoots by about 0.03 days each time.

The Gregorian calendar corrects the overshoot with two more rules. The result is three conditions, and they are applied in this order:

RuleEffectExample
Divisible by 400leap2000, 2400
Otherwise divisible by 100not leap1900, 2100
Otherwise divisible by 4leap2024, 2028
Otherwisenot leap2023, 2025

The order matters. 2000 is divisible by 400, by 100 and by 4, and only the first rule that applies decides. A program that tests "divisible by 100" before "divisible by 400" declares 2000 not to be a leap year, which it is.

Algorithm

1. Start

2. Read year

3. If year is divisible by 400, print leap year and go to step 7

4. If year is divisible by 100, print not a leap year and go to step 7

5. If year is divisible by 4, print leap year and go to step 7

6. Print not a leap year

7. Stop

Flowchart

Flowchart for the leap year test

Figure 6.1 Three decisions in order, and only the first one that applies decides

Three diamonds in a column, each one's No arm dropping into the next. Both output boxes are reached from more than one place, which is normal: two different paths can lead to the same answer.

The remainder operator

The whole program rests on one operator. % gives the remainder after a division of whole numbers.

#include <stdio.h>

int main(void)
{
    printf("2024 %% 4   = %d\n", 2024 % 4);
    printf("2023 %% 4   = %d\n", 2023 % 4);
    printf("1900 %% 100 = %d\n", 1900 % 100);
    printf("2000 %% 400 = %d\n", 2000 % 400);

    return 0;
}
2024 % 4   = 0
2023 % 4   = 3
1900 % 100 = 0
2000 % 400 = 0

A remainder of 0 means the division came out exactly, which is what "divisible by" means. So "year is divisible by 4" is written year % 4 == 0.

Two things about %. It works on whole numbers only: 7.5 % 2 does not compile, because a remainder of a fractional division is not defined in C. And to print a literal per cent sign inside a printf string you write %%, as the listing above does, because a single % starts a format specifier.

munotes.in20

Practical 1(c): Is This a Leap Year?

The program

#include <stdio.h>

int main(void)
{
    int year;

    printf("Enter a year: ");
    scanf("%d", &year);

    if (year % 400 == 0)
        printf("%d is a leap year\n", year);
    else if (year % 100 == 0)
        printf("%d is not a leap year\n", year);
    else if (year % 4 == 0)
        printf("%d is a leap year\n", year);
    else
        printf("%d is not a leap year\n", year);

    return 0;
}
1900
Enter a year: 1900 is not a leap year

The four cases, all run

One run proves one case. Here are all four in one program, so that nothing is being taken on trust.

#include <stdio.h>

int main(void)
{
    int years[] = {2000, 1900, 2024, 2023};

    for (int i = 0; i < 4; i++) {
        int y = years[i];
        int leap = (y % 400 == 0) || (y % 4 == 0 && y % 100 != 0);

        printf("%d: %s\n", y, leap ? "leap" : "not leap");
    }

    return 0;
}
2000: leap
1900: not leap
2024: leap
2023: not leap

That listing also shows the second way of writing the test, as one condition rather than a chain:

(year % 400 == 0) || (year % 4 == 0 && year % 100 != 0)

&& is and: both sides must be true. || is or: either side will do. != is not equal to. Read the expression aloud and it is the rule in English: divisible by 400, or else divisible by 4 but not by 100.

Both forms are correct and either is accepted. The chain of else if is easier to trace in a viva because it matches the flowchart box for box; the single condition is shorter and is what you will see in most code.

What beginners get wrong

Testing only year % 4 == 0. This is right for 2024, 2023, 2025 and every year a student is likely to try, and wrong for 1900, 2100 and 2200. It is the reason MU can set this practical at all.

Getting the order backwards. Testing % 100 before % 400 makes 2000 not a leap year.

Writing year % 4 = 0. One = assigns, two compare. With -Wall the compiler warns; without it the program is wrong and silent.

Using % on a float. It does not compile. The remainder operator is for whole numbers.

Mixing up && and ||. (y % 4 == 0 || y % 100 != 0) is true for almost every year, including 2023, because 2023 is not divisible by 100.

munotes.in21

Practical 1(c): Is This a Leap Year?

Quick revision

  • Three rules in order: divisible by 400 is leap; else divisible by 100 is not; else divisible by 4 is leap; else not.
  • % is the remainder operator, whole numbers only.
  • "Divisible by n" is year % n == 0.
  • %% prints a literal per cent sign inside printf.
  • One condition form: (y % 400 == 0) || (y % 4 == 0 && y % 100 != 0).
  • && is and, || is or, != is not equal to, == is equal to.
  • 2000 is leap, 1900 is not, 2024 is leap, 2023 is not.

What goes in your journal

Aim, the seven-step algorithm, the flowchart with all three diamonds, the program, and the output. Run it four times, once for each case in the table, and write all four outputs in the journal: an examiner who sees only 2024 in a leap year write-up has been given a reason to ask whether you know about 1900.

Test yourself

1. Is 1900 a leap year? No. It is divisible by 100 and not by 400, so the second rule applies and it is an ordinary year of 365 days.

2. Why is testing year % 4 == 0 alone not enough? Because the Gregorian calendar skips the extra day in century years that are not divisible by 400. That test declares 1900 and 2100 to be leap years, and they are not.

3. What does 2023 % 4 give?

  1. 2023 is 505 fours and 3 left over.

4. Write the leap year test as one expression. (year % 400 == 0) || (year % 4 == 0 && year % 100 != 0)

5. Why does printf("%d %% 4", n) print a per cent sign? Because %% is the escape for a literal per cent inside a format string. A single % would be read as the start of a format specifier.

6. How many days does February have in 2024 and in 1900? 29 in 2024, which is a leap year, and 28 in 1900, which is not.

Contents This chapter on its own page

munotes.in22

Chapter Seven

Practical 2(a): The Roots of a Quadratic Equation

Syllabus topic Module 1, Practical 2(a): "Write a program to calculate roots of a quadratic equation."

Aim

To calculate the roots of a quadratic equation.

The mathematics, in two lines

A quadratic equation is any equation of the form:

a x x + b * x + c = 0

with a not zero. Its two roots are given by the quadratic formula:

x = (-b + square root of (b b - 4 a c)) / (2 a)

= (-b - square root of (b b - 4 a c)) / (2 a)

The part under the square root sign has a name of its own, the discriminant, usually written D:

D = b b - 4 a * c

It is the whole program, because its sign decides what kind of roots there are.

DiscriminantThe roots areWhy
D greater than 0real and differentthe square root is a real number, added and subtracted
D equal to 0real and equalthe square root is 0, so both formulas give the same value
D less than 0complexthe square root of a negative number is not real

That table is the answer to "explain the three cases", which is the viva question this practical always carries.

Algorithm

1. Start

2. Read a, b and c

3. If a is 0, print that it is not a quadratic equation and go to step 9

4. Set D to b into b minus 4 into a into c

5. If D is greater than 0, set root1 to (minus b plus the square root of D) divided by 2a, set root2 to (minus b minus the square root of D) divided by 2a, print both, and go to step 9

6. If D is 0, set root1 to minus b divided by 2a, print it, and go to step 9

7. Set real to minus b divided by 2a and set imaginary to the square root of minus D divided by 2a

8. Print the two complex roots

9. Stop

The library the square root lives in

sqrt is not part of the C language. It is a function in the standard mathematics library, declared in math.h, so a program that uses it must include that header:

#include <math.h>

On many Linux systems the compiler must also be told to link the mathematics library, with -lm at the end of the command:

cc -std=c17 -Wall -Wextra roots.c -o roots -lm

If your compiler answers with "undefined reference to sqrt" the code is correct and the -lm is missing. On macOS and in most IDE toolchains the library is linked automatically and the flag is unnecessary. Add it anyway; on a system that does not need it, it does nothing.

sqrt takes a double and returns a double. Handing it a negative number is the one thing it cannot do, which is why the program tests the sign of D before calling it.

munotes.in23

Practical 2(a): The Roots of a Quadratic Equation

The program

#include <stdio.h>
#include <math.h>

int main(void)
{
    float a, b, c, d, root1, root2, real, imag;

    printf("Enter a, b and c: ");
    scanf("%f %f %f", &a, &b, &c);

    if (a == 0) {
        printf("a is zero, so this is not a quadratic equation\n");
    } else {
        d = b * b - 4 * a * c;
        printf("Discriminant = %.2f\n", d);

        if (d > 0) {
            root1 = (-b + sqrt(d)) / (2 * a);
            root2 = (-b - sqrt(d)) / (2 * a);
            printf("Roots are real and different: %.2f and %.2f\n", root1, root2);
        } else if (d == 0) {
            root1 = -b / (2 * a);
            printf("Roots are real and equal: %.2f and %.2f\n", root1, root1);
        } else {
            real = -b / (2 * a);
            imag = sqrt(-d) / (2 * a);
            printf("Roots are complex: %.2f + %.2fi and %.2f - %.2fi\n",
                   real, imag, real, imag);
        }
    }

    return 0;
}
1 -7 12
Enter a, b and c: Discriminant = 1.00
Roots are real and different: 4.00 and 3.00

Check it by hand. With a as 1, b as minus 7 and c as 12, D is 49 minus 48, which is 1. The square root of 1 is 1, so the roots are (7 plus 1) over 2 and (7 minus 1) over 2, which are 4 and 3. And indeed x squared minus 7x plus 12 factorises as (x minus 4)(x minus 3).

All three cases, run

#include <stdio.h>
#include <math.h>

int main(void)
{
    float sets[3][3] = { {1, -7, 12}, {1, -4, 4}, {1, 2, 5} };

    for (int i = 0; i < 3; i++) {
        float a = sets[i][0], b = sets[i][1], c = sets[i][2];
        float d = b * b - 4 * a * c;

        printf("a=%.0f b=%.0f c=%.0f  D=%.0f  ", a, b, c, d);

        if (d > 0)
            printf("real and different: %.2f, %.2f\n",
                   (-b + sqrt(d)) / (2 * a), (-b - sqrt(d)) / (2 * a));
        else if (d == 0)
            printf("real and equal: %.2f\n", -b / (2 * a));
        else
            printf("complex: %.2f +- %.2fi\n", -b / (2 * a), sqrt(-d) / (2 * a));
    }

    return 0;
}
a=1 b=-7 c=12  D=1  real and different: 4.00, 3.00
a=1 b=-4 c=4  D=0  real and equal: 2.00
a=1 b=2 c=5  D=-16  complex: -1.00 +- 2.00i

The middle case is x squared minus 4x plus 4, which is (x minus 2) squared, so both roots are 2. The last is x squared plus 2x plus 5, whose roots are minus 1 plus 2i and minus 1 minus 2i.

munotes.in24

Practical 2(a): The Roots of a Quadratic Equation

Why a == 0 is tested first

Because dividing by 2 * a when a is zero is division by zero. With integers that crashes the program; with floats C gives you inf or nan, which prints as inf or nan and looks like a bug in your formula.

More importantly it is not a quadratic equation at all. With a of 0 the equation is b * x + c = 0, a linear equation with the single root minus c over b. Saying so is worth a mark; printing nan is worth none.

What beginners get wrong

Forgetting math.h. Without it the compiler does not know what sqrt returns and, in C17, refuses to compile at all.

Calling sqrt(d) before checking the sign of D. The square root of a negative number is nan, and every calculation that touches a nan gives nan.

Writing 2a for 2 * a. C has no implied multiplication. 2a is not an expression.

Writing b b as b ^ 2. In C, ^ is the bitwise exclusive-or operator, not a power. Use b b, or pow(b, 2) from math.h, and prefer b * b because it is faster and exact.

Using int for the coefficients. The roots of a quadratic are rarely whole numbers. Use float.

Quick revision

  • D = b b - 4 a * c, the discriminant.
  • D greater than 0: real and different. D equal to 0: real and equal. D less than 0: complex.
  • Roots: (-b + sqrt(D)) / (2 a) and (-b - sqrt(D)) / (2 a).
  • When D is negative, the real part is -b / (2 a) and the imaginary part is sqrt(-D) / (2 a).
  • sqrt needs #include <math.h>, and on Linux the compiler may need -lm.
  • Test a == 0 first: with a of 0 it is a linear equation, not a quadratic.
  • ^ is not a power operator in C.

What goes in your journal

Aim, the nine-step algorithm, the flowchart with one decision for a == 0 and two more for the sign of D, the program, and three separate runs, one for each case, with the coefficients and the output of each. A journal that shows only the real and different case has demonstrated a third of the practical.

Test yourself

1. What is the discriminant, and what does it tell you? b b - 4 a * c. Its sign decides whether the roots are real and different, real and equal, or complex.

2. For x squared minus 4x plus 4, what are the roots? Both are 2. The discriminant is 16 minus 16, which is 0, so the roots are real and equal.

munotes.in25

Practical 2(a): The Roots of a Quadratic Equation

3. Which header declares sqrt, and what may the compiler need besides? math.h. On many Linux systems the mathematics library must also be linked with -lm.

4. Why must a == 0 be tested before anything else? Because the formula divides by 2 * a, and because with a of 0 the equation is linear and has one root, not two.

5. What does b ^ 2 compute in C? The bitwise exclusive-or of b and 2, not b squared. C has no power operator.

6. The program printed nan for both roots. What happened? sqrt was called with a negative discriminant. The sign of D must be tested before the square root is taken.

Contents This chapter on its own page

munotes.in26

Chapter Eight

Practical 2(b): A Menu Driven Calculator, with switch

Syllabus topic Module 1, Practical 2(b): "Write a menu driven program using switch case to perform add / subtract / multiply / divide based on the users choice."

Aim

To write a menu driven program using switch case to perform addition, subtraction, multiplication or division based on the user's choice.

What "menu driven" means

The program prints a list of things it can do, each with a number beside it. The user types a number, the program does that one thing, and then the menu comes back. It keeps coming back until the user chooses to stop.

Two ideas make that work and you need both. switch picks one of several courses of action from a single value. A loop brings the menu back afterwards. Neither alone is a menu.

The switch statement

switch (expression) {
case constant1:
    statements, then break;
case constant2:
    statements, then break;
default:
    statements
}

The expression is worked out once, and control jumps to the case whose constant matches it. If none matches, control goes to default, and if there is no default the whole statement does nothing.

Four rules, and each one is a viva question.

The expression must be an integer type. An int, a char, or an enumeration constant. You cannot switch on a float and you cannot switch on a string.

Each case label must be a constant. case 3: is fine, case n: is not, because the compiler must know every label's value while it is compiling.

No two labels may be the same. case 1: twice does not compile.

break leaves the switch. Without it, execution carries straight on into the next case, which is called fall-through.

Fall-through, run rather than described

#include <stdio.h>

int main(void)
{
    for (int choice = 1; choice <= 3; choice++) {
        printf("choice %d: ", choice);

        switch (choice) {
        case 1:
            printf("one ");
        case 2:
            printf("two ");
            break;
        case 3:
            printf("three ");
        }

        printf("\n");
    }

    return 0;
}
choice 1: one two
choice 2: two
choice 3: three

Choice 1 printed two words. There is no break after case 1, so once control entered at case 1 it kept going until it met one. Case 2 has a break, so it stopped there.

This is not a defect in C. Fall-through is deliberately useful when several values want the same treatment:

case 'a':
case 'e':
case 'i':
case 'o':
case 'u':
    printf("vowel\n");
    break;

Five labels, one body, no repetition. But an accidental fall-through is one of the commonest bugs in a first-semester program, and it is invisible in the source: nothing is missing that the eye expects to see. Write break at the end of every case, and when you mean to fall through, say so in a comment.

The program

#include <stdio.h>

int main(void)
{
    int choice;
    float x, y;

    do {
        printf("\n1 Add  2 Subtract  3 Multiply  4 Divide  5 Exit\n");
        printf("Enter your choice: ");
        scanf("%d", &choice);

        if (choice >= 1 && choice <= 4) {
            printf("Enter two numbers: ");
            scanf("%f %f", &x, &y);

            switch (choice) {
            case 1:
                printf("%.2f + %.2f = %.2f\n", x, y, x + y);
                break;
            case 2:
                printf("%.2f - %.2f = %.2f\n", x, y, x - y);
                break;
            case 3:
                printf("%.2f * %.2f = %.2f\n", x, y, x * y);
                break;
            case 4:
                if (y == 0)
                    printf("Division by zero is not allowed\n");
                else
                    printf("%.2f / %.2f = %.2f\n", x, y, x / y);
                break;
            }
        } else if (choice != 5) {
            printf("Please choose a number between 1 and 5\n");
        }
    } while (choice != 5);

    printf("Calculator closed\n");
    return 0;
}
munotes.in27

Practical 2(b): A Menu Driven Calculator, with switch

1
12 5
4
9 0
7
5

1 Add  2 Subtract  3 Multiply  4 Divide  5 Exit
Enter your choice: Enter two numbers: 12.00 + 5.00 = 17.00

1 Add  2 Subtract  3 Multiply  4 Divide  5 Exit
Enter your choice: Enter two numbers: Division by zero is not allowed

1 Add  2 Subtract  3 Multiply  4 Divide  5 Exit
Enter your choice: Please choose a number between 1 and 5

1 Add  2 Subtract  3 Multiply  4 Divide  5 Exit
Enter your choice: Calculator closed

Four passes through the loop: an addition, a division by zero that was refused, an invalid choice of 7 that was rejected, and 5 to stop.

Why do while and not while

do {
    ...
} while (choice != 5);

A do while loop runs its body first and tests afterwards, so the body always runs at least once. That is exactly what a menu needs: the menu has to be shown before the user can possibly have made a choice.

Write it as an ordinary while and the condition tests choice before anything has put a value in it, so the loop is deciding on whatever rubbish happened to be in that memory. The compiler warns about the uninitialised read, and that warning is the language telling you the loop is the wrong shape.

Why division by zero is refused rather than attempted

Dividing an int by zero crashes the program outright. Dividing a float by zero does not: C gives inf, or nan if the top is zero as well, and the program carries on printing meaningless results.

Neither is acceptable in a program that is going to be marked. The one line that tests y == 0 before the division is the difference between a program that handles a bad input and one that is defeated by it, and examiners look for it.

switch against if-else-if

switchif-else-if
Testsone value against several constantsany conditions at all
Value may beint, char, enumanything, including floats and ranges
Labelsconstants onlyfull expressions
Ranges such as marks > 75cannot be writtennatural
Falls through without breakyesnever
Reads better whenthere are several fixed choices, as in a menuthe conditions differ in kind
munotes.in28

Practical 2(b): A Menu Driven Calculator, with switch

What beginners get wrong

Leaving out break. The next case runs too. Nothing warns you.

Switching on a float. switch (price) does not compile.

Writing case choice:. A case label must be a constant known at compile time.

Testing the choice inside switch but reading the numbers outside it. Then an invalid choice still asks for two numbers. Read the operands only when the choice is one that needs them, as the program above does.

Forgetting default. The program above uses an if to reject bad input before the switch, which is equally good; what is not acceptable is a switch with neither, because an invalid choice then does nothing at all and the user cannot tell whether the program worked.

Quick revision

  • switch (n) { case 1: ... break; default: ... }.
  • The switch expression must be an integer type; case labels must be constants; no two labels may be equal.
  • Without break, control falls through into the next case.
  • Deliberate fall-through groups several labels onto one body.
  • do while tests at the bottom, so the body runs at least once. A menu needs that.
  • Test for division by zero before dividing, not after.
  • switch compares one value against constants; if-else-if can test anything.

What goes in your journal

Aim, an algorithm whose steps include the loop and the five choices, a flowchart with the menu inside the loop and a decision for each case, the program, and one output showing at least three passes: a successful operation, a refused division by zero, and the exit. A run that shows only one addition has not demonstrated that the program is menu driven.

Test yourself

1. What happens if break is left out of a case? Execution falls through into the next case and keeps going until it meets a break or the end of the switch.

2. Can you switch on a float? No. The controlling expression must have an integer type: int, char or an enumeration.

3. Why is do while the right loop for a menu? Because it tests at the bottom, so the menu is always displayed at least once before the user's choice is examined.

4. Give one case where fall-through is wanted. Grouping several labels onto one body, such as the five vowels all leading to one printf.

5. What does 9.0 / 0 give in a float calculation, and why is that not good enough? inf. It does not crash, so the program carries on and prints a meaningless result, which is harder to notice than a crash.

munotes.in29

Practical 2(b): A Menu Driven Calculator, with switch

6. What is the purpose of default? It is the case that runs when no label matches, so the program can tell the user their choice was not valid.

Contents This chapter on its own page

munotes.in30

Chapter Nine

Practical 2(c): Patterns of Asterisks

Syllabus topic Module 1, Practical 2(c): "Write a program to print the pattern of asterisks."

Aim

To print a pattern of asterisks.

The one idea: rows outside, columns inside

Every one of these patterns is made by two loops, one inside the other.

The outer loop runs once per row. The inner loop runs once per character in that row. When the inner loop finishes, you print a newline and the outer loop starts the next row.

for (row = 1; row <= n; row++) {
    for (col = 1; col <= how_many_in_this_row; col++)
        printf("*");
    printf("\n");
}

That skeleton never changes. The only thing that changes from pattern to pattern is the expression how_many_in_this_row, and finding it is the whole exercise.

Finding the rule

Write the pattern out, number the rows, and count.

RowAsterisksSpaces before
110
220
330
440
550

The asterisk count is the row number, so the inner loop runs row times. That is the right-angled triangle, and it is the pattern MU's wording points at.

For the inverted triangle the counts are 5, 4, 3, 2, 1, so with n of 5 the count is n - row + 1. For the pyramid each row has an odd number of asterisks, 1, 3, 5, 7, 9, which is 2 * row - 1, and it is pushed right by n - row spaces to make the left edge slope.

The program

#include <stdio.h>

int main(void)
{
    int n, row, col;

    printf("Enter the number of rows: ");
    scanf("%d", &n);

    for (row = 1; row <= n; row++) {
        for (col = 1; col <= row; col++)
            printf("*");
        printf("\n");
    }

    return 0;
}
5
Enter the number of rows: *
**
***
****
*****

The first asterisk sits on the same line as the prompt, because the prompt has no \n of its own. On your screen the 5 you typed is between them.

Five patterns, all printed

#include <stdio.h>

int main(void)
{
    int n = 5, row, col;

    printf("1. Right angled triangle\n");
    for (row = 1; row <= n; row++) {
        for (col = 1; col <= row; col++)
            printf("*");
        printf("\n");
    }

    printf("\n2. Inverted triangle\n");
    for (row = 1; row <= n; row++) {
        for (col = 1; col <= n - row + 1; col++)
            printf("*");
        printf("\n");
    }

    printf("\n3. Right aligned triangle\n");
    for (row = 1; row <= n; row++) {
        for (col = 1; col <= n - row; col++)
            printf(" ");
        for (col = 1; col <= row; col++)
            printf("*");
        printf("\n");
    }

    printf("\n4. Pyramid\n");
    for (row = 1; row <= n; row++) {
        for (col = 1; col <= n - row; col++)
            printf(" ");
        for (col = 1; col <= 2 * row - 1; col++)
            printf("*");
        printf("\n");
    }

    printf("\n5. Hollow square\n");
    for (row = 1; row <= n; row++) {
        for (col = 1; col <= n; col++) {
            if (row == 1 || row == n || col == 1 || col == n)
                printf("*");
            else
                printf(" ");
        }
        printf("\n");
    }

    return 0;
}
munotes.in31

Practical 2(c): Patterns of Asterisks

1. Right angled triangle
*
**
***
****
*****

2. Inverted triangle
*****
****
***
**
*

3. Right aligned triangle
    *
   **
  ***
 ****
*****

4. Pyramid
    *
   ***
  *****
 *******
*********

5. Hollow square
*****
*   *
*   *
*   *
*****

Five shapes, one skeleton, five different expressions inside it. Compare pattern 3 with pattern 4: the only difference is that the second inner loop runs row times instead of 2 * row - 1 times. That is what is meant by finding the rule.

The spaces are printed, not implied

Pattern 3 has a loop that prints a space. This is the step students leave out, and then wonder why the triangle will not move to the right.

The screen has no idea where you want a character to go. If you want four blanks before an asterisk, four blanks must be printed. A printf(" ") in a loop is how that is done, and the number of them is worked out exactly like the number of asterisks: write the rows out, count, and find the expression.

The pyramid's two rules, side by side

RowSpaces n - rowAsterisks 2 * row - 1Total width
1415
2336
3257
4178
5099

The asterisk counts go up by two each row and the spaces come down by one, which is what makes both edges slope at the same angle.

What beginners get wrong

Putting printf("\n") inside the inner loop. Then every asterisk is on a line of its own, and n rows become the whole pattern in a column.

Starting row at 0 but writing the rule as though it started at 1. Either is fine, but the rule must match. With row from 0 to n-1 the triangle's inner loop runs row + 1 times.

Not resetting the inner counter. In C the inner for sets col = 1 every time it starts, so this cannot happen; it is a real bug in languages where the loop variable is set outside.

Printing a tab instead of a space. A tab moves to the next tab stop, not one column, so the shape collapses.

Trying to draw the pattern with one loop. It cannot be done with one loop and a counter without effectively simulating two.

Quick revision

  • Outer loop for rows, inner loop for the characters in a row, printf("\n") after the inner loop.
  • Right angled triangle: inner loop runs row times.
  • Inverted triangle: n - row + 1 times.
  • Right aligned: n - row spaces, then row asterisks.
  • Pyramid: n - row spaces, then 2 * row - 1 asterisks.
  • Hollow square: print an asterisk when row or col is 1 or n, a space otherwise.
  • Spaces must be printed; the screen will not indent for you.
munotes.in32

Practical 2(c): Patterns of Asterisks

What goes in your journal

Aim, the algorithm with both loops written as steps, the flowchart with the inner loop drawn inside the outer one, the program, and the printed pattern copied exactly, asterisk for asterisk. Write the counting table for your pattern beside it: it is the working, and it shows the examiner that you derived the rule instead of remembering it.

Test yourself

1. How many times does the inner loop run in total for a right angled triangle of 5 rows? 1 plus 2 plus 3 plus 4 plus 5, which is 15.

2. What goes wrong if printf("\n") is put inside the inner loop? Every asterisk goes onto its own line, so the pattern comes out as a single column.

3. For a pyramid of n rows, how many asterisks are in row r? 2 * r - 1, so row 1 has one and row 5 has nine.

4. How do you move a triangle to the right? Print spaces before the asterisks, with an inner loop of their own that runs n - row times.

5. Write the inner loop condition for an inverted triangle with n rows. col <= n - row + 1

6. Which characters are asterisks in a hollow square? Those in the first row, the last row, the first column or the last column. Everything else is a space.

Contents This chapter on its own page

munotes.in33

Chapter Ten

Practical 3(a): Reversing the Digits of a Number

Syllabus topic Module 1, Practical 3(a): "Write a program using while loop to reverse the digits of a number."

Aim

To reverse the digits of a number using a while loop.

The technique, done by hand first

Take 4271. Peel the last digit off, and build the answer from the pieces as they come.

Stepn beforen % 10rev beforerev * 10 + n % 10n after
14271101427
2427711742
3422171724
44417217240

Two operations do all the work.

n % 10 gives the last digit. The remainder after dividing by ten is whatever is in the units place. 4271 divided by 10 is 427 with 1 left over, and that 1 is the last digit.

n / 10 removes the last digit. This is integer division, and here the truncation that was a trap in [Practical 1(a): Simple Interest] is exactly what you want: 4271 divided by 10 is 427.1 in arithmetic and 427 in C, and 427 is the number with its last digit gone.

rev * 10 + digit puts the digit on the end of the answer. Multiplying by ten shifts everything one place left and leaves a zero in the units place, and adding the digit fills it.

The loop stops when n reaches 0, because there is then nothing left to peel.

Algorithm

1. Start

2. Read n

3. Set rev to 0

4. While n is greater than 0, repeat steps 5 and 6

5. Set rev to rev into 10 plus the remainder of n divided by 10

6. Set n to n divided by 10

7. Print rev

8. Stop

Flowchart

The reverse digits program as a flowchart

Figure 10.1 The test is at the top, and the loop's arrow returns above it

Why a while loop, and not a for loop

A for loop is the right shape when you know at the start how many times the body will run: five rows, ten students, n terms. You write the count into the loop's own head, and reading the head tells you the answer.

Here you do not know how many times. It runs once per digit, and the number of digits is not known until the digits have been counted, which is the same work the loop is doing. What you do know is when to stop: when n reaches zero. That is a condition, not a count, and a condition at the top of a loop is a while.

Both compile. Both work. But MU asks for the while loop, and she asks for it because it is the honest shape for this problem.

The program

#include <stdio.h>

int main(void)
{
    int n, original, rev = 0;

    printf("Enter a number: ");
    scanf("%d", &n);
    original = n;

    while (n > 0) {
        rev = rev * 10 + n % 10;
        n = n / 10;
    }

    printf("%d reversed is %d\n", original, rev);
    return 0;
}
munotes.in34

Practical 3(a): Reversing the Digits of a Number

4271
Enter a number: 4271 reversed is 1724

original exists only so that the number can be printed at the end. The loop destroys n as it works, and by the time it finishes n is zero. Keeping a copy before a loop consumes a value is a habit worth forming now.

Three inputs that show what the program really does

#include <stdio.h>

int main(void)
{
    int cases[] = {4271, 1200, 7};

    for (int i = 0; i < 3; i++) {
        int n = cases[i], rev = 0;

        while (n > 0) {
            rev = rev * 10 + n % 10;
            n = n / 10;
        }

        printf("%d reversed is %d\n", cases[i], rev);
    }

    return 0;
}
4271 reversed is 1724
1200 reversed is 21
7 reversed is 7

The middle one is worth a paragraph. 1200 reversed should read 0021, and the program prints 21. Nothing has gone wrong: 0021 and 21 are the same number, and a number has no leading zeros to print. If the zeros must appear, the answer is not a number at all but a string of characters, and it is handled with the string techniques of [Practical 6(a): Extracting Part of a String]. Saying that in a viva is worth more than the program.

Negative numbers, and what this program does with them

while (n > 0) is false immediately for a negative number, so the loop never runs and the answer is 0. That is a real limitation, and there are two honest things to do about it.

Say so, or handle it:

int sign = (n < 0) ? -1 : 1;
n = n * sign;              /* work with the positive value */
... the loop as before ...
rev = rev * sign;          /* put the sign back */

Either is acceptable in a first-semester practical. What is not acceptable is a program that prints 0 for minus 4271 and a student who does not know why.

Two more things this loop can do

Once you can peel digits off a number, three of the standard first-year programs are the same loop with a different body.

ProgramBody of the loop
Reverse the digitsrev = rev * 10 + n % 10
Sum of the digitssum = sum + n % 10
Count the digitscount = count + 1

All three end with n = n / 10 and all three stop at n > 0. A palindrome number check is then the reverse program plus one comparison, which is why examiners set them together.

munotes.in35

Practical 3(a): Reversing the Digits of a Number

What beginners get wrong

Writing rev = rev + n % 10. That gives the sum of the digits, not the reverse. The multiply by ten is what makes the position.

Forgetting n = n / 10. The loop then peels the same last digit forever and the program hangs. If a program never finishes, look for the statement that was supposed to move the loop on.

Not saving the original. After the loop n is 0, so a program that prints n at the end prints 0.

Using % on a float. It does not compile. This program is about whole numbers throughout.

Declaring rev without setting it to 0. An uninitialised variable holds whatever was in that memory, so the answer has a random number added to the front of it. Initialise every accumulator.

Quick revision

  • n % 10 is the last digit; n / 10 is the number without it.
  • rev = rev * 10 + n % 10 appends a digit to the answer.
  • Loop while n > 0; it stops when there is nothing left to peel.
  • Set rev to 0 before the loop, and keep a copy of the original.
  • while when you know the stopping condition, for when you know the count.
  • Reversing 1200 gives 21, because a number has no leading zeros.
  • The same loop gives the sum of digits and the count of digits.

What goes in your journal

Aim, the eight-step algorithm, the flowchart with the loop arrow returning above the test, the program, and the output. Write out the trace table from the top of this chapter for your own number: four rows showing n, the digit, and rev at each step. It is the clearest evidence that you know what the loop does, and it takes two minutes.

Test yourself

1. What does n % 10 give, and what does n / 10 give? The last digit, and the number with its last digit removed.

2. Why is rev multiplied by 10 before the digit is added? To move the digits already collected one place to the left, leaving the units place free for the new digit.

3. What does the program print for 1200, and is that wrong? 21, and it is not wrong. 0021 and 21 are the same number, and leading zeros cannot be printed from an integer.

4. Your program never stops. What is the likeliest cause? n = n / 10 is missing, so the condition n > 0 never becomes false.

5. Why a while loop rather than a for loop here? Because the number of repetitions is not known in advance. What is known is the stopping condition, and a condition tested at the top of a loop is a while.

munotes.in36

Practical 3(a): Reversing the Digits of a Number

6. How would you turn this into a sum of digits program? Replace the body with sum = sum + n % 10, keeping the same condition and the same n = n / 10.

Contents This chapter on its own page

munotes.in37

Chapter Eleven

Practical 3(b): The Factorial of a Number

Syllabus topic Module 1, Practical 3(b): "Write a program to calculate the factorial of a given number."

Aim

To calculate the factorial of a given number.

What a factorial is

The factorial of a whole number n, written n with an exclamation mark after it, is the product of every whole number from 1 up to n.

5! = 1 2 3 4 5 = 120

Two facts about it are asked in vivas and are easy to forget.

0! is 1, not 0. It is defined that way because the factorial counts the number of ways of arranging n things in order, and there is exactly one way to arrange nothing. Every formula that uses factorials, the combinations formula among them, breaks if 0! is anything else.

Factorial is not defined for negative numbers. A program asked for the factorial of minus 5 must say so, not print 1.

Algorithm

1. Start

2. Read n

3. If n is less than 0, print that the factorial is not defined and go to step 8

4. Set fact to 1

5. For i from 1 to n, repeat step 6

6. Set fact to fact into i

7. Print fact

8. Stop

Step 4 sets fact to 1 and not to 0. The loop multiplies, and anything multiplied by zero is zero, so an accumulator that is going to be multiplied starts at 1. An accumulator that is going to be added to starts at 0. Getting that backwards is the single commonest mistake in this practical.

The program

#include <stdio.h>

int main(void)
{
    int n, i;
    unsigned long long fact = 1;

    printf("Enter a number: ");
    scanf("%d", &n);

    if (n < 0) {
        printf("Factorial is not defined for a negative number\n");
    } else {
        for (i = 1; i <= n; i++)
            fact = fact * i;

        printf("%d! = %llu\n", n, fact);
    }

    return 0;
}
6
Enter a number: 6! = 720

Notice that the loop runs from 1 to n inclusive, and that when n is 0 it does not run at all, leaving fact at its starting value of 1, which is the right answer. The program handles 0 correctly without a special case, and that is worth saying in the journal.

The types, which this practical is really about

unsigned long long is the widest whole number type C guarantees, and the standard requires it to hold at least 18446744073709551615. Its printf specifier is %llu.

That is not decoration. Factorials grow faster than almost anything else a first-year program computes, and an int runs out very quickly.

TypeLargest value it must holdLargest factorial it can hold
int32767 by the standard, 2147483647 in practice12! = 479001600
unsigned int4294967295 in practice12!
long long922337203685477580720!
unsigned long long1844674407370955161520!
munotes.in38

Practical 3(b): The Factorial of a Number

So int fails at 13 and unsigned long long fails at 21. There is no C type that holds 21!, and a program that needs it must store the answer differently, in an array of digits.

The overflow, run

#include <stdio.h>

int main(void)
{
    unsigned int small = 1;
    unsigned long long big = 1;

    for (unsigned int i = 1; i <= 13; i++) {
        small = small * i;
        big = big * i;
    }

    printf("13! in an unsigned int:       %u\n", small);
    printf("13! in an unsigned long long: %llu\n", big);

    return 0;
}
13! in an unsigned int:       1932053504
13! in an unsigned long long: 6227020800

The true value is 6227020800. An unsigned int on this machine holds 32 bits, so it can count to 4294967295 and then starts again at zero. 6227020800 minus 4294967296 is 1932053504, which is exactly what it printed.

Two things follow, and the second is the one that matters.

Unsigned arithmetic wraps, and the standard says so. An unsigned value that grows too large is reduced modulo one more than the largest value it can hold. It is defined, predictable, and wrong for this purpose.

Signed overflow is not defined at all. If small had been a plain int, the C standard would place no requirement on what happens: any value, or none. That is why this demonstration is written in unsigned arithmetic. Never write a program that relies on a signed integer overflowing.

Guarding the input

A program that will be marked should refuse what it cannot compute rather than printing a wrapped value:

if (n < 0)
    printf("Factorial is not defined for a negative number\n");
else if (n > 20)
    printf("%d! is too large for this program\n", n);
else
    ... the loop ...

Three cases, all of them answered. That is a complete program; the one that prints 1932053504 for 13 without comment is not.

for against while, for this loop

for (i = 1; i <= n; i++)while
Suitsa known number of repetitionsa known stopping condition
Start, test and stepall three in the loop head, visible at oncescattered across three places
This practicalnatural: exactly n multiplicationsworks, but the count has to be managed by hand
The previous practicalthe digit count is unknown, so it does not fitnatural

Both loops can do either job. The question is which one lets a reader see what is happening, and in a counted loop it is the for.

What beginners get wrong

Starting fact at 0. Every product is then 0.

Writing for (i = 1; i < n; i++). One multiplication short: 5! comes out as 24.

Using int and believing 13!. The value wraps, and with a signed int the behaviour is not even defined.

munotes.in39

Practical 3(b): The Factorial of a Number

Printing an unsigned long long with %d. The wrong specifier prints rubbish. Use %llu.

Returning 1 for a negative input. The factorial of a negative number does not exist; say so.

Forgetting that 0! is 1. The loop gets this right by itself; a student answering a viva often does not.

Quick revision

  • n! = 1 2 ... * n. 0! is 1. Negative factorials do not exist.
  • A multiplying accumulator starts at 1; an adding accumulator starts at 0.
  • for (i = 1; i <= n; i++) fact = fact * i;
  • int holds up to 12!, unsigned long long up to 20!. Nothing in C holds 21!.
  • %llu prints an unsigned long long.
  • Unsigned overflow wraps and is defined; signed overflow is undefined behaviour.
  • 5! is 120, 6! is 720, 10! is 3628800, 13! is 6227020800.

What goes in your journal

Aim, the eight-step algorithm, a flowchart with the loop and the negative-input decision, the program, and the output. Run it three times: a normal value, 0, and a negative number, and write all three outputs down. Add one line to the conclusion about the limit of the type you used, with the value at which it fails. That one line is what separates this write-up from every other one in the class.

Test yourself

1. What is 0!, and why?

  1. There is exactly one way to arrange nothing, and every formula built on factorials requires it.

2. Why is fact initialised to 1 and not to 0? Because the loop multiplies. Starting at 0 makes every product 0.

3. What is the largest factorial an int can hold on a typical machine? 12!, which is 479001600. 13! is 6227020800 and does not fit in 32 bits.

4. What does the C standard say about an int that overflows? Nothing. Signed integer overflow is undefined behaviour, so the program may print anything. Unsigned overflow, by contrast, is defined to wrap.

5. Which format specifier prints an unsigned long long? %llu.

6. Your 5! came out as 24. What is wrong? The loop condition is i < n instead of i <= n, so the final multiplication by 5 never happens.

Contents This chapter on its own page

munotes.in40

Chapter Twelve

Practical 3(c): The Fibonacci Series

Syllabus topic Module 1, Practical 3(c): "Write a program to print the Fibonacci series."

Aim

To print the Fibonacci series.

What the series is

Each term is the sum of the two before it. The first two terms are given, and everything after that follows.

0, 1, 1, 2, 3, 5, 8, 13, 21, 34, 55, ...

Written as a rule, with the terms numbered from 0:

F(0) = 0

F(1) = 1

F(n) = F(n - 1) + F(n - 2) for n >= 2

Check the rule against the list. F(2) is 1 plus 0, which is 1. F(3) is 1 plus 1, which is 2. F(4) is 2 plus 1, which is 3.

Where the series starts, and why it matters in an examination

Some books begin 0, 1, 1, 2 and some begin 1, 1, 2, 3. Both are in print and neither is wrong. What is wrong is a program whose output does not match the definition the student wrote above it.

So do this: write the first two terms in your algorithm as values, print the series, and make sure the printed series begins with those two values. Then whichever convention your examiner has in mind, your work is internally consistent, which is what is actually being marked. This book starts at 0, because that is the form in which the rule F(n) = F(n-1) + F(n-2) needs no adjustment.

The technique: three variables and a shuffle

You never need the whole series in memory. At any moment the program is holding just two numbers, the last two terms, and it makes the next one from them.

Stepabnext = a + bprinted
start010, 1
11111
21222
32333
43555

After each term is printed, a takes the old value of b and b takes the new term. The order of those two assignments matters: overwrite a first, because once b has been overwritten the old value of b is gone.

Algorithm

1. Start

2. Read n, the number of terms

3. Set a to 0 and b to 1

4. Print a and b

5. For i from 3 to n, repeat steps 6, 7 and 8

6. Set next to a plus b

7. Print next

8. Set a to b, then set b to next

9. Stop

The program

#include <stdio.h>

int main(void)
{
    int n, i, a = 0, b = 1, next;

    printf("How many terms? ");
    scanf("%d", &n);

    printf("Fibonacci series: ");

    if (n >= 1)
        printf("%d ", a);
    if (n >= 2)
        printf("%d ", b);

    for (i = 3; i <= n; i++) {
        next = a + b;
        printf("%d ", next);
        a = b;
        b = next;
    }

    printf("\n");
    return 0;
}
munotes.in41

Practical 3(c): The Fibonacci Series

10
How many terms? Fibonacci series: 0 1 1 2 3 5 8 13 21 34

The two if statements before the loop are not padding. Asked for one term, the program must print 0 and nothing else; asked for two, 0 and 1. Without those guards a request for one term prints two, which is a wrong answer to the question that was asked.

The same thing without the shuffle

There is a second form, a line shorter, that some colleges teach:

#include <stdio.h>

int main(void)
{
    int a = 0, b = 1, next, i;

    for (i = 1; i <= 10; i++) {
        printf("%d ", a);
        next = a + b;
        a = b;
        b = next;
    }

    printf("\n");
    return 0;
}
0 1 1 2 3 5 8 13 21 34

Here every term is printed at the top of the loop, so there are no special cases for one term and two, and the loop simply runs n times. It is the neater program. The first version is shown as well because it is the one that matches the algorithm most colleges expect in the journal, step for step.

What beginners get wrong

Writing a = b; b = next; in the wrong order, or using one variable too few. With only a and b and no next, assigning a = a + b destroys the value of a that the next line needs.

Starting the loop at 1 when the first two terms have already been printed. Then twelve terms come out when ten were asked for. If the first two are printed before the loop, the loop starts at 3.

Printing the series but never the term count asked for. Read n and honour it.

Assuming int is enough. F(47) is 2971215073, which is past the largest int. For a first-semester practical of ten or twenty terms this never bites, but it is a fair viva question, and the answer is to use long long.

Confusing this with factorial. Both are classic loops; one adds the two previous values, the other multiplies by a counter.

Quick revision

  • Each term is the sum of the two before it. F(0) is 0, F(1) is 1.
  • 0, 1, 1, 2, 3, 5, 8, 13, 21, 34, 55.
  • Keep two variables, a and b, and a third for the new term.
  • After printing, a = b then b = next, in that order.
  • Guard the cases n equal to 1 and n equal to 2 if the first terms are printed before the loop.
  • F(47) overflows a 32-bit int; use long long for long series.
munotes.in42

Practical 3(c): The Fibonacci Series

What goes in your journal

Aim, the nine-step algorithm, the flowchart with the loop and the two assignments in its body, the program, and the output for at least ten terms. Add the trace table from this chapter, four or five rows of a, b and next; it takes a minute and it is the proof that the shuffle is understood.

Test yourself

1. What are the first eight terms? 0, 1, 1, 2, 3, 5, 8, 13.

2. Why is a third variable needed? Because a and b must both be updated from values that the update itself destroys. next holds the new term while a takes the old b.

3. What goes wrong if b = next; is written before a = b;? b is overwritten first, so a is then given the new term instead of the old b, and the series goes wrong from the third term onward.

4. Asked for one term, what should the program print? 0, and nothing else.

5. State the rule for F(n). F(n) = F(n - 1) + F(n - 2) for n of 2 or more, with F(0) = 0 and F(1) = 1.

6. Why can the series not be printed far with an int? Because the terms grow quickly: F(47) is 2971215073, which is larger than a 32-bit signed int can hold.

Contents This chapter on its own page

munotes.in43

Chapter Thirteen

Practical 4(a): The Area of a Square, Using a Function

Syllabus topic Module 1, Practical 4(a): "Write a program to print area of square using function."

Aim

To print the area of a square using a function.

What a function is, and why a program has any

A function is a named piece of a program that does one job. You give it some values, it does its work, and it hands a value back.

main is a function. printf is a function. A program is nothing but functions, and until now yours have had exactly one.

Three reasons to write your own, and only the third is the real one.

It is written once and used many times. The area calculation appears in one place instead of four.

It has a name, so the program reads like a description of itself. area(side) says what is happening; s * s buried inside a printf does not.

It can be reasoned about on its own. You can be sure area is right by looking at three lines, without reading the rest of the program. That is the reason large programs are possible at all.

The six words, kept apart

This vocabulary is asked in vivas constantly, and students mix up the last two.

WordWhat it isIn this program
Declaration, or prototypetells the compiler the name, the return type and the parameter typesfloat area(float side);
Definitionthe declaration plus the body that does the workfloat area(float side) { ... }, with its body
Callthe place where the function is usedarea(s)
Return typethe type of the value handed backfloat
Parameterthe variable named in the definition, which receives the valueside
Argumentthe value actually passed at the calls, or 4.5

The last two in one line: a parameter is written in the function, an argument is written at the call. The argument's value is copied into the parameter when the call happens.

The program

#include <stdio.h>

float area(float side);

int main(void)
{
    float s, a;

    printf("Enter the side of the square: ");
    scanf("%f", &s);

    a = area(s);

    printf("Side = %.2f\n", s);
    printf("Area = %.2f\n", a);
    printf("A square of side 3 has area %.2f\n", area(3));

    return 0;
}

float area(float side)
{
    return side * side;
}
4.5
Enter the side of the square: Side = 4.50
Area = 20.25
A square of side 3 has area 9.00

Check it: 4.5 by 4.5 is 20.25.

Reading the program in the order the compiler does

float area(float side); near the top is the declaration. It ends in a semicolon and has no body. It exists so that when the compiler reaches area(s) inside main, it already knows that area takes one float and gives back a float.

Leave it out and put the definition after main, and C17 refuses to compile: calling a function the compiler has not seen is not allowed in modern C. Put the whole definition above main instead and no declaration is needed, because the definition is also a declaration. Both arrangements are correct; this book declares first and defines last, because that keeps main at the top where a reader looks for it.

munotes.in44

Practical 4(a): The Area of a Square, Using a Function

a = area(s); is the call. The value in s is copied into the parameter side, the body runs, and the value that return produces becomes the value of the expression area(s), which is then assigned to a.

return side * side; computes and hands back. return does two things at once: it supplies the value, and it ends the function immediately. Any statement after it in the same block never runs.

area(3) inside the last printf shows that a call is an expression. It has a value, so it can go anywhere a value can go. The argument here is a constant rather than a variable, which is allowed: the argument is a value, not a place.

What void means, in both positions

void print_line(void);

Before the name, void means the function returns nothing. Such a function is called for what it does, not for what it gives back, and it ends either at its closing brace or at a bare return; with no value.

Inside the brackets, void means the function takes nothing. That is why main is written int main(void).

Empty brackets, float area(), are not the same thing in C as (void), and C23 finally made them mean the same. Write (void) and the question never arises.

Where the parameter lives

The parameter side exists only while the function is running. It is created when the call happens, it holds a copy of the argument, and it is destroyed when the function returns.

#include <stdio.h>

void spoil(float side);

int main(void)
{
    float s = 4.5;

    spoil(s);
    printf("Back in main, s is still %.2f\n", s);

    return 0;
}

void spoil(float side)
{
    side = 99;
    printf("Inside the function, side is %.2f\n", side);
}
Inside the function, side is 99.00
Back in main, s is still 4.50

The function changed its own copy and main never noticed. That is call by value, it is how every C function call works unless you go out of your way, and it is the subject of [Practical 7: Swapping Two Numbers, by Value and by Reference].

What beginners get wrong

Putting a semicolon after the definition's head. float area(float side); followed by a body is a declaration and then a stray block. The definition's head has no semicolon.

munotes.in45

Practical 4(a): The Area of a Square, Using a Function

Forgetting the return type. In C17 a function with no stated return type does not compile.

Declaring the function inside main. It is legal but pointless and confusing. Declare at file scope, above main.

Writing return; in a function that promises a value, or return x; in a void function. Both are errors.

Expecting a changed parameter to reach the caller. It does not. The parameter is a copy.

Calling the function without using its value. area(s); on a line of its own computes the area and throws it away.

Quick revision

  • Declaration: name, return type, parameter types, ends in a semicolon, no body.
  • Definition: the same head with a body, and no semicolon after the head.
  • Call: area(s), an expression with a value.
  • Parameter is in the function; argument is at the call.
  • return supplies the value and ends the function at once.
  • void before the name means no value returned; void in the brackets means no parameters taken.
  • Arguments are passed by value: the function works on a copy.

What goes in your journal

Aim, an algorithm that names the function as a step ("call area with s and store the result"), a flowchart, the program with the declaration, the definition and the call all visible, and the output. In the conclusion define parameter and argument in one sentence each; it is the viva question this practical always carries.

Test yourself

1. What is the difference between a parameter and an argument? The parameter is the variable named in the function definition; the argument is the value supplied at the call and copied into the parameter.

2. Why is a declaration needed when the definition comes after main? Because the compiler reads the file from the top, and in C17 a call to a function it has not yet seen is an error.

3. What does return do, apart from supplying a value? It ends the function immediately. Statements after it in the same block are never reached.

4. What does void mean in int main(void)? That main takes no parameters.

5. A function changed its parameter. Does the caller's variable change? No. The argument's value was copied into the parameter, so the function changed only its own copy.

6. Is area(3) legal when the parameter is a float? Yes. The declaration tells the compiler a float is expected, so the constant 3 is converted to 3.0 before the call.

Contents This chapter on its own page

munotes.in46

Chapter Fourteen

Practical 4(b): A Recursive Function

Syllabus topic Module 1, Practical 4(b): "Write a program using recursive function."

Aim

To write a program using a recursive function.

What recursion is

A function is recursive when it calls itself.

That sounds circular, and it would be, except for one thing: each call is made with a smaller problem than the one it was given, and there is a smallest problem that is answered outright without calling anything. That smallest case is the base case, and the calls that reduce the problem are the recursive case.

Every recursive function has both. A recursive function without a base case does not stop.

Factorial, defined recursively

The loop definition says 5! is 1 times 2 times 3 times 4 times 5. The recursive definition says something shorter:

0! = 1

n! = n * (n - 1)!

Two lines. The first is the base case; the second is the recursive case. Read the second aloud: the factorial of n is n multiplied by the factorial of one less than n. That is the whole program.

The program, printing its own call stack

#include <stdio.h>

unsigned long long fact(int n);

int main(void)
{
    unsigned long long answer = fact(5);

    printf("5! = %llu\n", answer);
    return 0;
}

unsigned long long fact(int n)
{
    unsigned long long result;

    printf("fact(%d) is called\n", n);

    if (n <= 1) {
        printf("fact(%d) returns 1, the base case\n", n);
        return 1;
    }

    result = n * fact(n - 1);
    printf("fact(%d) returns %llu\n", n, result);
    return result;
}
fact(5) is called
fact(4) is called
fact(3) is called
fact(2) is called
fact(1) is called
fact(1) returns 1, the base case
fact(2) returns 2
fact(3) returns 6
fact(4) returns 24
fact(5) returns 120
5! = 120

Read that output twice. It is the best explanation of recursion there is, because the program wrote it about itself.

The first five lines go down. Each call needs the answer to a smaller call before it can do its own multiplication, so it stops and waits. Five calls are now in progress at once, none of them finished.

The sixth line is the base case. fact(1) needs nothing from anybody. It answers 1 and returns.

The last five lines come back up, in the opposite order. Now that fact(1) has answered, fact(2) can finish its multiplication and return 2; then fact(3) can finish, and so on. The calls unwind in exactly the reverse of the order they were made.

What the machine is actually doing

Every call to a function is given its own small area of memory, called a stack frame, holding that call's parameters and local variables. The frames are stacked: the newest one is on top, and a function that returns takes its frame away.

DepthFramenwaiting for
1fact(5)5fact(4)
2fact(4)4fact(3)
3fact(3)3fact(2)
4fact(2)2fact(1)
5fact(1)1nothing, it is the base case
munotes.in47

Practical 4(b): A Recursive Function

Five frames exist at the deepest point, and each holds its own n. That is why fact(3) still knows its n is 3 when control comes back to it after a long detour.

No base case, and what happens

unsigned long long bad(int n)
{
    return n * bad(n - 1);      /* nothing stops this */
}

Called with 5, it calls itself with 4, 3, 2, 1, 0, minus 1, minus 2, forever. Each call adds a frame, the stack runs out of room, and the program is killed by the operating system with a message such as "segmentation fault" or "stack overflow".

This is not shown running, because a crash proves nothing a sentence cannot. What it is worth knowing is the diagnosis: a C program that dies immediately with a segmentation fault and has a recursive function in it has almost always lost its base case, or has a base case the recursion can step past. if (n == 0) is a weaker base case than if (n <= 0), because a call with minus 1 slips straight through it.

Recursion against iteration

RecursionIteration, the loop
The definitionreads like the mathematicsreads like the procedure
Memoryone stack frame per level, all at oncea fixed handful of variables
Speedslower: every level is a function callfaster
Depthlimited by the stack, a few thousand levelsunlimited
Dangera missing base case crashes the programa wrong condition loops forever
Suitsproblems defined in terms of themselves: trees, Towers of Hanoi, quicksortcounting, totalling, scanning

Anything that can be written one way can be written the other. For factorial, the loop is the better program; the recursion is the better explanation. That is the honest answer in a viva, and it is better than "recursion is elegant".

Two more recursive functions worth knowing

#include <stdio.h>

int fib(int n);
int gcd(int a, int b);

int main(void)
{
    for (int i = 0; i < 8; i++)
        printf("%d ", fib(i));
    printf("\n");

    printf("gcd(48, 18) = %d\n", gcd(48, 18));
    return 0;
}

int fib(int n)
{
    if (n < 2)
        return n;
    return fib(n - 1) + fib(n - 2);
}

int gcd(int a, int b)
{
    if (b == 0)
        return a;
    return gcd(b, a % b);
}
0 1 1 2 3 5 8 13
gcd(48, 18) = 6

fib is the definition from [Practical 3(c): The Fibonacci Series] written directly as code, and it is also a warning: it recomputes the same values over and over, so fib(40) takes a noticeable time where the loop version is instant. gcd is Euclid's algorithm, and recursion suits it perfectly because each step really is the same problem on smaller numbers.

munotes.in48

Practical 4(b): A Recursive Function

What beginners get wrong

No base case, or one the recursion steps over. Use <= rather than == where the argument might overshoot.

Forgetting that the function must make progress. fact(n) calling fact(n) is not recursion, it is a hang. Every call must be on a smaller problem.

Expecting a local variable to be shared. Each call has its own. That is the point of the stack frame.

Using recursion where a loop is plainly better, and then being unable to say why in the viva. Say: same result, more memory, slower, but closer to the definition.

Thinking a recursive function needs a loop inside it. It does not. The repetition is the calling.

Quick revision

  • A recursive function calls itself on a smaller problem.
  • It must have a base case, answered without recursion, and the recursive case must make progress towards it.
  • n! = n * (n - 1)! with 0! = 1.
  • Each call gets its own stack frame with its own parameters and locals.
  • The calls go down to the base case, then unwind in reverse order.
  • No base case means a stack overflow and a crash.
  • Recursion costs memory and speed; it buys a definition that reads like the mathematics.

What goes in your journal

Aim, the two-line recursive definition as the algorithm, a flowchart with the decision for the base case and a box that says "call fact(n minus 1)", the program, and the output. Copy the trace, all eleven lines: it is the evidence that you know the order in which the calls return, which is the one thing about recursion a viva always tests.

Test yourself

1. What are the two parts every recursive function must have? A base case, which returns without recursing, and a recursive case, which calls the function on a smaller problem.

2. In what order do the calls of fact(5) finish? fact(1) first, then fact(2), fact(3), fact(4) and fact(5) last. They unwind in the reverse of the order they were made.

3. What is a stack frame? The area of memory given to one call of a function, holding its parameters and local variables. Each call has its own.

4. What happens when the base case is missing? The function calls itself forever, the stack runs out of memory, and the program crashes.

5. Give one problem where recursion is clearly the better choice. The greatest common divisor by Euclid's algorithm, or the Towers of Hanoi, where each step is genuinely the same problem on a smaller input.

munotes.in49

Practical 4(b): A Recursive Function

6. Why is if (n == 0) a weaker base case than if (n <= 0)? Because a call made with a negative number steps straight past it and the recursion never stops.

Contents This chapter on its own page

munotes.in50

Chapter Fifteen

Practical 4(c): sqrt and abs, and the Headers They Need

Syllabus topic Module 1, Practical 4(c): "Write a program to square root, abs() value using function."

Aim

To use the library functions for square root and absolute value.

Library functions, and what a header is

printf was never written by you, and neither is sqrt. They come from the standard library, a collection of functions that every C compiler supplies.

A header file is how your program learns that they exist. It holds the declarations, in exactly the form of [Practical 4(a): The Area of a Square, Using a Function]: the name, the return type and the parameter types, with no body. #include <math.h> copies those declarations into your file before compilation, so the compiler knows that sqrt takes a double and returns one.

The bodies are somewhere else, already compiled, and the linker joins them to your program at the last step.

FunctionHeaderTakesReturns
sqrt(x)math.hdoubledouble, the square root
pow(x, y)math.htwo doublesdouble, x raised to y
fabs(x)math.hdoubledouble, the absolute value
abs(n)stdlib.hintint, the absolute value
labs(n)stdlib.hlonglong

Two headers, and that is the whole difficulty of this practical: abs is in stdlib.h with the general utilities, and fabs is in math.h with the mathematics.

What absolute value means

The absolute value of a number is its size with the sign thrown away. The absolute value of minus 17 is 17, and of 17 is also 17. It answers "how far from zero", which is why it is what you use for a difference that should not be negative.

The program

#include <stdio.h>
#include <stdlib.h>
#include <math.h>

int main(void)
{
    int n = -17;
    double x = -4.7, y = 2.0;

    printf("abs(%d) = %d\n", n, abs(n));
    printf("fabs(%.1f) = %.1f\n", x, fabs(x));
    printf("sqrt(%.1f) = %.4f\n", y, sqrt(y));
    printf("sqrt(144) = %.1f\n", sqrt(144.0));
    printf("pow(2, 10) = %.0f\n", pow(2.0, 10.0));

    return 0;
}
abs(-17) = 17
fabs(-4.7) = 4.7
sqrt(2.0) = 1.4142
sqrt(144) = 12.0
pow(2, 10) = 1024

The square root of 2 is an irrational number, so %.4f shows four places of a value that never ends. sqrt(144.0) is exactly 12, and the .0 is there only because a double is being printed.

The trap: abs on a floating point value

#include <stdio.h>
#include <stdlib.h>

int main(void)
{
    double x = -4.7;

    printf("abs of -4.7 came out as %d\n", abs(x));
    return 0;
}
abs of -4.7 came out as 4

Four, not 4.7. abs takes an int, so minus 4.7 was converted to minus 4 before the function ever saw it, and the fraction was thrown away exactly as in [Practical 1(a): Simple Interest].

The compiler does warn about this, and the warning is only in -Wall. A student compiling without warnings gets 4 and no explanation. That is the single best argument for the flags this book has used since [Your First Program: Writing It, Compiling It and Running It].

munotes.in51

Practical 4(c): sqrt and abs, and the Headers They Need

The rule to remember: abs for whole numbers, fabs for anything with a decimal point. The f is for floating point, not for function.

Writing them yourself

MU's wording is "using function", and a fair reading is either the library's function or your own. Both are worth knowing, and writing your own is the better viva answer because it shows what the library one does.

#include <stdio.h>

int my_abs(int n);
double my_sqrt(double x);

int main(void)
{
    printf("my_abs(-17) = %d\n", my_abs(-17));
    printf("my_sqrt(2) = %.4f\n", my_sqrt(2.0));
    printf("my_sqrt(144) = %.4f\n", my_sqrt(144.0));

    return 0;
}

int my_abs(int n)
{
    return (n < 0) ? -n : n;
}

double my_sqrt(double x)
{
    double guess = x / 2.0;
    int i;

    if (x <= 0)
        return 0;

    for (i = 0; i < 30; i++)
        guess = (guess + x / guess) / 2.0;

    return guess;
}
my_abs(-17) = 17
my_sqrt(2) = 1.4142
my_sqrt(144) = 12.0000

my_abs is the conditional operator from [Practical 1(b): The Greatest of Three Numbers, with the Conditional Operator], doing the only thing absolute value ever does.

my_sqrt is the Babylonian method, which is about four thousand years old. Guess an answer; the true root lies between the guess and x divided by the guess, so the average of those two is a better guess; repeat. It doubles the number of correct digits each time, so thirty rounds is far more than enough for any double.

What beginners get wrong

Including only math.h and calling abs. It is in stdlib.h.

Calling abs on a double. The value is truncated to an int first.

Writing sqrt(-9). The result is nan, and every later calculation that touches it becomes nan too. Test the sign first.

Expecting sqrt to return an int. It returns a double, and printing it with %d prints rubbish.

Using pow(x, 2) for a square. It works, but x * x is faster and exact, where pow goes through floating point and can return 8.999999 for what should be 9.

Forgetting -lm on Linux. The code is right and the linker cannot find the library. See [Practical 2(a): The Roots of a Quadratic Equation].

Quick revision

  • sqrt, pow and fabs are in math.h. abs and labs are in stdlib.h.
  • abs is for whole numbers, fabs for floating point. The f is for float.
  • sqrt takes and returns a double; print it with %f.
  • sqrt of a negative number is nan.
  • Absolute value is the distance from zero: (n < 0) ? -n : n.
  • Prefer x * x to pow(x, 2).
  • A header supplies declarations; the linker supplies the bodies.
munotes.in52

Practical 4(c): sqrt and abs, and the Headers They Need

What goes in your journal

Aim, the algorithm, the program, and the output. Write the two-header table into the conclusion, because "which header is abs in" is asked in more vivas than any other question about this practical. If you have room, add your own my_abs; it is three lines and it demonstrates that you know what the library is doing.

Test yourself

1. Which header declares abs, and which declares fabs? stdlib.h declares abs; math.h declares fabs.

2. What does abs(-4.7) give, and why?

  1. The argument is converted to an int first, so the fractional part is lost before the function runs.

3. What does sqrt return for a negative argument? nan, which stands for not a number. Any later arithmetic involving it also gives nan.

4. Why is x * x better than pow(x, 2)? It is faster and exact, where pow computes through floating point and may return a value very slightly off.

5. What is in a header file? Declarations: names, return types and parameter types. The compiled bodies of the functions are elsewhere and are joined on by the linker.

6. Write a one-line absolute value without using any library. (n < 0) ? -n : n

Contents This chapter on its own page

munotes.in53

Chapter Sixteen

Practical 4(d): The goto Statement

Syllabus topic Module 1, Practical 4(d): "Write a program using goto statement ."

Aim

To write a program using the goto statement.

What it is

goto transfers control to a label elsewhere in the same function. A label is a name followed by a colon, written before a statement.

    goto finish;
    ...
finish:
    printf("done\n");

That is the whole feature. No condition, no return, no bookkeeping: control simply continues from the label.

Three rules bound it, and each is a fair viva question.

The label must be in the same function. You cannot jump from main into another function or out of one. A label's name is known only inside the function that contains it, which also means two functions may both have a label called finish.

A label must be attached to a statement. A label immediately before a closing brace was an error in C17 and was permitted in C23. If your compiler refuses end: }, write end: ; with an empty statement, and it will accept it.

You may not jump into the scope of a variable length array. Jumping forward past an ordinary declaration is allowed, but the variable is then not initialised, which is its own kind of trouble.

The one use that is defensible

Breaking out of nested loops. break leaves one loop, the one it is in, and there is no break 2 in C. So with two loops the choices are a flag variable tested in both conditions, or a goto.

#include <stdio.h>

int main(void)
{
    int i, j, target = 35;

    for (i = 1; i <= 9; i++) {
        for (j = 1; j <= 9; j++) {
            if (i * j == target)
                goto found;
        }
    }

    printf("%d is not in the nine times table\n", target);
    return 0;

found:
    printf("%d = %d * %d\n", target, i, j);
    return 0;
}
35 = 5 * 7

The jump goes forward, out of both loops, to a label at the end of the function. It is easy to follow: there is one label, it is below the code that jumps to it, and it is the exit.

Compare the version without goto:

#include <stdio.h>

int main(void)
{
    int i, j, target = 35, found = 0;

    for (i = 1; i <= 9 && !found; i++) {
        for (j = 1; j <= 9 && !found; j++) {
            if (i * j == target)
                found = 1;
        }
    }

    if (found)
        printf("%d = %d * %d\n", target, i - 1, j - 1);
    else
        printf("%d is not in the nine times table\n", target);

    return 0;
}
35 = 5 * 7

Same answer, and look at the price: an extra variable, two extra conditions, and i - 1 and j - 1 in the printout because both loops incremented once more before their conditions were retested. Those two subtractions are a bug waiting to happen, and they are exactly the kind of bug the goto version cannot have.

munotes.in54

Practical 4(d): The goto Statement

The use that is not defensible

A goto that jumps backwards makes a loop, and this is what gave the statement its reputation.

#include <stdio.h>

int main(void)
{
    int i = 1;

again:
    printf("%d ", i);
    i++;
    if (i <= 5)
        goto again;

    printf("\n");
    return 0;
}
1 2 3 4 5

It works. It is also a for loop with its three parts scattered over four lines and no word anywhere saying that this is a loop at all. A reader has to find the label, find every goto that targets it, and reconstruct the loop in their head.

Now imagine that with three labels and seven jumps in a hundred-line function, which is what real programs looked like before structured programming. The code becomes a tangle that cannot be read top to bottom, and the name for it, from Edsger Dijkstra's 1968 letter to the editor of the Communications of the ACM, is spaghetti code. That letter is why goto is banned by almost every coding standard in use today.

What to say about it in a viva

Not "goto is bad". Say this:

Any program written with goto can be written without it, using only sequence, selection and repetition. This is the structured program theorem, proved by Corrado Bohm and Giuseppe Jacopini in 1966, and it is why the language does not need goto.

C keeps it for the cases where the alternatives are worse: leaving nested loops, and jumping to one block of cleanup code from several error checks inside one function. Both are jumps forward to a single exit.

A backward goto should always be a loop instead. while, do while and for say what they are; a label does not.

goto against break, continue and return

Jumps toLeavesUse it for
breakjust after the enclosing loop or switchone levelending a loop early
continuethe next iteration's test or stepnothingskipping the rest of one iteration
returnthe callerthe whole functionthe answer is known
gotoany label in the same functionas many levels as you likeleaving nested loops, or one shared exit

What beginners get wrong

Jumping into a loop from outside it. The loop's counter is then never initialised.

Jumping forward over a declaration. The variable exists but holds nothing meaningful.

Using goto to make a menu repeat. Use do while, as [Practical 2(b): A Menu Driven Calculator, with switch] does. It is one line shorter and says what it means.

munotes.in55

Practical 4(d): The goto Statement

Reusing one label for two purposes. A label is not a function; control never comes back.

Putting the label at the end with nothing after it. In C17 a label needs a statement; write end: ; if there is nothing to do.

Quick revision

  • goto label; transfers control to label: in the same function.
  • A label is a name and a colon, attached to a statement.
  • It cannot jump between functions.
  • Defensible use: breaking out of nested loops, and a single cleanup exit.
  • A backward goto is a loop written badly; use while, do while or for.
  • Bohm and Jacopini, 1966: sequence, selection and repetition are enough for any program.
  • Dijkstra's 1968 letter is the origin of the objection, and of the phrase spaghetti code.

What goes in your journal

Aim, the algorithm, the program, and the output. Two programs if you have room: the nested loop search with goto and the same search with a flag, because the comparison is the practical's real content. In the conclusion write one sentence on why goto is permitted in C and avoided in practice.

Test yourself

1. What is a label? A name followed by a colon, placed before a statement, marking a destination for goto.

2. Can goto jump from one function into another? No. A label is visible only inside the function that contains it.

3. Give one use of goto that is generally accepted. Leaving two or more nested loops at once, because break leaves only the innermost.

4. What is the difference between break and goto? break leaves exactly one enclosing loop or switch and always lands just after it. goto can leave any number of them and lands wherever the label is.

5. What does the structured program theorem say? That any program can be written using only sequence, selection and repetition, so goto is never strictly necessary. Bohm and Jacopini proved it in 1966.

6. Why is a backward goto worse than a forward one? Because it creates a loop that is not announced as a loop, so a reader must reconstruct the control flow instead of reading it.

Contents This chapter on its own page

munotes.in56

Chapter Seventeen

Practical 5(a): Ten Students' Roll Numbers and Names

Syllabus topic Module 1, Practical 5(a): "Write a program to print rollno and names of 10 students using array."

Aim

To print the roll numbers and names of ten students using an array.

What an array is

An array is a set of variables of the same type, stored one after another in memory, sharing one name and told apart by a number.

int roll[10];

That declares ten int variables at once. They are roll[0], roll[1], and so on up to roll[9]. The number in the square brackets is the subscript or index.

Ten variables, one name. Without arrays, ten students would need ten declarations, ten scanf calls and ten printf calls, and a hundred students would be impossible. With an array, ten and a hundred differ by one digit.

The first subscript is 0, and the last is n minus 1

This is the rule that catches everybody, and it is worth understanding rather than memorising.

The subscript is not "which one" but how far from the start. roll[0] is the element at the beginning, no distance along. roll[3] is three places along, which is the fourth element. So an array of ten has subscripts 0 to 9, and roll[10] is not part of it.

StudentSubscriptWritten
1st0roll[0]
2nd1roll[1]
10th9roll[9]
does not exist10roll[10]

That is also why a loop over an array is written for (i = 0; i < n; i++) and not i <= n. Read the condition as "while i is still inside", and i < n is exactly that.

An array of names is an array of arrays

A name is not one character, it is several, so ten names need ten sets of characters.

char name[10][20];

Read the brackets left to right: ten things, each of which is twenty characters. name[0] is the first student's name, name[0][0] is the first letter of it.

Twenty is a choice, and it must be big enough for the longest name plus one. The extra one is for the terminating null character, written '\0', which C puts at the end of every string to mark where it stops. A nineteen-letter name needs twenty characters of room.

The program

#include <stdio.h>

int main(void)
{
    int roll[10];
    char name[10][20];
    int i, n = 10;

    for (i = 0; i < n; i++) {
        printf("Roll number and name of student %d: ", i + 1);
        scanf("%d %19s", &roll[i], name[i]);
    }

    printf("\nRoll  Name\n");
    printf("----  --------------------\n");

    for (i = 0; i < n; i++)
        printf("%-4d  %s\n", roll[i], name[i]);

    return 0;
}
101 Asha
102 Ravi
103 Meena
104 Imran
105 Neha
106 Vikram
107 Sunita
108 Rahul
109 Farah
110 Kiran
Roll number and name of student 1: Roll number and name of student 2: Roll number and name of student 3: Roll number and name of student 4: Roll number and name of student 5: Roll number and name of student 6: Roll number and name of student 7: Roll number and name of student 8: Roll number and name of student 9: Roll number and name of student 10:
Roll  Name
----  --------------------
101   Asha
102   Ravi
103   Meena
104   Imran
105   Neha
106   Vikram
107   Sunita
108   Rahul
109   Farah
110   Kiran
munotes.in57

Practical 5(a): Ten Students' Roll Numbers and Names

Four details in that program

&roll[i] has an ampersand, and name[i] does not. roll[i] is one int, so scanf needs its address, exactly as in [Practical 1(a): Simple Interest]. But name[i] is already an array of characters, and the name of an array is the address of its first element. Putting an & in front of it is a common mistake; leaving it off the integer is a worse one.

%19s and not %s. The number limits scanf to nineteen characters, leaving room for the terminating null in a space of twenty. Plain %s will happily read a fiftieth character into memory that belongs to something else, which is the defect behind a great many real security failures. Always give %s a width.

%s stops at the first space. Type "Asha Kulkarni" and scanf("%19s") reads "Asha" and leaves " Kulkarni" waiting, where the next scanf("%d") will trip over it. This practical therefore uses single-word names. Reading a name with a space in it needs a different call, and it is taught in [Practical 9: A Structure, and Two Records of It], where MU's own example requires it.

%-4d left-aligns in a field of four. The minus sign means align left; without it the column is right-aligned. It is how a table is made to line up, and it costs nothing.

Arrays are not bounds checked, and nothing warns you

int roll[10];

roll[10] = 999;        /* there is no roll[10] */
roll[-1] = 42;         /* nor a roll[-1] */

Both of those compile. C does not check that a subscript is inside the array, at compile time or at run time. The statements write over whatever bytes happen to sit next to the array, which might be another variable, might be nothing important, and might be part of the machinery that lets the function return.

The C standard calls it undefined behaviour, which means the program may print a wrong answer, may crash, or may appear to work for a year and then fail. So the safety is entirely yours:

Every loop over an array uses i < n, never i <= n. That one habit prevents most of them.

munotes.in58

Practical 5(a): Ten Students' Roll Numbers and Names

How many elements is it, really

An array does not remember how long it is, but the compiler knows, and sizeof asks it.

#include <stdio.h>

int main(void)
{
    int roll[10];
    char name[10][20];

    printf("one int is %zu bytes\n", sizeof(int));
    printf("roll is %zu bytes, so %zu elements\n",
           sizeof roll, sizeof roll / sizeof roll[0]);
    printf("name is %zu bytes: %zu names of %zu characters\n",
           sizeof name, sizeof name / sizeof name[0], sizeof name[0]);

    return 0;
}
one int is 4 bytes
roll is 40 bytes, so 10 elements
name is 200 bytes: 10 names of 20 characters

sizeof array / sizeof array[0] is the standard way of asking how many elements an array has: the total number of bytes divided by the bytes in one element. %zu is the format specifier for the type sizeof produces, and using %d for it is a warning with -Wall.

Two cautions. sizeof(int) is 4 on this machine and on almost every machine you will meet, but the C standard only requires an int to hold values from minus 32767 to 32767, so a very small system may use 2. Never write 4 where sizeof(int) belongs.

And this trick works only where the array itself is visible. Pass an array to a function and what arrives is an address, so sizeof inside the function gives the size of a pointer, not of the array. That is why every function in this book that takes an array also takes its length as a second parameter.

What beginners get wrong

Starting at 1 or ending at n. The subscripts are 0 to n minus 1.

Putting & before an array name in scanf. The array name is already an address.

Leaving & off a single variable in scanf. The value goes nowhere.

Using %s with no width. A long input writes past the end of the array.

Typing a name with a space. %s stops at the space and the rest of the line confuses the next read.

Declaring the array too small for the data. There is no error; the extra values land on something else.

Quick revision

  • int roll[10]; gives ten ints, roll[0] to roll[9].
  • The subscript is the distance from the start, so the first is 0 and the last is n minus 1.
  • char name[10][20]; is ten strings of up to nineteen characters plus the terminating null.
  • A string ends with '\0', and room must be left for it.
  • An array name is the address of its first element, so scanf needs no & for it.
  • %19s limits the read; plain %s does not and is dangerous.
  • %s stops at whitespace.
  • C does not check array bounds. i < n in every loop.
  • sizeof a / sizeof a[0] is the number of elements.
munotes.in59

Practical 5(a): Ten Students' Roll Numbers and Names

What goes in your journal

Aim, an algorithm with two loops, one to read and one to print, a flowchart showing both, the program, and the output table with all ten students. In the conclusion write the two facts this practical exists to teach: that the subscripts run from 0 to n minus 1, and that an array name in scanf needs no ampersand.

Test yourself

1. int a[5]; Which subscripts are valid? 0, 1, 2, 3 and 4. a[5] is past the end.

2. Why does scanf("%d", &roll[i]) need an & when scanf("%19s", name[i]) does not? roll[i] is a single int, so its address must be given. name[i] is an array, and an array name already is the address of its first element.

3. What does char name[10][20]; declare? Ten strings, each with room for twenty characters, which is nineteen letters and the terminating null.

4. What happens if you write to a[10] in an array of ten? Nothing is checked. The write lands on whatever is next in memory, and the behaviour is undefined: the program may work, may print nonsense, or may crash.

5. Why i < n and not i <= n? Because the last valid subscript is n minus 1. i <= n runs one step past the end of the array.

6. How do you find the number of elements in an array? sizeof a / sizeof a[0], the total size divided by the size of one element.

Contents This chapter on its own page

munotes.in60

Chapter Eighteen

Practical 5(b): Sorting an Array

Syllabus topic Module 1, Practical 5(b): "Write a program to sort the elements of array in ascending or descending order"

Aim

To sort the elements of an array in ascending or descending order.

Bubble sort, in one sentence

Walk through the array comparing each pair of neighbours, and swap them when they are in the wrong order. Repeat until a whole walk passes with no swap.

It is called bubble sort because on each walk the largest value left unsorted travels all the way to the right, the way a bubble rises.

Watching one pass

Start with 5, 1, 4, 2, 8 and sort ascending. Compare each neighbouring pair, left to right.

CompareArray beforeSwap?Array after
5 and 15 1 4 2 8yes1 5 4 2 8
5 and 41 5 4 2 8yes1 4 5 2 8
5 and 21 4 5 2 8yes1 4 2 5 8
5 and 81 4 2 5 8no1 4 2 5 8

Four comparisons for five elements, and the largest value, 8, has reached the end. It is now in its final place, so the next pass need only go as far as the second to last element. That is the whole reason the inner loop's limit shrinks by one each pass.

Algorithm

1. Start

2. Read n, then read n elements into the array a

3. Read the choice: 1 for ascending, 2 for descending

4. For i from 0 to n minus 2, repeat steps 5 and 6

5. For j from 0 to n minus i minus 2, repeat step 6

6. If the pair a[j] and a[j+1] is in the wrong order for the chosen direction, swap them

7. Print the array

8. Stop

The swap, and why it takes three statements

temp = a[j];
a[j] = a[j + 1];
a[j + 1] = temp;

Two statements are not enough. Writing a[j] = a[j+1]; first destroys the value in a[j], and the second statement then has nothing to put back. temp holds the value that is about to be overwritten. This is the same reasoning as the third variable in [Practical 3(c): The Fibonacci Series].

The program

#include <stdio.h>

int main(void)
{
    int a[50], n, i, j, temp, choice;

    printf("How many elements? ");
    scanf("%d", &n);

    for (i = 0; i < n; i++) {
        printf("Element %d: ", i + 1);
        scanf("%d", &a[i]);
    }

    printf("1 for ascending, 2 for descending: ");
    scanf("%d", &choice);

    for (i = 0; i < n - 1; i++) {
        for (j = 0; j < n - i - 1; j++) {
            int wrong_order = (choice == 1) ? (a[j] > a[j + 1])
                                            : (a[j] < a[j + 1]);
            if (wrong_order) {
                temp = a[j];
                a[j] = a[j + 1];
                a[j + 1] = temp;
            }
        }
    }

    printf("Sorted: ");
    for (i = 0; i < n; i++)
        printf("%d ", a[i]);
    printf("\n");

    return 0;
}
munotes.in61

Practical 5(b): Sorting an Array

5
5 1 4 2 8
1
How many elements? Element 1: Element 2: Element 3: Element 4: Element 5: 1 for ascending, 2 for descending: Sorted: 1 2 4 5 8

One line does both directions:

wrong_order = (choice == 1) ? (a[j] > a[j + 1]) : (a[j] < a[j + 1]);

Ascending means a pair is wrong when the left is bigger. Descending means it is wrong when the left is smaller. Everything else about the sort is identical, which is worth saying out loud in a viva: the direction of a sort is one comparison.

Every pass, printed

#include <stdio.h>

void show(const char *label, int a[], int n);

int main(void)
{
    int a[] = {5, 1, 4, 2, 8};
    int n = 5, i, j, temp, comparisons = 0, swaps = 0;

    show("start ", a, n);

    for (i = 0; i < n - 1; i++) {
        for (j = 0; j < n - i - 1; j++) {
            comparisons++;
            if (a[j] > a[j + 1]) {
                temp = a[j];
                a[j] = a[j + 1];
                a[j + 1] = temp;
                swaps++;
            }
        }
        printf("pass %d", i + 1);
        show("", a, n);
    }

    printf("%d comparisons, %d swaps\n", comparisons, swaps);
    return 0;
}

void show(const char *label, int a[], int n)
{
    printf("%s: ", label);
    for (int i = 0; i < n; i++)
        printf("%d ", a[i]);
    printf("\n");
}
start : 5 1 4 2 8
pass 1: 1 4 2 5 8
pass 2: 1 2 4 5 8
pass 3: 1 2 4 5 8
pass 4: 1 2 4 5 8
10 comparisons, 4 swaps

Two things to notice. The array was already sorted after pass 2, and passes 3 and 4 did nothing but compare. And the count is what the theory predicts.

How many comparisons

The outer loop runs n minus 1 times. The inner loop runs n minus 1 times, then n minus 2, and so on down to 1.

(n - 1) + (n - 2) + ... + 2 + 1 = n * (n - 1) / 2

For n of 5 that is 5 times 4 divided by 2, which is 10, and the program counted 10. For n of 100 it is 4950, and for n of 1000 it is 499500. The work grows with the square of the number of elements, which is written O(n squared) and is the answer to "what is the time complexity of bubble sort".

nComparisons
510
1045
1004950
1000499500
munotes.in62

Practical 5(b): Sorting an Array

That is why bubble sort is taught and not used. It is the easiest sort to write and to explain, and for anything past a few hundred elements it is far too slow.

The improvement worth knowing

If a whole pass makes no swap, the array is sorted and the remaining passes are wasted. One flag stops them:

for (i = 0; i < n - 1; i++) {
    int swapped = 0;
    for (j = 0; j < n - i - 1; j++)
        if (a[j] > a[j + 1]) { ...swap...; swapped = 1; }
    if (!swapped)
        break;
}

With that, an already sorted array costs one pass instead of n minus 1. It is a fair viva question and it is two lines.

What beginners get wrong

Swapping with two statements. The first value is lost. Use temp.

Writing j < n - 1 instead of j < n - i - 1. It still sorts, because the tail is already in order, but it compares pairs that are known to be right, so the count doubles for nothing.

Writing j <= n - i - 1. Then a[j + 1] is a[n - i], and on the first pass that is a[n], one past the end of the array.

Comparing a[i] with a[i + 1] in the inner loop. The inner loop's counter is j. Mixing them is the commonest typing error in this practical.

Sorting in the wrong direction and swapping the printing instead. Printing an ascending array backwards is not a descending sort, and an examiner will ask you to sort strings next, where the trick collapses.

Quick revision

  • Bubble sort compares neighbours and swaps them when they are in the wrong order.
  • After pass i, the last i elements are in their final places.
  • Outer loop i < n - 1; inner loop j < n - i - 1.
  • A swap needs three statements and a temp.
  • Ascending: swap when a[j] > a[j + 1]. Descending: swap when a[j] < a[j + 1].
  • Comparisons: n * (n - 1) / 2, which is O(n squared).
  • A swapped flag lets an already sorted array finish in one pass.

What goes in your journal

Aim, the algorithm with both loops, a flowchart with the inner loop inside the outer, the program, and two outputs: the same data sorted ascending and sorted descending. Add the pass-by-pass table for five elements and the comparison count. The count is what the viva asks about, and a journal that shows it has answered the question before it is asked.

Test yourself

1. Why does the inner loop's limit shrink on each pass? Because after pass i the largest i elements have reached their final places at the end, so there is no need to compare them again.

munotes.in63

Practical 5(b): Sorting an Array

2. How many comparisons does bubble sort make for 10 elements? 45, which is 10 times 9 divided by 2.

3. What is the time complexity of bubble sort? O(n squared): the work grows with the square of the number of elements.

4. Why does a swap need a third variable? Because assigning one element to the other destroys the value that the second assignment needs.

5. What single change sorts in descending order instead? Reverse the comparison: swap when a[j] < a[j + 1] instead of when a[j] > a[j + 1].

6. How can an already sorted array be detected? Set a flag when any swap happens in a pass. If a whole pass makes no swap, the array is sorted and the loop can stop.

Contents This chapter on its own page

munotes.in64

Chapter Nineteen

Practical 6(a): Extracting Part of a String

Syllabus topic Module 1, Practical 6(a): "Write a program to extract the portion of a character string and print the extracted part."

Aim

To extract a portion of a character string and print the extracted part.

What a string is in C

There is no string type in C. A string is an array of characters with a marker at the end saying where it stops, and that marker is a character whose value is zero, written '\0' and called the null character or the terminating null.

So the word Mumbai occupies seven characters, not six:

Subscript0123456
CharacterMumbai\0

Every library function that works on strings finds the end by looking for that zero. A character array without one is not a string, and passing it to printf("%s") makes the function run off the end of the array printing whatever it finds until it happens on a zero byte.

That is the single most important fact in this chapter and the next two.

Counting from zero, and what the user is counting from

C counts positions from 0. A person asked "where does the portion start" will answer counting from 1.

This book keeps the C convention in the program and says so to the user, because that is the honest choice: a program that quietly subtracts 1 somewhere is a program whose behaviour a student cannot explain. Whichever you choose, say which in the prompt and be consistent, and write it in the journal.

Extracting by hand

#include <stdio.h>

int main(void)
{
    char source[100], part[100];
    int start, length, i;

    printf("Enter a string: ");
    scanf("%99[^\n]", source);

    printf("Start position (counting from 0) and length: ");
    scanf("%d %d", &start, &length);

    for (i = 0; i < length && source[start + i] != '\0'; i++)
        part[i] = source[start + i];

    part[i] = '\0';

    printf("Source: %s\n", source);
    printf("Extracted: %s\n", part);
    printf("Length extracted: %d\n", i);

    return 0;
}
Mumbai University
7 10
Enter a string: Start position (counting from 0) and length: Source: Mumbai University
Extracted: University
Length extracted: 10

Position 7 is the U, because M is 0, u is 1, m is 2, b is 3, a is 4, i is 5, the space is 6, and U is 7. Ten characters from there is the whole of University.

The four things that program does

scanf("%99[^\n]", source) reads a whole line, spaces included. The square brackets are a scan set: [^\n] means "any characters except a newline", so the read stops at the end of the line rather than at the first space. The 99 limits it to ninety-nine characters, leaving the hundredth for the terminating null. This is how you read a line with scanf, and it is what [Practical 5(a): Ten Students' Roll Numbers and Names] promised would come later.

munotes.in65

Practical 6(a): Extracting Part of a String

The loop copies character by character. source[start + i] is the character i places past the starting point.

The condition has two halves. i < length stops when enough characters have been taken. source[start + i] != '\0' stops at the end of the source, which is what saves the program when the user asks for more characters than there are.

part[i] = '\0'; is the line students forget. Without it, part is a character array with no end marker, and printf("%s", part) prints the extracted text followed by whatever rubbish lies after it in memory. The loop leaves i holding the number of characters copied, which is exactly the subscript the null belongs in.

The same thing with the library

string.h has a function for it, and MU's reference list names the library, so it is worth knowing both.

#include <stdio.h>
#include <string.h>

int main(void)
{
    char source[] = "Mumbai University";
    char part[20];

    strncpy(part, source + 7, 10);
    part[10] = '\0';

    printf("Extracted: %s\n", part);
    printf("Its length is %zu\n", strlen(part));

    return 0;
}
Extracted: University
Its length is 10

source + 7 is the address of the eighth character. Adding a number to an array name moves along it, which is the same arithmetic the subscript was doing.

strncpy copies at most n characters. The n is a safety limit, and it is why strncpy is preferred to strcpy, which copies until it meets a null and will happily run past the end of the destination.

part[10] = '\0'; is still needed, and this is the trap. strncpy copies the terminating null only if it fits within the n characters. Here ten characters were asked for and ten were copied, so no null was copied with them, and without that line part has no end. It is the same forgotten line as in the loop version, wearing a library's clothes.

The rule: after strncpy, put the null in yourself. It costs one line and it removes a whole class of bug.

The functions this practical touches

FunctionWhat it doesWatch out for
strlen(s)how many characters before the nulldoes not count the null itself
strcpy(d, s)copies s into d, null and allno limit; overruns d if s is longer
strncpy(d, s, n)copies at most n charactersmay leave no null; add one
strcat(d, s)joins s onto the end of dd must have room for both
strcmp(a, b)comparesreturns an ordering, not a truth value

All of them need #include <string.h>, and all of them are taught where MU sets them: strlen and strcmp in [Practical 6(c): strlen and strcmp].

What beginners get wrong

Forgetting the terminating null in the extracted part. The output has rubbish after it.

munotes.in66

Practical 6(a): Extracting Part of a String

Using %s in scanf for a string with a space. It stops at the space. Use %99[^\n].

Counting positions from 1 in the program and from 0 in the head, or the other way round. Choose one, say which in the prompt, and keep to it.

Asking for more characters than the string has. Without the second half of the loop condition, the copy runs past the end of the source.

Thinking strlen counts the null. It does not. strlen("Mumbai") is 6, and the array needs 7.

Declaring the destination too small. char part[5] with ten characters copied into it writes over its neighbours, and nothing warns you.

Quick revision

  • A string is a character array ending in '\0'.
  • strlen counts characters before the null; the array needs one more.
  • scanf("%99[^\n]", s) reads a whole line including spaces.
  • Extract by hand: copy source[start + i] for i from 0, then set part[i] = '\0'.
  • source + 7 is the address of the character at subscript 7.
  • strncpy(d, s, n) copies at most n characters and may leave no null; add it yourself.
  • The loop condition needs both i < length and a test for the end of the source.

What goes in your journal

Aim, the algorithm with the copy loop and the null written as its own step, a flowchart, the program, and the output for at least two runs: a portion from the middle, and a request for more characters than the string holds. In the conclusion write the one sentence this practical exists for: a string is a character array that ends in a null, and a copy that does not add one is not a string.

Test yourself

1. How many bytes does the string "Mumbai" occupy? Seven: six letters and the terminating null.

2. What does part[i] = '\0'; do, and what happens without it? It marks the end of the extracted string. Without it, printf("%s", part) keeps printing past the end until it meets a zero byte somewhere in memory.

3. Why is scanf("%s", s) wrong for reading a full name? It stops at the first whitespace character, so only the first word is read.

4. What is the danger in strncpy? It copies the terminating null only if the null falls within the n characters. Copying exactly n characters leaves the destination with no end marker.

5. What does source + 7 mean? The address of the character at subscript 7, that is, the eighth character of the string.

6. strlen("Mumbai University") gives what?

  1. Sixteen letters and one space, and the null is not counted.

Contents This chapter on its own page

munotes.in67

Chapter Twenty

Practical 6(b): Is This String a Palindrome?

Syllabus topic Module 1, Practical 6(b): "Write a program to find the given string is palindrome or not."

Aim

To check whether a given string is a palindrome.

What a palindrome is

A palindrome is a word, a number or a phrase that reads the same forwards and backwards. madam, level, radar, 121.

The test follows straight from the definition: the first character must equal the last, the second must equal the second from last, and so on until the two ends meet in the middle.

The technique: two subscripts, walking inward

Take madam, whose characters sit at subscripts 0 to 4.

Stepijs[i]s[j]Same?
104mmyes
213aayes
322i is no longer less than j, so stop

Two comparisons for five characters, and then the subscripts meet. For a word of even length, say abba, they cross without meeting: i is 2 and j is 1, and i < j is false. Either way the condition i < j is what stops the loop, and it is right for both.

Notice there is no need to go all the way. Comparing the first with the last already checks both ends, so the loop does about half as many comparisons as there are characters.

Algorithm

1. Start

2. Read the string s

3. Set i to 0 and j to the length of s minus 1

4. Set flag to 1

5. While i is less than j, repeat steps 6 and 7

6. If s[i] is not equal to s[j], set flag to 0 and leave the loop

7. Increase i by 1 and decrease j by 1

8. If flag is 1, print that it is a palindrome, otherwise print that it is not

9. Stop

The program

#include <stdio.h>
#include <string.h>

int main(void)
{
    char s[100];
    int i, j, flag = 1;

    printf("Enter a string: ");
    scanf("%99[^\n]", s);

    i = 0;
    j = strlen(s) - 1;

    while (i < j) {
        if (s[i] != s[j]) {
            flag = 0;
            break;
        }
        i++;
        j--;
    }

    if (flag)
        printf("%s is a palindrome\n", s);
    else
        printf("%s is not a palindrome\n", s);

    return 0;
}
madam
Enter a string: madam is a palindrome

strlen(s) - 1 is the subscript of the last character, not of the null. strlen("madam") is 5, and the last letter is at subscript 4.

The break leaves the loop the moment a mismatch is found. There is no point comparing the rest; one difference settles it.

Several strings, tested

#include <stdio.h>
#include <string.h>

int is_palindrome(const char *s);

int main(void)
{
    const char *tests[] = {"madam", "level", "abba", "mumbai", "a", ""};

    for (int k = 0; k < 6; k++)
        printf("%-8s %s\n", tests[k][0] ? tests[k] : "(empty)",
               is_palindrome(tests[k]) ? "palindrome" : "not a palindrome");

    return 0;
}

int is_palindrome(const char *s)
{
    int i = 0, j = (int) strlen(s) - 1;

    while (i < j) {
        if (s[i] != s[j])
            return 0;
        i++;
        j--;
    }

    return 1;
}
munotes.in68

Practical 6(b): Is This String a Palindrome?

madam    palindrome
level    palindrome
abba     palindrome
mumbai   not a palindrome
a        palindrome
(empty)  palindrome

The last two rows are the edge cases an examiner reaches for, and both come out as palindromes.

A single character: i is 0 and j is 0, so i < j is false, the loop never runs, and the flag is still 1. That is the right answer, and the program gets it without a special case.

The empty string: strlen("") is 0, so j is minus 1, i < j is false at once, and the answer is again yes. This one is a convention rather than a discovery. A string with no characters reads the same in both directions because there is nothing to read, so it is a palindrome vacuously, in the same way that 0! is 1. If your examiner wants an empty input refused instead, that is one if before the loop; what matters is that you know which answer your program gives and why.

Spaces and capital letters: the decision this chapter makes

Is Madam a palindrome? Compared exactly, no: M and m are different characters, because a capital M is 77 and a small m is 109. Is nurses run a palindrome? Compared exactly, no, because of the space.

There is no single right answer, and a student who has not decided will be caught out.

This book compares exactly, as typed. That is what MU's wording asks for: the given string, as given. A program that silently strips spaces is answering a different question from the one it was asked.

But the other version is two if statements away, and it is worth knowing:

#include <stdio.h>
#include <string.h>
#include <ctype.h>

int main(void)
{
    char s[] = "Madam, I'm Adam";
    int i = 0, j = (int) strlen(s) - 1, flag = 1;

    while (i < j) {
        if (!isalpha((unsigned char) s[i])) { i++; continue; }
        if (!isalpha((unsigned char) s[j])) { j--; continue; }

        if (tolower((unsigned char) s[i]) != tolower((unsigned char) s[j])) {
            flag = 0;
            break;
        }
        i++;
        j--;
    }

    printf("\"%s\" is %sa palindrome when letters alone are compared\n",
           s, flag ? "" : "not ");

    return 0;
}
"Madam, I'm Adam" is a palindrome when letters alone are compared

isalpha and tolower come from ctype.h. The cast to unsigned char is not decoration: those functions are defined for values of an unsigned char and for the end-of-file marker, and passing a plain char that happens to be negative is undefined behaviour. It costs nothing and it is correct.

munotes.in69

Practical 6(b): Is This String a Palindrome?

The number palindrome, which is a different program

An examiner may ask for a palindrome number rather than a string. That is not this program with %d: it is the reverse-digits loop from [Practical 3(a): Reversing the Digits of a Number] followed by one comparison.

reverse the digits of n into rev
if (rev == original) it is a palindrome

Both are worth having in the journal, because the two questions sound identical and the programs share nothing.

What beginners get wrong

Setting j to strlen(s). That subscript holds the null character, so the first comparison is against '\0' and every string comes out as not a palindrome.

Looping while i <= j. At the middle character of an odd-length word it compares that character with itself, which is harmless but wasteful, and on an even-length word it compares a pair that has already been compared.

Comparing the whole string with its reverse using ==. In C, == on two arrays compares addresses, not contents, and two different arrays never have the same address. Use strcmp, which is [Practical 6(c): strlen and strcmp].

Forgetting to reset the flag between two tests when the program checks several strings in one run.

Assuming Madam is a palindrome. Not when compared exactly. Decide, and say so.

Quick revision

  • A palindrome reads the same forwards and backwards.
  • Two subscripts: i from 0, j from strlen(s) - 1, loop while i < j.
  • Compare s[i] with s[j], then i++ and j--.
  • One mismatch settles it; leave the loop at once.
  • strlen(s) - 1 is the last character; strlen(s) is the null.
  • A single character is a palindrome.
  • Capitals and spaces matter unless you deliberately ignore them with tolower and isalpha from ctype.h.
  • A palindrome number is the reverse-digits loop plus one comparison, not this program.

What goes in your journal

Aim, the nine-step algorithm, a flowchart with the loop and the mismatch decision, the program, and two runs: one palindrome and one that is not. Write the walking table for your own word, three or four rows of i, j and the two characters. In the conclusion state what your program does about capitals and spaces; that is the question you will be asked.

Test yourself

1. Why is j set to strlen(s) - 1? Because strlen counts the characters and the subscripts start at 0, so the last character sits one place before the count. strlen(s) is the subscript of the terminating null.

2. Why i < j rather than i <= j? Because when i and j meet, that character is the middle one and comparing it with itself proves nothing. All the pairs have already been checked.

munotes.in70

Practical 6(b): Is This String a Palindrome?

3. Is Madam a palindrome? Not when the characters are compared exactly, because a capital M and a small m are different characters. It is if the comparison is made case insensitive with tolower.

4. How many comparisons does a word of ten characters need? Five. Each comparison settles a pair, and ten characters make five pairs.

5. Why can you not test s == reverse with ==? Because == on arrays compares their addresses, which are always different for two separate arrays. Contents are compared with strcmp.

6. How would you check whether a number is a palindrome? Reverse its digits with the loop from the reversing practical, and compare the reversed value with a saved copy of the original.

Contents This chapter on its own page

munotes.in71

Chapter Twenty-One

Practical 6(c): strlen and strcmp

Syllabus topic Module 1, Practical 6(c): "Write a program to using strlen(), strcmp() function ."

Aim

To use the strlen() and strcmp() functions.

Both come from one header

#include <string.h>

Without it the compiler does not know these functions exist, and in C17 that is an error, not a warning.

strlen: how long is this string

strlen(s) counts the characters in s up to but not including the terminating null.

#include <stdio.h>
#include <string.h>

int main(void)
{
    char a[20] = "Mumbai";
    char b[] = "";
    char c[] = "BSc IT";

    printf("strlen(\"%s\") = %zu\n", a, strlen(a));
    printf("strlen(\"%s\") = %zu  (the empty string)\n", b, strlen(b));
    printf("strlen(\"%s\") = %zu  (the space counts)\n", c, strlen(c));

    printf("sizeof a = %zu, strlen(a) = %zu\n", sizeof a, strlen(a));

    return 0;
}
strlen("Mumbai") = 6
strlen("") = 0  (the empty string)
strlen("BSc IT") = 6  (the space counts)
sizeof a = 20, strlen(a) = 6

The last line is the distinction this function is always asked about.

sizeof astrlen(a)
Answershow big is the boxhow much is in it
Counts the nullyesno
Decidedwhen the program is compiledwhen the function runs
Is it a functionno, an operatoryes
For char a[20] = "Mumbai"206

strlen returns a size_t, an unsigned type, whose format specifier is %zu. Two consequences follow, and the second bites.

Print it with %zu, not %d. With -Wall the compiler says so.

strlen(s) - 1 on an empty string is not minus 1. It is an enormous positive number, because an unsigned type cannot be negative and subtracting one from zero wraps to the largest value. A loop written for (i = strlen(s) - 1; i >= 0; i--) with i unsigned never ends. Cast to int when you need to go backwards, as [Practical 6(b): Is This String a Palindrome?] does.

strcmp: which of these two comes first

strcmp(a, b) compares two strings and returns a number that tells you their order:

Return valueMeaning
0the strings are identical
negativea comes before b
positivea comes after b

The comparison is character by character from the left. At the first pair that differs, the character codes decide, and the rest of both strings is irrelevant.

#include <stdio.h>
#include <string.h>

int main(void)
{
    printf("strcmp(\"apple\", \"apple\")  = %d\n", strcmp("apple", "apple"));
    printf("strcmp(\"apple\", \"banana\") = %d\n", strcmp("apple", "banana"));
    printf("strcmp(\"banana\", \"apple\") = %d\n", strcmp("banana", "apple"));
    printf("strcmp(\"Apple\", \"apple\")  = %d\n", strcmp("Apple", "apple"));
    printf("strcmp(\"app\", \"apple\")    = %d\n", strcmp("app", "apple"));

    return 0;
}
strcmp("apple", "apple")  = 0
strcmp("apple", "banana") = -1
strcmp("banana", "apple") = 1
strcmp("Apple", "apple")  = -1
strcmp("app", "apple")    = -1

Read those five lines carefully, because each one answers a viva question.

Identical strings give 0. Not 1. This is the whole trap and it has its own section below.

munotes.in72

Practical 6(c): strlen and strcmp

apple against banana is negative, because at the first character a is 97 and b is 98.

Apple against apple is negative, because a capital A is 65 and a small a is 97. Every capital letter comes before every small letter in ASCII, which is why an ordinary alphabetical sort puts Zebra before apple.

app against apple is negative. The first three characters match; then app has its terminating null, which is 0, against l, which is 108. A string that is a prefix of another always comes first.

The standard promises only the sign of the value, not its size. This library returns minus 1 and 1; another may return the difference of the character codes, minus 1 and 1 or minus 25 and 25. Test the sign, never the value. A program that says if (strcmp(a, b) == -1) is relying on something no standard guarantees.

The trap: 0 means equal

#include <stdio.h>
#include <string.h>

int main(void)
{
    char a[] = "ravi", b[] = "ravi";

    if (strcmp(a, b))
        printf("WRONG: the program thinks they are different\n");
    else
        printf("RIGHT: they are the same\n");

    if (strcmp(a, b) == 0)
        printf("The correct test: they are the same\n");

    return 0;
}
RIGHT: they are the same
The correct test: they are the same

The first if gives the right answer here for the wrong reason, and that is why it is dangerous. In C, a condition is true when it is not zero. strcmp returns 0 for equal strings, so if (strcmp(a, b)) is true exactly when the strings are different. Writing it that way and reading it as "if a equals b" gets the answer backwards on every pair that differs, and right on every pair that matches, which is the worst possible failure pattern for testing.

Always write it out: if (strcmp(a, b) == 0) for equal, != 0 for different.

And the other half of the same lesson:

if (a == b)              /* WRONG for strings */

That compares two addresses. Two separate arrays are at two different addresses, so it is false even when the contents match. == works on numbers and characters; strings need strcmp.

Both functions, in one program that does something

#include <stdio.h>
#include <string.h>

int main(void)
{
    char first[50], second[50];
    int result;

    printf("Enter the first string: ");
    scanf("%49[^\n]", first);
    printf("Enter the second string: ");
    scanf(" %49[^\n]", second);

    printf("\"%s\" has %zu characters\n", first, strlen(first));
    printf("\"%s\" has %zu characters\n", second, strlen(second));

    result = strcmp(first, second);

    if (result == 0)
        printf("The two strings are equal\n");
    else if (result < 0)
        printf("\"%s\" comes before \"%s\"\n", first, second);
    else
        printf("\"%s\" comes after \"%s\"\n", first, second);

    return 0;
}
banana
apple
Enter the first string: Enter the second string: "banana" has 6 characters
"apple" has 5 characters
"banana" comes after "apple"
munotes.in73

Practical 6(c): strlen and strcmp

The space at the start of the second format string, " %49[^\n]", skips the newline left behind by the first read. Without it the second read sees that newline immediately, matches nothing, and leaves second untouched. This is the commonest reason a program that reads two lines appears to skip the second one.

The other comparison functions

FunctionWhat it does
strcmp(a, b)compares the whole strings
strncmp(a, b, n)compares only the first n characters
strcasecmp(a, b)compares ignoring case, on most systems

strncmp is genuinely useful: strncmp(word, "com", 3) == 0 asks whether a word begins with those three letters. strcasecmp is not in the C standard, though nearly every system has it; if a program must be portable, lower both strings yourself with tolower from ctype.h.

What beginners get wrong

if (strcmp(a, b)) read as "if equal". It is true when they differ.

Testing strcmp(a, b) == -1. Only the sign is guaranteed.

Using == on two strings. That compares addresses.

Printing strlen with %d. It returns a size_t; use %zu.

Believing strlen counts the null. It does not; the array needs one more byte than strlen reports.

Calling strlen inside a loop condition. for (i = 0; i < strlen(s); i++) recounts the whole string on every single iteration. Count once into a variable before the loop.

Forgetting #include <string.h>.

Quick revision

  • Both are in string.h.
  • strlen(s) counts characters before the null and returns a size_t, printed with %zu.
  • sizeof is the size of the array; strlen is the length of the string in it.
  • strcmp(a, b) returns 0 if equal, negative if a comes first, positive if b does.
  • Only the sign is guaranteed, never the value.
  • Equal is strcmp(a, b) == 0. A bare if (strcmp(a, b)) means different.
  • == on strings compares addresses and is always the wrong test.
  • A leading space in a scanf format skips the leftover newline.

What goes in your journal

Aim, the algorithm, the program using both functions, and the output for two pairs: two equal strings and two that differ. Write the three return values of strcmp into the conclusion as a small table. If there is room, add the sizeof against strlen line with its two numbers; both are asked in vivas and both fit on one line.

Test yourself

1. What does strcmp return when the strings are identical? 0.

2. char a[20] = "Mumbai"; What are sizeof a and strlen(a)? 20 and 6. sizeof is the size of the array, strlen the length of the string in it.

munotes.in74

Practical 6(c): strlen and strcmp

3. Why is if (strcmp(a, b)) the wrong way to test for equality? Because a non-zero value is true in C, and strcmp returns non-zero exactly when the strings are different. The test means "if they differ".

4. Is strcmp("Apple", "apple") negative or positive, and why? Negative. A capital A is 65 and a small a is 97, so the capital comes first.

5. Why should you not test strcmp(a, b) == -1? Because the standard guarantees only the sign of the result. A different library may return minus 25 for the same pair.

6. Why is for (i = 0; i < strlen(s); i++) a poor loop? Because strlen is called on every iteration and walks the whole string each time. Store the length in a variable before the loop.

Contents This chapter on its own page

munotes.in75

Chapter Twenty-Two

Practical 7: Swapping Two Numbers, by Value and by Reference

Syllabus topic Module 1, Practical 7: "Write a program to swap two numbers using a function. Pass the values to be swapped to this function using call-by-value method and call-by-reference method."

Aim

To swap two numbers using a function, passing the values by the call by value method and by the call by reference method.

The two ways a value reaches a function

Call by value copies. The function is given a copy of the argument, works on the copy, and the caller's variable is untouched. This is what C does by default, and it is what happened in [Practical 4(a): The Area of a Square, Using a Function].

Call by reference passes the address of the caller's variable instead of its value. The function then reaches through that address and changes the original.

C has only one calling mechanism, and it is call by value. What is called call by reference in C is call by value where the value being copied happens to be an address. That is worth saying precisely, because it is exactly what an examiner is listening for.

Pointers, in the smallest useful amount

A pointer is a variable that holds the address of another variable.

WrittenRead asMeans
int *p;p is a pointer to intdeclares p, which can hold the address of an int
p = &a;p gets the address of a& is the address-of operator
*pthe thing p points at is the dereference operator, and p is another name for a
*p = 7;put 7 where p pointschanges a itself

The does two different jobs and that is what confuses people. In the declaration int p; it says what kind of variable p is. In an expression *p it goes and fetches what p points at. Same character, two jobs, and the context tells them apart.

You have already used & in every scanf since [Practical 1(a): Simple Interest]. scanf("%d", &a) passes the address of a so that scanf can change it. scanf is call by reference, and you have been writing it for six practicals.

Both versions, in one program

#include <stdio.h>

void swap_by_value(int x, int y);
void swap_by_reference(int *x, int *y);

int main(void)
{
    int a = 10, b = 20;

    printf("Before any call:        a = %d, b = %d\n", a, b);

    swap_by_value(a, b);
    printf("After call by value:    a = %d, b = %d\n", a, b);

    swap_by_reference(&a, &b);
    printf("After call by reference: a = %d, b = %d\n", a, b);

    return 0;
}

void swap_by_value(int x, int y)
{
    int temp = x;

    x = y;
    y = temp;

    printf("  inside the function:  x = %d, y = %d\n", x, y);
}

void swap_by_reference(int *x, int *y)
{
    int temp = *x;

    *x = *y;
    *y = temp;

    printf("  inside the function:  *x = %d, *y = %d\n", *x, *y);
}
Before any call:        a = 10, b = 20
  inside the function:  x = 20, y = 10
After call by value:    a = 10, b = 20
  inside the function:  *x = 20, *y = 10
After call by reference: a = 20, b = 10
munotes.in76

Practical 7: Swapping Two Numbers, by Value and by Reference

Those five lines are the whole practical, and line 2 against line 3 is the point.

Inside swap_by_value the swap worked. x became 20 and y became 10. The function did its job perfectly, on its own two copies.

Back in main, nothing had changed. a was still 10. The copies were destroyed when the function returned, and they were the only things that were swapped.

Inside swap_by_reference the swap also worked, and this time x and y were not copies: they were other names for main's own a and b.

Back in main, a and b had exchanged values. The function reached through the addresses it was given.

Why the copies are genuinely separate

#include <stdio.h>

void examine(int x, int *p);

int main(void)
{
    int a = 10;

    examine(a, &a);
    return 0;
}

void examine(int x, int *p)
{
    printf("the copy holds %d, and the pointer points at %d\n", x, *p);

    x = 99;
    printf("after changing the copy, the original is still %d\n", *p);

    *p = 77;
    printf("after changing through the pointer, the original is %d\n", *p);
    printf("but the copy is still %d\n", x);
}
the copy holds 10, and the pointer points at 10
after changing the copy, the original is still 10
after changing through the pointer, the original is 77
but the copy is still 99

One function, one variable in main, and two routes to it. Changing x moved nothing but x. Changing *p moved the original, and the copy did not notice.

The three lines of a swap, in both forms

By valueBy reference
Parametersint x, int yint x, int y
Savetemp = x;temp = *x;
Movex = y;x = y;
Restorey = temp;*y = temp;
Called asswap(a, b);swap(&a, &b);
Effect on the callernonethe values are exchanged

Every x has become *x, and the call has gained two ampersands. Nothing else changed.

Call by value against call by reference

Call by valueCall by reference
What is passeda copy of the valuethe address of the variable
The function can change the originalnoyes
Parameter is declaredint xint *x
Call is writtenf(a)f(&a)
Cost of a large itemthe whole thing is copiedone address, whatever the size
Safetythe caller's data cannot be damagedthe function can damage it
Use it forinputs the function only readsvalues the function must change, and large structures
munotes.in77

Practical 7: Swapping Two Numbers, by Value and by Reference

Neither is better. A function that only reads its input should take it by value, because then a reader of the call knows nothing can change. A function that must change something has no choice.

Where you have already met both

printf("%d", a) is call by value. It only reads.

scanf("%d", &a) is call by reference. It must change a, so it is given the address.

An array passed to a function is always by reference, in effect, because an array name is the address of its first element. That is why the sorting function in a larger program can sort the caller's array, and why char * parameters in [Practical 6(b): Is This String a Palindrome?] could read the caller's string.

What beginners get wrong

Calling swap(a, b) on the pointer version. The compiler refuses: an int is not an int *.

Calling swap(&a, &b) on the value version. The compiler refuses the other way round.

Writing x = y instead of x = y inside the pointer version. That swaps the two pointers, which are themselves copies, so nothing happens in the caller. It compiles and it looks right, and it is the classic wrong answer to this practical.

Forgetting temp. Two statements lose a value, as in [Practical 5(b): Sorting an Array].

Declaring int x, y; and expecting two pointers. The attaches to the name, not to the type, so that declares one pointer and one plain int. Write int x, y;.

Saying "C supports call by reference". Say instead: C passes everything by value, and call by reference is achieved by passing an address by value.

Quick revision

  • Call by value copies the argument; the function cannot change the caller's variable.
  • Call by reference passes an address; the function changes the original through it.
  • int p; declares a pointer. &a is the address of a. p is what p points at.
  • In a declaration says "pointer"; in an expression says "fetch what is there".
  • Swap by reference: temp = x; x = y; y = temp; called as swap(&a, &b).
  • int x, y; declares two pointers. int* x, y; does not.
  • scanf is call by reference and always has been; that is what its & is for.
  • Strictly, C has only call by value. Passing an address by value is how call by reference is done.

What goes in your journal

Aim, an algorithm for each of the two functions, a flowchart, one program containing both functions, and the full output with all five lines. The output is the whole demonstration: the "inside the function" line of the value version proves that the swap worked, and the line after it proves that it did not reach main. In the conclusion write one sentence saying why, in terms of copies and addresses.

munotes.in78

Practical 7: Swapping Two Numbers, by Value and by Reference

Test yourself

1. What is passed in call by value, and what in call by reference? A copy of the value, and the address of the variable.

2. In int p;, what does the mean, and what does it mean in *p = 7;? In the declaration it says p is a pointer. In the statement it fetches the thing p points at, so 7 is stored in that variable.

3. Why does the call by value swap fail? Because the function swaps its own two copies, which are destroyed when it returns. The caller's variables were never touched.

4. Which of printf and scanf uses call by reference, and why? scanf, because it must store a value into the caller's variable, and to do that it needs the address.

5. What does int x, y; declare? One pointer to int, called x, and one ordinary int called y. The binds to the name.

6. Does C support call by reference? Not directly. Every argument is passed by value. Call by reference is obtained by passing an address by value and dereferencing it inside the function.

Contents This chapter on its own page

munotes.in79

Chapter Twenty-Three

Practical 8(a): Reading a Matrix of m Rows and n Columns

Syllabus topic Module 1, Practical 8(a): "Write a program to read a matrix of size m*n."

Aim

To read a matrix of size m by n.

What a two dimensional array is

int a[3][4];

Read the brackets left to right: three things, each of which is four ints. So it is three rows of four columns, twelve ints in all, and an element is named by two subscripts:

a[i][j]        /* row i, column j */

Both subscripts start at 0, exactly as in [Practical 5(a): Ten Students' Roll Numbers and Names]. For int a[3][4] the rows are 0 to 2 and the columns are 0 to 3.

col 0col 1col 2col 3
row 0a[0][0]a[0][1]a[0][2]a[0][3]
row 1a[1][0]a[1][1]a[1][2]a[1][3]
row 2a[2][0]a[2][1]a[2][2]a[2][3]

Row first, then column. Always. a[2][3] is the element in row 2, column 3, and a[3][2] is outside this array altogether.

Row major order

Memory is a single line of bytes; it has no rows. A two dimensional array is stored row by row, the whole of row 0, then the whole of row 1, and so on. That arrangement is called row major order, and C uses it.

#include <stdio.h>

int main(void)
{
    int a[3][4];

    printf("one int          %zu bytes\n", sizeof(int));
    printf("one row, a[0]    %zu bytes, so %zu ints\n",
           sizeof a[0], sizeof a[0] / sizeof a[0][0]);
    printf("the whole array  %zu bytes, so %zu rows\n",
           sizeof a, sizeof a / sizeof a[0]);

    return 0;
}
one int          4 bytes
one row, a[0]    16 bytes, so 4 ints
the whole array  48 bytes, so 3 rows

Three rows of sixteen bytes make forty-eight, laid out end to end. Two consequences follow.

a[0] is itself an array, of four ints, which is why sizeof a[0] has an answer at all. A two dimensional array in C is an array of arrays.

A loop that runs along a row is faster than one that runs down a column, because consecutive elements of a row are next to each other in memory. It does not matter at this size; it matters a great deal in a large program, and it is a good viva answer.

The program

#include <stdio.h>

int main(void)
{
    int a[10][10];
    int m, n, i, j;

    printf("Enter the number of rows and columns: ");
    scanf("%d %d", &m, &n);

    if (m < 1 || m > 10 || n < 1 || n > 10) {
        printf("Rows and columns must be between 1 and 10\n");
        return 0;
    }

    for (i = 0; i < m; i++)
        for (j = 0; j < n; j++) {
            printf("Element [%d][%d]: ", i, j);
            scanf("%d", &a[i][j]);
        }

    printf("\nThe %d by %d matrix:\n", m, n);

    for (i = 0; i < m; i++) {
        for (j = 0; j < n; j++)
            printf("%5d", a[i][j]);
        printf("\n");
    }

    return 0;
}
munotes.in80

Practical 8(a): Reading a Matrix of m Rows and n Columns

2 3
1 2 3
4 5 6
Enter the number of rows and columns: Element [0][0]: Element [0][1]: Element [0][2]: Element [1][0]: Element [1][1]: Element [1][2]:
The 2 by 3 matrix:
    1    2    3
    4    5    6

Three things in that program

The array is declared 10 by 10 and only m by n of it is used. That is the ordinary way this practical is written, and the reason is in the next section.

The size check is not optional. With m of 20 the reading loop writes past the end of the array, which is undefined behaviour and may corrupt anything. Two comparisons prevent it.

%5d prints in a field five wide, which is what makes the columns line up. Without it, 5 and 100 take different amounts of room and the matrix comes out ragged. This is the same field width idea as %-4d in the students practical, without the minus because numbers look right when they are right aligned.

Why not declare the array after reading m and n

C99 added variable length arrays, which allow exactly that:

int m, n;
scanf("%d %d", &m, &n);
int a[m][n];           /* a variable length array */

It compiles in C17 and it works. There are two reasons this book does not use it.

It is optional. C11 made variable length arrays a feature a compiler may choose not to support, and the macro __STDC_NO_VLA__ exists to say so. Microsoft's compiler, which many college laboratories use behind an IDE, does not support them.

There is nowhere to report a failure. A fixed array of 10 by 10 that is asked for 20 by 20 can print a message. A variable length array asked for 100000 by 100000 simply fails, usually by crashing, and there is no way to test for it.

So: declare generously, check the input, and use the part you need. If your college teaches the variable length form, use it and say in your journal that you know why the fixed form is the safer one.

Reading in the other order

The program above reads row by row, which is how a matrix is written on paper. Reading column by column takes one change, the loops swapped:

for (j = 0; j < n; j++)
    for (i = 0; i < m; i++)
        scanf("%d", &a[i][j]);

The subscripts stay a[i][j], because that is still row i column j. Only the order in which the elements are visited has changed. Examiners ask for this to see whether a student understands that the loops and the subscripts are separate decisions.

munotes.in81

Practical 8(a): Reading a Matrix of m Rows and n Columns

What beginners get wrong

Writing a[i, j]. In C the comma is an operator; a[i, j] means a[j], which is a whole row. Two sets of brackets, always.

Mixing up the subscripts. a[j][i] is the transpose, and it is a different matrix.

Declaring int a[m][n] before m and n have been read. At that point they hold rubbish.

Not checking m and n against the declared size. The reading loop then writes outside the array.

Printing without a field width. The matrix comes out ragged and the journal looks careless.

Forgetting printf("\n") after each row. The whole matrix prints on one line.

Quick revision

  • int a[3][4]; is three rows of four columns, twelve ints.
  • a[i][j] is row i, column j. Both subscripts start at 0.
  • C stores a matrix in row major order: all of row 0, then all of row 1.
  • a[0] is itself an array, so a matrix is an array of arrays.
  • Declare generously, read m and n, check them, and use the part you need.
  • Variable length arrays exist in C99 and are optional since C11; not every compiler has them.
  • %5d aligns the columns; printf("\n") ends each row.

What goes in your journal

Aim, the algorithm with both nested loops, a flowchart showing the inner loop inside the outer, the program, and the output for at least a 2 by 3 matrix so that the difference between rows and columns is visible. A 3 by 3 example hides a subscript mix-up, because the transpose of a square matrix is the same shape; a rectangular one exposes it immediately.

Test yourself

1. int a[3][4]; How many ints, and what is the last valid element? Twelve, and the last is a[2][3].

2. What is row major order? Storing a two dimensional array one whole row after another in memory. C uses it.

3. Why is a[i, j] wrong? Because the comma is an operator, so the expression evaluates to j and a[j] is a whole row, not an element. Use a[i][j].

4. Why declare int a[10][10] and read smaller sizes into it? Because the size must be known where the array is declared unless variable length arrays are used, and those are optional in C11 and absent from some compilers. A fixed size also allows the program to reject an input that is too large.

5. What does %5d do? Prints the integer right aligned in a field five characters wide, so columns line up.

6. How do you read the matrix column by column instead? Swap the two loops so that j is the outer one. The subscripts stay a[i][j].

Contents This chapter on its own page

munotes.in82

Chapter Twenty-Four

Practical 8(b): Multiplying Two Matrices in a Function

Syllabus topic Module 1, Practical 8(b): "Write a program to multiply two matrices using a function."

Aim

To multiply two matrices using a function.

When two matrices can be multiplied at all

Not every pair can. The rule is short and it is the first thing an examiner asks.

The number of columns in the first matrix must equal the number of rows in the second.

If A is m by n and B is p by q, then A times B exists only when n equals p, and the answer is m by q.

ABProduct exists?Answer is
2 by 33 by 2yes, 3 equals 32 by 2
3 by 22 by 4yes, 2 equals 23 by 4
2 by 32 by 3no, 3 is not 2does not exist

A program that does not test this will read two matrices, run its loops over memory it does not own, and print nonsense. The test is one if.

How one element of the answer is worked out

The element in row i, column j of the product is made from row i of A and column j of B: multiply them element by element and add up the results.

C[i][j] = A[i][0] B[0][j] + A[i][1] B[1][j] + ... + A[i][n-1] * B[n-1][j]

Worked, with A of 2 by 3 and B of 3 by 2:

AB
12378
456910
1112

C[0][0] takes row 0 of A, which is 1, 2, 3, against column 0 of B, which is 7, 9, 11:

1 7 + 2 9 + 3 * 11 = 7 + 18 + 33 = 58

C[0][1] takes the same row of A against column 1 of B, which is 8, 10, 12:

1 8 + 2 10 + 3 * 12 = 8 + 20 + 36 = 64

C[1][0] and C[1][1] use row 1 of A, which is 4, 5, 6:

4 7 + 5 9 + 6 * 11 = 28 + 45 + 66 = 139

= 4 8 + 5 10 + 6 * 12 = 32 + 50 + 72 = 154

So the product is 58, 64 on the first row and 139, 154 on the second.

The three loops, and what each one is for

for (i = 0; i < m; i++)            /* which row of the answer */
    for (j = 0; j < q; j++) {      /* which column of the answer */
        c[i][j] = 0;
        for (k = 0; k < n; k++)    /* walk along the row and down the column */
            c[i][j] += a[i][k] * b[k][j];
    }

The outer two loops visit every element of the answer. The innermost loop computes one such element.

munotes.in83

Practical 8(b): Multiplying Two Matrices in a Function

Look at the subscripts in the innermost line. a[i][k] walks along row i of A as k grows; b[k][j] walks down column j of B. They move together, which is exactly the pairing the formula asks for. Getting a[k][i] or b[j][k] there is the commonest error in this practical, and the result is the product of transposes, which is a different matrix.

c[i][j] = 0; before the innermost loop is not optional. The element is an accumulator and it must start empty, for the reason given in [Practical 3(b): The Factorial of a Number].

Passing a matrix to a function

A one dimensional array is passed as an address and the function does not need to know how long it is. A two dimensional array is different: to work out where a[i][j] lives, the compiler must know how many columns there are in a row, because the rows are laid end to end.

So the parameter is written with the first size left out and every later size given:

void multiply(int a[][10], int b[][10], int c[][10], int m, int n, int q);

int a[][10] means "an array of rows, each row being ten ints". The number of rows may be left blank because it is never needed for the address arithmetic; the 10 may not.

That is why every matrix in this program is declared [10][10] and the used part is m by n: the second dimension must be a fixed number that the function and the caller agree on. Passing a 10 by 10 array to a function expecting int a[][20] does not compile, and it should not.

The program

#include <stdio.h>

void read_matrix(int a[][10], int rows, int cols, const char *name);
void print_matrix(int a[][10], int rows, int cols, const char *name);
void multiply(int a[][10], int b[][10], int c[][10], int m, int n, int q);

int main(void)
{
    int a[10][10], b[10][10], c[10][10];
    int m, n, p, q;

    printf("Rows and columns of the first matrix: ");
    scanf("%d %d", &m, &n);
    printf("Rows and columns of the second matrix: ");
    scanf("%d %d", &p, &q);

    if (n != p) {
        printf("Cannot multiply: the first has %d columns "
               "and the second has %d rows\n", n, p);
        return 0;
    }

    read_matrix(a, m, n, "A");
    read_matrix(b, p, q, "B");

    multiply(a, b, c, m, n, q);

    print_matrix(a, m, n, "A");
    print_matrix(b, p, q, "B");
    print_matrix(c, m, q, "A times B");

    return 0;
}

void read_matrix(int a[][10], int rows, int cols, const char *name)
{
    int i, j;

    printf("Enter %d values for matrix %s: ", rows * cols, name);

    for (i = 0; i < rows; i++)
        for (j = 0; j < cols; j++)
            scanf("%d", &a[i][j]);
}

void print_matrix(int a[][10], int rows, int cols, const char *name)
{
    int i, j;

    printf("\nMatrix %s (%d by %d)\n", name, rows, cols);

    for (i = 0; i < rows; i++) {
        for (j = 0; j < cols; j++)
            printf("%6d", a[i][j]);
        printf("\n");
    }
}

void multiply(int a[][10], int b[][10], int c[][10], int m, int n, int q)
{
    int i, j, k;

    for (i = 0; i < m; i++)
        for (j = 0; j < q; j++) {
            c[i][j] = 0;
            for (k = 0; k < n; k++)
                c[i][j] += a[i][k] * b[k][j];
        }
}
munotes.in84

Practical 8(b): Multiplying Two Matrices in a Function

2 3
3 2
1 2 3 4 5 6
7 8 9 10 11 12
Rows and columns of the first matrix: Rows and columns of the second matrix: Enter 6 values for matrix A: Enter 6 values for matrix B:
Matrix A (2 by 3)
     1     2     3
     4     5     6

Matrix B (3 by 2)
     7     8
     9    10
    11    12

Matrix A times B (2 by 2)
    58    64
   139   154

The four numbers are the four worked out by hand above.

The mismatched case, refused

#include <stdio.h>

int main(void)
{
    int n = 3, p = 2;

    if (n != p)
        printf("Cannot multiply: first has %d columns, second has %d rows\n",
               n, p);
    else
        printf("The product is defined\n");

    return 0;
}
Cannot multiply: first has 3 columns, second has 2 rows

That is the whole guard, and a program that has it is a complete answer where one that does not is half of one.

Two facts about matrix multiplication worth a viva mark

It is not commutative. A times B is not generally B times A, and very often B times A does not even exist: with A of 2 by 3 and B of 3 by 2, A times B is 2 by 2 and B times A is 3 by 3. Two different matrices, from the same pair.

It is associative. (A times B) times C equals A times (B times C), whenever the sizes allow.

What beginners get wrong

Not testing that the columns of A match the rows of B.

Writing a[k][i] or b[j][k] in the innermost line. The subscripts are a[i][k] * b[k][j], in that order.

Forgetting c[i][j] = 0;. The element starts with whatever was in that memory and every answer is wrong by a random amount.

Putting c[i][j] = 0; in the wrong place. Inside the k loop it wipes the running total on every step and leaves only the last term.

Declaring the function parameter as int a[][]. Both dimensions blank does not compile: the column count is needed.

Assuming the answer has the same shape as the inputs. A of m by n times B of n by q gives m by q.

munotes.in85

Practical 8(b): Multiplying Two Matrices in a Function

Quick revision

  • A of m by n times B of p by q exists only when n equals p, and the answer is m by q.
  • C[i][j] is row i of A against column j of B, multiplied term by term and added.
  • Three loops: i over the answer's rows, j over its columns, k along the row and down the column.
  • c[i][j] = 0; goes between the j loop and the k loop.
  • The inner statement is c[i][j] += a[i][k] * b[k][j];.
  • A function parameter is int a[][10]: the first size may be blank, the rest may not.
  • Matrix multiplication is associative but not commutative.

What goes in your journal

Aim, the algorithm with all three loops, a flowchart, the program with the three functions, and the output showing both input matrices and the product. Work one element out by hand beside the output, the way C[0][0] is worked out in this chapter: it is the proof that you know where the numbers came from, and it is two lines of arithmetic.

Test yourself

1. When can two matrices be multiplied? When the number of columns of the first equals the number of rows of the second.

2. A is 3 by 5 and B is 5 by 2. What size is the product? 3 by 2.

3. What are the three loops for? The outer two visit each element of the answer, by row and by column. The innermost adds up the products of one row of A with one column of B.

4. Where does c[i][j] = 0; go, and why? Immediately inside the j loop and before the k loop, because the element is an accumulator and must be empty before the summing begins, but only once per element.

5. Why must a function parameter be int a[][10] rather than int a[][]? Because the compiler works out the address of a[i][j] from the number of columns in a row. The row count is never needed; the column count always is.

6. Is A times B the same as B times A? No. Matrix multiplication is not commutative, and for non square matrices the second product often does not exist at all.

Contents This chapter on its own page

munotes.in86

Chapter Twenty-Five

Practical 9: A Structure, and Two Records of It

Syllabus topic Module 1, Practical 9: "Write a program to print the structure using Title Author Subject Book ID. Print the details of two students."

Aim

To print a structure with the members Title, Author, Subject and Book ID, and to print the details of two records.

MU's wording, and what this chapter does about it

Her printed task names four members, Title, Author, Subject and Book ID, and then says "Print the details of two students." Those four members describe a book, not a student.

Both readings are covered here, because it costs one extra structure and it means that whichever your examiner has in mind, your journal answers it. Ask your own teacher which they want; if they have no preference, write the book structure, since those are the four members MU actually printed.

What a structure is

An array holds many items of the same type. A structure holds several items of different types under one name.

A book has a title, which is text, and an identity number, which is a number. No array can hold both. A structure can:

struct book {
    char title[40];
    char author[30];
    char subject[20];
    int  book_id;
};

The things inside are called members or fields. struct book is now a type, in the same way that int is a type, and variables can be declared of it:

struct book b1, b2;
struct book library[100];

The semicolon after the closing brace is required. Leaving it off produces a confusing error on the line after, because the compiler is still waiting for the declaration to end.

Reaching a member

The dot operator gets at one member of a structure variable:

b1.book_id = 101;
strcpy(b1.title, "The C Language");
printf("%s\n", b1.title);

b1.title is an ordinary char array and b1.book_id is an ordinary int. They behave exactly as they would if they had been declared on their own, and every rule about strings from [Practical 6(a): Extracting Part of a String] still applies to b1.title.

You cannot write b1.title = "The C Language";. An array name is not something that can be assigned to. Use strcpy, or set the members when the variable is created:

struct book b1 = {"The C Language", "Dennis Ritchie", "Programming", 101};

That form is called an initialiser list, and the values go in the order the members were declared.

The program

#include <stdio.h>
#include <string.h>

struct book {
    char title[40];
    char author[30];
    char subject[20];
    int  book_id;
};

void print_book(struct book b);

int main(void)
{
    struct book b1 = {"The C Language", "Dennis Ritchie",
                      "Programming", 101};
    struct book b2;

    strcpy(b2.title, "Database Systems");
    strcpy(b2.author, "Ramez Elmasri");
    strcpy(b2.subject, "Databases");
    b2.book_id = 102;

    printf("Details of two books\n");
    printf("--------------------\n");
    print_book(b1);
    print_book(b2);

    printf("One structure occupies %zu bytes\n", sizeof(struct book));

    return 0;
}

void print_book(struct book b)
{
    printf("Book ID : %d\n", b.book_id);
    printf("Title   : %s\n", b.title);
    printf("Author  : %s\n", b.author);
    printf("Subject : %s\n", b.subject);
    printf("\n");
}
munotes.in87

Practical 9: A Structure, and Two Records of It

Details of two books
--------------------
Book ID : 101
Title   : The C Language
Author  : Dennis Ritchie
Subject : Programming

Book ID : 102
Title   : Database Systems
Author  : Ramez Elmasri
Subject : Databases

One structure occupies 96 bytes

Both ways of filling a structure are shown: b1 with an initialiser list at the point of declaration, and b2 member by member with strcpy afterwards.

The size is worth a second look, because it is not what the members add up to. 40 plus 30 plus 20 plus 4 is 94, and the program reported 96.

The two extra bytes are padding. The three character arrays end at byte 90, and an int on this machine is happiest starting at an address that is a multiple of 4, so the compiler leaves bytes 90 and 91 unused and puts book_id at byte 92. That takes the structure to 96.

So never add the members up by hand and call it the size. Ask sizeof, which is what it is for, and remember that the answer may differ between compilers and between machines.

The students reading, with two records read from the user

#include <stdio.h>

struct student {
    int  roll;
    char name[30];
    char course[15];
    float marks;
};

int main(void)
{
    struct student s[2];
    int i;

    for (i = 0; i < 2; i++) {
        printf("Student %d roll number: ", i + 1);
        scanf("%d", &s[i].roll);

        printf("Student %d name: ", i + 1);
        scanf(" %29[^\n]", s[i].name);

        printf("Student %d course: ", i + 1);
        scanf(" %14[^\n]", s[i].course);

        printf("Student %d marks: ", i + 1);
        scanf("%f", &s[i].marks);
    }

    printf("\nRoll  Name             Course      Marks\n");
    printf("----  ---------------  ----------  -----\n");

    for (i = 0; i < 2; i++)
        printf("%-4d  %-15s  %-10s  %5.2f\n",
               s[i].roll, s[i].name, s[i].course, s[i].marks);

    return 0;
}
101
Asha Kulkarni
BSc IT
78.5
102
Ravi Deshmukh
BSc IT
65.25
Student 1 roll number: Student 1 name: Student 1 course: Student 1 marks: Student 2 roll number: Student 2 name: Student 2 course: Student 2 marks:
Roll  Name             Course      Marks
----  ---------------  ----------  -----
101   Asha Kulkarni    BSc IT      78.50
102   Ravi Deshmukh    BSc IT      65.25

struct student s[2]; is an array of structures, which is how any real record keeping program is built: s[0] and s[1] are whole students, and s[i].name is one member of one of them.

Reading a name that has a space in it

scanf(" %29[^\n]", s[i].name) is the call that [Practical 5(a): Ten Students' Roll Numbers and Names] promised. Three parts, and all three are needed.

%[^\n] is a scan set meaning "every character except a newline", so the read stops at the end of the line rather than at the first space. That is what lets "Asha Kulkarni" arrive whole.

munotes.in88

Practical 9: A Structure, and Two Records of It

The 29 limits it to twenty-nine characters, leaving the thirtieth for the terminating null.

The leading space skips whatever whitespace is waiting, and there is always some: the Enter you pressed after the roll number is still in the buffer. Without that space the scan set sees the newline at once, matches nothing, and the name is left empty. This is the single commonest defect in a first-semester program that mixes numbers and names.

Structure against array

ArrayStructure
Holdsmany items of one typeseveral items of different types
Items are named bya subscripta member name
Reached witha[3]b.title
Can be assigned wholenoyes: b2 = b1; copies every member
Passed to a functionas an address, so changes are seenby value, so the function gets a copy
Suitsa list of the same thingone thing with several properties

The fourth row is worth dwelling on. b2 = b1; is legal and copies all four members including the character arrays, where copying two plain arrays needs strcpy or a loop. Structures may be assigned; arrays may not.

typedef, so you can stop writing struct

typedef struct {
    char title[40];
    int  book_id;
} Book;

Book b1, b2;                 /* no `struct` keyword needed */

typedef gives a type a second, shorter name. It is common in real code and costs nothing to know.

What beginners get wrong

Forgetting the semicolon after the closing brace of the structure.

Writing b1.title = "something";. An array cannot be assigned. Use strcpy.

Using -> on a structure variable. The arrow is for a pointer to a structure: p->title means (*p).title. With a plain variable it is the dot.

Declaring the structure inside main and then using it in a function. Declare the structure at file scope, above main, so every function can see it.

Reading a name with %s. It stops at the first space, so "Asha Kulkarni" becomes "Asha".

Leaving out the leading space in " %29[^\n]". The name comes out empty.

Adding the member sizes up by hand and expecting sizeof to agree. Padding may make the structure larger.

Quick revision

  • A structure groups items of different types under one name; an array groups items of one type.
  • struct book { ... }; and the semicolon after the brace is required.
  • b1.title is the dot operator; p->title is for a pointer to a structure.
  • Initialise with struct book b1 = {...}; or fill members later with strcpy and assignment.
  • A character array member cannot be assigned with =; use strcpy.
  • A whole structure can be assigned: b2 = b1; copies every member.
  • struct student s[2]; is an array of structures; s[i].name is one member of one record.
  • scanf(" %29[^\n]", name) reads a line with spaces; the leading space skips the waiting newline.
  • sizeof may exceed the sum of the members because of padding.
munotes.in89

Practical 9: A Structure, and Two Records of It

What goes in your journal

Aim, the algorithm, the program, and the output showing two complete records. Write the structure declaration out separately in the journal with each member labelled with its type; that declaration is what the viva is about. In the conclusion give the one sentence distinction between a structure and an array.

Test yourself

1. What is the difference between an array and a structure? An array holds many items of the same type, named by a subscript. A structure holds several items of possibly different types, named by member names.

2. What is wrong with b1.title = "The C Language";? title is an array, and an array cannot be assigned to. Use strcpy(b1.title, "The C Language").

3. When is -> used instead of .? When the variable on the left is a pointer to a structure. p->title is short for (*p).title.

4. Can one structure be assigned to another? Yes. b2 = b1; copies all the members, including the character arrays.

5. Why does scanf(" %29[^\n]", name) need the leading space? Because the newline left behind by the previous read is still waiting, and without the space the scan set would stop on it at once and read nothing.

6. Why might sizeof(struct book) be larger than the sum of its members? Because the compiler may insert padding bytes so that members begin at addresses the machine handles efficiently.

Contents This chapter on its own page

munotes.in90

Chapter Twenty-Six

Practical 10: Designing the Bank Management System

Syllabus topic Module 1, Practical 10: "Create a mini project on "Bank management system". The program should be menu driven."

Aim

To design a menu driven mini project for a bank management system.

Why this chapter exists

Practical 10 is the only one MU calls a project, and a project is not a longer program. It is a program with a design behind it, and the design is what is marked when the examiner looks at your journal and asks "why did you do it this way?".

So this chapter decides four things, in the order a real design decides them, and the next chapter writes the code. Nothing here is invented afterwards to explain a program that already exists.

Step 1: what is the data

One account has an identity, an owner and an amount of money. Those are three things of different types, so they belong in a structure ([Practical 9: A Structure, and Two Records of It]).

MemberTypeWhy
numberintthe account number the customer quotes
namechar[30]the holder's name, which may contain a space
balancefloatrupees and paise, so not an int

A bank has many accounts, all of them the same kind of thing, so they belong in an array. An array of structures:

struct account bank[100];
int count = 0;

count is how many accounts actually exist, which is not the same as the size of the array. Every loop over the accounts runs to count, never to 100, and every new account is added at subscript count before count goes up by one. That pair, the array and the count of how much of it is in use, is the commonest arrangement in C and it is worth recognising.

Step 2: what can the user do

Five operations plus a way out. Each is one menu item and each becomes one function.

ChoiceOperationWhat it needsWhat it does
1Create accountname, opening depositallots the next number, stores the record
2Depositaccount number, amountadds to that balance
3Withdrawaccount number, amounttakes from that balance if allowed
4Balance enquiryaccount numberprints one account
5Display allnothingprints every account
6Exitnothingends the program

Each operation is a void function taking no parameters, because the array is at file scope and every one of them works on it. That is the simplest arrangement that works for a first-semester project, and it is what most colleges expect.

Step 3: what must be refused

This is the part that separates a project from an exercise, and it is where the marks are. A program that accepts anything is not a bank.

RuleWhyWhat the program does
The opening deposit must be at least 500a minimum balance is a real bank rulerefuses to open the account
A deposit must be more than zeroa deposit of minus 100 is a withdrawal in disguiserefuses
A withdrawal must be more than zerosame reasonrefuses
A withdrawal must leave at least 500the minimum balance againrefuses and says what the balance is
The account number must existotherwise the program writes to a record that is not theresays there is no such account
The menu choice must be 1 to 6anything else is a mistakesays so and shows the menu again
munotes.in91

Practical 10: Designing the Bank Management System

Six rules and six messages. Write them down before writing the program, because each one is an if and a return, and they are much harder to insert afterwards.

Notice that every rule says what the program does as well as what it forbids. "Refuses" is not enough: the user must be told why, or they will try the same thing again.

Step 4: how the parts fit together

main
  the menu loop
      switch on the choice
          create_account
          deposit        \
          withdraw        >  each of these first calls find_account
          enquiry        /
          display_all

One function is shared by three others. find_account(number) searches the array and returns the subscript of the account, or minus 1 if there is none.

int find_account(int number)
{
    for (i = 0; i < count; i++)
        if (bank[i].number == number)
            return i;
    return -1;
}

Minus 1 is chosen as the "not found" answer because it is not a valid subscript, so it can never be mistaken for one. Deposit, withdraw and enquiry all begin the same way: find the account, and if the answer is minus 1, print a message and return.

Writing that search once instead of three times is the whole argument for functions from [Practical 4(a): The Area of a Square, Using a Function], appearing in a program large enough for it to matter.

The algorithm for the whole program

1. Start

2. Set count to 0 and the next account number to 1001

3. Display the menu and read the choice

4. If the choice is 1, create an account: read the name and the opening deposit, refuse if it is below 500, otherwise store the record and increase count

5. If the choice is 2, deposit: read the number, find the account, refuse if not found or if the amount is not positive, otherwise add it to the balance

6. If the choice is 3, withdraw: read the number, find the account, refuse if not found, if the amount is not positive, or if the balance would fall below 500, otherwise subtract it

7. If the choice is 4, find the account and print it

8. If the choice is 5, print every account from 0 to count minus 1

9. If the choice is not between 1 and 6, print that the choice is not valid

10. If the choice is not 6, go back to step 3

11. Stop

munotes.in92

Practical 10: Designing the Bank Management System

What this design leaves out, and why saying so is worth marks

The accounts disappear when the program ends. They are in memory, not in a file. Saving them would need file handling, which MU does not set in Semester 1, so it is out of scope. Say this in your conclusion: it shows you know the limit of your own program rather than having failed to notice it.

There is no password and no date. A real system has both. This one is a mini project, and adding them without being asked makes the program longer without making it better.

float loses fractions of a paisa. A float keeps about seven significant digits, so a balance in the crores is no longer exact to the paisa. Real banking software uses integer paise, or a decimal type, for exactly this reason. In a mini project a float is the right choice and knowing why it would not be in a real one is a very good viva answer.

What beginners get wrong

Writing the program first and the design afterwards. The journal then shows a design that describes the code instead of a design the code came from, and it always shows.

Looping to 100 instead of to count. The program then reads accounts that do not exist.

Letting find_account return the account number instead of the subscript. Then the caller cannot reach the record, and returning 0 for "not found" collides with the first real subscript.

Putting the validation in main. Each rule belongs in the function that owns the operation, so that the rule and the change are never separated.

Forgetting to increase count. Every new account then overwrites the last one.

Quick revision

  • One account is a structure: number, name, balance. The bank is an array of them plus a count.
  • count is how many accounts exist; loops run to count, not to the array's size.
  • Six menu items: create, deposit, withdraw, enquiry, display all, exit.
  • One shared function, find_account, returns a subscript or minus 1.
  • Minus 1 cannot be a subscript, which is why it is the "not found" answer.
  • Six rules to enforce, each with its own message to the user.
  • The data is lost when the program ends; say so rather than leaving it unsaid.

What goes in your journal

This chapter is the front half of the Practical 10 write-up, and it is unusual in that almost all of it is writing rather than code. Include the aim, the structure definition with each member's type and purpose, the menu table, the table of rules, the diagram of which function calls which, and the eleven-step algorithm. The program itself goes in the next entry.

munotes.in93

Practical 10: Designing the Bank Management System

An examiner who reads that design before reading your program already knows you understood the task. That is worth more than a longer program.

Test yourself

1. Why is an account a structure rather than three arrays? Because the three values describe one thing and are of different types. Keeping them in one record means they cannot be accidentally separated.

2. What does count hold, and why not use the array's size? The number of accounts that actually exist. The array is declared large enough for the most the program will hold, and most of it is unused.

3. Why does find_account return minus 1 when there is no such account? Because minus 1 is not a valid subscript, so the caller can never confuse it with a real one. Returning 0 would collide with the first account.

4. Name three things the program must refuse. An opening deposit below the minimum, a withdrawal that would take the balance below the minimum, and an operation on an account number that does not exist.

5. What happens to the accounts when the program ends? They are lost. They are held in memory only, and saving them would require file handling.

6. Why is float acceptable here but not in a real bank? Because a float keeps only about seven significant digits, so large balances stop being exact to the paisa. Real systems store money as whole paise or in a decimal type.

Contents This chapter on its own page

munotes.in94

Chapter Twenty-Seven

Practical 10: The Bank Management System, Written and Run

Syllabus topic Module 1, Practical 10: "Create a mini project on "Bank management system". The program should be menu driven."

Aim

To write and run a menu driven mini project for a bank management system.

The program

The design is [Practical 10: Designing the Bank Management System] and is not repeated here. This is that design turned into C, function for function.

#include <stdio.h>
#include <string.h>

#define MAX_ACCOUNTS 100

struct account {
    int   number;
    char  name[30];
    float balance;
};

static struct account bank[MAX_ACCOUNTS];
static int count = 0;
static int next_number = 1001;

int  find_account(int number);
void create_account(void);
void deposit(void);
void withdraw(void);
void enquiry(void);
void display_all(void);

int main(void)
{
    int choice = 0;

    do {
        printf("\nBANK MANAGEMENT SYSTEM\n");
        printf("1 Create account   2 Deposit   3 Withdraw\n");
        printf("4 Balance enquiry  5 Display all   6 Exit\n");
        printf("Enter your choice: ");

        if (scanf("%d", &choice) != 1) {
            printf("\nInput ended\n");
            break;
        }

        switch (choice) {
        case 1: create_account(); break;
        case 2: deposit();        break;
        case 3: withdraw();       break;
        case 4: enquiry();        break;
        case 5: display_all();    break;
        case 6: printf("Thank you for banking with us\n"); break;
        default: printf("Please choose a number between 1 and 6\n");
        }
    } while (choice != 6);

    return 0;
}

int find_account(int number)
{
    int i;

    for (i = 0; i < count; i++)
        if (bank[i].number == number)
            return i;

    return -1;
}

void create_account(void)
{
    float amount;

    if (count == MAX_ACCOUNTS) {
        printf("The bank is full\n");
        return;
    }

    printf("Name of the account holder: ");
    scanf(" %29[^\n]", bank[count].name);

    printf("Opening deposit: ");
    scanf("%f", &amount);

    if (amount < 500) {
        printf("The opening deposit must be at least 500. Account not opened\n");
        return;
    }

    bank[count].number = next_number;
    bank[count].balance = amount;
    next_number++;
    count++;

    printf("Account %d opened for %s with a balance of %.2f\n",
           bank[count - 1].number, bank[count - 1].name,
           bank[count - 1].balance);
}

void deposit(void)
{
    int number, i;
    float amount;

    printf("Account number: ");
    scanf("%d", &number);

    i = find_account(number);
    if (i == -1) {
        printf("There is no account %d\n", number);
        return;
    }

    printf("Amount to deposit: ");
    scanf("%f", &amount);

    if (amount <= 0) {
        printf("A deposit must be more than zero\n");
        return;
    }

    bank[i].balance += amount;
    printf("Deposited %.2f. The balance of account %d is now %.2f\n",
           amount, bank[i].number, bank[i].balance);
}

void withdraw(void)
{
    int number, i;
    float amount;

    printf("Account number: ");
    scanf("%d", &number);

    i = find_account(number);
    if (i == -1) {
        printf("There is no account %d\n", number);
        return;
    }

    printf("Amount to withdraw: ");
    scanf("%f", &amount);

    if (amount <= 0) {
        printf("A withdrawal must be more than zero\n");
        return;
    }

    if (bank[i].balance - amount < 500) {
        printf("Refused. The balance of %.2f would fall below the "
               "minimum of 500.00\n", bank[i].balance);
        return;
    }

    bank[i].balance -= amount;
    printf("Withdrew %.2f. The balance of account %d is now %.2f\n",
           amount, bank[i].number, bank[i].balance);
}

void enquiry(void)
{
    int number, i;

    printf("Account number: ");
    scanf("%d", &number);

    i = find_account(number);
    if (i == -1) {
        printf("There is no account %d\n", number);
        return;
    }

    printf("Account %d  %s  balance %.2f\n",
           bank[i].number, bank[i].name, bank[i].balance);
}

void display_all(void)
{
    int i;

    if (count == 0) {
        printf("There are no accounts yet\n");
        return;
    }

    printf("\nNumber  Name                     Balance\n");
    printf("------  -----------------------  ----------\n");

    for (i = 0; i < count; i++)
        printf("%-6d  %-23s  %10.2f\n",
               bank[i].number, bank[i].name, bank[i].balance);
}
munotes.in95

Practical 10: The Bank Management System, Written and Run

1
Asha Kulkarni
5000
1
Ravi Deshmukh
2000
2
1001
1500
3
1002
1800
3
1001
1000
4
1001
5
7
6

BANK MANAGEMENT SYSTEM
1 Create account   2 Deposit   3 Withdraw
4 Balance enquiry  5 Display all   6 Exit
Enter your choice: Name of the account holder: Opening deposit: Account 1001 opened for Asha Kulkarni with a balance of 5000.00

BANK MANAGEMENT SYSTEM
1 Create account   2 Deposit   3 Withdraw
4 Balance enquiry  5 Display all   6 Exit
Enter your choice: Name of the account holder: Opening deposit: Account 1002 opened for Ravi Deshmukh with a balance of 2000.00

BANK MANAGEMENT SYSTEM
1 Create account   2 Deposit   3 Withdraw
4 Balance enquiry  5 Display all   6 Exit
Enter your choice: Account number: Amount to deposit: Deposited 1500.00. The balance of account 1001 is now 6500.00

BANK MANAGEMENT SYSTEM
1 Create account   2 Deposit   3 Withdraw
4 Balance enquiry  5 Display all   6 Exit
Enter your choice: Account number: Amount to withdraw: Refused. The balance of 2000.00 would fall below the minimum of 500.00

BANK MANAGEMENT SYSTEM
1 Create account   2 Deposit   3 Withdraw
4 Balance enquiry  5 Display all   6 Exit
Enter your choice: Account number: Amount to withdraw: Withdrew 1000.00. The balance of account 1001 is now 5500.00

BANK MANAGEMENT SYSTEM
1 Create account   2 Deposit   3 Withdraw
4 Balance enquiry  5 Display all   6 Exit
Enter your choice: Account number: Account 1001  Asha Kulkarni  balance 5500.00

BANK MANAGEMENT SYSTEM
1 Create account   2 Deposit   3 Withdraw
4 Balance enquiry  5 Display all   6 Exit
Enter your choice:
Number  Name                     Balance
------  -----------------------  ----------
1001    Asha Kulkarni               5500.00
1002    Ravi Deshmukh               2000.00

BANK MANAGEMENT SYSTEM
1 Create account   2 Deposit   3 Withdraw
4 Balance enquiry  5 Display all   6 Exit
Enter your choice: Please choose a number between 1 and 6

BANK MANAGEMENT SYSTEM
1 Create account   2 Deposit   3 Withdraw
4 Balance enquiry  5 Display all   6 Exit
Enter your choice: Thank you for banking with us

What that session shows

Nine passes through the menu, and each one was chosen to exercise a different part of the design.

PassChoiceWhat was tested
11opening an account, and the first number allotted
21a second account, and the number going up
32a deposit into an account that exists
43a withdrawal REFUSED, because it would breach the minimum balance
53a withdrawal allowed
64a balance enquiry
75the full list
87an invalid menu choice
96exit
munotes.in96

Practical 10: The Bank Management System, Written and Run

The fourth pass is the important one. Account 1002 holds 2000 and 1800 was asked for, which would leave 200, below the minimum of 500. The program refused and said what the balance was. A run that shows only successful operations has not demonstrated that the rules exist.

Everything from Module 1, in one program

This is why MU sets the mini project last.

TechniqueWhere it is used hereWhere it was taught
if and comparison operatorsevery validation rule[Practical 1(c): Is This a Leap Year?]
switch with defaultthe menu[Practical 2(b): A Menu Driven Calculator, with switch]
do whilethe menu loop[Practical 2(b): A Menu Driven Calculator, with switch]
forfind_account and display_all[Practical 3(b): The Factorial of a Number]
functions with parameters and a return valuefind_account[Practical 4(a): The Area of a Square, Using a Function]
void functionsthe five operations[Practical 4(a): The Area of a Square, Using a Function]
arrays and a used countbank and count[Practical 5(a): Ten Students' Roll Numbers and Names]
reading a name with spacesscanf(" %29[^\n]", ...)[Practical 9: A Structure, and Two Records of It]
structures and arrays of structuresstruct account bank[100][Practical 9: A Structure, and Two Records of It]
field widths in printfthe account table[Practical 5(a): Ten Students' Roll Numbers and Names]

Nothing in it is new. That is the point: a project is the existing tools put together, not a new subject.

Three details worth defending in a viva

static on the global variables. static struct account bank[100]; at file scope means the name is visible in this file only. In a one file program it changes nothing you can see, and it is the right habit: a name that does not need to be visible elsewhere should not be.

if (scanf("%d", &choice) != 1) break; guards against the input running out or the user typing letters. scanf returns the number of items it successfully read, and testing that return value is the difference between a program that ends cleanly and one that loops forever printing its menu. Most first-semester programs ignore it; this is a good thing to be asked about.

bank[count - 1] in the success message of create_account. count has already been increased by then, so the record just written is at count - 1. Writing the message before the increase would be simpler; it is written this way here deliberately, because the order of those two lines is exactly the kind of thing an examiner points at and asks you to explain.

munotes.in97

Practical 10: The Bank Management System, Written and Run

How to extend it, if you have time

None of these is required, and adding them badly is worse than not adding them. If you do add one, say in the journal why.

Save to a file. fopen, fwrite of the array, fclose, and read it back at the start. This makes the accounts survive the program ending, and it is the natural next step; MU sets file handling in the second semester.

A transaction list. A second structure holding the account number, the date, the type and the amount, in its own array, appended to by deposit and withdraw.

Search by name. strcmp in a loop, from [Practical 6(c): strlen and strcmp].

Close an account. Move the last record into the gap and reduce count by one, which is the standard way of deleting from an array when the order does not matter.

What beginners get wrong

Writing the whole thing in main. It will be two hundred lines with no shape, and a viva question about any part of it means reading the whole thing.

Testing only the happy path. Every refusal in the design must appear in the output you write up.

Using count and the array size interchangeably.

Reading the name with %s. "Asha Kulkarni" becomes "Asha".

Leaving out the default in the switch. An invalid choice then does nothing at all and the user cannot tell whether the program is working.

Changing the balance before the checks. Every validation must come before the assignment, so that a refused operation leaves the record exactly as it was.

Quick revision

  • One structure for an account, one array of them, one count of how many are in use.
  • do while around the menu; switch with a default inside it.
  • One shared find_account returning a subscript or minus 1.
  • Validate first, change afterwards. A refused operation must leave the data untouched.
  • scanf returns how many items it read; test it.
  • Field widths in printf are what make the account table line up.
  • The data lives in memory only and is lost when the program ends.

What goes in your journal

Aim, the program in full, and the complete session, both the successful operations and the refused ones. The design went in the previous entry; refer to it rather than repeating it.

Write a short conclusion of three or four lines: what the project demonstrates, which Module 1 techniques it uses, and the one limitation you know about, which is that the accounts are not saved. A conclusion that names its own limitation is the mark of somebody who understands their program.

Test yourself

1. Why is the menu loop a do while? Because the menu must be displayed before the user can have made a choice, so the body has to run at least once before the condition is tested.

munotes.in98

Practical 10: The Bank Management System, Written and Run

2. What does find_account return, and what does minus 1 mean? The subscript of the matching account, or minus 1 when there is none. Minus 1 is safe because it is not a valid subscript.

3. Why must the validation come before the balance is changed? So that a refused operation leaves the record exactly as it was. Changing first and checking afterwards would need the change undone.

4. What does scanf return, and why test it? The number of input items it read successfully. Testing it catches the input ending or the user typing something that is not a number, which would otherwise leave the program looping.

5. Why is the new account printed as bank[count - 1]? Because count has already been increased, so the record just written is one place below the new value.

6. How would you make the accounts survive the program ending? By writing the array to a file with fopen and fwrite before exit, and reading it back at the start.

7. Which single change would let the bank hold more than a hundred accounts? Raising the constant the array is declared with. Every loop already runs to count rather than to a fixed number, so nothing else has to change.

Contents This chapter on its own page

munotes.in99

Module II

munotes.in

Chapter Twenty-Eight

Getting Into MySQL, and What a Database Is

Syllabus topic Major Practical 1, Module 2, particulars row 5: "30 Hours(DBMS - Practical)". The mechanics every one of her ten DBMS practicals is performed with

In one line

A database is an organised collection of related data. A database management system, or DBMS, is the program that stores it, protects it, and answers questions about it. SQL is the language you ask those questions in.

In the wording you can use in a viva: a DBMS is software that defines, creates, maintains and controls access to a database, providing data definition, data manipulation, data integrity, data security, concurrent access and recovery.

Why not just keep the data in files

Module 1 ended with a bank whose accounts vanished when the program stopped. The obvious answer is to write them to a file, and that answer runs into six problems that a DBMS exists to solve.

Problem with plain filesWhat a DBMS does about it
The same fact is stored in several files and they disagreeone copy, referred to from everywhere
Every program must know the file's exact layoutprograms ask in SQL and never see the layout
Nothing stops a nonsense value being writtenconstraints refuse it
Two programs writing at once corrupt the filetransactions and locking
Anybody who can read the file can read everythingprivileges, granted per user and per table
A crash halfway through leaves the file half writtenrecovery, and the whole change is undone

Those six are the standard answer to "advantages of a DBMS over a file system", and each of them is a practical in this module.

The words you need before anything else

WordWhat it means
Table, or relationa grid of data about one kind of thing, such as books
Row, record or tupleone of them, such as one book
Column, field or attributeone property, such as the title
Primary keythe column whose value identifies a row uniquely
Foreign keya column holding a primary key value from another table, which is how tables are joined
Schemathe structure: the tables, their columns and their types
Instancethe data actually in them at this moment
SQLStructured Query Language, the language all of this is written in
Queryone question or one instruction written in SQL

The three pairs in rows two to four are worth knowing in both forms. The first word of each pair is what a person says; the second is what a textbook says; the third is what the relational model says.

Getting to a prompt

MU names no particular DBMS. Her reference list names MySQL twice, and her Course Objectives are plain SQL, so this book uses MySQL, and every listing in it was run on one.

mysql -u root -p

-u root is the user name, -p asks for the password. Your college laboratory will have its own user and password; ask, and write them on the inside cover of your journal, not on a loose sheet. If the laboratory uses phpMyAdmin or MySQL Workbench instead of the command line, the same statements are typed into a query window and the results appear in a grid; nothing in this book changes.

munotes.in100

Getting Into MySQL, and What a Database Is

Once in, the prompt reads mysql> and it is waiting.

The one rule about semicolons

Every statement ends with a semicolon. The client does not send a statement until it sees one.

This is the first thing that goes wrong for everybody. You type SELECT * FROM book and press Enter, and the prompt changes to -> and nothing happens. The client is not stuck; it is waiting for the rest of the statement. Type a semicolon and press Enter, and it runs.

To abandon a half typed statement, type \c and press Enter.

Which server is this

SELECT VERSION() AS server_version;
+----------------+
| server_version |
+----------------+
| 9.6.0          |
+----------------+

That is the version this book was written and checked on. Yours will print its own number, and nothing in this module depends on the difference. Where a chapter does turn on a version, it says so in its own words.

Making the database this book uses

Everything in Module 2 works on one small library: authors, the books they wrote, the members of the library, and the loans. It is the same library whose book record MU's own Module 1 Practical 9 asked for, with the members Title, Author, Subject and Book ID.

First, check that the name is free. SHOW DATABASES lists every database on the server, and LIKE narrows the list:

SHOW DATABASES LIKE 'library%';

Nothing came back at all: not a table with no rows in it, just nothing. So no database whose name begins with library exists yet, and the name is free. An empty result is not an error; it is an answer.

CREATE DATABASE librarydb;

Nothing is printed because nothing was asked for. Silence after a statement that changes something is normal, and it means it worked.

Creating a database does not put you inside it. USE does that:

USE librarydb;
SELECT DATABASE() AS now_using;
+-----------+
| now_using |
+-----------+
| librarydb |
+-----------+

From here until the end of this session, every table you name is looked for in librarydb. SELECT DATABASE() is how you ask which one you are in, and it is worth typing whenever a statement complains that a table does not exist: half the time you are in the wrong database.

What is in it so far

SHOW TABLES;

Nothing again, which is correct: a new database has no tables. Filling it is [Practical 2: Viewing Databases, Creating One and Listing Its Tables] onward.

munotes.in101

Getting Into MySQL, and What a Database Is

How the chapters from here on work

Every chapter of Module 2 after the three ER chapters begins with the library already built and its rows already in place, and says so in its own opening. You do not rebuild it for each practical. If you want to start again from scratch at any point:

DROP DATABASE librarydb;
CREATE DATABASE librarydb;
USE librarydb;

and then run the table creation of [Practical 2: Creating a Table, and Choosing Its Data Types] again.

Upper case, lower case, and layout

SQL keywords may be written in any case: select, SELECT and SeLeCt are the same word. Table and column names are case sensitive on Linux and not on Windows, which is a real source of confusion when a program written in a laboratory is run at home.

The convention this book follows, and which every examiner recognises:

  • Keywords in capitals: SELECT, FROM, WHERE, ORDER BY.
  • Names in small letters: book, member_name.
  • One clause per line when a query is longer than a line.

It is a convention, not a rule, and it exists so that a long query can be read.

What beginners get wrong

Forgetting the semicolon. The prompt turns into -> and waits.

Thinking CREATE DATABASE also selects it. It does not; USE does.

Working in the wrong database. SELECT DATABASE(); tells you which one.

Expecting output from a statement that changes something. Silence means success.

Typing a query into the shell instead of into mysql. The shell has never heard of SELECT.

Writing the password on the command line after -p. It is visible to everyone on the machine. Let it ask.

Quick revision

  • A database is organised related data; a DBMS manages it; SQL is the language.
  • Advantages over files: no duplication, programs do not know the layout, constraints, concurrency, security, recovery.
  • Table equals relation; row equals record equals tuple; column equals field equals attribute.
  • A primary key identifies a row; a foreign key points at another table's primary key.
  • Every statement ends with a semicolon; \c abandons one.
  • SHOW DATABASES;, CREATE DATABASE name;, USE name;, SELECT DATABASE();, SHOW TABLES;.
  • Keywords are not case sensitive; table names may be, depending on the operating system.

What goes in your journal

This chapter is not one of MU's ten practicals, so it has no entry of its own. Write the connection command, the database name, and the semicolon rule on the inside cover with the compile command from Module 1. You will need all three every week.

Test yourself

1. What is the difference between a database and a DBMS? The database is the data. The DBMS is the software that stores, protects and gives access to it.

munotes.in102

Getting Into MySQL, and What a Database Is

2. Give three advantages of a DBMS over keeping data in ordinary files. Data is not duplicated and cannot disagree with itself; constraints refuse invalid values; access can be granted per user and per table. Concurrency and recovery are two more.

3. What are the three names for one row of a table? Row, record and tuple.

4. You typed a SELECT and the prompt changed to ->. What happened? The semicolon was missing, so the client is still waiting for the end of the statement.

5. What does USE librarydb; do that CREATE DATABASE librarydb; does not? It makes that database the current one, so that later statements find its tables. Creating a database does not select it.

6. A statement printed nothing at all. Did it fail? No. A statement that changes something rather than asking something prints nothing when it succeeds. An error would have printed an error.

Contents This chapter on its own page

munotes.in103

Chapter Twenty-Nine

Practical 1: The ER Diagram: Entities, Attributes and Keys

Syllabus topic Module 2, Practical 1: "Conceptual Designing using ER Diagrams (Identifying entities, attributes, keys and relationships between entities, cardinalities, generalization, specialization etc.)"

Aim

To design a database conceptually using an ER diagram, identifying entities, attributes and keys.

What an ER diagram is for

Before a single table is created, somebody has to decide what the database is about. An entity relationship diagram is how that decision is written down and shown to people who do not write SQL.

It was proposed by Peter Chen in 1976, and its shapes have barely changed since:

ShapeWhat it stands for
Rectanglean entity set: a kind of thing the database records
Ellipsean attribute: a property of a thing
Diamonda relationship set: how two kinds of thing are connected
Linejoins a shape to what it belongs to
Underlinemarks the key attribute

An ER diagram is a conceptual design. It has no tables, no data types and no SQL in it, and that is deliberate: it can be shown to a librarian who knows nothing about databases and they can tell you whether it is right. Turning it into tables comes later, in [Practical 8: Turning the ER Model Into Tables].

Entity, entity set, attribute

An entity is one particular thing: the copy of The C Language on the shelf, with book id 1.

An entity set is all the things of that kind: every book in the library. The rectangle in the diagram is the entity set, and a common examination slip is to draw one rectangle per book.

An attribute is a property every entity in the set has: a book has a title, a subject, a price and a date it was added.

A BOOK entity set with its attributes

Figure 29.1 An entity set is a rectangle; each attribute is an ellipse joined to it; the key is underlined

The kinds of attribute, which examiners ask for by name

KindMeansExampleDrawn as
Simplecannot be broken down usefullypriceplain ellipse
Compositemade of parts that are useful on their ownaddress, made of street, city, pinan ellipse with smaller ones hanging off it
Single valuedexactly one per entitybook_idplain ellipse
Multivaluedseveral per entitya member's phone numbersdouble ellipse
Derivedworked out from others, not storedage, from the date of birthdashed ellipse
Keyidentifies the entity uniquelybook_idunderlined

The distinction between stored and derived is the one that carries a mark. Date of birth is stored; age is derived, because storing it would mean it is wrong from tomorrow.

Keys

A key is an attribute, or a combination of attributes, whose value is different for every entity in the set. It is how one row is told apart from all the others.

TermWhat it means
Super keyany set of attributes that identifies a row uniquely, even with spare attributes in it
Candidate keya super key with nothing spare: remove any attribute and it stops being unique
Primary keythe candidate key chosen to be the one used, and underlined in the diagram
Alternate keya candidate key that was not chosen
Composite keya key made of more than one attribute together
Foreign keyan attribute holding the primary key value of another entity set
munotes.in104

Practical 1: The ER Diagram: Entities, Attributes and Keys

Worked on the library's members. Suppose each member has member_id, an email address, a name and a city.

  • {member_id} is a candidate key: no two members share one.
  • {email} is another candidate key, if the library insists on different addresses.
  • {member_id, name} is a super key but not a candidate key, because name is spare.
  • {name} is not a key at all: two members may both be called Asha Kulkarni.
  • Choose member_id as the primary key; email is then an alternate key.

Choose a primary key that is short, never changes, and is never unknown. That is why member_id beats email: a member may change their email, and the value would then have to be changed everywhere it has been used.

The four-step method

This is what to do in an examination when handed a paragraph of English.

Step 1: underline every noun. Each is a candidate for an entity set or an attribute. It is an entity set if it has properties of its own, or if there can be many of them for one owner, or if more than one other thing refers to it. Otherwise it is an attribute.

Step 2: underline every verb that joins two nouns. Each is a candidate relationship set.

Step 3: for every entity set, find a key. If there is none, it is a weak entity set, which is covered in [Practical 1: Relationships and Cardinality].

Step 4: for every relationship, ask how many on each side. That is the cardinality, and it is the next chapter.

Worked example: the library, from a paragraph

The library keeps books. Each book has a number, a title, a subject, a price and the date it was added, and is written by one author. An author has a name and a country and may have written several books, or none yet. Members join the library on a date and have a name, a course and a city. A member may borrow many books and a book may be borrowed by many members over time, and for each borrowing the library records the date it was issued and the date it came back.

Step 1: the nouns.

NounEntity set or attributeWhy
libraryneitherit is the whole database, not a thing in it
bookentity setit has five properties of its own
number, title, subject, price, date addedattributes of BOOKno properties of their own
authorentity setit has a name and a country, and many books point at it
name, countryattributes of AUTHOR
memberentity setit has properties, and there are many
name, course, city, date joinedattributes of MEMBER
borrowingsee step 2it is what the verb produces
munotes.in105

Practical 1: The ER Diagram: Entities, Attributes and Keys

author is the interesting decision. If a book simply carried an author's name as text, author would be an attribute. It is an entity set here because the paragraph gives it a property of its own, the country, and because one author may have several books: storing the country beside every book would repeat it, and two rows could then disagree about where the same author is from. That repetition is the signal that a noun should be an entity set of its own.

Step 2: the verbs.

VerbRelationship setBetween
is written bywritesAUTHOR and BOOK
may borrowborrowsMEMBER and BOOK

Step 3: the keys. BOOK is identified by its number, so book_id is the primary key. AUTHOR gets author_id and MEMBER gets member_id. None of the three has a natural key that is short, unchanging and always known, which is why each is given a number of its own.

Step 4: where do issued_on and returned_on go? Not on BOOK, because a book is borrowed many times and would need many dates. Not on MEMBER, for the same reason. They belong to the borrowing itself, so they are attributes of the relationship, drawn as ellipses hanging off the diamond.

The library as an ER diagram

Figure 29.2 The whole library: three entity sets, two relationships, and two attributes that belong to a relationship rather than to either entity set

That last point is the one most often got wrong and most often asked about. An attribute that needs two entities named before it means anything belongs to the relationship between them. A mark in a marks database is the same: it needs a student and a subject.

What it does NOT mean

A rectangle is not a table. It becomes one later, usually, but the diagram is deliberately free of tables so that it can be discussed with people who do not know SQL.

A rectangle is not one thing. It is the whole set of them. One rectangle marked BOOK stands for every book in the library.

An ellipse is not a column. Multivalued and composite attributes do not become single columns at all; what happens to them is [Practical 8: Turning the ER Model Into Tables].

munotes.in106

Practical 1: The ER Diagram: Entities, Attributes and Keys

What beginners get wrong

Drawing one rectangle per entity instead of per entity set.

Making everything an entity set. A price has no properties of its own and nothing else refers to it, so it is an attribute.

Making everything an attribute. Then the same author's country is written beside every one of their books, and two of them can disagree.

Storing a derived value. Age instead of date of birth is wrong by tomorrow.

Choosing a primary key that can change. An email or a phone number is a poor primary key for that reason.

Hanging a relationship's attribute off an entity set. issued_on on BOOK cannot work, because a book is borrowed more than once.

Quick revision

  • ER diagram: rectangle is an entity set, ellipse an attribute, diamond a relationship set, underline the key.
  • Entity is one thing; entity set is all of them.
  • Attributes: simple, composite, single valued, multivalued, derived, key.
  • Derived attributes are not stored; a dashed ellipse marks them.
  • Super key has spares; candidate key has none; primary key is the chosen candidate; the rest are alternate keys.
  • A primary key should be short, unchanging and never unknown.
  • Nouns become entity sets or attributes; verbs become relationships.
  • An attribute needing two entities named belongs to the relationship between them.

What goes in your journal

Aim, the problem statement as you were given it, the four steps with your own noun and verb tables, the list of entity sets with their attributes and their keys underlined, and the diagram drawn with a ruler. Write one line under the diagram justifying one decision, such as why AUTHOR is an entity set and not an attribute. That single line is what a viva asks about.

Test yourself

1. What is the difference between an entity and an entity set? An entity is one particular thing; an entity set is the whole collection of things of that kind. The rectangle stands for the set.

2. What shape is used for a relationship, and what for an attribute? A diamond for a relationship set, an ellipse for an attribute.

3. Distinguish a super key, a candidate key and a primary key. A super key identifies a row uniquely but may contain spare attributes. A candidate key is a super key with nothing spare. The primary key is the candidate key chosen for use; the others become alternate keys.

4. Why is age a derived attribute? Because it can be worked out from the date of birth. Storing it would make it wrong as soon as the person's birthday passes.

5. Where do issued_on and returned_on belong, and why? To the borrows relationship, because they mean nothing until both a member and a book are named, and one book is borrowed many times.

munotes.in107

Practical 1: The ER Diagram: Entities, Attributes and Keys

6. When should a noun be made an entity set rather than an attribute? When it has properties of its own, when there can be several of it for one owner, or when putting it inside another entity set would repeat it and let two copies disagree.

Contents This chapter on its own page

munotes.in108

Chapter Thirty

Practical 1: Relationships and Cardinality

Syllabus topic Module 2, Practical 1: "... relationships between entities, cardinalities ..."

Aim

To identify the relationships between entities and their cardinalities.

What a relationship is

A relationship connects entities. writes connects an author to a book; borrows connects a member to a book.

A relationship set is all the connections of that kind, and it is the diamond in the diagram. The degree of a relationship set is how many entity sets it joins:

DegreeNameExample
2binarya member borrows a book
3ternarya doctor prescribes a drug to a patient
1unary, or recursivean employee supervises an employee

Almost every relationship you will be asked to draw is binary. The unary case surprises students: one entity set joined to itself, such as a student being the class representative of other students, and it is drawn with two lines from the diamond back to the same rectangle, each labelled with the role.

Cardinality: the two questions

The cardinality ratio says how many entities of one set may be joined to one entity of the other. There are three ratios, and two questions decide which one applies.

Question one: for ONE entity on the left, how many on the right? Question two: for ONE entity on the right, how many on the left?

Ask them in exactly that order, out loud, with a concrete example. The answers are "one" or "many", and the two answers together give the ratio.

The three cardinality ratios

Figure 30.1 One to one, one to many and many to many, with the letter on each line saying how many of that side

AnswersRatioWrittenExample from a library
one, oneone to one1:1a member has one library card and a card belongs to one member
one, manyone to many1:Nan author writes many books, a book has one author
many, manymany to manyM:Na member borrows many books, a book is borrowed by many members

Reading the letters the right way round

This is where marks are lost, because the letter goes on the line beside the entity set it counts.

On the writes relationship in the library diagram, the line to AUTHOR carries a 1 and the line to BOOK carries an N. Read it as: for one author there are N books, and for one book there is 1 author.

The test that settles any argument: cover one side of the diagram with your hand and read the letter on the other. If you cannot say the sentence out loud, the letters are the wrong way round.

Working the library's two relationships

writes, between AUTHOR and BOOK.

  • For one author, how many books? Many. An author may have written several, and one has written none yet.
  • For one book, how many authors? One, in this design. The paragraph said "is written by one author".
  • So it is 1:N, one author to many books.
munotes.in109

Practical 1: Relationships and Cardinality

borrows, between MEMBER and BOOK.

  • For one member, how many books? Many, over time.
  • For one book, how many members? Many, over time, though only one at a time.
  • So it is M:N.

Notice what decided the second one: the phrase over time. If the question had been "how many members hold this book right now", the answer would be one and the relationship would be 1:N. A cardinality is only meaningful once you have said at what moment you are counting, and saying so is worth a mark.

Participation: must every entity take part

Cardinality says how many. Participation says whether an entity may sit out.

ParticipationMeansDrawn as
Totalevery entity of the set must take part in at least one relationshipa double line
Partialan entity may take part in nonea single line

In the library: every book must have an author, so BOOK's participation in writes is total and gets a double line. An author need not have written any book yet, so AUTHOR's participation is partial and gets a single line.

Total participation is also called an existence dependency, and it is what becomes a NOT NULL on a foreign key when the diagram is turned into tables.

Weak entity sets

Some things cannot be identified on their own.

A weak entity set has no key of its own. Its entities are identified only in combination with another entity set, called the owner or identifying entity set, and the relationship between them is the identifying relationship.

FeatureDrawn as
Weak entity setdouble rectangle
Identifying relationshipdouble diamond
Partial key, or discriminatordashed underline
Participation of the weak setalways total, so a double line

The library's example: if the library recorded each physical copy of a book as copy 1, copy 2, copy 3, then copy number 2 means nothing on its own, because every book has a copy 2. A copy is identified only as "copy 2 of book 1001". COPY is a weak entity set, book_id plus copy_no identifies it, and copy_no alone is its discriminator.

A weak entity set is always totally participating in its identifying relationship, because a copy cannot exist without a book.

What beginners get wrong

Putting the letters on the wrong lines. Cover one side and read the other.

Deciding a cardinality without saying at what moment. A book is borrowed by many members over time and by one at a time.

Confusing cardinality with participation. Cardinality is how many; participation is whether any at all. A relationship can be 1:N with total participation on both sides, or on neither.

munotes.in110

Practical 1: Relationships and Cardinality

Drawing an arrow the wrong way. Where arrows are used instead of letters, the arrow points at the side that has one. That is the opposite of many students' guess.

Calling a relationship with an attribute a weak entity set. A weak entity set has no key. A relationship with an attribute simply has an attribute.

Forgetting that a relationship may have attributes at all.

Quick revision

  • Degree: binary joins two entity sets, ternary three, unary one to itself.
  • Cardinality ratios: 1:1, 1:N and M:N.
  • Two questions: for one on the left how many on the right, and for one on the right how many on the left.
  • The letter sits beside the entity set it counts.
  • Participation: total is a double line and means every entity must take part; partial is a single line.
  • Weak entity set: no key of its own, double rectangle, identifying relationship a double diamond, discriminator dashed underlined, participation always total.
  • A cardinality is only meaningful once the moment of counting is stated.

What goes in your journal

Aim, the list of relationships you found, and for each of them the two questions written out with their answers and the ratio that follows. The two questions are the working, and a journal that shows only the final letters has left out what the practical is teaching. Then the diagram with every line labelled and every total participation drawn double.

Test yourself

1. What are the three cardinality ratios? One to one, one to many and many to many, written 1:1, 1:N and M:N.

2. Which two questions decide the ratio? For one entity on the left, how many on the right, and for one entity on the right, how many on the left.

3. What is the difference between cardinality and participation? Cardinality says how many entities of one set may join to one of the other. Participation says whether an entity may take part in none at all.

4. How is total participation drawn, and what does it become in a table? A double line, and it becomes a NOT NULL on the foreign key.

5. What is a weak entity set, and how is it identified? An entity set with no key of its own. It is identified by its owner's key together with its own discriminator, and it is drawn as a double rectangle joined by a double diamond.

6. Is the relationship between a member and a book one to many or many to many? Many to many over time, because a member borrows many books and a book is borrowed by many members. It is one to many only if the question is about who holds it at this moment.

Contents This chapter on its own page

munotes.in111

Chapter Thirty-One

Practical 1: Generalization and Specialization

Syllabus topic Module 2, Practical 1: "... generalization, specialization etc."

Aim

To apply generalization and specialization to an ER design.

The problem both of them solve

Suppose the library has two kinds of member: students, who have a course, and staff, who have a department. Both kinds have a member id, a name and a date of joining.

Two entity sets, and most of their attributes are the same:

STUDENTSTAFF
member_idmember_id
namename
joined_onjoined_on
coursedepartment

That repetition is the smell. If a third kind of member appears later, the shared attributes are written a third time, and a change to one of them has to be made in three places.

Generalization: bottom up

Generalization takes two or more entity sets that share attributes and pulls the shared ones up into a new, more general entity set.

STUDENT and STAFF become subclasses of a new MEMBER, which holds member_id, name and joined_on. STUDENT keeps only course; STAFF keeps only department.

Generalization: MEMBER with STUDENT and STAFF beneath it

Figure 31.1 The superclass holds what they share; each subclass adds only what is its own

The triangle marked ISA is the notation, and it is read "is a": a student is a member. The superclass is above it, the subclasses below.

Generalization is bottom up. You start with the specific sets, notice what they share, and invent the general one.

Specialization: top down

Specialization is the same picture arrived at from the other end. You start with MEMBER, notice that some members have a course and others have a department, and split the set into subclasses.

GeneralizationSpecialization
Directionbottom uptop down
You start withseveral specific entity setsone general entity set
You are looking forwhat they have in commonhow they differ
Resulta new superclassnew subclasses
Drawn asthe same ISA trianglethe same ISA triangle

The picture is identical; only the thinking that produced it differs. That is the whole distinction, and it is exactly what a viva asks. Do not look for a difference in the diagram, because there is none.

Attribute inheritance

A subclass inherits every attribute of its superclass, and every relationship the superclass takes part in.

So STUDENT has member_id, name, joined_on and course, even though only course is drawn beside it. And if MEMBER borrows books, a student borrows books without the relationship being drawn again.

That is what makes the arrangement worth having: say a thing once, and everything below inherits it.

The two constraints, which are asked by name

Disjoint or overlapping. May one entity belong to more than one subclass at the same time?

  • Disjoint, sometimes marked d: no. A member is a student or a member of staff, not both.
  • Overlapping, marked o: yes. A person might be both a student and a member of staff, at a college that employs its own research students.
munotes.in112

Practical 1: Generalization and Specialization

Total or partial. Must every entity of the superclass belong to some subclass?

  • Total: yes, drawn with a double line into the triangle. Every member is either a student or staff, and there is no third kind.
  • Partial: no, a single line. There may be members who are neither, such as an alumnus.

The two constraints are independent, so there are four combinations, and a complete answer names one from each pair: "disjoint and total", "overlapping and partial", and so on.

Aggregation, the other thing the triangle is confused with

Aggregation treats a whole relationship as though it were an entity set, so that another relationship can be attached to it. It is drawn as a box round the diamond and the two rectangles it joins.

The library's example: a member borrows a book, and the library wants to record which member of staff approved a particular borrowing. approved_by connects a member of staff to the borrowing, not to the member and not to the book separately, so the borrows relationship is boxed and the new diamond is attached to the box.

It is not generalization and it is not a weak entity set. Examiners ask students to tell the three apart, and the answer is:

What it isDrawn as
Generalizationa subclass is a kind of superclassISA triangle
Weak entity setan entity that cannot be identified alonedouble rectangle, double diamond
Aggregationa relationship treated as an entity seta box round a relationship

Where this goes next

A superclass and its subclasses can be turned into tables in three ways, and choosing between them is part of [Practical 8: Turning the ER Model Into Tables]:

One table per entity set. MEMBER, STUDENT and STAFF, with the subclass tables holding the key and their own attributes only. This is the general answer and works for every combination of constraints.

One table per subclass only. STUDENT and STAFF each hold all the inherited attributes as well. This works only when the specialization is total and disjoint; otherwise an entity belonging to no subclass has nowhere to live.

One table for everything. A single MEMBER table with course and department both in it, and one of them null for each row, plus a column saying which kind it is. Simple, and it fills the table with nulls.

What beginners get wrong

Looking for a difference between the two diagrams. There is none. The difference is in how you got there.

Repeating the inherited attributes beside the subclasses. They are inherited; drawing them again is a mistake, not extra detail.

munotes.in113

Practical 1: Generalization and Specialization

Confusing disjoint with total. Disjoint is about belonging to two subclasses at once; total is about belonging to none.

Using an ISA triangle for a relationship. A student is a member; a member borrows a book. The first is ISA, the second is a diamond. "Is a" against "does something to" is the test.

Specialising on a value that changes every day. If the subclass a row belongs to changes often, a column holding the kind is simpler than two tables.

Quick revision

  • Generalization is bottom up: pull shared attributes into a new superclass.
  • Specialization is top down: split a superclass into subclasses.
  • The diagram is the same for both: an ISA triangle.
  • Subclasses inherit every attribute and every relationship of the superclass.
  • Disjoint means an entity is in at most one subclass; overlapping means it may be in several.
  • Total means every superclass entity is in some subclass; partial means it need not be.
  • Aggregation boxes a relationship so another relationship can attach to it.
  • Three ways to make tables from it: one per entity set, one per subclass, or one for everything.

What goes in your journal

Aim, the two subclasses with the attributes they share listed separately from the attributes that are their own, the diagram with the ISA triangle, and a line naming both constraints, such as "disjoint and total". In the conclusion write the one sentence distinguishing generalization from specialization, because that sentence is the viva question for this half of the practical.

Test yourself

1. What is the difference between generalization and specialization? The direction of thinking. Generalization starts with specific entity sets and pulls out what they share into a new superclass. Specialization starts with one general entity set and divides it. The resulting diagram is the same.

2. What does the ISA triangle mean? That each subclass is a kind of the superclass: a student is a member.

3. What does a subclass inherit? Every attribute of its superclass and every relationship the superclass takes part in.

4. What is the difference between a disjoint and an overlapping specialization? Disjoint means an entity may belong to at most one subclass. Overlapping means it may belong to more than one at the same time.

5. What does total specialization mean? That every entity of the superclass must belong to at least one subclass. It is drawn with a double line into the triangle.

6. How is aggregation different from generalization? Aggregation treats a whole relationship as an entity set so that another relationship can be joined to it, and is drawn as a box. Generalization arranges entity sets into a kind-of hierarchy with a triangle.

Contents This chapter on its own page

munotes.in114

Chapter Thirty-Two

Practical 2: Viewing Databases, Creating One and Listing Its Tables

Syllabus topic Module 2, Practical 2: "Viewing all databases", "Creating a Database", "Viewing all Tables in a Database"

Aim

To view all the databases on the server, to create a database, and to view all the tables in it.

Viewing all databases

SHOW DATABASES;

That lists every database the user you logged in as is allowed to see. On a college server the list is long; on your own machine it is short. Four of them are there on every MySQL installation and are not yours to touch:

information_schema holds a description of every database, table and column on the server.

mysql holds the server's own accounts and privileges.

performance_schema holds measurements of what the server is doing.

sys holds easier-to-read views over performance_schema.

information_schema is worth remembering rather than avoiding: it is how a program finds out what tables exist, and it is used at the end of this chapter.

Because the full list differs from machine to machine, narrow it when you are looking for something in particular:

SHOW DATABASES LIKE 'library%';

Nothing came back, so nothing beginning with library exists here yet. LIKE uses % to mean "any characters", which is the same pattern language as the WHERE ... LIKE of [Practical 4: Simple Queries].

Creating a database

CREATE DATABASE librarydb;

Nothing is printed, which means it worked.

Run it a second time and the server refuses, because the name is taken:

CREATE DATABASE librarydb;
ERROR 1007 (HY000): Can't create database 'librarydb'; database exists

That error has a number, 1007, and a five character SQLSTATE, HY000. Both are worth noticing: a program that talks to MySQL tests the number, not the English, because the English changes between versions and languages.

When you do not care whether it already exists, say so:

CREATE DATABASE IF NOT EXISTS librarydb;

That prints nothing and does nothing, because the database is already there. IF NOT EXISTS turns the error into a warning, which is what you want in a script that may be run twice.

What the server actually created

You asked for a name and nothing else, and the server filled in two more things.

SELECT default_character_set_name AS charset,
       default_collation_name     AS collation
FROM information_schema.schemata
WHERE schema_name = 'librarydb';
+---------+--------------------+
| charset | collation          |
+---------+--------------------+
| utf8mb4 | utf8mb4_0900_ai_ci |
+---------+--------------------+

A character set decides what characters may be stored. A collation decides how they sort and compare.

utf8mb4 is the character set that can hold every character in Unicode, Devanagari and emoji included, in up to four bytes each. It is the sensible default and it is what modern MySQL chooses. The older utf8 in MySQL is a three byte version that cannot hold all of Unicode, which is why the name with mb4 exists at all.

The ai_ci on the end of the collation name means accent insensitive, case insensitive, so Mumbai and mumbai compare equal. That is why WHERE title = 'learn sql' finds Learn SQL later in this book, and it surprises students who expect SQL to behave like C.

munotes.in115

Practical 2: Viewing Databases, Creating One and Listing Its Tables

SHOW CREATE DATABASE librarydb; prints the same two facts as one long line, in the form of the statement that would recreate the database. It is the quicker thing to type and the harder thing to read.

To choose them yourself rather than take what you are given:

CREATE DATABASE librarydb
  CHARACTER SET utf8mb4
  COLLATE utf8mb4_0900_ai_ci;

Selecting it, and confirming which one you are in

USE librarydb;
SELECT DATABASE() AS now_using;
+-----------+
| now_using |
+-----------+
| librarydb |
+-----------+

USE is one of the few statements the interactive client accepts without a semicolon. Write the semicolon anyway, so that the same line also works inside a script. SELECT DATABASE() is the question "where am I", and it is the first thing to type when a statement says a table does not exist.

Viewing all the tables in it

SHOW TABLES;

Nothing, because the database is new. Once tables exist, the same statement lists them, and [Practical 2: Creating a Table, and Choosing Its Data Types] makes some.

You can look inside another database without leaving this one:

SHOW TABLES FROM othername;

And you can narrow the list exactly as with databases:

SHOW TABLES LIKE 'b%';

The same question asked of information_schema

Everything SHOW tells you is also stored as ordinary rows in information_schema, and it can be queried like any table.

SELECT schema_name AS found
FROM information_schema.schemata
WHERE schema_name = 'librarydb';
+-----------+
| found     |
+-----------+
| librarydb |
+-----------+

That matters for two reasons. It is how a program asks the question, because a program can put a WHERE on a query and cannot put one on SHOW. And it is a fair viva question: SHOW is a convenience for a person; information_schema is the same information as data.

Removing a database

DROP DATABASE librarydb;

There is no confirmation and no undo. Every table and every row in it is gone. On a shared college server, check twice which database you are naming, and never type it with a wildcard or from memory. DROP DATABASE IF EXISTS name; is the form that does not complain when there is nothing to drop.

The statements in this practical

StatementWhat it does
SHOW DATABASES;lists every database you may see
SHOW DATABASES LIKE 'pattern';the same list, narrowed
CREATE DATABASE name;makes one
the same with IF NOT EXISTSmakes one, and says nothing when it is already there
SHOW CREATE DATABASE name;the statement that would recreate it, character set and all
USE name;selects it for the rest of the session
SELECT DATABASE();which one am I in
SHOW TABLES;lists the tables in the current database
SHOW TABLES FROM name;lists another database's tables
DROP DATABASE name;deletes it and everything in it
munotes.in116

Practical 2: Viewing Databases, Creating One and Listing Its Tables

What beginners get wrong

Expecting CREATE DATABASE to select the database. It does not; USE does.

Running SHOW TABLES with no database selected. The server answers ERROR 1046 (3D000): No database selected.

Being surprised that a second CREATE DATABASE fails. Use IF NOT EXISTS when you do not care.

Typing DROP DATABASE on a shared server without checking the name.

Assuming table names behave the same everywhere. On Linux they are case sensitive; on Windows they are not, so a query that works in the laboratory may fail at home.

Treating information_schema, mysql, performance_schema or sys as spare space. They belong to the server.

Quick revision

  • SHOW DATABASES; lists them; four system databases are always there.
  • CREATE DATABASE name; makes one, and IF NOT EXISTS stops the error on a second run.
  • SHOW CREATE DATABASE name; shows the character set and collation it was given.
  • utf8mb4 holds every Unicode character; the collation ending ai_ci compares without regard to accents or case.
  • USE name; selects; SELECT DATABASE(); asks which is selected.
  • SHOW TABLES; lists the current database's tables; SHOW TABLES FROM other; lists another's.
  • information_schema holds the same information as queryable rows.
  • DROP DATABASE has no undo.

What goes in your journal

Aim, every statement with the server's reply written underneath it exactly as it appeared, including the error from the second CREATE DATABASE. That error is the evidence that you tried it rather than copied it, and an examiner notices.

Test yourself

1. Which four databases exist on every MySQL server? information_schema, mysql, performance_schema and sys.

2. What does CREATE DATABASE print when it succeeds? Nothing. Silence after a statement that changes something means it worked.

3. What is the difference between CREATE DATABASE x; and CREATE DATABASE IF NOT EXISTS x;? The first is an error if x already exists. The second does nothing and carries on.

4. What is a collation? The set of rules for comparing and sorting the characters of a character set. A collation ending ai_ci ignores accents and case, so Mumbai and mumbai compare equal.

5. How do you list the tables of a database you are not currently in? SHOW TABLES FROM thatdatabase;

6. Why would a program query information_schema instead of using SHOW? Because information_schema holds the same information as ordinary rows, so a query can filter, join and sort it. SHOW cannot take a WHERE.

Contents This chapter on its own page

munotes.in117

Chapter Thirty-Three

Practical 2: Creating a Table, and Choosing Its Data Types

Syllabus topic Module 2, Practical 2: "Creating Tables (With and Without Constraints)", the without half

Aim

To create tables without constraints, choosing an appropriate data type for each column.

The statement

CREATE TABLE tablename (
    column1  type,
    column2  type,
    ...
);

A comma after every column except the last. The round brackets are required. The whole thing ends with a semicolon like any other statement.

Choosing a type

Every column has a type, and the type decides three things at once: what may be stored, how much room it takes, and what the column can be compared and sorted against. Getting it wrong is not a style question. A date kept as text cannot be subtracted from another date, and a number kept as text sorts 10 before 9.

Whole numbers

TypeRange, signedBytes
TINYINTminus 128 to 1271
SMALLINTminus 32768 to 327672
MEDIUMINTabout minus 8 million to 8 million3
INTabout minus 2.1 billion to 2.1 billion4
BIGINTabout minus 9.2 quintillion upward8

Add UNSIGNED and the range moves: a TINYINT UNSIGNED holds 0 to 255. Use it where a negative value is meaningless, such as a count.

INT(11) appears in old examples and in old textbooks. The number in brackets was only a display width and never a limit on the value, and MySQL 8.0.17 deprecated it. Write INT, not INT(11).

Numbers with a fractional part

TypeHow it is storedUse it for
DECIMAL(p, s)exactly, as digits: p digits in all, s of them after the pointmoney, and anything that must be exact
FLOAT, DOUBLEapproximately, in binarymeasurements, scientific quantities

DECIMAL(7,2) holds up to 99999.99 exactly. That is the right type for a price.

FLOAT and DOUBLE are the SQL versions of the C types from [Practical 1(a): Simple Interest], and they carry the same surprise: one tenth cannot be written exactly in binary, so a sum of prices held in a FLOAT can come out a paisa wrong. Money goes in DECIMAL. This is the single most useful thing in this chapter.

Text

TypeLengthPaddedUse it for
CHAR(n)always n charactersyes, with spacescodes that are always the same length
VARCHAR(n)up to n, as long as the value needsnonames, titles, anything varying
TEXTup to about 65 thousand charactersnolong passages

CHAR(20) holding "Asha" occupies twenty characters; the sixteen spaces are stored and then thrown away when the value is read back. VARCHAR(20) holding "Asha" occupies four characters plus one byte recording the length.

So use CHAR only when every value really is the same length, such as a two letter state code or a fixed-format identifier. For a title or a name, VARCHAR.

The n in VARCHAR(n) is a maximum, not a reservation. VARCHAR(200) costs nothing extra for a short value, so there is no reason to make it uncomfortably tight, and every reason not to make it so wide that a wrong value goes unnoticed.

munotes.in118

Practical 2: Creating a Table, and Choosing Its Data Types

Dates and times

TypeHoldsWritten as
DATEa date'2025-07-21'
TIMEa time of day'14:30:00'
DATETIMEboth'2025-07-21 14:30:00'
TIMESTAMPboth, and converts to the session's time zonethe same
YEARa year2025

The format is YYYY-MM-DD, always, whatever your country writes on paper. '21-07-2025' is not a date to MySQL, and the effort of writing them her way is repaid the moment you want the difference between two of them in [Practical 5: Date Functions].

The tables, created

CREATE TABLE author (
    author_id    INT,
    author_name  VARCHAR(18),
    country      VARCHAR(10)
);

CREATE TABLE book (
    book_id    INT,
    title      VARCHAR(20),
    subject    VARCHAR(12),
    price      DECIMAL(7,2),
    added_on   DATE,
    author_id  INT
);

Nothing printed, so both were created. Not one constraint anywhere: any column may be left empty, two books may have the same book_id, and author_id may hold 999 when no such author exists. That is what "without constraints" means, and it is why the next chapter exists.

Looking at what was made

DESCRIBE book;
+-----------+--------------+------+-----+---------+-------+
| Field     | Type         | Null | Key | Default | Extra |
+-----------+--------------+------+-----+---------+-------+
| book_id   | int          | YES  |     | NULL    |       |
| title     | varchar(20)  | YES  |     | NULL    |       |
| subject   | varchar(12)  | YES  |     | NULL    |       |
| price     | decimal(7,2) | YES  |     | NULL    |       |
| added_on  | date         | YES  |     | NULL    |       |
| author_id | int          | YES  |     | NULL    |       |
+-----------+--------------+------+-----+---------+-------+

DESCRIBE, which may be shortened to DESC, lists the columns with their types. Read the six columns of its answer:

Column of the answerMeans
Fieldthe column's name
Typeits data type
Nullwhether it may be left empty, and YES everywhere here
KeyPRI, UNI or MUL if it takes part in a key, and empty here
Defaultwhat goes in when no value is given
Extraanything else, such as auto_increment

Four of those six are about constraints, and all four are empty or permissive in this table. Come back to this output after the next chapter and every one of them will have something in it.

SHOW TABLES;

It printed the two names under a heading reading Tables_in_librarydb, which is why the heading is not reproduced here: it carries whatever your database is called. The same list, asked for as ordinary data, is:

SELECT table_name AS tables_here
FROM information_schema.tables
WHERE table_schema = DATABASE()
ORDER BY table_name;
munotes.in119

Practical 2: Creating a Table, and Choosing Its Data Types

+-------------+
| tables_here |
+-------------+
| author      |
| book        |
+-------------+

Two tables, where [Practical 2: Viewing Databases, Creating One and Listing Its Tables] found none.

The types the server really gave the columns

SELECT column_name AS col, data_type AS type,
       character_maximum_length AS max_len, is_nullable AS nullable
FROM information_schema.columns
WHERE table_name = 'book' AND table_schema = DATABASE()
ORDER BY ordinal_position;
+-----------+---------+---------+----------+
| col       | type    | max_len | nullable |
+-----------+---------+---------+----------+
| book_id   | int     |    NULL | YES      |
| title     | varchar |      20 | YES      |
| subject   | varchar |      12 | YES      |
| price     | decimal |    NULL | YES      |
| added_on  | date    |    NULL | YES      |
| author_id | int     |    NULL | YES      |
+-----------+---------+---------+----------+

Every column is nullable, which is the default and which the next chapter changes.

What beginners get wrong

CHAR where VARCHAR belongs. Names and titles vary in length.

FLOAT for money. Use DECIMAL.

A date stored as VARCHAR. It then cannot be compared, subtracted or sorted properly, and '09-01-2025' sorts before '10-12-2024'.

Writing INT(11). The bracketed number was only a display width and is deprecated.

A comma after the last column. The server answers with a syntax error pointing at the closing bracket.

Making every text column VARCHAR(255) without thinking. It works, and it means the table no longer documents what it holds.

Using a keyword as a column name, such as order or desc. It has to be written in backticks everywhere afterwards; choose a different name.

Quick revision

  • CREATE TABLE name (column type, column type, ...); with no comma after the last.
  • Whole numbers: TINYINT, SMALLINT, MEDIUMINT, INT, BIGINT, and UNSIGNED to move the range.
  • Exact fractions: DECIMAL(p, s). Money always.
  • Approximate: FLOAT and DOUBLE, never for money.
  • CHAR(n) is fixed length and padded; VARCHAR(n) is up to n and is not.
  • Dates are YYYY-MM-DD, always.
  • DESCRIBE table; lists the columns, their types, and whether each may be null.
  • Without constraints, every column may be empty and every value may repeat.

What goes in your journal

Aim, a table of the columns you chose with the type of each and one line saying why, then the CREATE TABLE statements, then the DESCRIBE output. The column of reasons is what is being marked: an examiner wants to see that price is DECIMAL because it is money and title is VARCHAR because titles vary in length.

Test yourself

1. What is the difference between CHAR(10) and VARCHAR(10)? CHAR(10) always occupies ten characters and pads short values with spaces. VARCHAR(10) occupies only what the value needs, up to ten, plus a byte for the length.

munotes.in120

Practical 2: Creating a Table, and Choosing Its Data Types

2. Which type should a price use, and why not FLOAT? DECIMAL, because it stores digits exactly. FLOAT stores an approximation in binary, so sums of money can come out a paisa wrong.

3. In what format must a DATE literal be written? 'YYYY-MM-DD', for example '2025-07-21'.

4. What does DESCRIBE book; tell you? The name and type of every column, whether it may be null, what key it takes part in, its default value, and anything extra such as auto increment.

5. What does INT(11) mean? Only a display width, never a limit on the value. It is deprecated; write INT.

6. What can go wrong in a table created with no constraints? Any column may be left empty, values that should be unique may repeat, and a column that should refer to another table may hold a value that is not there.

Contents This chapter on its own page

munotes.in121

Chapter Thirty-Four

Practical 2: The Constraints, and What Each One Refuses

Syllabus topic Module 2, Practical 2: "Creating Tables (With and Without Constraints)", the with half

Aim

To create tables with constraints, and to demonstrate what each constraint refuses.

What a constraint is

A constraint is a rule about what may be stored, enforced by the server rather than by your program. Once it is in place, no query, no program and no user can put a row in that breaks it.

That last sentence is the whole argument for constraints, and it is the answer to "why not check in the program instead". A program can be bypassed: somebody types a statement at the prompt, or a second program is written next year by somebody who does not know the rule. A constraint cannot be bypassed. It lives with the data.

The six MU's practicals need:

ConstraintRefuses
NOT NULLa missing value in that column
PRIMARY KEYa duplicate, and a missing value
UNIQUEa duplicate
DEFAULTnothing; it supplies a value when none is given
CHECKa value the condition rejects
FOREIGN KEYa value that is not in the other table

AUTO_INCREMENT is not a constraint but belongs with them, because it is the usual way a primary key is filled in.

The tables, with their constraints

CREATE TABLE author (
    author_id    INT          AUTO_INCREMENT PRIMARY KEY,
    author_name  VARCHAR(18)  NOT NULL,
    email        VARCHAR(30)  UNIQUE,
    country      VARCHAR(10)  DEFAULT 'India'
);

CREATE TABLE book (
    book_id    INT          AUTO_INCREMENT PRIMARY KEY,
    title      VARCHAR(20)  NOT NULL,
    subject    VARCHAR(12)  NOT NULL,
    price      DECIMAL(7,2) CHECK (price > 0),
    added_on   DATE         DEFAULT (CURRENT_DATE),
    author_id  INT,
    FOREIGN KEY (author_id) REFERENCES author(author_id)
);

Two tables and seven constraints. Each of them is now proved.

INSERT INTO author (author_name, email, country) VALUES
  ('Dennis Ritchie', 'dmr@example.com',  'USA'),
  ('Vikram Vaswani', 'vv@example.com',   'India');

INSERT INTO book (title, subject, price, added_on, author_id) VALUES
  ('The C Language', 'Programming', 395.00, '2023-06-12', 1),
  ('MySQL Reference', 'Databases',  499.00, '2024-01-19', 2);

NOT NULL

NOT NULL says the column must have a value in every row.

INSERT INTO book (title, subject, price, author_id)
VALUES (NULL, 'Databases', 300.00, 1);
ERROR 1048 (23000): Column 'title' cannot be null

Error 1048. The column was named in the statement and given NULL, and the server refused before anything was written.

NULL is not zero and it is not an empty string. It means "no value here", and it is the absence of information rather than a particular piece of information. A price of 0 is a free book; a price of NULL is a book whose price nobody has recorded. Keeping those two apart is what NULL is for, and it is why [Practical 4: Simple Queries] spends a section on it.

PRIMARY KEY

A PRIMARY KEY does two things at once: the value must be unique, and it may not be NULL. A table may have only one.

munotes.in122

Practical 2: The Constraints, and What Each One Refuses

INSERT INTO book (book_id, title, subject, price, author_id)
VALUES (1, 'Another Book', 'Programming', 250.00, 1);
ERROR 1062 (23000): Duplicate entry '1' for key 'book.PRIMARY'

Error 1062, and it names the value it refused, 1, and the key it broke, PRIMARY.

UNIQUE

UNIQUE refuses a duplicate but permits NULL, and permits more than one NULL, because two unknown values are not known to be the same.

INSERT INTO author (author_name, email)
VALUES ('Another Writer', 'dmr@example.com');
ERROR 1062 (23000): Duplicate entry 'dmr@example.com' for key 'author.email'

The same error number as a duplicate primary key, 1062, but the key it names is email rather than PRIMARY. Read the key name in a 1062: it tells you which rule you broke.

PRIMARY KEYUNIQUE
Duplicatesrefusedrefused
NULLrefusedallowed
How many per tableoneas many as you like
Purposeidentifies the rowenforces a business rule

DEFAULT

DEFAULT supplies a value when the statement does not.

INSERT INTO author (author_name, email) VALUES ('New Writer', 'nw@example.com');

SELECT author_name, country FROM author WHERE author_name = 'New Writer';
+-------------+---------+
| author_name | country |
+-------------+---------+
| New Writer  | India   |
+-------------+---------+

country was never mentioned and came out as India, because that is what the column's DEFAULT says.

DEFAULT (CURRENT_DATE) on book.added_on does the same with today's date, which is why this book cannot print that particular result: it would be different tomorrow. A default that is an expression rather than a constant goes in round brackets, which MySQL has allowed since 8.0.13.

CHECK

CHECK refuses any row for which its condition is false.

INSERT INTO book (title, subject, price, author_id)
VALUES ('Free Book', 'Programming', -50.00, 1);
ERROR 3819 (HY000): Check constraint 'book_chk_1' is violated.

Error 3819, and it names the constraint. MySQL invented the name book_chk_1 because none was given; naming it yourself makes the error readable:

price DECIMAL(7,2) CONSTRAINT price_must_be_positive CHECK (price > 0)

A version warning that matters in a college laboratory. CHECK was accepted and silently ignored by MySQL until 8.0.16. On an older server the row above is inserted with a price of minus 50 and nothing is said. If your laboratory's server is older than that, the constraint has to be written as a trigger or enforced in the program, and saying so in your journal is worth more than pretending it worked.

FOREIGN KEY

A FOREIGN KEY says that a value in this column must exist as a primary key in another table. It is what stops a book pointing at an author who does not exist, and it is the constraint that makes a database more than a heap of tables.

INSERT INTO book (title, subject, price, author_id)
VALUES ('Orphan Book', 'Databases', 300.00, 99);
munotes.in123

Practical 2: The Constraints, and What Each One Refuses

ERROR 1452 (23000): Cannot add or update a child row: a foreign key constraint fails (`librarydb`.`book`, CONSTRAINT `book_ibfk_1` FOREIGN KEY (`author_id`) REFERENCES `author` (`author_id`))

Error 1452. There is no author 99, so there can be no book by author 99.

It works in the other direction too. An author who has books cannot be deleted, because that would leave those books pointing at nothing:

DELETE FROM author WHERE author_id = 1;
ERROR 1451 (23000): Cannot delete or update a parent row: a foreign key constraint fails (`librarydb`.`book`, CONSTRAINT `book_ibfk_1` FOREIGN KEY (`author_id`) REFERENCES `author` (`author_id`))

Error 1451, the other half of the same rule. Notice that the two errors are different numbers: 1452 is a child row with no parent, 1451 is a parent row with children.

The technical name for what these two errors protect is referential integrity: every foreign key value either is NULL or matches an existing primary key.

What to do instead of refusing

You can tell the server what should happen to the children when a parent goes:

WrittenEffect on the child rows when the parent is deleted
nothing, or ON DELETE RESTRICTthe delete is refused, as above
ON DELETE CASCADEthey are deleted too
ON DELETE SET NULLtheir foreign key becomes NULL, so the column must allow it

ON UPDATE takes the same three, for when the parent's key value changes.

Choose deliberately. CASCADE on a library's authors would delete every book by an author you removed, which is almost certainly not what a librarian wants; SET NULL leaves the books and forgets who wrote them, which may be exactly right.

AUTO_INCREMENT

AUTO_INCREMENT fills a column with the next number when none is given. A table may have one such column and it must be a key.

INSERT INTO author (author_name, email) VALUES ('Fourth Writer', 'f@example.com');

SELECT author_id, author_name FROM author ORDER BY author_id;
+-----------+----------------+
| author_id | author_name    |
+-----------+----------------+
|         1 | Dennis Ritchie |
|         2 | Vikram Vaswani |
|         4 | New Writer     |
|         5 | Fourth Writer  |
+-----------+----------------+

Nobody supplied any author_id after the first two, and the numbers went on by themselves.

Now look for 3. There is no author 3, and no author was ever deleted. The number was used up by the insert that broke the UNIQUE on email earlier in this chapter: MySQL allotted the next value before it checked the constraint, the row was refused, and the number went with it.

That is the first of two facts about AUTO_INCREMENT that are asked and that surprise people. The numbers are not reused. A failed insert, a rolled back transaction or a deleted row leaves a gap, and the counter never goes back, because a number that has been handed out may already have been written down somewhere else.

munotes.in124

Practical 2: The Constraints, and What Each One Refuses

And the numbers are not a count. The largest id is 5 and there are four authors. Never answer "how many rows are there" by looking at the last id; ask COUNT(*), which is [Practical 4: Aggregate Functions, GROUP BY and HAVING].

Adding a constraint to a table that already exists

Constraints do not have to be decided at creation.

ALTER TABLE book ADD CONSTRAINT price_positive CHECK (price > 0);
ALTER TABLE book DROP CHECK price_positive;

That is [Practical 3: Altering a Table That Already Holds Data], and it is why naming a constraint when you create it is worth the extra words: an unnamed one is book_chk_1, and dropping it means looking that name up first.

The error numbers, collected

NumberMeans
1048a NULL in a NOT NULL column
1062a duplicate in a PRIMARY KEY or UNIQUE; the message names which
1452a foreign key value with no matching parent row
1451a parent row that still has children
3819a CHECK condition was false

Write that table into your journal. Every one of the five appears in this chapter's output, and being able to say what a number means is exactly the kind of thing a viva asks.

What beginners get wrong

Thinking NULL is zero or an empty string. It is the absence of a value.

Expecting UNIQUE to refuse NULLs. It does not, and it allows several of them.

Declaring two PRIMARY KEY columns separately. A table has one primary key; if it must be made of two columns, declare it once as PRIMARY KEY (a, b).

Relying on CHECK on a server older than 8.0.16. It is accepted and ignored there.

Creating the child table before the parent. The REFERENCES has nothing to point at and the statement fails.

A foreign key whose type does not match the key it references. INT must reference INT, and an UNSIGNED mismatch is refused too.

Assuming AUTO_INCREMENT values are consecutive. They are unique, not gapless.

Quick revision

  • A constraint is enforced by the server, so no program can bypass it.
  • NOT NULL: the column must have a value. Error 1048.
  • PRIMARY KEY: unique and not null, one per table. Error 1062.
  • UNIQUE: no duplicates, but NULLs are allowed and may repeat. Error 1062, naming the column.
  • DEFAULT: supplies a value when none is given; an expression goes in brackets.
  • CHECK: refuses rows where the condition is false. Error 3819. Ignored before MySQL 8.0.16.
  • FOREIGN KEY: the value must exist in the referenced table. Error 1452 inserting, 1451 deleting the parent.
  • ON DELETE CASCADE or SET NULL change what happens to the children.
  • AUTO_INCREMENT fills in the next number, never reuses one, and needs a key column.
munotes.in125

Practical 2: The Constraints, and What Each One Refuses

What goes in your journal

Aim, the two CREATE TABLE statements with every constraint, and then one failed statement per constraint with the server's error copied exactly, error number and all. This is the practical where the failures are the answer. A journal showing only successful inserts has demonstrated that the tables exist, not that the constraints work.

Test yourself

1. What two rules does a PRIMARY KEY enforce? The value must be unique, and it may not be NULL.

2. How does UNIQUE differ from PRIMARY KEY? UNIQUE allows NULL, and allows more than one NULL. A table may have many UNIQUE columns and only one primary key.

3. What is referential integrity? The rule that every foreign key value either is NULL or matches an existing primary key in the referenced table, so no row can point at something that does not exist.

4. What happens when you delete an author who has books, and how could that be changed? The delete is refused with error 1451. Declaring the foreign key ON DELETE CASCADE would delete the books too, and ON DELETE SET NULL would leave them with a NULL author.

5. On which versions of MySQL is CHECK enforced? From 8.0.16 onward. Before that it was parsed and ignored.

6. An insert was refused by a constraint. What happens to the AUTO_INCREMENT number it was allotted? It is lost, and the next successful insert takes the following number. The counter never goes back, so gaps are normal and mean nothing is wrong.

7. Why enforce a rule with a constraint rather than in the program? Because the constraint travels with the data and cannot be bypassed by another program, or by somebody typing a statement at the prompt.

Contents This chapter on its own page

munotes.in126

Chapter Thirty-Five

Practical 2: Inserting, Updating and Deleting Rows

Syllabus topic Module 2, Practical 2: "Inserting/Updating/Deleting Records in a Table"

Aim

To insert, update and delete records in a table.

The three verbs, and the family they belong to

SQL's statements fall into families, and knowing which is which is a standing viva question.

FamilyStands forStatements
DDLData Definition LanguageCREATE, ALTER, DROP, TRUNCATE, RENAME
DMLData Manipulation LanguageINSERT, UPDATE, DELETE
DQLData Query LanguageSELECT
DCLData Control LanguageGRANT, REVOKE
TCLTransaction Control LanguageCOMMIT, ROLLBACK, SAVEPOINT

This practical is the whole of DML. DDL is [Practical 3: Altering a Table That Already Holds Data], DQL is [Practical 4: Simple Queries], and DCL and TCL are the two chapters of Practical 10.

Some books put SELECT inside DML. Either answer is accepted; say which classification you are using and be consistent.

INSERT, three ways

With every column, in order

INSERT INTO author VALUES (6, 'Byron Gottfried', 'USA');

SELECT * FROM author WHERE author_id = 6;
+-----------+-----------------+---------+
| author_id | author_name     | country |
+-----------+-----------------+---------+
|         6 | Byron Gottfried | USA     |
+-----------+-----------------+---------+

No column list, so the values must be given for every column, in the order the table declares them. It is the shortest form and the most fragile: add a column to the table tomorrow and every statement of this shape stops working.

With a column list

INSERT INTO author (author_name, author_id) VALUES ('Behrouz Forouzan', 7);

SELECT * FROM author WHERE author_id = 7;
+-----------+------------------+---------+
| author_id | author_name      | country |
+-----------+------------------+---------+
|         7 | Behrouz Forouzan | NULL    |
+-----------+------------------+---------+

Naming the columns lets you give them in any order and leave some out, and country was left out and came back NULL. Prefer this form. It survives a change to the table and it says what it means.

Several rows at once

INSERT INTO member (member_id, member_name, course, city, joined_on) VALUES
  (7, 'Sunita Rane',  'BSc IT', 'Mumbai', '2025-02-11'),
  (8, 'Rahul Jadhav', 'BSc CS', 'Panvel', '2025-02-14');

SELECT member_id, member_name, city FROM member WHERE member_id > 6;
+-----------+--------------+--------+
| member_id | member_name  | city   |
+-----------+--------------+--------+
|         7 | Sunita Rane  | Mumbai |
|         8 | Rahul Jadhav | Panvel |
+-----------+--------------+--------+

One statement, two rows, one set of brackets each. It is faster than two statements because the server does the work once, and it is all or nothing: if either row breaks a constraint, neither is inserted.

From a query

INSERT INTO old_members (member_id, member_name)
SELECT member_id, member_name FROM member WHERE joined_on < '2025-01-01';

No VALUES at all: the rows come from a SELECT. It is how a table is filled from another table, and it is worth knowing that it exists.

UPDATE

UPDATE table SET column = value, column = value WHERE condition;
UPDATE book SET price = 425.00 WHERE book_id = 1;

SELECT book_id, title, price FROM book WHERE book_id = 1;
munotes.in127

Practical 2: Inserting, Updating and Deleting Rows

+---------+----------------+--------+
| book_id | title          | price  |
+---------+----------------+--------+
|       1 | The C Language | 425.00 |
+---------+----------------+--------+

Several columns at once, separated by commas:

UPDATE member
SET city = 'Dombivli', course = 'BSc IT'
WHERE member_id = 5;

SELECT member_id, member_name, course, city FROM member WHERE member_id = 5;
+-----------+-------------+--------+----------+
| member_id | member_name | course | city     |
+-----------+-------------+--------+----------+
|         5 | Neha Patil  | BSc IT | Dombivli |
+-----------+-------------+--------+----------+

The new value may be worked out from the old one, which is how a percentage rise is applied:

UPDATE book SET price = price * 1.10 WHERE subject = 'Databases';

SELECT book_id, title, price FROM book WHERE subject = 'Databases';
+---------+------------------+--------+
| book_id | title            | price  |
+---------+------------------+--------+
|       2 | Database Systems | 792.55 |
|       3 | MySQL Reference  | 548.90 |
|       6 | Learn SQL        | 308.00 |
+---------+------------------+--------+

price = price * 1.10 reads oddly to a programmer and is ordinary in SQL: on the right of the = the column means its current value, and on the left it means where the new value goes. It is applied to every row the WHERE matches, each using its own old price.

The mistake this chapter exists to prevent

CREATE TABLE book_copy AS SELECT * FROM book;

CREATE TABLE ... AS SELECT makes a table with the same columns and the same rows. Now the damage can be done where it does not matter:

UPDATE book_copy SET price = 100.00;

SELECT book_id, title, price FROM book_copy ORDER BY book_id;
+---------+-------------------+--------+
| book_id | title             | price  |
+---------+-------------------+--------+
|       1 | The C Language    | 100.00 |
|       2 | Database Systems  | 100.00 |
|       3 | MySQL Reference   | 100.00 |
|       4 | Data Structures   | 100.00 |
|       5 | Computer Networks | 100.00 |
|       6 | Learn SQL         | 100.00 |
|       7 | Operating Systems | 100.00 |
|       8 | Discrete Maths    | 100.00 |
+---------+-------------------+--------+

Every book now costs 100. There was no WHERE, so every row matched, and MySQL did exactly as it was told without a question or a warning.

There is no undo for that outside a transaction. [Practical 10: COMMIT and ROLLBACK] is the chapter that gives you one, and it is the reason transactions exist.

Three habits, and they cost nothing:

Write the WHERE first. Type WHERE book_id = 1 before you type the SET.

Run it as a SELECT first. SELECT * FROM book WHERE ... with the same condition shows you exactly which rows are about to change. If it returns forty rows and you expected one, you have just saved yourself.

munotes.in128

Practical 2: Inserting, Updating and Deleting Rows

On a real database, START TRANSACTION first. Then a wrong answer is one ROLLBACK away.

DELETE

DELETE FROM book_copy WHERE book_id = 8;

SELECT COUNT(*) AS rows_left FROM book_copy;
+-----------+
| rows_left |
+-----------+
|         7 |
+-----------+

Seven left of eight. DELETE removes whole rows; there is no such thing as deleting one column's value, and setting it to NULL with an UPDATE is what that would mean.

DELETE without a WHERE empties the table, exactly as UPDATE without one changes every row.

A constraint can refuse a delete, and in this database it does:

DELETE FROM book WHERE book_id = 1;
ERROR 1451 (23000): Cannot delete or update a parent row: a foreign key constraint fails (`librarydb`.`loan`, CONSTRAINT `loan_ibfk_1` FOREIGN KEY (`book_id`) REFERENCES `book` (`book_id`))

Book 1 has been borrowed, so a row in loan points at it. Error 1451, from [Practical 2: The Constraints, and What Each One Refuses], and it is the database protecting itself: deleting the book would leave a loan of a book that does not exist.

The three verbs compared

INSERTUPDATEDELETE
Acts onnew rowsexisting rowsexisting rows
Needs a WHEREnoyes, in practiceyes, in practice
Without a WHEREnot applicablechanges every rowempties the table
Can a constraint refuse ityesyesyes
Undone by ROLLBACKyes, inside a transactionyesyes

What beginners get wrong

An UPDATE or DELETE with no WHERE. It is the commonest serious mistake in this module.

UPDATE book SET price = 400 WHERE price = NULL. Nothing matches: NULL is never equal to anything, not even to NULL. The test is WHERE price IS NULL.

Quoting numbers, or not quoting text. Text and dates go in single quotes; numbers do not.

Double quotes for a string. MySQL accepts them, the SQL standard does not, and a query written with them may not run on another server. Use single quotes.

Inserting a child row before its parent. A book by author 99 is refused until author 99 exists.

Expecting DELETE to remove a column. It removes rows. Removing a column is ALTER TABLE ... DROP COLUMN.

Thinking UPDATE with no matching rows is an error. It is not; it changes nothing and reports zero rows affected.

Quick revision

  • DDL creates and changes structure; DML changes data; DQL asks; DCL grants; TCL commits.
  • INSERT INTO t VALUES (...) needs every column in order; INSERT INTO t (cols) VALUES (...) is safer.
  • Several rows go in one INSERT, separated by commas, and it is all or nothing.
  • INSERT INTO t (cols) SELECT ... fills a table from a query.
  • UPDATE t SET col = value WHERE cond; and the new value may use the old one.
  • Without a WHERE, UPDATE changes every row and DELETE empties the table.
  • Run the WHERE as a SELECT first.
  • WHERE col = NULL never matches; use IS NULL.
  • Text and dates in single quotes; numbers bare.
munotes.in129

Practical 2: Inserting, Updating and Deleting Rows

What goes in your journal

Aim, then one statement of each kind with the table printed before and after it, so the change is visible. The before and after are the evidence; a journal with only the statements has not shown that they did anything.

Include the refused DELETE with its error, and write one line on what a missing WHERE would have done. That line is the whole safety lesson of the practical.

Test yourself

1. Which family do INSERT, UPDATE and DELETE belong to? DML, the Data Manipulation Language.

2. Why is INSERT INTO t (col1, col2) VALUES (...) better than INSERT INTO t VALUES (...)? Because it does not depend on the number or the order of the table's columns, so it still works when the table changes, and it says which value goes where.

3. What does UPDATE book SET price = 100; do? Sets every book's price to 100, because there is no WHERE to limit it.

4. Why does WHERE price = NULL match nothing? Because NULL is not equal to anything, including another NULL. The correct test is IS NULL.

5. How do you make a copy of a table to practise on? CREATE TABLE copy AS SELECT * FROM original;

6. Your DELETE was refused with error 1451. What does that mean? Another table has rows whose foreign key points at the row you tried to delete, so deleting it would break referential integrity.

7. What is the safest way to check an UPDATE before running it? Run a SELECT with the same WHERE clause and look at how many rows come back.

Contents This chapter on its own page

munotes.in130

Chapter Thirty-Six

Practical 3: Altering a Table That Already Holds Data

Syllabus topic Module 2, Practical 3: "Altering a Table"

Aim

To alter a table: to add, modify, rename and drop columns and constraints.

Why a table ever needs altering

Because the design was decided before anybody used it. A library that has run for a year discovers that it wants to record a book's edition, that titles are longer than twenty characters, and that subject should never have been allowed to be empty.

ALTER TABLE makes those changes without losing the rows, which is what makes it worth a practical of its own. Creating a new table and copying is the alternative, and it means downtime and a chance to lose data.

ALTER TABLE is DDL, like CREATE. In MySQL that has one consequence that catches everybody, and it is at the end of this chapter.

Adding a column

ALTER TABLE book ADD COLUMN edition INT;

SELECT book_id, title, edition FROM book WHERE book_id <= 3;
+---------+------------------+---------+
| book_id | title            | edition |
+---------+------------------+---------+
|       1 | The C Language   |    NULL |
|       2 | Database Systems |    NULL |
|       3 | MySQL Reference  |    NULL |
+---------+------------------+---------+

The new column arrived on every existing row, holding NULL, because no value was supplied for rows that were written before the column existed. That is the only thing it could hold, and it is why a column added to a populated table cannot simply be NOT NULL.

Unless it is given a default:

ALTER TABLE book ADD COLUMN pages INT NOT NULL DEFAULT 0;

SELECT book_id, title, pages FROM book WHERE book_id <= 3;
+---------+------------------+-------+
| book_id | title            | pages |
+---------+------------------+-------+
|       1 | The C Language   |     0 |
|       2 | Database Systems |     0 |
|       3 | MySQL Reference  |     0 |
+---------+------------------+-------+

Now the existing rows have 0 rather than NULL, and the column can be NOT NULL honestly.

By default a new column goes at the end. To put it elsewhere:

ALTER TABLE book ADD COLUMN isbn VARCHAR(13) AFTER title;

DESCRIBE book;
+-----------+--------------+------+-----+---------+-------+
| Field     | Type         | Null | Key | Default | Extra |
+-----------+--------------+------+-----+---------+-------+
| book_id   | int          | NO   | PRI | NULL    |       |
| title     | varchar(20)  | NO   |     | NULL    |       |
| isbn      | varchar(13)  | YES  |     | NULL    |       |
| subject   | varchar(12)  | NO   |     | NULL    |       |
| price     | decimal(7,2) | YES  |     | NULL    |       |
| added_on  | date         | YES  |     | NULL    |       |
| author_id | int          | YES  | MUL | NULL    |       |
| edition   | int          | YES  |     | NULL    |       |
| pages     | int          | NO   |     | 0       |       |
+-----------+--------------+------+-----+---------+-------+
munotes.in131

Practical 3: Altering a Table That Already Holds Data

isbn sits third, immediately after title, where AFTER title asked for it. FIRST puts a column at the very front. Neither changes anything a query can see, because a query names its columns; it is for the convenience of whoever reads DESCRIBE.

Changing a column's type

ALTER TABLE book MODIFY COLUMN title VARCHAR(60);

SELECT column_name AS col, data_type AS type, character_maximum_length AS max_len
FROM information_schema.columns
WHERE table_schema = DATABASE() AND table_name = 'book' AND column_name = 'title';
+-------+---------+---------+
| col   | type    | max_len |
+-------+---------+---------+
| title | varchar |      60 |
+-------+---------+---------+

Twenty has become sixty. Widening a column is always safe, because every value that fitted in the old type fits in the new one.

Narrowing is not:

ALTER TABLE book MODIFY COLUMN title VARCHAR(5);
ERROR 1265 (01000): Data truncated for column 'title' at row 1

Error 1265. The C Language is fourteen characters and will not fit in five, so the server refused rather than cutting the value short. That refusal is the server's strict mode at work, and it is what you want: an ALTER that silently truncated every title would be far worse than one that failed.

The same applies to type changes. Turning a VARCHAR into an INT succeeds only if every value in it is a number.

Renaming a column

Two ways, and the difference is asked.

ALTER TABLE book RENAME COLUMN edition TO edition_no;

SELECT column_name AS col
FROM information_schema.columns
WHERE table_schema = DATABASE() AND table_name = 'book'
  AND column_name LIKE 'edition%';
+------------+
| col        |
+------------+
| edition_no |
+------------+

RENAME COLUMN changes only the name, and is available from MySQL 8.0. The older way changes the name and the type together, so the type has to be restated even when it is not changing:

ALTER TABLE book CHANGE COLUMN edition_no edition SMALLINT;

SELECT column_name AS col, data_type AS type
FROM information_schema.columns
WHERE table_schema = DATABASE() AND table_name = 'book'
  AND column_name LIKE 'edition%';
+---------+----------+
| col     | type     |
+---------+----------+
| edition | smallint |
+---------+----------+
ClauseChanges the nameChanges the typeNeeds the type restated
MODIFY COLUMN old newtypenoyesyes
CHANGE COLUMN old new newtypeyesyesyes
RENAME COLUMN old TO newyesnono

MODIFY for the type, RENAME COLUMN for the name, CHANGE when you want both. Forgetting to restate the type in a CHANGE is the classic mistake: CHANGE COLUMN edition_no edition alone does not compile, because the type is not optional there.

Dropping a column

ALTER TABLE book DROP COLUMN isbn;

SELECT column_name AS col
FROM information_schema.columns
WHERE table_schema = DATABASE() AND table_name = 'book'
ORDER BY ordinal_position;
+-----------+
| col       |
+-----------+
| book_id   |
| title     |
| subject   |
| price     |
| added_on  |
| author_id |
| edition   |
| pages     |
+-----------+
munotes.in132

Practical 3: Altering a Table That Already Holds Data

The column and every value in it are gone, on every row, and there is no undo outside a transaction. Check twice.

Adding and dropping constraints

ALTER TABLE book ADD CONSTRAINT price_positive CHECK (price > 0);

Naming the constraint is what makes it droppable without looking the name up:

INSERT INTO book (book_id, title, subject, price, author_id)
VALUES (99, 'Bad Price', 'Programming', -10.00, 1);
ERROR 3819 (HY000): Check constraint 'price_positive' is violated.
ALTER TABLE book DROP CHECK price_positive;

The same statement adds a unique constraint, a primary key or a foreign key:

ALTER TABLE author ADD CONSTRAINT author_email_unique UNIQUE (email);
ALTER TABLE book   ADD CONSTRAINT book_author_fk
                   FOREIGN KEY (author_id) REFERENCES author(author_id);
ALTER TABLE book   DROP FOREIGN KEY book_author_fk;

Note that the words after DROP differ by kind: DROP CHECK, DROP FOREIGN KEY, DROP INDEX, and DROP PRIMARY KEY with no name at all, since there is only one.

Adding a constraint the data already breaks

This is the situation the chapter's title is about, and it is the one an examiner sets.

ALTER TABLE member MODIFY COLUMN city VARCHAR(9) NOT NULL;
ERROR 1138 (22004): Invalid use of NULL value

Error 1138. One member has a NULL city, so the column cannot become NOT NULL while that row is there. A constraint must be true of the data before it can be added.

The fix is to mend the data first:

UPDATE member SET city = 'Unknown' WHERE city IS NULL;

ALTER TABLE member MODIFY COLUMN city VARCHAR(9) NOT NULL;

SELECT column_name AS col, is_nullable AS nullable
FROM information_schema.columns
WHERE table_schema = DATABASE() AND table_name = 'member'
  AND column_name = 'city';
+------+----------+
| col  | nullable |
+------+----------+
| city | NO       |
+------+----------+

Now it is true of every row, so the server accepts it. That order, mend the data and then add the constraint, is the whole answer to this practical's hardest viva question.

The one that catches everybody: DDL commits

In MySQL, a DDL statement causes an implicit commit. The transaction you were in is committed before the ALTER runs, and a ROLLBACK afterwards cannot undo either.

START TRANSACTION;
UPDATE book SET price = 1;        -- can still be rolled back
ALTER TABLE book ADD COLUMN x INT; -- commits the UPDATE, silently
ROLLBACK;                          -- too late: the price is 1 for good

So on a real database, do the data work and the structure work in separate steps, and never assume an ALTER inside a transaction can be undone. This is one of the differences between MySQL and PostgreSQL that catches people moving between them, and it is worth knowing by name: MySQL's DDL is not transactional.

munotes.in133

Practical 3: Altering a Table That Already Holds Data

What beginners get wrong

Expecting an added column to have values. It holds NULL unless a DEFAULT is given.

Adding NOT NULL to a column that already holds NULLs. Mend the data first.

Narrowing a column that holds longer values. Refused, and rightly.

Using CHANGE without restating the type. The type is required there.

Writing ALTER TABLE book DROP edition; and meaning the constraint, or the other way round. Say DROP COLUMN, DROP CHECK or DROP FOREIGN KEY.

Assuming an ALTER can be rolled back. DDL commits.

Altering a large table on a live system in the middle of the day. Some alterations copy the whole table.

Quick revision

  • ALTER TABLE t ADD COLUMN c type; and the new column is NULL on existing rows unless it has a DEFAULT.
  • AFTER col and FIRST place the new column; by default it goes last.
  • MODIFY COLUMN c newtype changes the type; RENAME COLUMN a TO b changes the name; CHANGE COLUMN a b type changes both and always needs the type.
  • Widening is safe; narrowing is refused when a value would not fit, error 1265.
  • ADD CONSTRAINT name ... and DROP CHECK, DROP FOREIGN KEY, DROP INDEX, DROP PRIMARY KEY.
  • A constraint can only be added if the data already satisfies it, error 1138 otherwise.
  • DDL causes an implicit commit in MySQL; an ALTER cannot be rolled back.

What goes in your journal

Aim, and then for each alteration the DESCRIBE before and the DESCRIBE after, so the change is visible. Include the two refusals, the narrowing and the NOT NULL, with their error numbers, and the UPDATE that made the second one succeed. Those three statements together are the practical: a structure change is limited by the data that is already there.

Test yourself

1. What value does a newly added column hold on existing rows? NULL, unless the column was added with a DEFAULT, in which case the default.

2. What is the difference between MODIFY, CHANGE and RENAME COLUMN? MODIFY changes the type only. RENAME COLUMN changes the name only. CHANGE changes both and always requires the type to be written out.

3. Why was MODIFY COLUMN title VARCHAR(5) refused? Because existing titles are longer than five characters, and truncating them silently would lose data. Error 1265.

4. How do you make a column NOT NULL when it already contains NULLs? Update those rows to a real value first, then run the ALTER.

5. Can an ALTER TABLE be rolled back in MySQL? No. DDL causes an implicit commit, which also commits whatever the transaction had done before it.

munotes.in134

Practical 3: Altering a Table That Already Holds Data

6. How do you drop a CHECK constraint you did not name? Find the name MySQL gave it, such as book_chk_1, in SHOW CREATE TABLE or information_schema.table_constraints, and drop that. Naming constraints when you create them avoids this.

Contents This chapter on its own page

munotes.in135

Chapter Thirty-Seven

Practical 3: Dropping, Truncating and Renaming

Syllabus topic Module 2, Practical 3: "Dropping/Truncating/Renaming Tables"

Aim

To drop, truncate and rename tables, and to distinguish between them.

The three, in one line each

DROP TABLE removes the table itself. Structure and rows, both gone.

TRUNCATE TABLE empties the table. Every row goes; the table, its columns and its constraints stay.

RENAME TABLE gives the table a different name. Nothing else changes.

And the fourth that belongs in the comparison:

DELETE FROM t with no WHERE also empties the table, and it is not the same as TRUNCATE. The differences are what this practical is really about.

Working on copies

Everything below is done on copies, so the library survives the chapter.

CREATE TABLE book_copy    AS SELECT * FROM book;
CREATE TABLE member_copy  AS SELECT * FROM member;
CREATE TABLE spare_table (id INT PRIMARY KEY AUTO_INCREMENT, note VARCHAR(20));
INSERT INTO spare_table (note) VALUES ('first'), ('second'), ('third');

CREATE TABLE ... AS SELECT copies the columns and the rows. It does not copy the primary key, the constraints or the indexes, which is worth knowing: a copy made this way is a table of the same data with none of the same rules.

TRUNCATE

SELECT COUNT(*) AS before_truncate FROM book_copy;
+-----------------+
| before_truncate |
+-----------------+
|               8 |
+-----------------+
TRUNCATE TABLE book_copy;

SELECT COUNT(*) AS after_truncate FROM book_copy;
+----------------+
| after_truncate |
+----------------+
|              0 |
+----------------+

Eight rows to none, and the table is still there:

DESCRIBE book_copy;
+-----------+--------------+------+-----+---------+-------+
| Field     | Type         | Null | Key | Default | Extra |
+-----------+--------------+------+-----+---------+-------+
| book_id   | int          | NO   |     | NULL    |       |
| title     | varchar(20)  | NO   |     | NULL    |       |
| subject   | varchar(12)  | NO   |     | NULL    |       |
| price     | decimal(7,2) | YES  |     | NULL    |       |
| added_on  | date         | YES  |     | NULL    |       |
| author_id | int          | YES  |     | NULL    |       |
+-----------+--------------+------+-----+---------+-------+

Every column still in place, which is the whole difference from DROP.

What TRUNCATE does to AUTO_INCREMENT

DELETE FROM spare_table;
INSERT INTO spare_table (note) VALUES ('after delete');

SELECT id, note FROM spare_table;
+----+--------------+
| id | note         |
+----+--------------+
|  4 | after delete |
+----+--------------+

Three rows were deleted and the next id was 4, because DELETE does not touch the counter. Now the other way:

TRUNCATE TABLE spare_table;
INSERT INTO spare_table (note) VALUES ('after truncate');

SELECT id, note FROM spare_table;
+----+----------------+
| id | note           |
+----+----------------+
|  1 | after truncate |
+----+----------------+

Back to 1. TRUNCATE resets the AUTO_INCREMENT counter and DELETE does not. That single difference is the one most often asked, and it is now demonstrated rather than claimed.

TRUNCATE and foreign keys

TRUNCATE TABLE member;
ERROR 1701 (42000): Cannot truncate a table referenced in a foreign key constraint (`librarydb`.`loan`, CONSTRAINT `loan_ibfk_2`)
munotes.in136

Practical 3: Dropping, Truncating and Renaming

Error 1701. Rows in loan point at member, so the server refuses to empty it, and it refuses even if no row would actually be orphaned, because TRUNCATE does not examine rows one at a time. DELETE FROM member would check each row and refuse only the ones that are referenced.

That is the second real difference: DELETE is row by row; TRUNCATE is wholesale. It is also why TRUNCATE is much faster on a large table.

DELETE against TRUNCATE

DELETE FROM tTRUNCATE TABLE t
FamilyDMLDDL
Removesrows, one at a timeall rows, wholesale
Can take a WHEREyesno
Resets AUTO_INCREMENTnoyes
Can be rolled backyes, inside a transactionno: DDL commits
Fires triggersyesno
Speed on a large tableslowfast
With a foreign key pointing at itrefuses only the referenced rowsrefuses outright

The row that costs marks is rollback. TRUNCATE is DDL, so it commits, and there is no undoing it. DELETE inside a transaction can be undone right up to the COMMIT, which is [Practical 10: COMMIT and ROLLBACK].

RENAME

RENAME TABLE book_copy TO book_archive;

SELECT table_name AS tables_here
FROM information_schema.tables
WHERE table_schema = DATABASE()
ORDER BY table_name;
+--------------+
| tables_here  |
+--------------+
| author       |
| book         |
| book_archive |
| loan         |
| member       |
| member_copy  |
| spare_table  |
+--------------+

book_copy is gone from the list and book_archive has appeared, with the same rows and the same structure. Renaming moves nothing: the table is not copied, so it is as fast on a large table as on a small one.

Two tables can be renamed in one statement, which is how a table is swapped for a new version without the old name ever being missing:

RENAME TABLE live TO old, staging TO live;

ALTER TABLE old_name RENAME TO new_name; does the same thing for one table, and is the form you will see more often.

A rename does not follow through into other objects. A view or a stored program that names the old table keeps naming it and breaks. So does a program of yours.

DROP

DROP TABLE book_archive;

SELECT table_name AS tables_here
FROM information_schema.tables
WHERE table_schema = DATABASE()
ORDER BY table_name;
+-------------+
| tables_here |
+-------------+
| author      |
| book        |
| loan        |
| member      |
| member_copy |
| spare_table |
+-------------+

Gone from the list entirely: the rows, the columns, the constraints, the indexes.

SELECT * FROM book_archive;
ERROR 1146 (42S02): Table 'librarydb.book_archive' doesn't exist

Error 1146. There is no such table any more.

A table that another table's foreign key points at cannot be dropped either:

munotes.in137

Practical 3: Dropping, Truncating and Renaming

DROP TABLE member;
ERROR 3730 (HY000): Cannot drop table 'member' referenced by a foreign key constraint 'loan_ibfk_2' on table 'loan'.

Error 3730, and the message names the constraint that is in the way. Drop the child table first, or drop the constraint.

DROP TABLE IF EXISTS t; does not complain when there is nothing to drop, which is what a script that may be run twice should use.

The four compared

DELETETRUNCATEDROP TABLERENAME TABLE
Rowsremoved, selectivelyall removedremovedkept
Table structurekeptkeptremovedkept
Table namekeptkeptgonechanged
FamilyDMLDDLDDLDDL
Rollbackyesnonono
AUTO_INCREMENTunchangedresetnot applicableunchanged

Learn that table. It is the answer to "differentiate between DELETE, TRUNCATE and DROP", which is asked in this practical's viva more reliably than anything else in Module 2.

What beginners get wrong

Thinking TRUNCATE can be rolled back. It is DDL and it commits.

Expecting DROP to keep the structure. That is TRUNCATE.

Putting a WHERE on TRUNCATE. It takes none. Use DELETE.

Expecting DELETE to reset AUTO_INCREMENT. It does not.

Dropping a parent table before its children. Refused, error 3730.

Renaming a table and forgetting the views and programs that name it. They break, and not until the next time they run.

Using DROP DATABASE when a table was meant. There is no confirmation.

Quick revision

  • DROP TABLE t; removes structure and rows. DROP TABLE IF EXISTS t; does not complain.
  • TRUNCATE TABLE t; empties it, resets AUTO_INCREMENT, cannot be rolled back, takes no WHERE, and is refused outright if a foreign key points at the table.
  • RENAME TABLE a TO b; or ALTER TABLE a RENAME TO b; changes only the name.
  • DELETE FROM t; is DML: row by row, rollback-able, fires triggers, leaves the counter alone.
  • CREATE TABLE c AS SELECT * FROM t; copies columns and rows but not keys, constraints or indexes.
  • Error 1146 is no such table; 1701 is truncate refused by a foreign key; 3730 is drop refused by one.

What goes in your journal

Aim, and then the four statements each with a SELECT COUNT(*) and a SHOW TABLES before and after, so that what survived each one is visible. Then the comparison table, which is what the practical is for.

Include the AUTO_INCREMENT demonstration: delete the rows, insert one, note the id; truncate, insert one, note the id. Two ids, 4 and 1, and they are the proof of the difference.

Test yourself

1. What is the difference between DELETE and TRUNCATE? DELETE is DML, removes rows one at a time, can take a WHERE, can be rolled back, fires triggers and leaves AUTO_INCREMENT alone. TRUNCATE is DDL, empties the table wholesale, takes no WHERE, cannot be rolled back, fires no triggers and resets the counter.

munotes.in138

Practical 3: Dropping, Truncating and Renaming

2. What does DROP TABLE leave behind? Nothing. The rows, the columns, the constraints and the indexes all go.

3. Why was TRUNCATE TABLE member refused? Because a foreign key in loan references member. TRUNCATE is refused outright whenever such a reference exists, without examining the rows.

4. After deleting all three rows and inserting one, what id does it get? And after truncating instead? 4 after the delete, because the counter is untouched. 1 after the truncate, because TRUNCATE resets it.

5. Does RENAME TABLE copy the data? No. Only the name changes, so it is as quick on a large table as on a small one.

6. What does CREATE TABLE copy AS SELECT * FROM t; not copy? The primary key, the other constraints and the indexes. It copies the columns and the rows only.

Contents This chapter on its own page

munotes.in139

Chapter Thirty-Eight

Practical 3: Backing Up a Database, and Restoring It

Syllabus topic Module 2, Practical 3: "Backing up / Restoring a Database"

Aim

To take a backup of a database and to restore it.

Why this one is different

Every other practical in this module is typed at the mysql> prompt. This one is not. mysqldump is a separate program that you run from the operating system's command line, in the same place you would type cd or dir.

So: leave the MySQL client first, with exit or quit. If you type mysqldump at the mysql> prompt you get a syntax error, and that is the commonest confusion in this practical.

What a backup of a database actually is

mysqldump does not copy the data files. It produces a text file full of SQL statements which, run in order, rebuild the database exactly: the CREATE DATABASE, every CREATE TABLE with its constraints, and an INSERT for the rows.

That is worth understanding rather than memorising, because three useful facts follow from it.

A backup can be read. Open it in any text editor and you can see exactly what will be restored. A backup you cannot inspect is a backup you are trusting rather than checking.

A backup can be edited. One table can be pulled out of it, or a name changed, before it is restored.

A backup can be restored anywhere. Onto a different machine, a different operating system, or a newer version of MySQL, because SQL is text and not a binary file format.

Taking the backup

mysqldump -u root -p --single-transaction --set-gtid-purged=OFF \
          --databases librarydb > librarydb.sql

It prints nothing when it works, and the file appears in whatever directory you were in.

PartWhat it is for
-u root -pthe user, and ask for the password
--single-transactiontake the whole dump as one consistent snapshot, so a change made while it runs cannot land in half of it
--databases librarydbthis database, and include the CREATE DATABASE in the file
> librarydb.sqlsend the output into a file instead of onto the screen

The > is the operating system's, not MySQL's. mysqldump writes its output to the screen, and > catches it in a file.

--single-transaction is worth a sentence of its own. Without it, mysqldump reads the tables one after another while other people are still using the database, so a member could be added between the dump of member and the dump of loan, and the backup would then hold a loan by a member who is not in it. With it, the whole dump sees one moment in time. Omit it and the program itself warns you.

Looking inside it

head -6 librarydb.sql
-- MySQL dump 10.13  Distrib 9.6.0, for macos26.2 (arm64)
--
-- Host: localhost    Database: librarydb
-- ------------------------------------------------------
-- Server version	9.6.0
munotes.in140

Practical 3: Backing Up a Database, and Restoring It

A header of comments, which begin with two hyphens. Further down is the part that matters:

DROP TABLE IF EXISTS `book`;
CREATE TABLE `book` (
  `book_id` int NOT NULL,
  `title` varchar(20) NOT NULL,
  `subject` varchar(12) NOT NULL,
  `price` decimal(7,2) DEFAULT NULL,
  `added_on` date DEFAULT NULL,
  `author_id` int DEFAULT NULL,
  PRIMARY KEY (`book_id`),
  KEY `author_id` (`author_id`),
  CONSTRAINT `book_ibfk_1` FOREIGN KEY (`author_id`) REFERENCES `author` (`author_id`)
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COLLATE=utf8mb4_0900_ai_ci;

LOCK TABLES `book` WRITE;
INSERT INTO `book` VALUES (1,'The C Language','Programming',395.00,'2023-06-12',1),(2,'Database Systems','Databases',720.50,'2023-07-01',2);
UNLOCK TABLES;

Three things to notice, and all three are viva answers.

DROP TABLE IF EXISTS comes first. A restore replaces what is there. Restoring onto a database that already has data does not merge; it overwrites.

The CREATE TABLE is the full one, with the primary key, the foreign key and the character set, not the shortened version anybody typed. It also names the storage engine, InnoDB, which is the engine that supports foreign keys and transactions.

All the rows of one table are in one INSERT. That is far faster to restore than one statement per row, and it is why a dump of a large table is one very long line.

Destroying it, so that the restore proves something

mysql -u root -p -e "DROP DATABASE librarydb;"
mysql -u root -p -e "SHOW DATABASES LIKE 'librarydb';"

Nothing came back. The database is gone, tables, rows and all.

-e runs one statement and exits, without opening the interactive prompt. It is how a single statement is run from a script.

Restoring

mysql -u root -p < librarydb.sql

The < is the opposite of >: it feeds the file into the program's input, so mysql reads and runs every statement in it, in order.

Notice which program is used. You dump with mysqldump and you restore with mysql. There is no mysqlrestore, and looking for one is the second commonest confusion in this practical. The reason is the one from the top of the chapter: a dump is a file of SQL statements, and the program that runs SQL statements is mysql.

Notice also that no database is named on the restore. It does not need one: the file itself begins with CREATE DATABASE librarydb and USE librarydb, because the dump was taken with --databases.

Proving the restore worked

mysql -u root -p --table librarydb \
      -e "SELECT (SELECT COUNT(*) FROM author) AS authors,
                 (SELECT COUNT(*) FROM book)   AS books,
                 (SELECT COUNT(*) FROM member) AS members,
                 (SELECT COUNT(*) FROM loan)   AS loans;"
+---------+-------+---------+-------+
| authors | books | members | loans |
+---------+-------+---------+-------+
|       5 |     8 |       6 |     7 |
+---------+-------+---------+-------+

Those are the same four counts the database had before it was destroyed. A backup nobody has restored is not a backup, and counting the rows afterwards is the cheapest possible check. Write both sets of counts in your journal, before and after, on the same page.

munotes.in141

Practical 3: Backing Up a Database, and Restoring It

The variations worth knowing

Each of these is mysqldump -u root -p followed by what is shown, and each sends its output into a file with >.

One database, with its CREATE DATABASE. The form used above, and the one to use by default.

mysqldump -u root -p --databases librarydb > f.sql

One database's tables, without the CREATE DATABASE. The difference is small and catches people: this file has no CREATE DATABASE and no USE, so the restore must be told where to put it.

mysqldump -u root -p librarydb > f.sql
mysql -u root -p librarydb < f.sql

Named tables only. List them after the database name.

mysqldump -u root -p librarydb book member > f.sql

The whole server, every database and the accounts with it.

mysqldump -u root -p --all-databases > f.sql

The structure and no rows. This is the one to use when you have to hand somebody the design of your database without the data in it, which is often exactly what a journal submission needs.

mysqldump -u root -p --no-data librarydb > f.sql

Only the rows a condition selects.

mysqldump -u root -p --where="subject='Databases'" librarydb book > f.sql

In phpMyAdmin

If your laboratory uses phpMyAdmin rather than the command line, the same two operations are the Export and Import tabs. Export with the Quick method and the SQL format produces the same file mysqldump would; Import takes it back. It is worth doing it once each way, because an examiner may ask for either, and because knowing that the graphical button runs mysqldump underneath is the answer to "how does Export work".

What beginners get wrong

Typing mysqldump at the mysql> prompt. It is an operating system command. Leave the client first.

Looking for a mysqlrestore program. There is none; restore with mysql.

Using > for the restore. mysql -u root -p > f.sql empties the file and hangs. Restore is <.

Restoring a dump taken without --databases and not naming the database. The statements have no database to go into and the server answers "No database selected".

Never testing a backup. Restore it somewhere and count the rows.

Leaving the password on the command line after -p. -pmypassword works and is visible to everyone on the machine. Let it ask.

Dumping a busy database without --single-transaction. The parts of the file may not agree with each other.

munotes.in142

Practical 3: Backing Up a Database, and Restoring It

Quick revision

  • mysqldump is an operating system command, not an SQL statement.
  • A dump is a text file of SQL statements that rebuild the database.
  • mysqldump -u root -p --single-transaction --databases db > file.sql.
  • Restore with mysql -u root -p < file.sql. There is no mysqlrestore.
  • --databases puts the CREATE DATABASE in the file; without it, name the database on the restore.
  • --all-databases for the whole server; --no-data for the structure only.
  • A restore overwrites: the dump begins with DROP TABLE IF EXISTS.
  • Prove a restore by counting the rows before and after.
  • In phpMyAdmin the same two operations are Export and Import.

What goes in your journal

Aim, the four commands in order, and the row counts twice: once before the database was dropped and once after it was restored. The two identical counts are the whole proof, and a write-up without them has shown that the commands were typed rather than that the backup worked.

If you have room, paste the first fifteen lines of the dump file. A student who can point at the DROP TABLE IF EXISTS in their own backup and say what it means has answered the viva already.

Test yourself

1. Which program takes the backup, and which restores it? mysqldump takes it; mysql restores it. There is no separate restore program, because the dump is a file of SQL statements.

2. What is actually inside a dump file? SQL statements: the CREATE DATABASE, a CREATE TABLE for every table with its constraints, and INSERT statements for the rows.

3. What does --single-transaction do and why does it matter? It takes the whole dump as one consistent snapshot, so changes made while the dump is running cannot appear in part of the file and not in the rest.

4. What is the difference between mysqldump db and mysqldump --databases db? With --databases the file contains CREATE DATABASE and USE, so it can be restored without naming a database. Without it, the file holds only the tables and the restore must name the database.

5. Why does > appear in the backup command and < in the restore? > sends a program's output into a file; < feeds a file into a program's input. mysqldump produces the file and mysql consumes it.

6. How do you prove a backup is good? Restore it, and compare the row counts of every table with what they were before.

7. What happens if you restore onto a database that already has data? It is overwritten. The dump begins each table with DROP TABLE IF EXISTS, so nothing is merged.

Contents This chapter on its own page

munotes.in143

Chapter Thirty-Nine

Practical 4: Simple Queries

Syllabus topic Module 2, Practical 4: "Simple Queries"

Aim

To write simple queries: to select columns, filter rows, sort and limit the result.

The shape of a SELECT

SELECT   columns
FROM     table
WHERE    condition
ORDER BY columns
LIMIT    how many;

Only the first two lines are required. The rest are in the order they must be written in, and writing them out of order is a syntax error.

Every column, and chosen columns

SELECT * FROM author;
+-----------+------------------+---------+
| author_id | author_name      | country |
+-----------+------------------+---------+
|         1 | Dennis Ritchie   | USA     |
|         2 | Ramez Elmasri    | USA     |
|         3 | Vikram Vaswani   | India   |
|         4 | Behrouz Forouzan | Iran    |
|         5 | Ashwin Pajankar  | India   |
+-----------+------------------+---------+

The asterisk means every column, in the order the table declares them.

It is useful while you are exploring and it is a poor habit in a query you are going to keep. Name the columns you want. Then the result cannot change shape when somebody adds a column tomorrow, the server reads less, and a person reading the query knows what it is for.

SELECT title, price FROM book;
+-------------------+--------+
| title             | price  |
+-------------------+--------+
| The C Language    | 395.00 |
| Database Systems  | 720.50 |
| MySQL Reference   | 499.00 |
| Data Structures   | 610.00 |
| Computer Networks | 540.75 |
| Learn SQL         | 280.00 |
| Operating Systems | 655.25 |
| Discrete Maths    | 430.00 |
+-------------------+--------+

Aliases: giving a column a better name

SELECT title AS book_title, price AS rupees FROM book LIMIT 3;
+------------------+--------+
| book_title       | rupees |
+------------------+--------+
| The C Language   | 395.00 |
| Database Systems | 720.50 |
| MySQL Reference  | 499.00 |
+------------------+--------+

AS renames a column in the result only; the table is untouched. The word AS may be left out, and should not be: SELECT title book_title works and reads like a mistake.

An alias with a space or a capital letter in it needs backticks: ` AS Book Title `.

DISTINCT: no repeats

SELECT subject FROM book;
+-------------+
| subject     |
+-------------+
| Programming |
| Databases   |
| Databases   |
| Programming |
| Networks    |
| Databases   |
| Systems     |
| Mathematics |
+-------------+
SELECT DISTINCT subject FROM book;
+-------------+
| subject     |
+-------------+
| Programming |
| Databases   |
| Networks    |
| Systems     |
| Mathematics |
+-------------+

Eight rows became five. DISTINCT applies to the whole row of the result, not to one column: SELECT DISTINCT subject, price gives every different combination of the two, which is usually not what somebody typing it meant.

WHERE: choosing rows

SELECT title, price FROM book WHERE subject = 'Databases';
munotes.in144

Practical 4: Simple Queries

+------------------+--------+
| title            | price  |
+------------------+--------+
| Database Systems | 720.50 |
| MySQL Reference  | 499.00 |
| Learn SQL        | 280.00 |
+------------------+--------+

The comparison operators:

OperatorMeans
=equal to. One equals sign, not two
<> or !=not equal to
<, >, <=, >=the usual comparisons
BETWEEN a AND bfrom a to b, both ends included
IN (a, b, c)equal to any one of these
LIKE 'pattern'matches a pattern
IS NULL, IS NOT NULLhas no value, has a value

= is the first surprise for anybody arriving from C, where = assigns and == compares. In SQL there is no assignment inside a WHERE, so one sign is enough.

SELECT title, price FROM book WHERE price BETWEEN 400 AND 600;
+-------------------+--------+
| title             | price  |
+-------------------+--------+
| MySQL Reference   | 499.00 |
| Computer Networks | 540.75 |
| Discrete Maths    | 430.00 |
+-------------------+--------+

BETWEEN includes both ends, so 400 and 600 would both be in. It is the same as price >= 400 AND price <= 600, and it is easier to read.

SELECT title, subject FROM book WHERE subject IN ('Networks', 'Systems');
+-------------------+----------+
| title             | subject  |
+-------------------+----------+
| Computer Networks | Networks |
| Operating Systems | Systems  |
+-------------------+----------+

IN is the same as a chain of ORs and is much shorter. It comes back in [Practical 7: Subqueries with IN], where the list is produced by a query instead of typed.

LIKE: matching a pattern

Two wildcards, and only two.

WildcardMatches
%any number of characters, including none
_exactly one character
SELECT title FROM book WHERE title LIKE 'C%';
+-------------------+
| title             |
+-------------------+
| Computer Networks |
+-------------------+
SELECT title FROM book WHERE title LIKE '%Systems';
+-------------------+
| title             |
+-------------------+
| Database Systems  |
| Operating Systems |
+-------------------+
SELECT member_name FROM member WHERE member_name LIKE '_a%';
+---------------+
| member_name   |
+---------------+
| Ravi Deshmukh |
+---------------+

The last one asks for names whose second letter is a: one character, then an a, then anything.

Whether LIKE 'c%' in small letters also finds Computer Networks depends on the column's collation, and in this database it does, because the collation ends ci for case insensitive, as [Practical 2: Viewing Databases, Creating One and Listing Its Tables] showed. On a server whose collation is case sensitive it would not. If a LIKE must ignore case whatever the collation, write WHERE LOWER(title) LIKE 'c%'.

AND, OR, NOT

SELECT title, subject, price FROM book
WHERE subject = 'Databases' AND price > 400;
+------------------+-----------+--------+
| title            | subject   | price  |
+------------------+-----------+--------+
| Database Systems | Databases | 720.50 |
| MySQL Reference  | Databases | 499.00 |
+------------------+-----------+--------+
munotes.in145

Practical 4: Simple Queries

AND binds tighter than OR, exactly as multiplication binds tighter than addition. So this:

WHERE subject = 'Databases' OR subject = 'Networks' AND price > 600

means "Databases, or else Networks costing more than 600", which is almost certainly not what was wanted. Bracket every mixture of AND and OR, even where you are sure:

WHERE (subject = 'Databases' OR subject = 'Networks') AND price > 600

NULL, and why it needs its own operators

One member has no city recorded. Ask for the members whose city is not Mumbai:

SELECT member_id, member_name, city FROM member WHERE city <> 'Mumbai';
+-----------+---------------+--------+
| member_id | member_name   | city   |
+-----------+---------------+--------+
|         2 | Ravi Deshmukh | Thane  |
|         4 | Imran Shaikh  | Kalyan |
|         6 | Vikram Rao    | Thane  |
+-----------+---------------+--------+

Neha Patil is missing from that answer, and her city is certainly not Mumbai. Here is why:

SELECT NULL = NULL     AS null_equals_null,
       NULL <> 'Mumbai' AS null_not_mumbai,
       NULL IS NULL    AS null_is_null;
+------------------+-----------------+--------------+
| null_equals_null | null_not_mumbai | null_is_null |
+------------------+-----------------+--------------+
|             NULL |            NULL |            1 |
+------------------+-----------------+--------------+

NULL is not a value; it is the absence of one. Any comparison with it gives NULL, which is neither true nor false, and a WHERE keeps only the rows for which the condition is true. So a row with a NULL city is excluded by city = 'Mumbai' and excluded by city <> 'Mumbai' alike.

That is what IS NULL and IS NOT NULL are for, and they are the only operators that work on it:

SELECT member_id, member_name FROM member WHERE city IS NULL;
+-----------+-------------+
| member_id | member_name |
+-----------+-------------+
|         5 | Neha Patil  |
+-----------+-------------+

To ask the question that was really meant, say so:

SELECT member_id, member_name, city FROM member
WHERE city <> 'Mumbai' OR city IS NULL;
+-----------+---------------+--------+
| member_id | member_name   | city   |
+-----------+---------------+--------+
|         2 | Ravi Deshmukh | Thane  |
|         4 | Imran Shaikh  | Kalyan |
|         5 | Neha Patil    | NULL   |
|         6 | Vikram Rao    | Thane  |
+-----------+---------------+--------+

COALESCE is the tidy way to put a stand-in value in the result:

SELECT member_name, COALESCE(city, 'not recorded') AS city FROM member LIMIT 3;
+---------------+--------+
| member_name   | city   |
+---------------+--------+
| Asha Kulkarni | Mumbai |
| Ravi Deshmukh | Thane  |
| Meena Iyer    | Mumbai |
+---------------+--------+

ORDER BY

SELECT title, price FROM book ORDER BY price DESC LIMIT 4;
+-------------------+--------+
| title             | price  |
+-------------------+--------+
| Database Systems  | 720.50 |
| Operating Systems | 655.25 |
| Data Structures   | 610.00 |
| Computer Networks | 540.75 |
+-------------------+--------+
munotes.in146

Practical 4: Simple Queries

ASC is ascending and is the default; DESC is descending. Sort by more than one column by listing them, and the second decides only where the first is equal:

SELECT subject, title FROM book ORDER BY subject ASC, title DESC;
+-------------+-------------------+
| subject     | title             |
+-------------+-------------------+
| Databases   | MySQL Reference   |
| Databases   | Learn SQL         |
| Databases   | Database Systems  |
| Mathematics | Discrete Maths    |
| Networks    | Computer Networks |
| Programming | The C Language    |
| Programming | Data Structures   |
| Systems     | Operating Systems |
+-------------+-------------------+

Without an ORDER BY, the order of a result is not promised. It usually comes out in a convenient order and that order can change when an index is added or the table grows. If the order matters, say so.

LIMIT

SELECT title, price FROM book ORDER BY price DESC LIMIT 3;
+-------------------+--------+
| title             | price  |
+-------------------+--------+
| Database Systems  | 720.50 |
| Operating Systems | 655.25 |
| Data Structures   | 610.00 |
+-------------------+--------+

The three most expensive books. LIMIT with two numbers skips some first:

SELECT title, price FROM book ORDER BY price DESC LIMIT 3, 2;
+-------------------+--------+
| title             | price  |
+-------------------+--------+
| Computer Networks | 540.75 |
| MySQL Reference   | 499.00 |
+-------------------+--------+

LIMIT 3, 2 means skip three and then take two, so those are the fourth and fifth most expensive. The clearer spelling is LIMIT 2 OFFSET 3, which means the same thing with the numbers the right way round.

LIMIT without ORDER BY is the commonest way of getting a wrong answer that looks right: "the top three" is meaningless until you have said top by what.

The order the server does it in

This is the part that explains several errors at once.

Written in this orderEvaluated in this order
SELECTFROM
FROMWHERE
WHEREGROUP BY
GROUP BYHAVING
HAVINGSELECT
ORDER BYORDER BY
LIMITLIMIT

SELECT is nearly last. Two consequences follow, and both are asked:

A column alias cannot be used in WHERE. SELECT price * 1.1 AS new_price FROM book WHERE new_price > 500 fails, because the WHERE runs before the alias exists. Repeat the expression, or use a subquery.

A column alias can be used in ORDER BY, because ORDER BY runs after SELECT.

What beginners get wrong

WHERE city = NULL. Never matches. Use IS NULL.

Forgetting that <> 'Mumbai' excludes the NULLs too.

Double quotes round a string. MySQL allows them; the standard does not. Use single quotes.

munotes.in147

Practical 4: Simple Queries

SELECT * in a query that will be kept. Name the columns.

Mixing AND and OR without brackets. AND binds tighter.

LIMIT with no ORDER BY, and calling the result the top three.

A column alias in the WHERE. The alias does not exist yet.

LIKE with no wildcard. LIKE 'Learn SQL' is just a slower =.

Quick revision

  • SELECT cols FROM t WHERE cond ORDER BY cols LIMIT n; and the clauses must be written in that order.
  • * is every column; name them instead in anything you keep.
  • AS renames a column in the result only.
  • DISTINCT removes duplicate rows of the result, not duplicate values of one column.
  • =, <>, <, >, BETWEEN a AND b inclusive, IN (list), LIKE, IS NULL.
  • % is any number of characters, _ is exactly one.
  • NULL compares equal to nothing, not even NULL. Use IS NULL and IS NOT NULL.
  • COALESCE(col, 'stand-in') supplies a value where there is none.
  • ORDER BY col DESC, and a second column decides ties.
  • LIMIT n and LIMIT skip, n, which is LIMIT n OFFSET skip.
  • Evaluation order: FROM, WHERE, GROUP BY, HAVING, SELECT, ORDER BY, LIMIT.

What goes in your journal

Aim, and then one query per feature with its result underneath: every column, chosen columns, an alias, DISTINCT, each comparison operator, both LIKE wildcards, IS NULL, ORDER BY and LIMIT. Ten or twelve short queries.

Include the pair that shows the NULL trap: WHERE city <> 'Mumbai' and then the same query with OR city IS NULL. Two results, one row different, and the difference is the practical's whole lesson.

Test yourself

1. Why does WHERE city = NULL return nothing? Because a comparison with NULL gives NULL rather than true, and WHERE keeps only rows where the condition is true. The test is IS NULL.

2. What is the difference between % and _ in a LIKE pattern? % matches any number of characters, including none. _ matches exactly one.

3. Is BETWEEN 400 AND 600 inclusive? Yes, both ends are included.

4. Why can a column alias be used in ORDER BY but not in WHERE? Because WHERE is evaluated before SELECT, where the alias is created, and ORDER BY is evaluated after it.

5. What does LIMIT 3, 2 mean? Skip the first three rows and return the next two. The same thing is written LIMIT 2 OFFSET 3.

6. What order does the server evaluate the clauses in? FROM, WHERE, GROUP BY, HAVING, SELECT, ORDER BY, LIMIT.

7. Why bracket a mixture of AND and OR? Because AND binds more tightly than OR, so an unbracketed mixture rarely means what it looks like.

Contents This chapter on its own page

munotes.in148

Chapter Forty

Practical 4: Aggregate Functions, GROUP BY and HAVING

Syllabus topic Module 2, Practical 4: "Simple Queries with Aggregate functions"

Aim

To write queries using aggregate functions.

What an aggregate function is

An ordinary expression works on one row at a time. An aggregate function works on many rows at once and returns a single value: how many, the smallest, the largest, the total, the average.

There are five, and MU's wording points at all of them.

FunctionReturns
COUNT(x)how many
SUM(x)the total
AVG(x)the average
MIN(x)the smallest
MAX(x)the largest
SELECT COUNT(*) AS books,
       MIN(price) AS cheapest,
       MAX(price) AS dearest,
       SUM(price) AS total,
       ROUND(AVG(price), 2) AS average
FROM book;
+-------+----------+---------+---------+---------+
| books | cheapest | dearest | total   | average |
+-------+----------+---------+---------+---------+
|     8 |   280.00 |  720.50 | 4130.50 |  516.31 |
+-------+----------+---------+---------+---------+

One row out of eight rows in. That is what an aggregate does: it collapses many rows into one.

ROUND(AVG(price), 2) is there because AVG returns as many decimal places as it needs, and a price wants two. ROUND is in [Practical 5: Math Functions].

The one rule about NULL, and the one exception

Every aggregate except COUNT(*) ignores NULLs.

SELECT COUNT(*)           AS all_loans,
       COUNT(loan_id)     AS loans_with_an_id,
       COUNT(returned_on) AS loans_returned
FROM loan;
+-----------+------------------+----------------+
| all_loans | loans_with_an_id | loans_returned |
+-----------+------------------+----------------+
|         7 |                7 |              5 |
+-----------+------------------+----------------+

Seven loans, seven loan ids, and only five return dates, because two books are still out and their returned_on is NULL.

WrittenCounts
COUNT(*)rows, whatever is in them
COUNT(column)rows where that column is not NULL
COUNT(DISTINCT column)different non-NULL values

That table answers "what is the difference between COUNT(*) and COUNT(column)", which is asked in this practical's viva more than any other question.

The same rule bites hardest on AVG, and this is worth being careful about:

SELECT SUM(returned_on IS NOT NULL) AS returned,
       COUNT(*) AS rows_in_table,
       AVG(DATEDIFF(returned_on, issued_on)) AS avg_days_out
FROM loan;
+----------+---------------+--------------+
| returned | rows_in_table | avg_days_out |
+----------+---------------+--------------+
|        5 |             7 |      14.2000 |
+----------+---------------+--------------+

AVG divided by five, the loans that have come back, not by seven. That is almost always what you want, and it is almost never what a student expects. An average over a column with NULLs in it is an average of the rows that have values.

GROUP BY

An aggregate over the whole table gives one number. GROUP BY gives one number per group.

SELECT subject, COUNT(*) AS how_many
FROM book
GROUP BY subject
ORDER BY how_many DESC, subject;
+-------------+----------+
| subject     | how_many |
+-------------+----------+
| Databases   |        3 |
| Programming |        2 |
| Mathematics |        1 |
| Networks    |        1 |
| Systems     |        1 |
+-------------+----------+

The server sorted the rows into piles by subject and counted each pile.

SELECT subject,
       COUNT(*) AS how_many,
       MIN(price) AS cheapest,
       ROUND(AVG(price), 2) AS average
FROM book
GROUP BY subject
ORDER BY subject;
munotes.in149

Practical 4: Aggregate Functions, GROUP BY and HAVING

+-------------+----------+----------+---------+
| subject     | how_many | cheapest | average |
+-------------+----------+----------+---------+
| Databases   |        3 |   280.00 |  499.83 |
| Mathematics |        1 |   430.00 |  430.00 |
| Networks    |        1 |   540.75 |  540.75 |
| Programming |        2 |   395.00 |  502.50 |
| Systems     |        1 |   655.25 |  655.25 |
+-------------+----------+----------+---------+

The rule about what may go in the SELECT

Every column in the SELECT must either be in the GROUP BY or be inside an aggregate function. Anything else has no single answer per group.

SELECT subject, title, COUNT(*) FROM book GROUP BY subject;
ERROR 1055 (42000): Expression #2 of SELECT list is not in GROUP BY clause and contains nonaggregated column 'librarydb.book.title' which is not functionally dependent on columns in GROUP BY clause; this is incompatible with sql_mode=only_full_group_by

Error 1055. Three books are Databases; which of their three titles should the one Databases row show? There is no answer, so the server refuses.

That refusal comes from a setting called ONLY_FULL_GROUP_BY, which has been on by default since MySQL 5.7. On an older server, or one where it has been switched off, the query runs and silently picks a title at random. If your laboratory's server accepts the query above, it is not enforcing the rule, and the answer you get is not to be trusted.

HAVING

WHERE filters rows, before grouping. HAVING filters groups, after.

SELECT subject, COUNT(*) AS how_many
FROM book
GROUP BY subject
HAVING COUNT(*) > 1;
+-------------+----------+
| subject     | how_many |
+-------------+----------+
| Programming |        2 |
| Databases   |        3 |
+-------------+----------+

Only the subjects with more than one book. That cannot be done with WHERE, because at the time WHERE runs there are no groups yet and so no counts:

SELECT subject, COUNT(*) FROM book WHERE COUNT(*) > 1 GROUP BY subject;
ERROR 1111 (HY000): Invalid use of group function

Error 1111, invalid use of group function, and the evaluation order from [Practical 4: Simple Queries] explains it exactly: WHERE runs before GROUP BY.

The two are often used together, and then each does its own job:

SELECT subject, COUNT(*) AS how_many, ROUND(AVG(price), 2) AS average
FROM book
WHERE price > 300
GROUP BY subject
HAVING COUNT(*) > 1
ORDER BY average DESC;
+-------------+----------+---------+
| subject     | how_many | average |
+-------------+----------+---------+
| Databases   |        2 |  609.75 |
| Programming |        2 |  502.50 |
+-------------+----------+---------+

Read it in the server's order: take the books costing more than 300, pile them by subject, throw away the piles with only one book in them, work out the count and the average for the piles that are left, and sort.

munotes.in150

Practical 4: Aggregate Functions, GROUP BY and HAVING

WHEREHAVING
Filtersrowsgroups
Runsbefore GROUP BYafter it
May use an aggregatenoyes
May use a column not in the GROUP BYyesno
Without a GROUP BYordinary filteringtreats the whole table as one group

Grouping by more than one column

SELECT m.course, b.subject, COUNT(*) AS borrowings
FROM loan l
JOIN member m ON m.member_id = l.member_id
JOIN book   b ON b.book_id   = l.book_id
GROUP BY m.course, b.subject
ORDER BY m.course, b.subject;
+--------+-------------+------------+
| course | subject     | borrowings |
+--------+-------------+------------+
| BSc CS | Networks    |          1 |
| BSc CS | Programming |          1 |
| BSc IT | Databases   |          3 |
| BSc IT | Programming |          2 |
+--------+-------------+------------+

One row per combination that actually occurs. Combinations nobody borrowed do not appear at all, which is a real limitation: a grouped query cannot show a zero for a group that has no rows. To show those, start from the table that has all the values and use an outer join, which is [Practical 6: The Outer Join].

The joins in that query are [Practical 6: The Inner Join]; they are used here because a grouping worth doing usually spans two tables.

What beginners get wrong

Putting a plain column in the SELECT that is not in the GROUP BY. Error 1055, and worse on a server that allows it.

An aggregate in the WHERE. Error 1111. It belongs in HAVING.

Expecting COUNT(column) to count every row. It skips the NULLs. COUNT(*) does not.

Expecting AVG to divide by the number of rows. It divides by the number of non-NULL values.

Using HAVING where WHERE would do. It works, and it makes the server group rows it is about to throw away.

Forgetting ORDER BY. A GROUP BY does not promise an order.

Reading a count of rows as a count of things. COUNT(*) on a join counts matched pairs; the number of distinct books in that result is COUNT(DISTINCT book_id).

Quick revision

  • Five aggregates: COUNT, SUM, AVG, MIN, MAX.
  • Every one except COUNT(*) ignores NULLs.
  • COUNT(*) counts rows, COUNT(col) counts non-NULL values, COUNT(DISTINCT col) counts different ones.
  • GROUP BY makes one result row per group.
  • Every SELECT column must be in the GROUP BY or inside an aggregate; ONLY_FULL_GROUP_BY enforces it. Error 1055.
  • WHERE filters rows before grouping; HAVING filters groups after.
  • An aggregate in a WHERE is error 1111.
  • A group with no rows does not appear at all.

What goes in your journal

Aim, then each of the five functions once over the whole table, then a GROUP BY with a count, then the same query with HAVING, and finally one query using WHERE, GROUP BY, HAVING and ORDER BY together. Six or seven queries with their results.

munotes.in151

Practical 4: Aggregate Functions, GROUP BY and HAVING

Add the COUNT(*) against COUNT(returned_on) pair and one line saying why the numbers differ. That pair is the evidence that you know the NULL rule, and it is the question that will be asked.

Test yourself

1. Which aggregate does not ignore NULLs? COUNT(*). It counts rows regardless of their contents; every other aggregate, including COUNT(column), skips NULLs.

2. What is the difference between WHERE and HAVING? WHERE filters individual rows before they are grouped. HAVING filters whole groups afterwards and may use aggregate functions.

3. Why is SELECT subject, title, COUNT(*) FROM book GROUP BY subject; refused? Because title is neither in the GROUP BY nor inside an aggregate, so there is no single title for a group of several books.

4. AVG over a column with NULLs: what is it the average of? Of the non-NULL values only. The NULL rows are not counted in the divisor.

5. Why can you not write WHERE COUNT(*) > 1? Because WHERE is evaluated before GROUP BY, so no groups and no counts exist yet. Use HAVING.

6. A subject with no books at all: does it appear in a GROUP BY subject result? No. A group with no rows produces no row. Showing it requires an outer join from a table that lists every subject.

Contents This chapter on its own page

munotes.in152

Chapter Forty-One

Practical 5: Date Functions

Syllabus topic Module 2, Practical 5: "Queries involving Date Functions"

Aim

To write queries involving date functions.

Why a date must be a DATE

A date kept in a VARCHAR is a piece of text. Text cannot be subtracted from text, and text sorts alphabetically, so '09-01-2025' comes before '10-12-2024'. A DATE column can be compared, subtracted, added to and sorted, and every function in this chapter works on it.

MySQL's format is YYYY-MM-DD, always, and it is the ISO 8601 order: largest unit first. That order is why a DATE sorts correctly even as text, and it is the reason the standard chose it.

Today, and now

SELECT CURDATE(), CURRENT_DATE, NOW(), CURTIME();

That listing really ran; its result is not printed here because it would be different tomorrow.

FunctionReturnsLooks like
CURDATE(), CURRENT_DATEtoday's date2025-07-21
NOW(), CURRENT_TIMESTAMPdate and time2025-07-21 14:30:00
CURTIME()the time of day14:30:00
SYSDATE()the time at the moment it is calledas NOW()

NOW() and SYSDATE() differ in a way worth one line: NOW() gives the time the statement started, so it is the same everywhere in one statement; SYSDATE() gives the time it was called, so two of them in one statement can disagree. Use NOW().

Pulling a date apart

SELECT issued_on,
       YEAR(issued_on)  AS yr,
       MONTH(issued_on) AS mth,
       DAY(issued_on)   AS dy
FROM loan WHERE loan_id <= 3;
+------------+------+------+------+
| issued_on  | yr   | mth  | dy   |
+------------+------+------+------+
| 2025-06-02 | 2025 |    6 |    2 |
| 2025-06-02 | 2025 |    6 |    2 |
| 2025-06-10 | 2025 |    6 |   10 |
+------------+------+------+------+

And the parts that have names rather than numbers:

SELECT issued_on,
       MONTHNAME(issued_on) AS month_name,
       DAYNAME(issued_on)   AS day_name,
       QUARTER(issued_on)   AS qtr
FROM loan WHERE loan_id = 1;
+------------+------------+----------+------+
| issued_on  | month_name | day_name | qtr  |
+------------+------------+----------+------+
| 2025-06-02 | June       | Monday   |    2 |
+------------+------------+----------+------+

DAYNAME is genuinely useful: a library that closes on Sundays can find loans issued on one without anybody working out the day by hand.

The difference between two dates

SELECT loan_id, issued_on, returned_on,
       DATEDIFF(returned_on, issued_on) AS days_kept
FROM loan
WHERE returned_on IS NOT NULL
ORDER BY loan_id;
+---------+------------+-------------+-----------+
| loan_id | issued_on  | returned_on | days_kept |
+---------+------------+-------------+-----------+
|       1 | 2025-06-02 | 2025-06-14  |        12 |
|       2 | 2025-06-02 | 2025-06-30  |        28 |
|       3 | 2025-06-10 | 2025-06-20  |        10 |
|       5 | 2025-07-05 | 2025-07-15  |        10 |
|       6 | 2025-07-08 | 2025-07-19  |        11 |
+---------+------------+-------------+-----------+

DATEDIFF(a, b) is a minus b, in whole days, and the order matters: the other way round gives a negative number. It counts days, not hours, so a loan issued and returned on the same day is 0.

munotes.in153

Practical 5: Date Functions

TIMESTAMPDIFF does the same in a unit you choose, which is how an age or a length of membership is worked out:

SELECT member_name, joined_on,
       TIMESTAMPDIFF(MONTH, joined_on, '2025-08-01') AS months_a_member
FROM member ORDER BY member_id LIMIT 4;
+---------------+------------+-----------------+
| member_name   | joined_on  | months_a_member |
+---------------+------------+-----------------+
| Asha Kulkarni | 2024-07-15 |              12 |
| Ravi Deshmukh | 2024-07-18 |              12 |
| Meena Iyer    | 2024-08-02 |              11 |
| Imran Shaikh  | 2024-08-09 |              11 |
+---------------+------------+-----------------+

TIMESTAMPDIFF(unit, a, b) is b minus a, which is the opposite way round from DATEDIFF. That is genuinely confusing and it is worth writing in your journal. The units are DAY, WEEK, MONTH, QUARTER, YEAR, HOUR, MINUTE and SECOND.

Adding and subtracting

SELECT issued_on,
       DATE_ADD(issued_on, INTERVAL 14 DAY) AS due_on,
       DATE_SUB(issued_on, INTERVAL 1 WEEK) AS a_week_before
FROM loan WHERE loan_id = 1;
+------------+------------+---------------+
| issued_on  | due_on     | a_week_before |
+------------+------------+---------------+
| 2025-06-02 | 2025-06-16 | 2025-05-26    |
+------------+------------+---------------+

INTERVAL 14 DAY is one thing, not two: a number and a unit written together. The units are the same list as for TIMESTAMPDIFF.

ADDDATE and SUBDATE are other names for the same two functions, and issued_on + INTERVAL 14 DAY also works. All three forms are accepted; use whichever your teacher uses.

The overdue query, which is what the practical is for

The library lends for fourteen days and charges two rupees a day after that.

SELECT l.loan_id, b.title,
       l.issued_on,
       DATE_ADD(l.issued_on, INTERVAL 14 DAY) AS due_on,
       l.returned_on,
       DATEDIFF(l.returned_on, DATE_ADD(l.issued_on, INTERVAL 14 DAY)) AS days_late
FROM loan l JOIN book b ON b.book_id = l.book_id
WHERE l.returned_on IS NOT NULL
ORDER BY l.loan_id;
+---------+------------------+------------+------------+-------------+-----------+
| loan_id | title            | issued_on  | due_on     | returned_on | days_late |
+---------+------------------+------------+------------+-------------+-----------+
|       1 | The C Language   | 2025-06-02 | 2025-06-16 | 2025-06-14  |        -2 |
|       2 | Database Systems | 2025-06-02 | 2025-06-16 | 2025-06-30  |        14 |
|       3 | MySQL Reference  | 2025-06-10 | 2025-06-24 | 2025-06-20  |        -4 |
|       5 | The C Language   | 2025-07-05 | 2025-07-19 | 2025-07-15  |        -4 |
|       6 | Learn SQL        | 2025-07-08 | 2025-07-22 | 2025-07-19  |        -3 |
+---------+------------------+------------+------------+-------------+-----------+

One loan is late and the rest are not, and the negative numbers are the days to spare. Turning that into a fine means charging nothing when the book was on time, which is what GREATEST does:

SELECT l.loan_id, b.title,
       GREATEST(DATEDIFF(l.returned_on,
                DATE_ADD(l.issued_on, INTERVAL 14 DAY)), 0) * 2 AS fine
FROM loan l JOIN book b ON b.book_id = l.book_id
WHERE l.returned_on IS NOT NULL
ORDER BY l.loan_id;
+---------+------------------+------+
| loan_id | title            | fine |
+---------+------------------+------+
|       1 | The C Language   |    0 |
|       2 | Database Systems |   28 |
|       3 | MySQL Reference  |    0 |
|       5 | The C Language   |    0 |
|       6 | Learn SQL        |    0 |
+---------+------------------+------+
munotes.in154

Practical 5: Date Functions

GREATEST(x, 0) is the larger of the two, so a negative number becomes 0 and nobody is paid for returning a book early. It is the SQL of the if a student would otherwise write in a program, and doing it in the query is the point of the practical.

For the books still out, the due date is compared with CURDATE():

SELECT loan_id, issued_on,
       GREATEST(DATEDIFF(CURDATE(),
                DATE_ADD(issued_on, INTERVAL 14 DAY)), 0) * 2 AS fine_so_far
FROM loan WHERE returned_on IS NULL;

That one really ran too, and its answer grows by two rupees every day, which is why it is not printed here.

Formatting a date for a person to read

SELECT issued_on,
       DATE_FORMAT(issued_on, '%d-%m-%Y')      AS indian,
       DATE_FORMAT(issued_on, '%d %M %Y')      AS long_form,
       DATE_FORMAT(issued_on, '%W, %e %b %y')  AS with_day
FROM loan WHERE loan_id = 1;
+------------+------------+--------------+------------------+
| issued_on  | indian     | long_form    | with_day         |
+------------+------------+--------------+------------------+
| 2025-06-02 | 02-06-2025 | 02 June 2025 | Monday, 2 Jun 25 |
+------------+------------+--------------+------------------+
CodeMeans
%dday, two digits
%eday, no leading zero
%mmonth, two digits
%bmonth, short name
%Mmonth, full name
%yyear, two digits
%Yyear, four digits
%Wweekday, full name
%H:%i:%shours, minutes, seconds

DATE_FORMAT produces text, not a date. Its result cannot be sorted as a date or subtracted from another date, so format at the very end, in the SELECT, and never store the formatted version.

Going the other way, STR_TO_DATE reads text into a date with the same codes:

SELECT STR_TO_DATE('21-07-2025', '%d-%m-%Y') AS a_real_date;
+-------------+
| a_real_date |
+-------------+
| 2025-07-21  |
+-------------+

That is what you use when somebody hands you a file of dates written the Indian way round.

Two more that answer real questions

SELECT LAST_DAY('2025-02-10') AS end_of_feb,
       LAST_DAY('2024-02-10') AS end_of_feb_leap,
       DAYOFYEAR('2025-03-01') AS day_number;
+------------+-----------------+------------+
| end_of_feb | end_of_feb_leap | day_number |
+------------+-----------------+------------+
| 2025-02-28 | 2024-02-29      |         60 |
+------------+-----------------+------------+

LAST_DAY knows about leap years, which is the leap year rule of [Practical 1(c): Is This a Leap Year?] solved once by the server so that nobody has to write it again.

What beginners get wrong

Storing a date in a VARCHAR. Nothing in this chapter then works.

Writing a date as '21-07-2025'. MySQL wants '2025-07-21'. Use STR_TO_DATE for input in another order.

Getting DATEDIFF's arguments the wrong way round. It is the first minus the second.

Expecting TIMESTAMPDIFF to take them in the same order. It does not: it is the second minus the first.

munotes.in155

Practical 5: Date Functions

Writing INTERVAL 14 DAYS. The unit is singular: DAY.

Sorting on a DATE_FORMAT result. It is text, so 09 January sorts before 10 December.

Forgetting that a date arithmetic with NULL gives NULL. A loan that has not come back has no days_kept, and the row shows NULL rather than being missing.

Quick revision

  • Dates are YYYY-MM-DD; store them in DATE, never in VARCHAR.
  • CURDATE() today, NOW() date and time, CURTIME() the time.
  • YEAR, MONTH, DAY, MONTHNAME, DAYNAME, QUARTER pull a date apart.
  • DATEDIFF(a, b) is a minus b in days.
  • TIMESTAMPDIFF(unit, a, b) is b minus a, in the unit named.
  • DATE_ADD(d, INTERVAL n DAY) and DATE_SUB; the unit is singular.
  • GREATEST(x, 0) turns a negative number into zero, which is how a fine stops at nothing.
  • DATE_FORMAT(d, '%d-%m-%Y') makes text; STR_TO_DATE(s, '%d-%m-%Y') reads it back.
  • LAST_DAY knows about leap years.

What goes in your journal

Aim, then one query per group: the parts of a date, a DATEDIFF, a DATE_ADD, a DATE_FORMAT and a STR_TO_DATE. Then the overdue query in full, with the due date, the days late and the fine in the same result.

The overdue query is the one to spend time on. It uses four ideas at once, it answers a question a real library asks, and it is the query an examiner will set.

Test yourself

1. In what format does MySQL expect a date literal? 'YYYY-MM-DD', year first.

2. DATEDIFF('2025-06-14', '2025-06-02') gives what?

  1. It is the first date minus the second, in whole days.

3. How do you add fourteen days to a date? DATE_ADD(d, INTERVAL 14 DAY), or equivalently d + INTERVAL 14 DAY.

4. Why should a DATE_FORMAT result never be stored in a column? Because it is text. It cannot be compared or subtracted as a date, and it sorts alphabetically rather than chronologically.

5. How would you charge two rupees a day for a late book and nothing for one returned on time? GREATEST(DATEDIFF(returned_on, due_on), 0) * 2, so that a negative number of days becomes zero.

6. Which is the second minus the first: DATEDIFF or TIMESTAMPDIFF? TIMESTAMPDIFF. DATEDIFF is the first minus the second.

Contents This chapter on its own page

munotes.in156

Chapter Forty-Two

Practical 5: String Functions

Syllabus topic Module 2, Practical 5: "Queries involving String Functions"

Aim

To write queries involving string functions.

Length: two functions, and they are not the same

SELECT title, LENGTH(title) AS bytes, CHAR_LENGTH(title) AS characters
FROM book WHERE book_id <= 3;
+------------------+-------+------------+
| title            | bytes | characters |
+------------------+-------+------------+
| The C Language   |    14 |         14 |
| Database Systems |    16 |         16 |
| MySQL Reference  |    15 |         15 |
+------------------+-------+------------+

For English text the two agree, and they part company the moment the text is not English:

SELECT LENGTH('मुंबई')      AS bytes,
       CHAR_LENGTH('मुंबई') AS characters;
+-------+------------+
| bytes | characters |
+-------+------------+
|    15 |          5 |
+-------+------------+

Five characters, fifteen bytes, because each Devanagari character takes three bytes in utf8mb4.

CHAR_LENGTH counts characters and is almost always what you want. LENGTH counts bytes and is what you want when you are worrying about storage. Using LENGTH to check that a name is not too long is a bug waiting for its first non-English input.

Changing case

SELECT title, UPPER(title) AS shouted, LOWER(subject) AS quiet
FROM book WHERE book_id = 1;
+----------------+----------------+-------------+
| title          | shouted        | quiet       |
+----------------+----------------+-------------+
| The C Language | THE C LANGUAGE | programming |
+----------------+----------------+-------------+

UCASE and LCASE are other names for the same two functions.

Their real use is making a comparison case insensitive whatever the column's collation says, which [Practical 4: Simple Queries] pointed at:

SELECT title FROM book WHERE LOWER(title) = 'learn sql';
+-----------+
| title     |
+-----------+
| Learn SQL |
+-----------+

Joining strings together

SELECT CONCAT(member_name, ' of ', course) AS who
FROM member WHERE member_id <= 3;
+-------------------------+
| who                     |
+-------------------------+
| Asha Kulkarni of BSc IT |
| Ravi Deshmukh of BSc IT |
| Meena Iyer of BSc CS    |
+-------------------------+

CONCAT gives NULL if any argument is NULL, and that catches everybody:

SELECT member_name,
       CONCAT(member_name, ', ', city)         AS plain_concat,
       CONCAT_WS(', ', member_name, city)      AS concat_ws
FROM member WHERE member_id IN (1, 5);
+---------------+-----------------------+-----------------------+
| member_name   | plain_concat          | concat_ws             |
+---------------+-----------------------+-----------------------+
| Asha Kulkarni | Asha Kulkarni, Mumbai | Asha Kulkarni, Mumbai |
| Neha Patil    | NULL                  | Neha Patil            |
+---------------+-----------------------+-----------------------+

Neha Patil has no city, so CONCAT produced NULL for her whole line and threw the name away with it. CONCAT_WS takes a separator as its first argument and skips the NULLs, which is why it exists. WS stands for "with separator".

The other repair is COALESCE(city, 'not recorded') from [Practical 4: Simple Queries], which puts a stand-in value in.

Taking part of a string

This is the SQL of [Practical 6(a): Extracting Part of a String], and there is one difference that matters.

SELECT title,
       SUBSTRING(title, 1, 3)  AS first_three,
       SUBSTRING(title, 3)     AS from_the_third,
       LEFT(title, 5)          AS leftmost,
       RIGHT(title, 4)         AS rightmost
FROM book WHERE book_id = 2;
munotes.in157

Practical 5: String Functions

+------------------+-------------+----------------+----------+-----------+
| title            | first_three | from_the_third | leftmost | rightmost |
+------------------+-------------+----------------+----------+-----------+
| Database Systems | Dat         | tabase Systems | Datab    | tems      |
+------------------+-------------+----------------+----------+-----------+

SUBSTRING(s, 1, 3) starts at position 1, and position 1 is the first character. In C, s[0] is the first character. The same paper teaches both conventions, in the same semester, and mixing them up is worth marks in either half.

C, in Module 1SQL, here
First characters[0]SUBSTRING(s, 1, 1)
Three characters from the starta loop from 0 to 2SUBSTRING(s, 1, 3)
Everything from the fourtha loop from 3SUBSTRING(s, 4)

SUBSTR and MID are other names for SUBSTRING. A negative starting position counts from the right, so SUBSTRING(title, -4) is the last four characters.

Finding something inside a string

SELECT title,
       INSTR(title, 'S')        AS first_capital_s,
       LOCATE('Sys', title)     AS where_sys,
       title LIKE '%Systems'    AS ends_with_systems
FROM book WHERE book_id IN (2, 7);
+-------------------+-----------------+-----------+-------------------+
| title             | first_capital_s | where_sys | ends_with_systems |
+-------------------+-----------------+-----------+-------------------+
| Database Systems  |               7 |        10 |                 1 |
| Operating Systems |              11 |        11 |                 1 |
+-------------------+-----------------+-----------+-------------------+

INSTR(s, sub) returns the position where sub starts, counting from 1, and 0 when it is not there. Zero is not a position, which is what makes it a safe answer for "not found", exactly as minus 1 was in [Practical 10: Designing the Bank Management System].

LOCATE(sub, s) is the same thing with its arguments the other way round, and it takes an optional third argument saying where to start looking.

Trimming and padding

SELECT CONCAT('[', TRIM('   Mumbai   '), ']')         AS trimmed,
       CONCAT('[', LTRIM('   Mumbai'), ']')           AS left_trimmed,
       CONCAT('[', RTRIM('Mumbai   '), ']')           AS right_trimmed,
       LPAD('7', 4, '0')                              AS padded,
       RPAD('BSc', 6, '.')                            AS dotted;
+----------+--------------+---------------+--------+--------+
| trimmed  | left_trimmed | right_trimmed | padded | dotted |
+----------+--------------+---------------+--------+--------+
| [Mumbai] | [Mumbai]     | [Mumbai]      | 0007   | BSc... |
+----------+--------------+---------------+--------+--------+

The square brackets are there so that the spaces can be seen; without them a trimmed and an untrimmed value look identical on the screen.

TRIM removes spaces from both ends and not from the middle. LPAD('7', 4, '0') makes a value four characters wide by adding zeros on the left, which is how a reference number is printed as 0007.

Replacing and reversing

SELECT REPLACE('BSc IT Semester 1', 'Semester', 'Sem') AS shortened,
       REVERSE('level')                                AS reversed,
       REPEAT('ab', 3)                                 AS repeated;
+--------------+----------+----------+
| shortened    | reversed | repeated |
+--------------+----------+----------+
| BSc IT Sem 1 | level    | ababab   |
+--------------+----------+----------+

REVERSE gives the palindrome test of [Practical 6(b): Is This String a Palindrome?] in one line:

munotes.in158

Practical 5: String Functions

SELECT 'level' AS word, ('level' = REVERSE('level')) AS is_palindrome,
       'mumbai' AS word2, ('mumbai' = REVERSE('mumbai')) AS is_palindrome2;
+-------+---------------+--------+----------------+
| word  | is_palindrome | word2  | is_palindrome2 |
+-------+---------------+--------+----------------+
| level |             1 | mumbai |              0 |
+-------+---------------+--------+----------------+

1 is true and 0 is false. MySQL has no separate boolean type: a comparison returns 1 or 0, and anything that is not 0 counts as true.

A worked example: initials from a name

SELECT member_name,
       CONCAT(LEFT(member_name, 1), '.',
              LEFT(SUBSTRING_INDEX(member_name, ' ', -1), 1), '.') AS initials
FROM member ORDER BY member_id LIMIT 4;
+---------------+----------+
| member_name   | initials |
+---------------+----------+
| Asha Kulkarni | A.K.     |
| Ravi Deshmukh | R.D.     |
| Meena Iyer    | M.I.     |
| Imran Shaikh  | I.S.     |
+---------------+----------+

SUBSTRING_INDEX(s, ' ', -1) takes everything after the last space, which is the surname. It is the most useful of the less known string functions: with 1 it takes everything before the first separator, with -1 everything after the last.

The functions in this practical

FunctionDoes
LENGTH(s)bytes
CHAR_LENGTH(s)characters
UPPER(s), LOWER(s)change case
CONCAT(a, b, ...)join; NULL if any argument is NULL
CONCAT_WS(sep, a, b, ...)join with a separator, skipping NULLs
SUBSTRING(s, from, len)part of a string, counting from 1
LEFT(s, n), RIGHT(s, n)the first or last n characters
INSTR(s, sub), LOCATE(sub, s)where it starts, or 0
TRIM(s), LTRIM(s), RTRIM(s)remove spaces from the ends
LPAD(s, n, c), RPAD(s, n, c)pad to a width
REPLACE(s, old, new)replace every occurrence
REVERSE(s)backwards
SUBSTRING_INDEX(s, sep, n)before the nth separator, or after it if n is negative

What beginners get wrong

Counting a string's characters from 0. SQL counts from 1.

Using LENGTH where CHAR_LENGTH was meant. They differ as soon as the text is not plain English.

Forgetting that CONCAT with a NULL gives NULL. Use CONCAT_WS or COALESCE.

Swapping the arguments of INSTR and LOCATE. INSTR(string, part), LOCATE(part, string).

Expecting TRIM to remove spaces from the middle. It works on the ends only; REPLACE(s, ' ', '') removes them all.

Putting a function on the column in a WHERE. WHERE LOWER(title) = 'learn sql' cannot use an index on title, so on a large table it reads every row. It is correct and it is slow, and knowing that is worth a mark.

Quick revision

  • LENGTH counts bytes, CHAR_LENGTH counts characters. Devanagari is three bytes a character in utf8mb4.
  • UPPER and LOWER, also called UCASE and LCASE.
  • CONCAT returns NULL if any argument is NULL; CONCAT_WS(sep, ...) skips NULLs.
  • SUBSTRING(s, start, length) and the first character is at 1, not 0.
  • LEFT, RIGHT, and a negative start in SUBSTRING counts from the right.
  • INSTR(s, sub) returns the position or 0; LOCATE(sub, s) is the same with the arguments reversed.
  • TRIM, LTRIM, RTRIM take spaces off the ends only.
  • LPAD, RPAD, REPLACE, REVERSE, REPEAT, SUBSTRING_INDEX.
  • A comparison returns 1 for true and 0 for false.
munotes.in159

Practical 5: String Functions

What goes in your journal

Aim, then one query per group with its result: the two lengths, the case functions, a CONCAT and the CONCAT_WS that repairs it, a SUBSTRING, an INSTR, a TRIM with brackets round it so the spaces show, and a REPLACE.

Write the counting-from-1 line into the conclusion in so many words, next to the C convention from Module 1. It is the single most useful sentence in this chapter and it is a viva question in both halves of the paper.

Test yourself

1. What is the difference between LENGTH and CHAR_LENGTH? LENGTH counts bytes and CHAR_LENGTH counts characters. They differ for any text outside the plain ASCII range; a Devanagari character is three bytes in utf8mb4.

2. From what position does SUBSTRING count? From 1. SUBSTRING(s, 1, 1) is the first character, where in C it would be s[0].

3. What does CONCAT('Asha', ', ', NULL) return? NULL. Any NULL argument makes the whole result NULL. CONCAT_WS skips NULLs instead.

4. What does INSTR return when the substring is not found? 0, which is not a valid position and so is a safe marker for "not there".

5. How would you print a number as four digits with leading zeros? LPAD(n, 4, '0').

6. Why can WHERE LOWER(title) = 'learn sql' be slow on a large table? Because the function is applied to the column, so an index on that column cannot be used and every row has to be read.

Contents This chapter on its own page

munotes.in160

Chapter Forty-Three

Practical 5: Math Functions

Syllabus topic Module 2, Practical 5: "Queries involving Math Functions"

Aim

To write queries involving math functions.

Rounding, three different ways

SELECT ROUND(516.3125, 2)    AS rounded,
       TRUNCATE(516.3125, 2) AS truncated,
       FLOOR(516.3125)       AS floored,
       CEIL(516.3125)        AS ceiled;
+---------+-----------+---------+--------+
| rounded | truncated | floored | ceiled |
+---------+-----------+---------+--------+
|  516.31 |    516.31 |     516 |    517 |
+---------+-----------+---------+--------+
FunctionDoes
ROUND(x, d)rounds to d decimal places, up or down, at the halfway point away from zero
TRUNCATE(x, d)cuts off after d decimal places, never rounding
FLOOR(x)the largest whole number not greater than x
CEIL(x), CEILING(x)the smallest whole number not less than x

For a positive number, TRUNCATE(x, 0) and FLOOR(x) give the same answer, which is why students think they are the same function. They are not, and a negative number proves it:

SELECT ROUND(-7.6)    AS rounded,
       TRUNCATE(-7.6, 0) AS truncated,
       FLOOR(-7.6)    AS floored,
       CEIL(-7.6)     AS ceiled;
+---------+-----------+---------+--------+
| rounded | truncated | floored | ceiled |
+---------+-----------+---------+--------+
|      -8 |        -7 |      -8 |     -7 |
+---------+-----------+---------+--------+

Read that row carefully.

TRUNCATE throws the fraction away, so minus 7.6 becomes minus 7: it moves towards zero.

FLOOR goes down, so minus 7.6 becomes minus 8: it moves away from zero for a negative number.

ROUND goes to the nearest, which is minus 8 here.

CEIL goes up, which is minus 7.

That table is the answer to "distinguish between ROUND, TRUNCATE, FLOOR and CEIL", and the negative example is what makes the answer convincing.

ROUND with no second argument rounds to a whole number. ROUND(x, -1) rounds to the nearest ten, which is occasionally exactly what a report wants.

Rounding a real column

SELECT subject,
       AVG(price)            AS raw_average,
       ROUND(AVG(price), 2)  AS to_paise,
       ROUND(AVG(price))     AS to_rupees
FROM book GROUP BY subject ORDER BY subject LIMIT 3;
+-------------+-------------+----------+-----------+
| subject     | raw_average | to_paise | to_rupees |
+-------------+-------------+----------+-----------+
| Databases   |  499.833333 |   499.83 |       500 |
| Mathematics |  430.000000 |   430.00 |       430 |
| Networks    |  540.750000 |   540.75 |       541 |
+-------------+-------------+----------+-----------+

AVG gives more decimal places than anybody wants; ROUND is what makes it a price again.

Absolute value, sign and remainder

SELECT ABS(-42)      AS absolute,
       SIGN(-42)     AS sign_of_it,
       SIGN(0)       AS sign_of_zero,
       MOD(17, 5)    AS remainder,
       17 % 5        AS same_thing,
       17 DIV 5      AS whole_part;
+----------+------------+--------------+-----------+------------+------------+
| absolute | sign_of_it | sign_of_zero | remainder | same_thing | whole_part |
+----------+------------+--------------+-----------+------------+------------+
|       42 |         -1 |            0 |         2 |          2 |          3 |
+----------+------------+--------------+-----------+------------+------------+

ABS is [Practical 4(c): sqrt and abs, and the Headers They Need] in SQL, and MOD is the % of [Practical 1(c): Is This a Leap Year?]. % and MOD are the same operator written two ways, and DIV is integer division: the whole part with the remainder thrown away.

munotes.in161

Practical 5: Math Functions

So the leap year rule of Module 1 is one expression here:

SELECT y AS year,
       (MOD(y, 400) = 0 OR (MOD(y, 4) = 0 AND MOD(y, 100) <> 0)) AS is_leap
FROM (SELECT 2000 AS y UNION SELECT 1900 UNION SELECT 2024
      UNION SELECT 2023) AS years
ORDER BY y;
+------+---------+
| year | is_leap |
+------+---------+
| 1900 |       0 |
| 2000 |       1 |
| 2023 |       0 |
| 2024 |       1 |
+------+---------+

The same four answers as the C program gave, which is the point: the rule is the rule, and the language only changes how it is written.

Powers and roots

SELECT POW(2, 10)   AS two_to_the_ten,
       POWER(3, 4)  AS three_to_the_four,
       SQRT(144)    AS root_of_144,
       ROUND(SQRT(2), 4) AS root_of_2,
       EXP(1)       AS e,
       ROUND(PI(), 5) AS pi;
+----------------+-------------------+-------------+-----------+-------------------+---------+
| two_to_the_ten | three_to_the_four | root_of_144 | root_of_2 | e                 | pi      |
+----------------+-------------------+-------------+-----------+-------------------+---------+
|           1024 |                81 |          12 |    1.4142 | 2.718281828459045 | 3.14159 |
+----------------+-------------------+-------------+-----------+-------------------+---------+

POW and POWER are the same function. SQRT of a negative number gives NULL rather than an error, which is worth knowing: the row survives and the value is missing.

The largest and smallest of several values

SELECT GREATEST(3, 17, 8)   AS biggest,
       LEAST(3, 17, 8)      AS smallest,
       GREATEST(-4, 0)      AS not_negative;
+---------+----------+--------------+
| biggest | smallest | not_negative |
+---------+----------+--------------+
|      17 |        3 |            0 |
+---------+----------+--------------+

GREATEST and LEAST are not MAX and MIN. These compare several values in one row; MAX and MIN from [Practical 4: Aggregate Functions, GROUP BY and HAVING] compare one value down many rows. Confusing the two pairs is a standing examination mistake, and the distinction is worth writing out:

GREATEST(a, b, c)MAX(col)
Comparesseveral columns or values, across one rowone column, across many rows
Isan ordinary functionan aggregate function
Needs a GROUP BYnoonly to group the rows
Resultone value per rowone value per group

GREATEST(x, 0) was used in [Practical 5: Date Functions] to stop a fine going negative, and that is its commonest real use.

A worked example: the fine, in whole rupees

The library charges two rupees a day, rounded up to a whole rupee, on a fourteen day loan.

SELECT l.loan_id, b.title,
       DATEDIFF(l.returned_on, DATE_ADD(l.issued_on, INTERVAL 14 DAY)) AS late,
       CEIL(GREATEST(DATEDIFF(l.returned_on,
            DATE_ADD(l.issued_on, INTERVAL 14 DAY)), 0) * 2) AS fine
FROM loan l JOIN book b ON b.book_id = l.book_id
WHERE l.returned_on IS NOT NULL
ORDER BY fine DESC, l.loan_id;
+---------+------------------+------+------+
| loan_id | title            | late | fine |
+---------+------------------+------+------+
|       2 | Database Systems |   14 |   28 |
|       1 | The C Language   |   -2 |    0 |
|       3 | MySQL Reference  |   -4 |    0 |
|       5 | The C Language   |   -4 |    0 |
|       6 | Learn SQL        |   -3 |    0 |
+---------+------------------+------+------+
munotes.in162

Practical 5: Math Functions

Three functions from three different chapters in one expression, which is what a real query looks like.

Formatting a number for a person

SELECT price,
       FORMAT(price, 2)          AS with_commas,
       CONCAT('Rs ', FORMAT(price, 2)) AS as_money
FROM book ORDER BY price DESC LIMIT 3;
+--------+-------------+-----------+
| price  | with_commas | as_money  |
+--------+-------------+-----------+
| 720.50 | 720.50      | Rs 720.50 |
| 655.25 | 655.25      | Rs 655.25 |
| 610.00 | 610.00      | Rs 610.00 |
+--------+-------------+-----------+

FORMAT(x, d) rounds to d places and puts separators in. Like DATE_FORMAT, it returns text, so it belongs at the very end of a query and never in a column.

Random numbers

RAND() gives a different value every time it is called, so its result is not printed in this book. It really runs:

SELECT book_id, title FROM book ORDER BY RAND() LIMIT 1;

ORDER BY RAND() LIMIT 1 picks a row at random, which is how a quiz question or a book of the day is chosen. On a large table it is slow, because every row gets a random number before the sort.

The functions in this practical

FunctionDoes
ROUND(x, d)round to d places
TRUNCATE(x, d)cut off at d places, towards zero
FLOOR(x)down to a whole number
CEIL(x)up to a whole number
ABS(x)size without the sign
SIGN(x)minus 1, 0 or 1
MOD(a, b), a % bremainder
a DIV bwhole part of the division
POW(a, b), POWER(a, b)a to the power b
SQRT(x)square root; NULL for a negative x
EXP(x), LOG(x), PI()the usual mathematics
GREATEST(...), LEAST(...)across one row
FORMAT(x, d)text, with separators
RAND()a random value between 0 and 1

What beginners get wrong

Thinking TRUNCATE and FLOOR are the same. They agree on positive numbers only.

Confusing GREATEST with MAX. One works across a row, the other down a column.

Storing a FORMAT result. It is text and cannot be summed.

Expecting SQRT(-1) to be an error. It is NULL.

Writing TRUNCATE(x) with no second argument. It needs both.

Rounding at every step of a calculation. Round once, at the end, or the small errors add up.

Using ORDER BY RAND() on a large table. It sorts the whole table.

Quick revision

  • ROUND(x, d) to the nearest; TRUNCATE(x, d) cuts towards zero; FLOOR down; CEIL up.
  • On a negative number all four differ: minus 7.6 gives minus 8, minus 7, minus 8 and minus 7.
  • ABS, SIGN, MOD(a, b) which is also a % b, and DIV for the whole part.
  • POW(a, b) and SQRT(x), which is NULL for a negative x.
  • GREATEST and LEAST work across one row; MAX and MIN work down many rows.
  • GREATEST(x, 0) is how a value is stopped from going negative.
  • FORMAT(x, d) makes text; keep it out of stored columns.
  • ORDER BY RAND() LIMIT 1 picks a random row and is slow on a large table.
munotes.in163

Practical 5: Math Functions

What goes in your journal

Aim, then one query holding all four rounding functions applied to the same positive number, and a second holding them applied to the same negative number. Those two rows side by side are this practical's whole content, and an examiner looking at them can see at once whether you know the difference.

Then ABS, MOD, POW and SQRT in one query, and the fine calculation as a worked example.

Test yourself

1. What do ROUND, TRUNCATE, FLOOR and CEIL give for minus 7.6? Minus 8, minus 7, minus 8 and minus 7.

2. When do TRUNCATE(x, 0) and FLOOR(x) agree? When x is positive or zero. For a negative x, TRUNCATE moves towards zero and FLOOR moves away from it.

3. What is the difference between GREATEST and MAX? GREATEST compares several values within one row and is an ordinary function. MAX compares one column across many rows and is an aggregate.

4. What does SQRT(-4) return? NULL, not an error.

5. How do you make a value that might be negative come out as zero instead? GREATEST(x, 0).

6. Why should FORMAT not be used in a stored column? Because it returns text with separators in it, which cannot be summed, compared as a number or sorted numerically.

Contents This chapter on its own page

munotes.in164

Chapter Forty-Four

Practical 6: The Inner Join

Syllabus topic Module 2, Practical 6: "Join Queries: Inner Join"

Aim

To write inner join queries.

Why a join is needed at all

The book table holds author_id and not the author's name, because [Practical 1: The ER Diagram: Entities, Attributes and Keys] made AUTHOR an entity set of its own. That was the right decision, and it leaves one question: how do you print a book with its author's name, when the two facts are in two tables?

A join is the answer. It builds rows of a result from rows of two tables, matched on a column they have in common.

First, the Cartesian product

Name two tables with no condition at all and you get every row of the first paired with every row of the second.

SELECT COUNT(*) AS pairs FROM book, author;
+-------+
| pairs |
+-------+
|    40 |
+-------+

Eight books and five authors make forty pairs. That is the Cartesian product, also called a cross join, and it is almost never what anybody wants: most of those forty pairs put a book beside an author who did not write it.

SELECT b.title, a.author_name
FROM book b, author a
ORDER BY b.book_id, a.author_id LIMIT 6;
+------------------+------------------+
| title            | author_name      |
+------------------+------------------+
| The C Language   | Dennis Ritchie   |
| The C Language   | Ramez Elmasri    |
| The C Language   | Vikram Vaswani   |
| The C Language   | Behrouz Forouzan |
| The C Language   | Ashwin Pajankar  |
| Database Systems | Dennis Ritchie   |
+------------------+------------------+

The first book appears five times, once beside each author. Only one of those five rows is true.

A join is a Cartesian product with a condition that keeps the true pairs. Once you have seen the forty rows, the condition is obvious: keep the pairs where the book's author_id is the author's author_id.

The inner join, written the modern way

SELECT b.title, a.author_name
FROM book b
INNER JOIN author a ON b.author_id = a.author_id
ORDER BY b.book_id;
+-------------------+------------------+
| title             | author_name      |
+-------------------+------------------+
| The C Language    | Dennis Ritchie   |
| Database Systems  | Ramez Elmasri    |
| MySQL Reference   | Vikram Vaswani   |
| Data Structures   | Behrouz Forouzan |
| Computer Networks | Behrouz Forouzan |
| Learn SQL         | Vikram Vaswani   |
| Operating Systems | Behrouz Forouzan |
| Discrete Maths    | Ramez Elmasri    |
+-------------------+------------------+

Eight rows, one per book, each beside its own author. Thirty-two of the forty pairs were thrown away by the ON.

The word INNER may be left out: JOIN alone means an inner join. Write it while you are learning, because it makes the contrast with the outer join of the next chapter visible.

munotes.in165

Practical 6: The Inner Join

The same thing written the older way

SELECT b.title, a.author_name
FROM book b, author a
WHERE b.author_id = a.author_id
ORDER BY b.book_id;
+-------------------+------------------+
| title             | author_name      |
+-------------------+------------------+
| The C Language    | Dennis Ritchie   |
| Database Systems  | Ramez Elmasri    |
| MySQL Reference   | Vikram Vaswani   |
| Data Structures   | Behrouz Forouzan |
| Computer Networks | Behrouz Forouzan |
| Learn SQL         | Vikram Vaswani   |
| Operating Systems | Behrouz Forouzan |
| Discrete Maths    | Ramez Elmasri    |
+-------------------+------------------+

The same eight rows. This is the form the older textbooks use, and it is still accepted everywhere.

JOIN ... ONcomma and WHERE
Introduced bySQL-92the original SQL
Join condition sitsin the ONmixed into the WHERE
Filtering condition sitsin the WHEREmixed into the WHERE
Forgetting the condition givesa syntax errora silent Cartesian product
Can it do an outer joinyesnot portably

The fourth row is why the modern form is better. Write FROM book b, author a and forget the WHERE, and you get forty rows and no complaint; write FROM book b JOIN author a and forget the ON, and the server tells you at once.

Use JOIN ... ON. Recognise the comma form, because your examiner may write it.

Aliases

book b and author a give each table a short name for the rest of the query. With AS or without, both work; book AS b is the same as book b.

They become compulsory as soon as both tables have a column of the same name. Both of these tables have author_id, so:

SELECT author_id FROM book b JOIN author a ON b.author_id = a.author_id;
ERROR 1052 (23000): Column 'author_id' in field list is ambiguous

Error 1052, ambiguous. The server cannot tell which author_id was meant, and neither could a reader. b.author_id or a.author_id settles it.

Qualify every column with its table's alias in a join, even where it is not ambiguous. It costs two characters and it makes the query readable.

USING, when the columns have the same name

SELECT b.title, a.author_name
FROM book b JOIN author a USING (author_id)
ORDER BY b.book_id LIMIT 3;
+------------------+----------------+
| title            | author_name    |
+------------------+----------------+
| The C Language   | Dennis Ritchie |
| Database Systems | Ramez Elmasri  |
| MySQL Reference  | Vikram Vaswani |
+------------------+----------------+

USING (col) is shorthand for ON b.col = a.col, available only when the column has the same name in both tables. It also merges the two columns into one in the result, so SELECT author_id after a USING is not ambiguous.

NATURAL JOIN goes further and joins on every column the two tables share, with no condition written at all. It is best avoided: it changes meaning silently the day somebody adds a column with a name that happens to match.

munotes.in166

Practical 6: The Inner Join

Filtering as well as joining

SELECT b.title, a.author_name, b.price
FROM book b JOIN author a ON b.author_id = a.author_id
WHERE a.country = 'India'
ORDER BY b.price DESC;
+-----------------+----------------+--------+
| title           | author_name    | price  |
+-----------------+----------------+--------+
| MySQL Reference | Vikram Vaswani | 499.00 |
| Learn SQL       | Vikram Vaswani | 280.00 |
+-----------------+----------------+--------+

The ON says how the tables go together; the WHERE says which of the joined rows to keep. Keeping the two clauses for their two jobs is what makes a long query readable, and on an outer join it changes the answer, which [Practical 6: The Outer Join] shows.

Joining three tables

SELECT m.member_name, b.title, l.issued_on
FROM loan l
JOIN member m ON m.member_id = l.member_id
JOIN book   b ON b.book_id   = l.book_id
ORDER BY l.loan_id;
+---------------+-------------------+------------+
| member_name   | title             | issued_on  |
+---------------+-------------------+------------+
| Asha Kulkarni | The C Language    | 2025-06-02 |
| Asha Kulkarni | Database Systems  | 2025-06-02 |
| Ravi Deshmukh | MySQL Reference   | 2025-06-10 |
| Ravi Deshmukh | Data Structures   | 2025-07-01 |
| Meena Iyer    | The C Language    | 2025-07-05 |
| Imran Shaikh  | Learn SQL         | 2025-07-08 |
| Neha Patil    | Computer Networks | 2025-07-21 |
+---------------+-------------------+------------+

Each JOIN brings in one more table with its own ON. Start from the table in the middle, the one holding the foreign keys, and join outward; here that is loan, which points at both member and book.

This is the query the many to many relationship of [Practical 1: Relationships and Cardinality] was drawn for. loan is what that M:N relationship became when it was turned into tables, and a join through it is how the relationship is read back.

Joining with an aggregate

SELECT a.author_name, COUNT(*) AS books, SUM(b.price) AS worth
FROM book b JOIN author a ON b.author_id = a.author_id
GROUP BY a.author_id, a.author_name
ORDER BY books DESC, a.author_name;
+------------------+-------+---------+
| author_name      | books | worth   |
+------------------+-------+---------+
| Behrouz Forouzan |     3 | 1806.00 |
| Ramez Elmasri    |     2 | 1150.50 |
| Vikram Vaswani   |     2 |  779.00 |
| Dennis Ritchie   |     1 |  395.00 |
+------------------+-------+---------+

Grouping by a.author_id, a.author_name rather than by the name alone is deliberate: two authors could share a name, and the id is what actually identifies them.

Ashwin Pajankar is not in that answer at all, because he has written no book and an inner join keeps only the rows that matched. Showing him with a count of 0 needs an outer join, and that is the next chapter.

munotes.in167

Practical 6: The Inner Join

A self join: two rows of the same table

SELECT b1.title AS book_one, b2.title AS book_two, b1.subject
FROM book b1
JOIN book b2 ON b1.subject = b2.subject AND b1.book_id < b2.book_id
ORDER BY b1.subject, b1.book_id;
+------------------+-----------------+-------------+
| book_one         | book_two        | subject     |
+------------------+-----------------+-------------+
| Database Systems | MySQL Reference | Databases   |
| Database Systems | Learn SQL       | Databases   |
| MySQL Reference  | Learn SQL       | Databases   |
| The C Language   | Data Structures | Programming |
+------------------+-----------------+-------------+

The same table joined to itself, under two different aliases, which is the only way the server can tell the two copies apart. The b1.book_id < b2.book_id does two jobs at once: it stops a book being paired with itself, and it keeps only one of each pair rather than both orders.

What beginners get wrong

Forgetting the join condition in the comma form. A silent Cartesian product, and eight books become forty rows.

Not qualifying an ambiguous column. Error 1052.

Joining on the wrong pair of columns. ON b.book_id = a.author_id runs and gives nonsense, because both are integers and the server has no idea that one is a book and the other an author.

Expecting an inner join to show the unmatched rows. It never does; that is what it means.

Writing NATURAL JOIN to save typing. It joins on every shared column name, which changes the day a column is added.

Counting rows of a join as things. A join of eight books and seven loans has as many rows as there are matches, not as there are books.

Quick revision

  • Two tables with no condition give the Cartesian product: every pair.
  • FROM a JOIN b ON a.col = b.col keeps the pairs that match. INNER is the default.
  • The older form is FROM a, b WHERE a.col = b.col, and forgetting the WHERE gives a silent Cartesian product.
  • Aliases shorten the query and are compulsory when a column name appears in both tables. Error 1052 otherwise.
  • USING (col) when the column has the same name in both; NATURAL JOIN is best avoided.
  • ON says how the tables join; WHERE says which joined rows to keep.
  • Three tables take two JOIN ... ON clauses; start from the table holding the foreign keys.
  • A self join joins a table to itself under two aliases.
  • An inner join drops every row that has no match on the other side.

What goes in your journal

Aim, and then the four queries in this order: the Cartesian product with its row count, the same pair of tables with an ON, the same again in the comma form, and a three table join. Then one join with a GROUP BY.

munotes.in168

Practical 6: The Inner Join

The first two are the ones to keep. Forty rows and then eight, from the same two tables, is the clearest demonstration of what a join does that anybody has found, and it is worth the half page.

Test yourself

1. What is a Cartesian product? Every row of one table paired with every row of the other, with no condition. Eight books and five authors give forty rows.

2. What does an inner join do with a row that has no match? It leaves it out of the result entirely.

3. Why is JOIN ... ON safer than the comma form? Because forgetting the ON is a syntax error, where forgetting the WHERE in the comma form silently produces the Cartesian product.

4. When must a column be qualified with its table alias? Whenever the same column name exists in more than one of the joined tables. Otherwise the server answers with error 1052, ambiguous column.

5. How many JOIN ... ON clauses does a query over three tables need? Two: one for each table brought in after the first.

6. What is a self join, and why does it need aliases? A join of a table to itself. The two copies need different aliases so that the query can say which of them each column belongs to.

7. Why does b1.book_id < b2.book_id appear in the self join? To stop a row pairing with itself, and to keep only one of each pair instead of both orders.

Contents This chapter on its own page

munotes.in169

Chapter Forty-Five

Practical 6: The Outer Join

Syllabus topic Module 2, Practical 6: "Join Queries: Outer Join"

Aim

To write outer join queries.

What an inner join leaves out

An inner join keeps only the rows that matched. Three rows of this library never match anything, and each of them is a real fact somebody might want to know:

  • Ashwin Pajankar has written no book.
  • Discrete Maths has never been borrowed.
  • Vikram Rao has never borrowed anything.

An outer join keeps the unmatched rows of one side, filling the other side's columns with NULL.

LEFT JOIN

SELECT a.author_name, b.title
FROM author a
LEFT JOIN book b ON b.author_id = a.author_id
ORDER BY a.author_id, b.book_id;
+------------------+-------------------+
| author_name      | title             |
+------------------+-------------------+
| Dennis Ritchie   | The C Language    |
| Ramez Elmasri    | Database Systems  |
| Ramez Elmasri    | Discrete Maths    |
| Vikram Vaswani   | MySQL Reference   |
| Vikram Vaswani   | Learn SQL         |
| Behrouz Forouzan | Data Structures   |
| Behrouz Forouzan | Computer Networks |
| Behrouz Forouzan | Operating Systems |
| Ashwin Pajankar  | NULL              |
+------------------+-------------------+

Nine rows, where the inner join of [Practical 6: The Inner Join] gave eight. The extra row is Ashwin Pajankar, with NULL where a title would be.

A LEFT JOIN keeps every row of the LEFT table, the one named in the FROM, whether or not it matched. Where there was no match, every column taken from the right table is NULL.

LEFT OUTER JOIN is the full spelling and means exactly the same; the word OUTER is optional and carries no meaning of its own.

Finding the rows that did not match

This is the pattern worth memorising, because it answers a whole family of questions.

SELECT a.author_name
FROM author a
LEFT JOIN book b ON b.author_id = a.author_id
WHERE b.book_id IS NULL;
+-----------------+
| author_name     |
+-----------------+
| Ashwin Pajankar |
+-----------------+

Read it in two steps. The LEFT JOIN keeps every author, matched or not. The WHERE b.book_id IS NULL then keeps only the ones where nothing matched, because a matched row would have a real book_id there.

The column tested must be one that cannot be NULL in a matched row, which is why b.book_id is used and not b.price. A book with no price recorded would slip through a test on b.price.

The same shape answers the other two questions:

SELECT b.title AS never_borrowed
FROM book b
LEFT JOIN loan l ON l.book_id = b.book_id
WHERE l.loan_id IS NULL;
+-------------------+
| never_borrowed    |
+-------------------+
| Operating Systems |
| Discrete Maths    |
+-------------------+
SELECT m.member_name AS never_borrowed_anything
FROM member m
LEFT JOIN loan l ON l.member_id = m.member_id
WHERE l.loan_id IS NULL;
+-------------------------+
| never_borrowed_anything |
+-------------------------+
| Vikram Rao              |
+-------------------------+
munotes.in170

Practical 6: The Outer Join

Counting, including the zeros

[Practical 6: The Inner Join] ended with a count per author that left out the author with no books. Now it can be done properly.

SELECT a.author_name, COUNT(b.book_id) AS books
FROM author a
LEFT JOIN book b ON b.author_id = a.author_id
GROUP BY a.author_id, a.author_name
ORDER BY books DESC, a.author_name;
+------------------+-------+
| author_name      | books |
+------------------+-------+
| Behrouz Forouzan |     3 |
| Ramez Elmasri    |     2 |
| Vikram Vaswani   |     2 |
| Dennis Ritchie   |     1 |
| Ashwin Pajankar  |     0 |
+------------------+-------+

Ashwin Pajankar is there with 0.

COUNT(b.book_id) and not COUNT(*). Here is why:

SELECT a.author_name, COUNT(*) AS wrong_count, COUNT(b.book_id) AS right_count
FROM author a
LEFT JOIN book b ON b.author_id = a.author_id
GROUP BY a.author_id, a.author_name
ORDER BY a.author_name LIMIT 2;
+------------------+-------------+-------------+
| author_name      | wrong_count | right_count |
+------------------+-------------+-------------+
| Ashwin Pajankar  |           1 |           0 |
| Behrouz Forouzan |           3 |           3 |
+------------------+-------------+-------------+

Ashwin Pajankar's group has one row in it, the row the LEFT JOIN manufactured with NULLs in it, so COUNT(*) counts that row and says 1. COUNT(b.book_id) skips it, because that column is NULL, and says 0.

That is the NULL rule of [Practical 4: Aggregate Functions, GROUP BY and HAVING] doing exactly what it promised, and it is the commonest wrong answer in this practical.

ON against WHERE, and why it matters here

On an inner join the two clauses are interchangeable. On an outer join they are not, and this is the hardest idea in the chapter.

SELECT a.author_name, b.title
FROM author a
LEFT JOIN book b ON b.author_id = a.author_id AND b.subject = 'Databases'
ORDER BY a.author_id, b.book_id;
+------------------+------------------+
| author_name      | title            |
+------------------+------------------+
| Dennis Ritchie   | NULL             |
| Ramez Elmasri    | Database Systems |
| Vikram Vaswani   | MySQL Reference  |
| Vikram Vaswani   | Learn SQL        |
| Behrouz Forouzan | NULL             |
| Ashwin Pajankar  | NULL             |
+------------------+------------------+
SELECT a.author_name, b.title
FROM author a
LEFT JOIN book b ON b.author_id = a.author_id
WHERE b.subject = 'Databases'
ORDER BY a.author_id, b.book_id;
+----------------+------------------+
| author_name    | title            |
+----------------+------------------+
| Ramez Elmasri  | Database Systems |
| Vikram Vaswani | MySQL Reference  |
| Vikram Vaswani | Learn SQL        |
+----------------+------------------+

Same tables, same condition, different answers.

In the ON, the condition decides what counts as a match. Authors with no Databases book are still kept, with NULL beside them, because the LEFT JOIN keeps every author whatever happens.

In the WHERE, the condition is applied after the join. The manufactured NULL rows have a NULL subject, NULL = 'Databases' is not true, and they are thrown away. The left join has been turned back into an inner join.

munotes.in171

Practical 6: The Outer Join

The rule in one line: a condition on the right hand table of a LEFT JOIN belongs in the ON, not in the WHERE unless you actually meant to drop the unmatched rows.

RIGHT JOIN

SELECT b.title, a.author_name
FROM book b
RIGHT JOIN author a ON b.author_id = a.author_id
ORDER BY a.author_id, b.book_id;
+-------------------+------------------+
| title             | author_name      |
+-------------------+------------------+
| The C Language    | Dennis Ritchie   |
| Database Systems  | Ramez Elmasri    |
| Discrete Maths    | Ramez Elmasri    |
| MySQL Reference   | Vikram Vaswani   |
| Learn SQL         | Vikram Vaswani   |
| Data Structures   | Behrouz Forouzan |
| Computer Networks | Behrouz Forouzan |
| Operating Systems | Behrouz Forouzan |
| NULL              | Ashwin Pajankar  |
+-------------------+------------------+

A RIGHT JOIN keeps every row of the right table, the one named after the JOIN. The nine rows are the same nine as the LEFT JOIN at the top of this chapter, because the two tables have simply changed places.

Every RIGHT JOIN can be written as a LEFT JOIN by swapping the tables, and almost all real code does, because a query reads better when the table you care about is the first one named. Knowing that A RIGHT JOIN B equals B LEFT JOIN A is a fair viva question.

The FULL OUTER JOIN MySQL does not have

A full outer join keeps the unmatched rows of both sides. MySQL does not support it, and this is a real finding rather than an omission in this book:

SELECT a.author_name, b.title
FROM author a FULL OUTER JOIN book b ON b.author_id = a.author_id;
ERROR 1064 (42000): You have an error in your SQL syntax; check the manual that corresponds to your MySQL server version for the right syntax to use near 'FULL OUTER JOIN book b ON b.author_id = a.author_id' at line 2

A syntax error. PostgreSQL and Oracle have it; MySQL does not.

The standard answer is a UNION of the two one-sided joins:

SELECT a.author_name, b.title
FROM author a LEFT JOIN book b ON b.author_id = a.author_id
UNION
SELECT a.author_name, b.title
FROM author a RIGHT JOIN book b ON b.author_id = a.author_id
ORDER BY author_name, title;
+------------------+-------------------+
| author_name      | title             |
+------------------+-------------------+
| Ashwin Pajankar  | NULL              |
| Behrouz Forouzan | Computer Networks |
| Behrouz Forouzan | Data Structures   |
| Behrouz Forouzan | Operating Systems |
| Dennis Ritchie   | The C Language    |
| Ramez Elmasri    | Database Systems  |
| Ramez Elmasri    | Discrete Maths    |
| Vikram Vaswani   | Learn SQL         |
| Vikram Vaswani   | MySQL Reference   |
+------------------+-------------------+

UNION stacks two results and removes the duplicates; UNION ALL keeps them and is faster where you know there are none. The two queries must have the same number of columns, in a compatible order.

munotes.in172

Practical 6: The Outer Join

Here every book has an author, so the right hand side adds nothing new and the answer is the same nine rows. On a pair of tables with orphans on both sides it would be longer than either half, which is the point of it.

The four joins compared

JoinKeeps
INNER JOINonly the rows that matched
LEFT JOINevery row of the left table, matched or not
RIGHT JOINevery row of the right table, matched or not
FULL OUTER JOINevery row of both; not in MySQL, use a UNION
CROSS JOINevery pair, with no condition at all

What beginners get wrong

Putting a condition on the right table in the WHERE. It silently turns the outer join back into an inner join.

COUNT(*) on an outer join. It counts the manufactured NULL row as 1. Count a column from the right table instead.

Testing a nullable column for the unmatched rows. Test the right table's primary key, which cannot be NULL in a matched row.

Expecting FULL OUTER JOIN to work in MySQL. It does not.

Reading a NULL in the result as missing data. In an outer join it means "no matching row", which is different: the data is not missing, the match is.

Swapping LEFT and RIGHT without swapping the tables. The answer changes completely.

Quick revision

  • LEFT JOIN keeps every row of the first table; unmatched rows get NULLs on the right.
  • RIGHT JOIN keeps every row of the second. A RIGHT JOIN B is B LEFT JOIN A.
  • OUTER is optional and adds nothing.
  • To find unmatched rows: LEFT JOIN then WHERE right.primary_key IS NULL.
  • Count with COUNT(right.col), never COUNT(*), or the empty groups come out as 1.
  • A condition on the right table belongs in the ON; in the WHERE it undoes the outer join.
  • MySQL has no FULL OUTER JOIN; use LEFT JOIN UNION RIGHT JOIN.
  • UNION removes duplicates, UNION ALL keeps them.

What goes in your journal

Aim, then the inner join and the left join over the same two tables one under the other, so the extra row and its NULLs are visible. Then the unmatched-rows pattern, then the count with COUNT(col) beside the same count with COUNT(*).

That last pair, 0 against 1 for the author with no books, is the single most useful thing in this chapter to have written down before a viva.

Test yourself

1. What does a LEFT JOIN keep that an INNER JOIN does not? Every row of the left table that had no match, with NULLs in the columns taken from the right table.

munotes.in173

Practical 6: The Outer Join

2. How do you list the rows of one table that have no match in another? LEFT JOIN the second table and then WHERE second.primary_key IS NULL.

3. Why must that test use the primary key rather than any column? Because a matched row could itself contain a NULL in an ordinary column, and would then be reported as unmatched.

4. Why does COUNT() give 1 for a group with no matching rows? Because the LEFT JOIN produced one row for that group, filled with NULLs, and COUNT() counts rows regardless of their contents.

5. What happens if a condition on the right table is put in the WHERE instead of the ON? The manufactured NULL rows fail the condition and are removed, so the outer join behaves as an inner join.

6. Does MySQL support FULL OUTER JOIN? No. The same result is obtained with a UNION of a LEFT JOIN and a RIGHT JOIN.

7. Rewrite A RIGHT JOIN B ON ... as a left join. B LEFT JOIN A ON ..., with the same condition.

Contents This chapter on its own page

munotes.in174

Chapter Forty-Six

Practical 7: Subqueries with IN

Syllabus topic Module 2, Practical 7: "Subqueries With IN clause"

Aim

To write subqueries using the IN clause.

What a subquery is

A subquery is a SELECT written inside another statement, in brackets. The inner query runs, its result is handed to the outer query, and the outer query finishes.

It is how a question with two parts is asked in one statement. "Which books have been borrowed?" has two parts: find the borrowed book ids, then find those books.

NameWhere it appearsReturns
Scalar subqueryanywhere a single value can goone row, one column
Row subquerycompared with a row of valuesone row, several columns
Table subqueryafter IN, EXISTS, or in the FROMmany rows
Correlated subqueryany of the above, referring to the outer queryre-runs for each outer row

The first three are in this chapter. Correlated subqueries are [Practical 7: Subqueries with EXISTS], where they belong.

IN with a list, and IN with a query

SELECT title, subject FROM book
WHERE subject IN ('Databases', 'Networks')
ORDER BY book_id;
+-------------------+-----------+
| title             | subject   |
+-------------------+-----------+
| Database Systems  | Databases |
| MySQL Reference   | Databases |
| Computer Networks | Networks  |
| Learn SQL         | Databases |
+-------------------+-----------+

That list was typed. Now let the database produce it:

SELECT title FROM book
WHERE book_id IN (SELECT book_id FROM loan)
ORDER BY book_id;
+-------------------+
| title             |
+-------------------+
| The C Language    |
| Database Systems  |
| MySQL Reference   |
| Data Structures   |
| Computer Networks |
| Learn SQL         |
+-------------------+

The inner query gives the book ids that appear in loan; the outer query keeps the books whose id is one of them. Six books have been borrowed and two have not.

The inner query must return exactly one column. More than one and the server refuses:

SELECT title FROM book WHERE book_id IN (SELECT book_id, member_id FROM loan);
ERROR 1241 (21000): Operand should contain 1 column(s)

Error 1241. The comparison is between one value and a list of single values, so a two column list is meaningless.

Reading a subquery from the inside out

SELECT member_name, city FROM member
WHERE member_id IN (
    SELECT member_id FROM loan
    WHERE issued_on >= '2025-07-01'
)
ORDER BY member_id;
+---------------+--------+
| member_name   | city   |
+---------------+--------+
| Ravi Deshmukh | Thane  |
| Meena Iyer    | Mumbai |
| Imran Shaikh  | Kalyan |
| Neha Patil    | NULL   |
+---------------+--------+

Run the inside on its own first, always, when a subquery does not give the answer you expected:

SELECT member_id FROM loan WHERE issued_on >= '2025-07-01';
+-----------+
| member_id |
+-----------+
|         2 |
|         3 |
|         4 |
|         5 |
+-----------+

Those are the ids the outer query was handed. Three rows, one of them repeated, which does not matter to IN: it asks whether the value is in the list, and a list with a repeat in it answers the same way.

munotes.in175

Practical 7: Subqueries with IN

Debugging a subquery means running the inner query by itself. It is the first thing to do and students rarely think of it.

NOT IN

SELECT title FROM book
WHERE book_id NOT IN (SELECT book_id FROM loan)
ORDER BY book_id;
+-------------------+
| title             |
+-------------------+
| Operating Systems |
| Discrete Maths    |
+-------------------+

The two books nobody has borrowed, which is the same answer the LEFT JOIN pattern of [Practical 6: The Outer Join] gave. Both are correct and both are used; which is faster depends on the tables and the indexes.

The NULL trap, which is what this practical is really for

SELECT title FROM book
WHERE author_id NOT IN (SELECT author_id FROM book WHERE subject = 'Databases');
+-------------------+
| title             |
+-------------------+
| The C Language    |
| Data Structures   |
| Computer Networks |
| Operating Systems |
+-------------------+

Now the same question asked where the inner list contains a NULL:

INSERT INTO author (author_id, author_name, country) VALUES (6, 'Unknown', NULL);
INSERT INTO book VALUES (9, 'Mystery Book', 'Programming', 200.00, '2025-01-01', NULL);
SELECT author_id FROM book WHERE subject = 'Programming';
+-----------+
| author_id |
+-----------+
|         1 |
|         4 |
|      NULL |
+-----------+

There is a NULL in that list now. Ask which books are not by one of those authors:

SELECT title FROM book
WHERE author_id NOT IN (SELECT author_id FROM book WHERE subject = 'Programming');

Nothing at all, and there certainly are books by other authors.

Here is why, and it follows from the NULL rule of [Practical 4: Simple Queries]. x NOT IN (1, 4, NULL) means x <> 1 AND x <> 4 AND x <> NULL. That last comparison is NULL, never true, so the whole AND can never be true. It can be false, when x really is 1 or 4, and otherwise it is NULL, and a WHERE keeps only what is true. So no row survives.

Three fixes, and the first is the one to use:

SELECT title FROM book
WHERE author_id NOT IN (
    SELECT author_id FROM book
    WHERE subject = 'Programming' AND author_id IS NOT NULL
);
+------------------+
| title            |
+------------------+
| Database Systems |
| MySQL Reference  |
| Learn SQL        |
| Discrete Maths   |
+------------------+

The second is NOT EXISTS, which does not have this problem at all and is [Practical 7: Subqueries with EXISTS]. The third is the LEFT JOIN pattern.

IN with a NULL in the list is harmless; NOT IN with a NULL in the list returns nothing. That sentence is worth writing in your journal exactly as it stands.

munotes.in176

Practical 7: Subqueries with IN

DELETE FROM book WHERE book_id = 9;
DELETE FROM author WHERE author_id = 6;

A scalar subquery: one value

SELECT title, price FROM book
WHERE price > (SELECT AVG(price) FROM book)
ORDER BY price DESC;
+-------------------+--------+
| title             | price  |
+-------------------+--------+
| Database Systems  | 720.50 |
| Operating Systems | 655.25 |
| Data Structures   | 610.00 |
| Computer Networks | 540.75 |
+-------------------+--------+

The inner query returns one number, so it can be used anywhere a number can go. This is the query MU's examiners set most often on this practical: the rows above the average, which cannot be written without a subquery, because an aggregate is not allowed in a WHERE.

If a scalar subquery returns more than one row, the server refuses:

SELECT title FROM book WHERE price > (SELECT price FROM book);
ERROR 1242 (21000): Subquery returns more than 1 row

Error 1242. Use IN, or ANY, or add a LIMIT 1 if one row really is what you meant.

ANY and ALL

SELECT title, price FROM book
WHERE price > ALL (SELECT price FROM book WHERE subject = 'Databases')
ORDER BY price DESC;
SELECT title, price FROM book
WHERE price > ANY (SELECT price FROM book WHERE subject = 'Databases')
ORDER BY price DESC;
+-------------------+--------+
| title             | price  |
+-------------------+--------+
| Database Systems  | 720.50 |
| Operating Systems | 655.25 |
| Data Structures   | 610.00 |
| Computer Networks | 540.75 |
| MySQL Reference   | 499.00 |
| Discrete Maths    | 430.00 |
| The C Language    | 395.00 |
+-------------------+--------+

> ALL means greater than every one of them, so greater than the largest. > ANY means greater than at least one of them, so greater than the smallest. SOME is another word for ANY.

And IN is exactly = ANY. That equivalence is a fair viva question and it makes both easier to remember.

A subquery in the FROM: a derived table

SELECT subject, books
FROM (SELECT subject, COUNT(*) AS books FROM book GROUP BY subject) AS counts
WHERE books > 1;
+-------------+-------+
| subject     | books |
+-------------+-------+
| Programming |     2 |
| Databases   |     3 |
+-------------+-------+

The inner query produces a result, and the outer query treats it as a table. A derived table must be given an alias, AS counts here, or the server refuses; it has no name of its own.

This is the other way to filter on an aggregate, and it is worth comparing with the HAVING of [Practical 4: Aggregate Functions, GROUP BY and HAVING], which does the same job more directly.

munotes.in177

Practical 7: Subqueries with IN

Where a subquery may appear

PlaceExample
In the WHEREWHERE price > (SELECT AVG(price) ...)
In the SELECT lista count of each book's loans, beside the book
In the FROMFROM (SELECT ...) AS alias
In the HAVINGHAVING COUNT(*) > (SELECT ...)
In INSERT, UPDATE, DELETEDELETE FROM book WHERE book_id NOT IN (...)

What beginners get wrong

NOT IN against a list that may contain NULL. The answer is silently empty. Exclude the NULLs, or use NOT EXISTS.

A subquery with two columns after IN. Error 1241.

A scalar subquery that returns several rows. Error 1242.

Forgetting the alias on a derived table. The server refuses it.

Not running the inner query on its own when the answer looks wrong. It takes ten seconds and it finds most of these.

Using a subquery where a join is clearer. Both are correct; on a large table the join is often faster, and either is accepted in an examination.

Quick revision

  • A subquery is a SELECT in brackets inside another statement.
  • IN (subquery) needs the inner query to return exactly one column.
  • NOT IN with a NULL anywhere in the list returns no rows at all.
  • The fix is AND col IS NOT NULL in the subquery, or NOT EXISTS.
  • A scalar subquery returns one row and one column and goes wherever a value goes.
  • WHERE price > (SELECT AVG(price) FROM book) is how rows above the average are found.
  • > ALL is greater than the largest; > ANY is greater than the smallest; IN is = ANY.
  • A subquery in the FROM is a derived table and must be given an alias.
  • Debug a subquery by running the inner query on its own.

What goes in your journal

Aim, then: IN with a typed list, IN with a subquery, the inner query run on its own so that the list it produces is visible, NOT IN, and the above-the-average scalar subquery.

Then the NULL trap, in three steps: the list with a NULL in it, the NOT IN that returns nothing, and the repaired version. Those three results together are the practical's real content, and an examiner who sees them knows you have understood rather than copied.

Test yourself

1. How many columns may the subquery after IN return? Exactly one.

2. Why does NOT IN return nothing when the list contains a NULL? Because x NOT IN (a, b, NULL) expands to x <> a AND x <> b AND x <> NULL, and the last comparison is NULL rather than true, so the whole condition can never be true.

munotes.in178

Practical 7: Subqueries with IN

3. Give two ways to repair such a query. Add AND col IS NOT NULL inside the subquery, or rewrite it with NOT EXISTS.

4. What is a scalar subquery? One that returns a single row with a single column, so it can be used anywhere a single value can be used.

5. How do you find the rows whose price is above the average? WHERE price > (SELECT AVG(price) FROM book). It must be a subquery because an aggregate cannot appear in a WHERE.

6. What is the difference between > ANY and > ALL? > ANY is true if the value beats at least one of the list, so it compares with the smallest. > ALL is true only if it beats every one, so it compares with the largest.

7. What must a subquery in the FROM clause be given? An alias, because a derived table has no name of its own.

Contents This chapter on its own page

munotes.in179

Chapter Forty-Seven

Practical 7: Subqueries with EXISTS

Syllabus topic Module 2, Practical 7: "Subqueries With EXISTS clause"

Aim

To write subqueries using the EXISTS clause.

What EXISTS asks

EXISTS (subquery) is true when the subquery returns at least one row, and false when it returns none.

That is the whole definition, and two things follow from it.

It does not care what the rows contain. Only whether there are any. SELECT 1, SELECT * and SELECT member_name inside an EXISTS behave identically, which is why the convention is to write SELECT 1 and stop anybody wondering.

It stops as soon as it finds one. The server does not have to build the whole inner result.

Correlation: the inner query refers to the outer one

An EXISTS is almost always correlated, meaning the inner query mentions a column from the outer query.

SELECT b.title
FROM book b
WHERE EXISTS (SELECT 1 FROM loan l WHERE l.book_id = b.book_id)
ORDER BY b.book_id;
+-------------------+
| title             |
+-------------------+
| The C Language    |
| Database Systems  |
| MySQL Reference   |
| Data Structures   |
| Computer Networks |
| Learn SQL         |
+-------------------+

The six books that have been borrowed, which is the same answer IN gave in [Practical 7: Subqueries with IN].

b.book_id inside the inner query is the outer query's book. So the inner query cannot be run on its own; it has no meaning until a particular book is chosen. Conceptually the server works through the books one at a time, and for each one asks "is there a loan of this book?".

Ordinary subqueryCorrelated subquery
Mentions the outer querynoyes
Can be run on its ownyesno
Conceptually runsonceonce per outer row
Usually written withIN, =, ANYEXISTS

"Once per outer row" is the meaning, not a promise about the work: MySQL's optimiser often rewrites a correlated subquery into a join. Say "conceptually" in a viva and you are exactly right.

NOT EXISTS

SELECT b.title
FROM book b
WHERE NOT EXISTS (SELECT 1 FROM loan l WHERE l.book_id = b.book_id)
ORDER BY b.book_id;
+-------------------+
| title             |
+-------------------+
| Operating Systems |
| Discrete Maths    |
+-------------------+

The two books nobody has borrowed. Three ways of asking that question now exist in this book:

Written asChapter
LEFT JOIN loan ..., then WHERE l.loan_id IS NULL[Practical 6: The Outer Join]
WHERE book_id NOT IN (SELECT ...)[Practical 7: Subqueries with IN]
WHERE NOT EXISTS (SELECT 1 ...)here

All three are correct and all three are accepted. The differences are worth one sentence each: the outer join is usually the fastest, NOT IN is the easiest to read and the one that breaks on NULLs, and NOT EXISTS is the one that is always right.

munotes.in180

Practical 7: Subqueries with EXISTS

Why NOT EXISTS is always right

[Practical 7: Subqueries with IN] showed NOT IN returning nothing at all because a NULL had got into the list. The same data, asked with NOT EXISTS:

INSERT INTO author (author_id, author_name, country) VALUES (6, 'Unknown', NULL);
INSERT INTO book VALUES (9, 'Mystery Book', 'Programming', 200.00, '2025-01-01', NULL);
SELECT b1.title
FROM book b1
WHERE b1.author_id NOT IN (SELECT b2.author_id FROM book b2
                           WHERE b2.subject = 'Programming');

Nothing, as before. Now with NOT EXISTS:

SELECT b1.title
FROM book b1
WHERE NOT EXISTS (SELECT 1 FROM book b2
                  WHERE b2.subject = 'Programming'
                    AND b2.author_id = b1.author_id)
ORDER BY b1.book_id;
+------------------+
| title            |
+------------------+
| Database Systems |
| MySQL Reference  |
| Learn SQL        |
| Discrete Maths   |
| Mystery Book     |
+------------------+

The real answer.

The reason is that EXISTS never compares anything with the list. It asks whether a row is there, and a row either is or is not; there is no third possibility for a NULL to occupy. NOT IN has to compare, and a comparison with NULL is neither true nor false.

When a subquery's column can be NULL, use NOT EXISTS. That is the practical's main finding.

DELETE FROM book WHERE book_id = 9;
DELETE FROM author WHERE author_id = 6;

EXISTS against IN, side by side

IN (subquery)EXISTS (subquery)
Asksis this value in that listdoes that query return any row
Inner query returnsone columnanything; SELECT 1 by convention
Correlatedusually notalmost always
With a NULL in the inner resultIN is fine, NOT IN returns nothingunaffected
Reads more naturally whenthe list is short and independentthe condition ties the two tables together

A question IN cannot ask

IN compares one value. When the match needs two columns at once, EXISTS is the natural way.

SELECT m.member_name, b.title
FROM member m JOIN book b ON 1 = 1
WHERE EXISTS (SELECT 1 FROM loan l
              WHERE l.member_id = m.member_id
                AND l.book_id   = b.book_id
                AND l.returned_on IS NULL)
ORDER BY m.member_id;
+---------------+-------------------+
| member_name   | title             |
+---------------+-------------------+
| Ravi Deshmukh | Data Structures   |
| Neha Patil    | Computer Networks |
+---------------+-------------------+

The member and book pairs where that member currently holds that book. Two columns have to agree at once, and EXISTS says so in two lines. MySQL does have a row constructor, WHERE (m.member_id, b.book_id) IN (SELECT ...), which can do it too, and it is far less readable.

The hardest one: for all

"Which members have borrowed every Databases book?" cannot be asked directly, because SQL has no "for all". It is asked by turning it inside out: a member for whom there is no Databases book that they have not borrowed.

munotes.in181

Practical 7: Subqueries with EXISTS

Two NOT EXISTS, one inside the other:

SELECT m.member_name
FROM member m
WHERE NOT EXISTS (
    SELECT 1 FROM book b
    WHERE b.subject = 'Databases'
      AND NOT EXISTS (
          SELECT 1 FROM loan l
          WHERE l.book_id = b.book_id AND l.member_id = m.member_id
      )
)
ORDER BY m.member_id;

Nobody, in this library, because no member has taken out all three Databases books.

Read it from the inside: the innermost query asks "did this member borrow this book?". The middle NOT EXISTS asks "is there a Databases book this member did not borrow?". The outer NOT EXISTS keeps the members for which there is no such book.

This is called relational division, it is the standard hard question of a database paper, and it is worth writing into your journal once even though it will not be needed often. Check it by making it true:

INSERT INTO loan VALUES (8, 2, 4, '2025-08-01', NULL),
                        (9, 3, 4, '2025-08-01', NULL),
                        (10, 6, 4, '2025-08-01', NULL);
SELECT m.member_name
FROM member m
WHERE NOT EXISTS (
    SELECT 1 FROM book b
    WHERE b.subject = 'Databases'
      AND NOT EXISTS (
          SELECT 1 FROM loan l
          WHERE l.book_id = b.book_id AND l.member_id = m.member_id
      )
)
ORDER BY m.member_id;
+--------------+
| member_name  |
+--------------+
| Imran Shaikh |
+--------------+

Imran Shaikh has now borrowed all three, and the query found him. A query that returns nothing has not been tested. Make it true, run it again, and you know it works.

What beginners get wrong

Writing an uncorrelated EXISTS. WHERE EXISTS (SELECT 1 FROM loan) is true for every row, because there is always at least one loan. The inner query must mention the outer one.

Selecting columns inside EXISTS and thinking they matter. They do not; only the existence of rows does.

Using NOT IN where the column can be NULL. Silently empty.

Trying to test two columns with IN. Use EXISTS.

Expecting a "for all" keyword. There is none; use two NOT EXISTS.

Not testing a query that returns nothing. Insert the data that should make it true and run it again.

Quick revision

  • EXISTS (subquery) is true when the subquery returns at least one row.
  • What the subquery selects does not matter; write SELECT 1.
  • An EXISTS is almost always correlated: the inner query mentions an outer column.
  • A correlated subquery cannot be run on its own and conceptually runs once per outer row.
  • NOT EXISTS is unaffected by NULLs, where NOT IN returns nothing.
  • Unmatched rows can be found three ways: outer join, NOT IN, NOT EXISTS.
  • "For all" is written as two nested NOT EXISTS, and is called relational division.
  • Test a query that returns nothing by making it true.
munotes.in182

Practical 7: Subqueries with EXISTS

What goes in your journal

Aim, then EXISTS and NOT EXISTS over the same pair of tables with their results, then the same question asked with NOT IN so that the two can be compared.

Include the NULL demonstration: the NOT IN that returns nothing and the NOT EXISTS that returns the right answer, on the same data. One line under them saying why is the whole practical.

If you have room, the two nested NOT EXISTS, with the run that returns nothing and the run after the extra loans that returns a name. The pair proves the query works, where either alone proves very little.

Test yourself

1. When is EXISTS true? When its subquery returns at least one row, whatever that row contains.

2. What is a correlated subquery? One that refers to a column of the outer query, so it cannot be run on its own and conceptually runs once for each row of the outer query.

3. Why is NOT EXISTS safe where NOT IN is not? Because EXISTS only asks whether rows are there and never compares a value with NULL, while NOT IN compares and a comparison with NULL is never true.

4. Does it matter what the subquery inside EXISTS selects? No. SELECT 1 is the convention precisely because the columns are ignored.

5. How is "every" expressed in SQL? As a double negative: there is no member of the set for which some required thing does not exist. Two nested NOT EXISTS.

6. Your query returned no rows. How do you know it is right? Insert data that should make it return something, and run it again. A query that has only ever returned nothing has not been tested.

Contents This chapter on its own page

munotes.in183

Chapter Forty-Eight

Practical 8: Turning the ER Model Into Tables

Syllabus topic Module 2, Practical 8: "Converting ER Model to Relational Model ... (Represent entities and relationships in Tabular form, Represent attributes as columns, identifying keys ...)"

Aim

To convert an ER model into a relational model: to represent entities and relationships in tabular form, attributes as columns, and to identify the keys.

The two models

The ER model is for thinking and for showing to people. Shapes, lines, no data types.

The relational model is what a database actually is. Everything in it is a relation, which is a table: a set of rows, each row a set of values, one per named column.

Converting from one to the other is mechanical. There are seven rules and they cover everything the three ER chapters drew.

ER constructBecomes
Strong entity seta table; its key attribute becomes the primary key
Simple attributea column
Composite attributeone column per part; the composite itself is not stored
Multivalued attributea table of its own, keyed by the owner's key plus the value
Derived attributenothing; it is computed when it is asked for
1:1 relationshipa foreign key on either side, usually the one with total participation
1:N relationshipa foreign key on the N side
M:N relationshipa table of its own, with both keys, and the pair as its primary key
Relationship attributea column of whichever table the relationship became
Weak entity seta table whose primary key is the owner's key plus its own discriminator
Generalizationone of three schemes, below

Learn that table. Reproducing it is a whole examination answer on this practical.

Rule 1: every strong entity set becomes a table

AUTHOR, BOOK and MEMBER each had a key of their own, so each becomes a table whose primary key is that attribute.

SELECT table_name AS made_from_an_entity_set
FROM information_schema.tables
WHERE table_schema = DATABASE() AND table_name IN ('author','book','member')
ORDER BY table_name;
+-------------------------+
| made_from_an_entity_set |
+-------------------------+
| author                  |
| book                    |
| member                  |
+-------------------------+

Those three tables came straight from the three rectangles. Each attribute ellipse became a column, and the underlined one became the primary key.

Rule 2: a 1:N relationship becomes a foreign key on the N side

writes was 1:N: one author, many books. So the foreign key goes on BOOK, which is the many side, and there is no writes table at all.

SELECT column_name AS col, referenced_table_name AS points_at
FROM information_schema.key_column_usage
WHERE table_schema = DATABASE() AND table_name = 'book'
  AND referenced_table_name IS NOT NULL;
+-----------+-----------+
| col       | points_at |
+-----------+-----------+
| author_id | author    |
+-----------+-----------+

The foreign key always goes on the many side. Putting it on the one side would need one author row to hold several book ids, and a column holds one value.

If the N side's participation was total, the foreign key is NOT NULL. If it was partial, it may be NULL.

munotes.in184

Practical 8: Turning the ER Model Into Tables

Rule 3: an M:N relationship becomes a table

borrows was M:N, so it cannot be a foreign key anywhere. It becomes a table of its own holding both keys, and that is exactly what loan is.

SELECT column_name AS col, referenced_table_name AS points_at
FROM information_schema.key_column_usage
WHERE table_schema = DATABASE() AND table_name = 'loan'
  AND referenced_table_name IS NOT NULL
ORDER BY column_name;
+-----------+-----------+
| col       | points_at |
+-----------+-----------+
| book_id   | book      |
| member_id | member    |
+-----------+-----------+

Two foreign keys, one to each side. The attributes of the relationship, issued_on and returned_on, became columns of this table, which is rule 9 in the list above, and it is where they belonged all along.

Its primary key is the interesting part

The pair (book_id, member_id) is the natural primary key of a relationship table, and it says: one member may borrow one book once. In a library that is wrong, because a member may borrow the same book again next year. So loan has a loan_id of its own instead.

That is a real decision and it is worth stating in the journal: use the pair as the primary key when the relationship can happen only once, and add a key of its own when it can repeat. A marks table, one mark per student per subject, takes the pair; a loans table does not.

Rule 4: a multivalued attribute becomes a table

A member may have several phone numbers. The ER diagram drew that as a double ellipse, and a column cannot hold several values.

CREATE TABLE member_phone (
    member_id INT,
    phone     VARCHAR(15),
    PRIMARY KEY (member_id, phone),
    FOREIGN KEY (member_id) REFERENCES member(member_id)
);

INSERT INTO member_phone VALUES (1, '9820000001'), (1, '2224440001'), (2, '9820000002');

SELECT m.member_name, p.phone
FROM member m JOIN member_phone p ON p.member_id = m.member_id
ORDER BY m.member_id, p.phone;
+---------------+------------+
| member_name   | phone      |
+---------------+------------+
| Asha Kulkarni | 2224440001 |
| Asha Kulkarni | 9820000001 |
| Ravi Deshmukh | 9820000002 |
+---------------+------------+

The primary key is the pair, because one member may have many numbers and the same number should not be recorded twice for the same member.

The wrong alternative, and it is what students reach for first, is phone1, phone2, phone3 as three columns of member. It fails in three ways: a member with four numbers cannot be recorded, most rows are full of NULLs, and finding who owns a given number means searching three columns.

Rule 5: a composite attribute becomes several columns

An address made of a street, a city and a pin code becomes three columns. The composite itself is not stored, because it is only a name for the group.

munotes.in185

Practical 8: Turning the ER Model Into Tables

street VARCHAR(40),
city   VARCHAR(20),
pin    CHAR(6)

Store the parts, not the whole, and assemble it with CONCAT_WS from [Practical 5: String Functions] when somebody wants it on one line. The reverse, one address column holding everything, makes "all the members in Thane" impossible to ask.

Rule 6: a derived attribute becomes nothing

Age is not stored; date of birth is, and the age is worked out with TIMESTAMPDIFF from [Practical 5: Date Functions] whenever it is asked for. A stored age is wrong from the next birthday.

Rule 7: a weak entity set takes its owner's key

A COPY of a book cannot be identified on its own: copy 2 means nothing until you say copy 2 of which book.

CREATE TABLE copy (
    book_id  INT,
    copy_no  INT,
    shelf    VARCHAR(8),
    PRIMARY KEY (book_id, copy_no),
    FOREIGN KEY (book_id) REFERENCES book(book_id) ON DELETE CASCADE
);

INSERT INTO copy VALUES (1, 1, 'A-12'), (1, 2, 'A-12'), (2, 1, 'B-03');

SELECT b.title, c.copy_no, c.shelf
FROM copy c JOIN book b ON b.book_id = c.book_id
ORDER BY c.book_id, c.copy_no;
+------------------+---------+-------+
| title            | copy_no | shelf |
+------------------+---------+-------+
| The C Language   |       1 | A-12  |
| The C Language   |       2 | A-12  |
| Database Systems |       1 | B-03  |
+------------------+---------+-------+

The primary key is the owner's key plus the discriminator, and ON DELETE CASCADE is right here where it usually is not: a copy of a book cannot outlive the book, which is exactly what a weak entity set means.

A 1:1 relationship

The foreign key may go on either side, and there are two considerations.

Put it on the side with total participation. If every member has a card but not every card has been issued, the foreign key goes on the card, where it can be NOT NULL.

Make it UNIQUE. That is what turns a 1:N foreign key into a 1:1 one; without it the table allows many cards per member and the diagram has been lost.

CREATE TABLE library_card (
    card_no   INT PRIMARY KEY,
    member_id INT NOT NULL UNIQUE,
    issued_on DATE,
    FOREIGN KEY (member_id) REFERENCES member(member_id)
);

Merging the two entity sets into one table is the other option, and is right when both sides are total: a member who must have a card and a card that must have a member are one thing described twice.

Generalization: three schemes

STUDENT and STAFF under MEMBER, from [Practical 1: Generalization and Specialization], can be laid out three ways.

CREATE TABLE student_sub (
    member_id INT PRIMARY KEY,
    course    VARCHAR(10),
    FOREIGN KEY (member_id) REFERENCES member(member_id)
);

CREATE TABLE staff_sub (
    member_id  INT PRIMARY KEY,
    department VARCHAR(12),
    FOREIGN KEY (member_id) REFERENCES member(member_id)
);

INSERT INTO student_sub VALUES (1, 'BSc IT'), (3, 'BSc CS');
INSERT INTO staff_sub   VALUES (2, 'Library');

SELECT m.member_name,
       COALESCE(st.course, sf.department, 'neither') AS kind
FROM member m
LEFT JOIN student_sub st ON st.member_id = m.member_id
LEFT JOIN staff_sub   sf ON sf.member_id = m.member_id
ORDER BY m.member_id LIMIT 4;
munotes.in186

Practical 8: Turning the ER Model Into Tables

+---------------+---------+
| member_name   | kind    |
+---------------+---------+
| Asha Kulkarni | BSc IT  |
| Ravi Deshmukh | Library |
| Meena Iyer    | BSc CS  |
| Imran Shaikh  | neither |
+---------------+---------+
SchemeTablesWorks whenCosts
One per entity setMEMBER, STUDENT, STAFFalwaysa join to see a whole entity
One per subclass onlySTUDENT, STAFF, each with all attributestotal and disjoint onlythe shared attributes are declared twice
One table for everythingMEMBER with course, department and a kind columnalways, best when the subclasses are smallmany NULLs

The first is the general answer and the one to give unless the question says otherwise.

The library, diagram and tables together

In the diagramIn the database
AUTHOR rectangleauthor table, author_id primary key
BOOK rectanglebook table, book_id primary key
MEMBER rectanglemember table, member_id primary key
writes diamond, 1:Nbook.author_id foreign key
borrows diamond, M:Nloan table, two foreign keys
issued_on on the diamondloan.issued_on column
member's phone numbers, double ellipsemember_phone table
COPY, weak entity setcopy table, key (book_id, copy_no)

That is the practical's answer, and it is what to put in the journal: the diagram on one page and this table on the next.

What beginners get wrong

Making a table for a 1:N relationship. It needs only a foreign key on the many side.

Putting the foreign key on the one side. A column holds one value.

Forgetting UNIQUE on a 1:1 foreign key. The table then allows 1:N and the diagram is lost.

Phone1, phone2, phone3. A multivalued attribute becomes a table.

Storing a composite attribute as one column. Nothing inside it can then be searched.

Storing a derived attribute. It goes stale.

Giving a weak entity set a key of its own. Its key is the owner's plus its discriminator, or it was not weak.

Forgetting the relationship's own attributes. issued_on has to land somewhere, and it is the relationship's table.

Quick revision

  • Entity set becomes a table; its key attribute becomes the primary key.
  • Simple attribute becomes a column; composite becomes several; derived becomes none.
  • Multivalued becomes a table keyed by the owner's key plus the value.
  • 1:N becomes a foreign key on the many side.
  • M:N becomes a table with both keys; the pair is the primary key unless the relationship can repeat.
  • 1:1 becomes a foreign key with UNIQUE, on the totally participating side.
  • A weak entity set's primary key is the owner's key plus its discriminator, with ON DELETE CASCADE.
  • A relationship's own attributes become columns of whatever the relationship became.
  • Generalization: one table per entity set is the general answer.
munotes.in187

Practical 8: Turning the ER Model Into Tables

What goes in your journal

Aim, the ER diagram you are converting, the seven rules as a table, and then the CREATE TABLE statements with the primary keys and the foreign keys written out in full.

Then a two column table, diagram construct on the left and what it became on the right, like the library one above. That mapping is the answer to "convert the following ER diagram to a relational schema", which is the question this practical sets.

Test yourself

1. What does a 1:N relationship become? A foreign key on the many side. No separate table is needed.

2. What does an M:N relationship become, and what is its primary key? A table of its own holding both foreign keys. The pair is the primary key when the relationship can happen only once; otherwise the table gets a key of its own.

3. What happens to a multivalued attribute? It becomes a separate table, keyed by the owner's primary key together with the value.

4. Where do a relationship's own attributes go? Into whatever table the relationship became, which for M:N is its own table and for 1:N is the table on the many side.

5. What is the primary key of a weak entity set's table? The owner's primary key together with the weak entity's discriminator.

6. What turns a 1:N foreign key into a 1:1 one? A UNIQUE constraint on the foreign key column.

7. Why is a derived attribute not stored? Because it can be computed from what is stored, and a stored copy goes out of date.

Contents This chapter on its own page

munotes.in188

Chapter Forty-Nine

Practical 8: Normalizing to Third Normal Form

Syllabus topic Module 2, Practical 8: "... apply Normalization on database ... normalization up to 3rd Normal Form"

Aim

To apply normalization to a database, up to third normal form.

What normalization is for

Normalization is the process of arranging columns into tables so that the same fact is stored in exactly one place.

The reason is not tidiness. It is that a fact stored twice can be changed in one place and not the other, and the database then contradicts itself. Those failures have names, and an examiner asks for them:

AnomalyWhat happens
Insertiona fact cannot be recorded because some unrelated fact is not known yet
Updatea fact is changed in one row and not in another, and the table disagrees with itself
Deletionremoving one fact removes another that nobody meant to remove

The bad table

Here is a loans table written the way somebody would write it in a spreadsheet: everything in one place.

CREATE TABLE bad_loan (
    loan_id      INT PRIMARY KEY,
    member_id    INT,
    member_name  VARCHAR(16),
    member_city  VARCHAR(9),
    book_id      INT,
    book_title   VARCHAR(20),
    author_name  VARCHAR(18),
    author_country VARCHAR(10),
    issued_on    DATE
);

INSERT INTO bad_loan VALUES
 (1, 1, 'Asha Kulkarni', 'Mumbai', 1, 'The C Language',  'Dennis Ritchie', 'USA',   '2025-06-02'),
 (2, 1, 'Asha Kulkarni', 'Mumbai', 2, 'Database Systems','Ramez Elmasri',  'USA',   '2025-06-02'),
 (3, 2, 'Ravi Deshmukh', 'Thane',  3, 'MySQL Reference', 'Vikram Vaswani', 'India', '2025-06-10');

SELECT loan_id, member_name, book_title, author_name FROM bad_loan;
+---------+---------------+------------------+----------------+
| loan_id | member_name   | book_title       | author_name    |
+---------+---------------+------------------+----------------+
|       1 | Asha Kulkarni | The C Language   | Dennis Ritchie |
|       2 | Asha Kulkarni | Database Systems | Ramez Elmasri  |
|       3 | Ravi Deshmukh | MySQL Reference  | Vikram Vaswani |
+---------+---------------+------------------+----------------+

Asha Kulkarni's name and city are already written twice. The three anomalies follow at once.

The update anomaly, run:

UPDATE bad_loan SET member_city = 'Thane' WHERE loan_id = 1;

SELECT loan_id, member_id, member_name, member_city FROM bad_loan
WHERE member_id = 1;
+---------+-----------+---------------+-------------+
| loan_id | member_id | member_name   | member_city |
+---------+-----------+---------------+-------------+
|       1 |         1 | Asha Kulkarni | Thane       |
|       2 |         1 | Asha Kulkarni | Mumbai      |
+---------+-----------+---------------+-------------+

One member, two cities. The table now asserts that member 1 lives in Thane and in Mumbai, and nothing in the database objects, because as far as it knows these are two unrelated rows.

The insertion anomaly. A new member who has not borrowed anything cannot be recorded at all, because every row of this table needs a loan_id and a book. The member's existence is hostage to a loan.

The deletion anomaly. Delete loan 3 and Ravi Deshmukh disappears from the database entirely, along with the fact that Vikram Vaswani is from India. Nobody meant to delete either.

DELETE FROM bad_loan WHERE loan_id = 3;

SELECT DISTINCT member_id, member_name FROM bad_loan;
+-----------+---------------+
| member_id | member_name   |
+-----------+---------------+
|         1 | Asha Kulkarni |
+-----------+---------------+
munotes.in189

Practical 8: Normalizing to Third Normal Form

Ravi Deshmukh is gone.

INSERT INTO bad_loan VALUES
 (3, 2, 'Ravi Deshmukh', 'Thane', 3, 'MySQL Reference', 'Vikram Vaswani', 'India', '2025-06-10');
UPDATE bad_loan SET member_city = 'Mumbai' WHERE member_id = 1;

Functional dependency, which is the language of the rest

X determines Y, written X to Y, means: knowing X tells you Y, and two rows with the same X must have the same Y.

In the bad table:

DependencyIn English
loan_id to everythingthe key determines the whole row
member_id to member_name, member_cityone member has one name and one city
book_id to book_title, author_name, author_countryone book has one title and one author
author_name to author_countryone author is from one country

Those four lines are the analysis, and the three normal forms are three rules about them. Write them down before touching the tables; the rest is mechanical.

Two more terms, both asked by name:

A prime attribute is one that is part of some candidate key. Everything else is non-prime.

A partial dependency is a non-prime attribute that depends on only part of a composite key. A transitive dependency is a non-prime attribute that depends on another non-prime attribute.

First normal form

A table is in 1NF when every value is atomic: one value per column per row, and no repeating groups.

The bad table is already in 1NF, and it is worth seeing what would not be:

loan_idmember_namebooks_borrowed
1Asha KulkarniThe C Language, Database Systems
2Ravi DeshmukhMySQL Reference

Two books in one cell. Nothing can be asked of it: "who borrowed Database Systems" needs a string search, the count of books needs the commas counted, and a title containing a comma breaks it.

The other failure of 1NF is a repeating group: book1, book2, book3 as three columns, which is the same mistake as phone1, phone2, phone3 in [Practical 8: Turning the ER Model Into Tables].

The fix in both cases is a row per value, which is exactly what the loan table already is.

Second normal form

A table is in 2NF when it is in 1NF and no non-prime attribute depends on only part of a composite key.

This form has something to say only when the primary key is composite. If the key is a single column, nothing can depend on part of it, and a 1NF table with a single column key is automatically in 2NF.

So take a table whose key really is composite: the loan keyed by the book and the member.

CREATE TABLE bad_loan2 (
    book_id     INT,
    member_id   INT,
    book_title  VARCHAR(20),
    member_name VARCHAR(16),
    issued_on   DATE,
    PRIMARY KEY (book_id, member_id)
);

INSERT INTO bad_loan2 VALUES
 (1, 1, 'The C Language',   'Asha Kulkarni', '2025-06-02'),
 (2, 1, 'Database Systems', 'Asha Kulkarni', '2025-06-02'),
 (3, 2, 'MySQL Reference',  'Ravi Deshmukh', '2025-06-10');

SELECT * FROM bad_loan2;
munotes.in190

Practical 8: Normalizing to Third Normal Form

+---------+-----------+------------------+---------------+------------+
| book_id | member_id | book_title       | member_name   | issued_on  |
+---------+-----------+------------------+---------------+------------+
|       1 |         1 | The C Language   | Asha Kulkarni | 2025-06-02 |
|       2 |         1 | Database Systems | Asha Kulkarni | 2025-06-02 |
|       3 |         2 | MySQL Reference  | Ravi Deshmukh | 2025-06-10 |
+---------+-----------+------------------+---------------+------------+

The key is (book_id, member_id). Now look at the dependencies:

  • book_title depends on book_id alone, which is part of the key.
  • member_name depends on member_id alone, which is the other part.
  • issued_on depends on the whole key, as it should.

Two partial dependencies, so the table is in 1NF and not in 2NF.

The fix is to move each partially dependent attribute to a table keyed by the part it actually depends on.

CREATE TABLE n2_book   (book_id INT PRIMARY KEY, book_title VARCHAR(20));
CREATE TABLE n2_member (member_id INT PRIMARY KEY, member_name VARCHAR(16));
CREATE TABLE n2_loan   (book_id INT, member_id INT, issued_on DATE,
                        PRIMARY KEY (book_id, member_id));

INSERT INTO n2_book   SELECT DISTINCT book_id, book_title   FROM bad_loan2;
INSERT INTO n2_member SELECT DISTINCT member_id, member_name FROM bad_loan2;
INSERT INTO n2_loan   SELECT book_id, member_id, issued_on   FROM bad_loan2;

SELECT b.book_title, m.member_name, l.issued_on
FROM n2_loan l
JOIN n2_book b   ON b.book_id   = l.book_id
JOIN n2_member m ON m.member_id = l.member_id
ORDER BY l.book_id;
+------------------+---------------+------------+
| book_title       | member_name   | issued_on  |
+------------------+---------------+------------+
| The C Language   | Asha Kulkarni | 2025-06-02 |
| Database Systems | Asha Kulkarni | 2025-06-02 |
| MySQL Reference  | Ravi Deshmukh | 2025-06-10 |
+------------------+---------------+------------+

The same three facts, and Asha Kulkarni's name is now stored once instead of twice. Nothing was lost: the join puts the original rows back exactly, which is what makes it a decomposition rather than a deletion.

Third normal form

A table is in 3NF when it is in 2NF and no non-prime attribute depends on another non-prime attribute.

That is the transitive dependency, and the library has one:

CREATE TABLE bad_book (
    book_id        INT PRIMARY KEY,
    book_title     VARCHAR(20),
    author_name    VARCHAR(18),
    author_country VARCHAR(10)
);

INSERT INTO bad_book VALUES
 (1, 'The C Language',    'Dennis Ritchie', 'USA'),
 (4, 'Data Structures',   'Behrouz Forouzan', 'Iran'),
 (5, 'Computer Networks', 'Behrouz Forouzan', 'Iran');

SELECT * FROM bad_book;
+---------+-------------------+------------------+----------------+
| book_id | book_title        | author_name      | author_country |
+---------+-------------------+------------------+----------------+
|       1 | The C Language    | Dennis Ritchie   | USA            |
|       4 | Data Structures   | Behrouz Forouzan | Iran           |
|       5 | Computer Networks | Behrouz Forouzan | Iran           |
+---------+-------------------+------------------+----------------+

book_id determines author_name, and author_name determines author_country. So author_country depends on the key through another non-key column, which is what transitive means: book id to author name to author country.

munotes.in191

Practical 8: Normalizing to Third Normal Form

Iran is written twice, and the update anomaly follows immediately:

UPDATE bad_book SET author_country = 'Canada' WHERE book_id = 4;

SELECT author_name, author_country FROM bad_book WHERE author_name = 'Behrouz Forouzan';
+------------------+----------------+
| author_name      | author_country |
+------------------+----------------+
| Behrouz Forouzan | Canada         |
| Behrouz Forouzan | Iran           |
+------------------+----------------+

One author, two countries.

The fix is to move the transitively dependent attribute to a table keyed by what it really depends on.

CREATE TABLE n3_author (author_name VARCHAR(18) PRIMARY KEY,
                        author_country VARCHAR(10));
CREATE TABLE n3_book   (book_id INT PRIMARY KEY, book_title VARCHAR(20),
                        author_name VARCHAR(18),
                        FOREIGN KEY (author_name) REFERENCES n3_author(author_name));

INSERT INTO n3_author VALUES ('Dennis Ritchie', 'USA'), ('Behrouz Forouzan', 'Iran');
INSERT INTO n3_book VALUES (1, 'The C Language', 'Dennis Ritchie'),
                           (4, 'Data Structures', 'Behrouz Forouzan'),
                           (5, 'Computer Networks', 'Behrouz Forouzan');

UPDATE n3_author SET author_country = 'Canada' WHERE author_name = 'Behrouz Forouzan';

SELECT b.book_title, a.author_name, a.author_country
FROM n3_book b JOIN n3_author a ON a.author_name = b.author_name
ORDER BY b.book_id;
+-------------------+------------------+----------------+
| book_title        | author_name      | author_country |
+-------------------+------------------+----------------+
| The C Language    | Dennis Ritchie   | USA            |
| Data Structures   | Behrouz Forouzan | Canada         |
| Computer Networks | Behrouz Forouzan | Canada         |
+-------------------+------------------+----------------+

One update, both books right. That is what normalization buys, and it is the paragraph to write in the conclusion.

In the real library database this is why AUTHOR is a table of its own, exactly as [Practical 1: The ER Diagram: Entities, Attributes and Keys] argued from the other direction. A well drawn ER diagram usually produces tables that are already in 3NF, which is a good thing to be able to say.

The three forms, on one line each

FormRequiresRemoves
1NFevery value atomic, no repeating groupslists and repeated columns
2NF1NF, and no partial dependency on part of a composite keyfacts about part of the key
3NF2NF, and no transitive dependency through a non-key columnfacts about a non-key column

The mnemonic examiners like: the key, the whole key, and nothing but the key. 1NF is about the data being atomic; 2NF says every non-key attribute depends on the whole key; 3NF says it depends on nothing but the key.

Beyond 3NF there is Boyce Codd normal form, and then 4NF and 5NF. MU's syllabus says up to third normal form, so they are outside this practical; knowing that they exist and that BCNF is a stricter 3NF is enough.

What beginners get wrong

Thinking 2NF applies to a table with a single column key. It cannot fail there.

Confusing partial with transitive. Partial is on part of the key; transitive is through a non-key column.

munotes.in192

Practical 8: Normalizing to Third Normal Form

Decomposing so far that the original rows cannot be rebuilt. A decomposition must be lossless: the join must give back exactly what was there.

Normalizing without writing the dependencies down. The dependencies are the working; the tables follow from them.

Believing normalization is always right. It costs joins, and a reporting database is often deliberately denormalized for speed. Say so in a viva: it shows judgment.

Forgetting to add the foreign keys after decomposing. Without them the pieces can drift apart, and the anomaly comes back by another road.

Quick revision

  • Normalization puts each fact in one place, to stop insertion, update and deletion anomalies.
  • X to Y means knowing X tells you Y.
  • A prime attribute is part of a candidate key; anything else is non-prime.
  • 1NF: atomic values, no repeating groups.
  • 2NF: 1NF and no non-prime attribute depending on part of a composite key.
  • 3NF: 2NF and no non-prime attribute depending on another non-prime attribute.
  • The key, the whole key, and nothing but the key.
  • A decomposition must be lossless: the join must rebuild the original rows.
  • A good ER design usually lands in 3NF already.

What goes in your journal

Aim, the unnormalized table with its rows, and the list of functional dependencies. Then three steps, each with the tables before and after and one sentence naming the dependency being removed.

Include at least one anomaly actually performed: the UPDATE that leaves one member in two cities is the best of them, because the result is visible in two rows of one printed table. Then the same update after normalization, touching one row and fixing everything. Those two results are the practical's whole argument.

Test yourself

1. What are the three anomalies normalization prevents? Insertion, update and deletion anomalies.

2. What does 1NF require? That every value is atomic, with one value per column per row and no repeating groups of columns.

3. What is a partial dependency, and which form removes it? A non-prime attribute depending on only part of a composite primary key. 2NF removes it.

4. What is a transitive dependency, and which form removes it? A non-prime attribute depending on another non-prime attribute rather than on the key directly. 3NF removes it.

5. Why can a table with a single column primary key not violate 2NF? Because there is no part of the key for an attribute to depend on; any dependency on the key is a dependency on the whole key.

6. What does lossless decomposition mean? That joining the decomposed tables back together returns exactly the original rows, with nothing lost and nothing invented.

7. Give the mnemonic for the first three normal forms. The key, the whole key, and nothing but the key.

Contents This chapter on its own page

munotes.in193

Chapter Fifty

Practical 9: Views

Syllabus topic Module 2, Practical 9: "Views: Creating Views (with and without check option), Dropping views, Selecting from a view"

Aim

To create views with and without the check option, to select from them, and to drop them.

What a view is

A view is a stored query with a name. It looks like a table and holds no data of its own: every time you select from it, the query underneath runs.

That single sentence answers most of the viva questions about views, and the rest of this chapter is what follows from it.

Creating one, and selecting from it

CREATE VIEW book_with_author AS
SELECT b.book_id, b.title, b.subject, b.price, a.author_name
FROM book b JOIN author a ON a.author_id = b.author_id;
SELECT title, author_name FROM book_with_author ORDER BY book_id LIMIT 4;
+------------------+------------------+
| title            | author_name      |
+------------------+------------------+
| The C Language   | Dennis Ritchie   |
| Database Systems | Ramez Elmasri    |
| MySQL Reference  | Vikram Vaswani   |
| Data Structures  | Behrouz Forouzan |
+------------------+------------------+

It is used exactly like a table. WHERE, ORDER BY, GROUP BY and joins all work on it:

SELECT author_name, COUNT(*) AS books, ROUND(AVG(price), 2) AS avg_price
FROM book_with_author
GROUP BY author_name
HAVING COUNT(*) > 1
ORDER BY books DESC, author_name;
+------------------+-------+-----------+
| author_name      | books | avg_price |
+------------------+-------+-----------+
| Behrouz Forouzan |     3 |    602.00 |
| Ramez Elmasri    |     2 |    575.25 |
| Vikram Vaswani   |     2 |    389.50 |
+------------------+-------+-----------+

That query is four lines because the join is already inside the view. That is what views are for: a complicated query is written once, given a name, and everybody else asks simple questions of it.

Why a database has views

ReasonWhat it means here
Simplicitya three table join becomes one name
Securityshow a user the columns they may see and no others
Consistencyeverybody's definition of "an overdue loan" is the same one
Independencethe tables underneath can change, and the view can be rewritten to keep its old shape

The security one is the reason MU pairs views with [Practical 10: Granting and Revoking Permissions]: a user can be given the right to read a view without any right over the tables it is built from, so the view is the permission.

A view is not a copy

SELECT COUNT(*) AS rows_in_view FROM book_with_author;
+--------------+
| rows_in_view |
+--------------+
|            8 |
+--------------+
INSERT INTO book VALUES (10, 'New Arrival', 'Programming', 350.00, '2025-08-01', 1);

SELECT COUNT(*) AS rows_in_view_now FROM book_with_author;
+------------------+
| rows_in_view_now |
+------------------+
|                9 |
+------------------+

Nobody touched the view and it changed, because the view is the query and the query was run again. A view is always current, and that is the difference between a view and a table made with CREATE TABLE ... AS SELECT, which is a copy taken once and stale from then on.

munotes.in194

Practical 9: Views

Selecting some columns only

CREATE VIEW member_public AS
SELECT member_id, member_name, course FROM member;
SELECT * FROM member_public ORDER BY member_id LIMIT 3;
+-----------+---------------+--------+
| member_id | member_name   | course |
+-----------+---------------+--------+
|         1 | Asha Kulkarni | BSc IT |
|         2 | Ravi Deshmukh | BSc IT |
|         3 | Meena Iyer    | BSc CS |
+-----------+---------------+--------+

city and joined_on are not in the view, so a user given access to member_public and nothing else cannot see them at all. This is column level security, done with a view.

An updatable view

Some views can be written to, and the change goes through to the table underneath.

CREATE VIEW cheap_books AS
SELECT book_id, title, subject, price FROM book WHERE price < 500;
SELECT book_id, title, price FROM cheap_books ORDER BY book_id;
+---------+-----------------+--------+
| book_id | title           | price  |
+---------+-----------------+--------+
|       1 | The C Language  | 395.00 |
|       3 | MySQL Reference | 499.00 |
|       6 | Learn SQL       | 280.00 |
|       8 | Discrete Maths  | 430.00 |
|      10 | New Arrival     | 350.00 |
+---------+-----------------+--------+
UPDATE cheap_books SET price = 310.00 WHERE book_id = 6;

SELECT book_id, title, price FROM book WHERE book_id = 6;
+---------+-----------+--------+
| book_id | title     | price  |
+---------+-----------+--------+
|       6 | Learn SQL | 310.00 |
+---------+-----------+--------+

The update went through the view into book.

A view is updatable only when there is a clear one to one correspondence between its rows and the rows of a table. A view is not updatable if it contains a DISTINCT, a GROUP BY or HAVING, an aggregate function, a UNION, or a subquery in the select list.

CREATE VIEW books_per_subject AS
SELECT subject, COUNT(*) AS books FROM book GROUP BY subject;
UPDATE books_per_subject SET books = 99 WHERE subject = 'Databases';
ERROR 1288 (HY000): The target table books_per_subject of the UPDATE is not updatable

Error 1288. There is no row of book that the number 3 belongs to; it was computed from three of them, and there is no way to push a change to it back.

A join view, and the surprise in it

A view over a join is updatable, provided the change touches only one of the joined tables. That is more permissive than most students expect, and it is worth seeing:

UPDATE book_with_author SET author_name = 'D. M. Ritchie' WHERE book_id = 1;

SELECT author_id, author_name FROM author WHERE author_id = 1;
+-----------+---------------+
| author_id | author_name   |
+-----------+---------------+
|         1 | D. M. Ritchie |
+-----------+---------------+

The statement named a book and changed an author. author_name belongs to author, so the update went there, and the name is now different for every book Dennis Ritchie wrote, not just for book 1.

munotes.in195

Practical 9: Views

That is correct behaviour and it is a trap worth knowing about: an update through a join view changes the table the column came from, not the rows the view showed. If that is not what you want, update the base table directly, where the statement says plainly what it is changing.

UPDATE author SET author_name = 'Dennis Ritchie' WHERE author_id = 1;

WITH CHECK OPTION, which is what MU names

Here is the problem the clause solves. The view shows books under 500. Nothing so far stops you putting a book through the view that does not belong in it:

INSERT INTO cheap_books VALUES (11, 'Expensive Book', 'Databases', 900.00);

SELECT book_id, title, price FROM book WHERE book_id = 11;
+---------+----------------+--------+
| book_id | title          | price  |
+---------+----------------+--------+
|      11 | Expensive Book | 900.00 |
+---------+----------------+--------+
SELECT book_id, title FROM cheap_books WHERE book_id = 11;

The row went into book and then vanished from the view, because 900 is not under 500. A user working only through cheap_books has inserted a row and cannot see it afterwards. That is the anomaly, and it is confusing enough to be worth a clause of its own.

CREATE VIEW cheap_books_checked AS
SELECT book_id, title, subject, price FROM book WHERE price < 500
WITH CHECK OPTION;
INSERT INTO cheap_books_checked VALUES (12, 'Another Costly One', 'Databases', 950.00);
ERROR 1369 (HY000): CHECK OPTION failed 'librarydb.cheap_books_checked'

Error 1369. WITH CHECK OPTION refuses any insert or update through the view that would produce a row the view cannot see. A row that does belong still goes in:

INSERT INTO cheap_books_checked VALUES (13, 'Pocket Guide', 'Programming', 120.00);

SELECT book_id, title, price FROM cheap_books_checked WHERE book_id = 13;
+---------+--------------+--------+
| book_id | title        | price  |
+---------+--------------+--------+
|      13 | Pocket Guide | 120.00 |
+---------+--------------+--------+

There are two forms of the clause, and they differ only when views are built on views:

WrittenChecks
WITH LOCAL CHECK OPTIONthis view's own condition
WITH CASCADED CHECK OPTIONthis view's condition and those of every view underneath it

CASCADED is the default when neither word is given, which is what the view above used.

Listing and dropping

SELECT table_name AS views_here
FROM information_schema.views
WHERE table_schema = DATABASE()
ORDER BY table_name;
+---------------------+
| views_here          |
+---------------------+
| book_with_author    |
| books_per_subject   |
| cheap_books         |
| cheap_books_checked |
| member_public       |
+---------------------+
DROP VIEW cheap_books_checked;

SELECT table_name AS views_left
FROM information_schema.views
WHERE table_schema = DATABASE()
ORDER BY table_name;
+-------------------+
| views_left        |
+-------------------+
| book_with_author  |
| books_per_subject |
| cheap_books       |
| member_public     |
+-------------------+
munotes.in196

Practical 9: Views

Dropping a view destroys no data. The view was only a stored query; the rows were always in the tables and are still there.

SELECT COUNT(*) AS books_still_here FROM book;
+------------------+
| books_still_here |
+------------------+
|               11 |
+------------------+

DROP VIEW IF EXISTS v; does not complain when the view is not there. CREATE OR REPLACE VIEW v AS ... redefines an existing view in one statement, which is how a view is changed.

Dropping a table that a view is built on is allowed, and the view then breaks the next time somebody selects from it, which is a genuine trap on a shared database.

View against table

TableView
Holds datayesno, it is a stored query
Occupies spaceyesonly the definition
Always currentit is the datayes, the query is re-run
Can be written toyesonly if it is updatable
DROP destroys datayesno
Can restrict columns or rows for securitynoyes

What beginners get wrong

Thinking a view stores a copy. It stores the query.

Expecting every view to be updatable. A view with GROUP BY, DISTINCT, an aggregate or a UNION is not.

Inserting through a view without WITH CHECK OPTION and then losing the row from sight.

Dropping a view and expecting the data to go. It does not; that is DROP TABLE.

Using DROP TABLE on a view. It is refused; a view is dropped with DROP VIEW.

Building a view on a view on a view. It works and it becomes very hard to see what will run.

Forgetting that the underlying table can be changed or dropped without the view being told.

Quick revision

  • A view is a stored query with a name; it holds no data.
  • CREATE VIEW v AS SELECT ...; and then select from v like a table.
  • Uses: simplicity, security, a shared definition, independence from the tables.
  • A view is always current, because the query runs each time.
  • Updatable only with a one to one row correspondence: no DISTINCT, GROUP BY, aggregate or UNION.
  • Without WITH CHECK OPTION, a row inserted through a view can fall outside it and vanish.
  • WITH CHECK OPTION refuses such a row. Error 1369.
  • LOCAL checks this view; CASCADED, the default, checks the views underneath too.
  • DROP VIEW v; removes the definition and no data. CREATE OR REPLACE VIEW redefines one.

What goes in your journal

Aim, the CREATE VIEW, a SELECT from it, and then the pair that is the whole practical: an insert through a view without the check option, and the same insert through a view with it.

munotes.in197

Practical 9: Views

Write the two results out in full. The first shows the row in the table and absent from the view; the second shows error 1369. Those two together explain the clause in a way no sentence does, and the question "what is the use of WITH CHECK OPTION" is what a viva asks about views.

Test yourself

1. What is a view? A named stored query. It looks like a table and holds no data; the query runs each time the view is used.

2. Give two reasons for using one. It hides the complexity of a long query behind a name, and it lets a user be given access to some columns or rows without access to the whole table.

3. When is a view not updatable? When its rows do not correspond one to one with rows of a table: if it contains DISTINCT, GROUP BY, HAVING, an aggregate function, a UNION, or a subquery in the select list.

4. What does WITH CHECK OPTION do? It refuses any insert or update made through the view that would produce a row the view's own condition excludes.

5. What happens without it? The row goes into the table and then disappears from the view, so the user who inserted it cannot see it.

6. Does DROP VIEW delete data? No. Only the stored query is removed; the tables and their rows are untouched.

7. What is the difference between LOCAL and CASCADED? LOCAL checks only this view's condition. CASCADED, which is the default, also checks the conditions of any views this one is built on.

Contents This chapter on its own page

munotes.in198

Chapter Fifty-One

Practical 10: Granting and Revoking Permissions

Syllabus topic Module 2, Practical 10: "DCL statements: Granting and revoking permissions"

Aim

To grant and revoke permissions on a database.

DCL, the fourth family

GRANT and REVOKE are the Data Control Language, the family from [Practical 2: Inserting, Updating and Deleting Rows] that decides who may do what.

Everything so far in this module has been done as root, who may do everything. That is right for learning and wrong for anything else. A library assistant who only issues books has no business being able to drop a table, and the way to arrange that is not to trust them: it is to give them exactly the privileges the job needs.

Creating a user

mysql -u root -p
CREATE USER 'libclerk'@'localhost' IDENTIFIED BY 'Clerk#123';

Two things about that statement.

A MySQL user is a name and a host together. 'libclerk'@'localhost' may connect only from this machine. 'libclerk'@'%' could connect from anywhere, and 'libclerk'@'192.168.1.%' from one network. They are different users even with the same name, and a login fails with "access denied" surprisingly often because the host part does not match where the connection came from.

A new user has no privileges at all. It can connect and do nothing, which is what GRANT USAGE ON . means when you see it later: usage is the privilege of having none.

Granting one privilege

GRANT SELECT ON librarydb.book TO 'libclerk'@'localhost';

Read it in three parts: what may be done, on what, and by whom.

PartWrittenMeans
PrivilegeSELECTmay read
Objectlibrarydb.bookthis one table
User'libclerk'@'localhost'this user from this host

The object may be narrowed or widened:

WrittenApplies to
.every database on the server
librarydb.*every table in one database
librarydb.bookone table
librarydb.book(title, price)two columns of one table, for SELECT, INSERT and UPDATE

The last row is worth knowing: privileges reach down to individual columns, which is the other way of doing what the view in [Practical 9: Views] did.

Proving it, from the other side

Now open a second connection as the new user. This is the part that cannot be faked, and it is what the practical is for.

mysql -u libclerk -p librarydb
+---------+------------------+
| book_id | title            |
+---------+------------------+
|       1 | The C Language   |
|       2 | Database Systems |
|       3 | MySQL Reference  |
+---------+------------------+

The SELECT on book worked. Everything else does not:

ERROR 1142 (42000) at line 1: INSERT command denied to user 'libclerk'@'localhost' for table 'book'
ERROR 1142 (42000) at line 1: SELECT command denied to user 'libclerk'@'localhost' for table 'member'

Error 1142 twice: once for the wrong action on a table the user may read, and once for the right action on a table the user may not touch at all. The message names the command, the user and the table, which is enough to see exactly what is missing.

munotes.in199

Practical 10: Granting and Revoking Permissions

The words "at line 1" appear because these statements were run with mysql -e "...", which makes the client treat them as a one line script. Typed at the mysql> prompt the same errors read ERROR 1142 (42000): INSERT command denied ... with no line number.

Looking at what a user has

SHOW GRANTS FOR 'libclerk'@'localhost';
+---------------------------------------------------------------+
| Grants for libclerk@localhost                                 |
+---------------------------------------------------------------+
| GRANT USAGE ON *.* TO `libclerk`@`localhost`                  |
| GRANT SELECT ON `librarydb`.`book` TO `libclerk`@`localhost`  |
+---------------------------------------------------------------+

Two lines. USAGE ON . is the user's right to exist and connect, and it carries no privilege of its own; it is there for every user. The second line is the grant that was made.

SHOW GRANTS; with no name shows your own.

Granting more, and taking it back

GRANT INSERT, UPDATE ON librarydb.book TO 'libclerk'@'localhost';

Several privileges in one statement, separated by commas. Back in the clerk's session, the insert that was refused a moment ago now succeeds and prints nothing, which is what a successful INSERT does.

REVOKE INSERT ON librarydb.book FROM 'libclerk'@'localhost';

And the same insert again:

ERROR 1142 (42000) at line 1: INSERT command denied to user 'libclerk'@'localhost' for table 'book'

REVOKE is GRANT written backwards, with FROM instead of TO. Notice that UPDATE was not revoked and is still there: revoking is as precise as granting.

Taking everything back on one object:

REVOKE ALL PRIVILEGES ON librarydb.book FROM 'libclerk'@'localhost';
ERROR 1044 (42000): Access denied for user 'libclerk'@'localhost' to database 'librarydb'

A different error number. 1142 means "you may not do that to this table"; 1044 means "you may not use this database at all", which is what is left when the user's last privilege in it has gone.

The privileges worth knowing

PrivilegeAllows
SELECTreading
INSERT, UPDATE, DELETEthe three DML verbs
CREATE, DROP, ALTERmaking and changing tables
INDEXcreating and dropping indexes
CREATE VIEW, SHOW VIEWmaking views and seeing their definitions
EXECUTErunning stored procedures
GRANT OPTIONpassing on one's own privileges to somebody else
ALL PRIVILEGESeverything on the named object

GRANT OPTION is the one to be careful with. Granted at the end of a statement:

GRANT SELECT ON librarydb.* TO 'libclerk'@'localhost' WITH GRANT OPTION;

the user may now give that same privilege to anybody else. On a college server that is how a privilege quietly spreads to people nobody intended, and it is a fair viva question: WITH GRANT OPTION delegates the power to grant, not just the privilege.

Removing the user

DROP USER 'libclerk'@'localhost';

The account and all its privileges go. It does not delete any data the user created; rows belong to the table, not to whoever inserted them.

munotes.in200

Practical 10: Granting and Revoking Permissions

The principle behind all of it

Least privilege: give each user exactly the privileges their work needs, and nothing more. Not because people are untrustworthy, but because a mistake or a stolen password can then do only as much damage as that account could do deliberately.

The library's three kinds of user fall out of it at once:

WhoNeeds
A readerSELECT on book and on a view of member showing only their own row
A clerkSELECT, INSERT, UPDATE on loan; SELECT on book and member
The librarianALL PRIVILEGES ON librarydb.*

Nobody needs ., and nobody needs DROP in the course of a normal day.

GRANT against REVOKE

GRANTREVOKE
Word before the userTOFROM
Effectadds privilegesremoves them
Applies touser and host togetherthe same
Takes effectat once for new connectionsat once, for the session in progress too
FamilyDCLDCL

What beginners get wrong

Forgetting the host part. 'libclerk' alone means 'libclerk'@'%', which is not the same user as 'libclerk'@'localhost'.

Using TO with REVOKE. It is FROM.

Granting ALL PRIVILEGES ON . to save time. That is a second root account.

Expecting a new user to be able to do anything. It has USAGE and nothing else.

Testing a grant in the administrator's own session. Nothing is proved until you connect as the new user.

Confusing 1044 with 1142. 1044 is the whole database refused; 1142 is one command on one table.

Forgetting WITH GRANT OPTION was given. The privilege can then travel without you.

Quick revision

  • GRANT and REVOKE are DCL.
  • A user is a name and a host: 'libclerk'@'localhost'.
  • CREATE USER 'u'@'h' IDENTIFIED BY 'password'; and the new user can do nothing.
  • GRANT privileges ON object TO user; and REVOKE privileges ON object FROM user;.
  • Objects: ., db.*, db.table, and db.table(columns).
  • SHOW GRANTS FOR 'u'@'h'; lists what a user has; USAGE ON . means none.
  • WITH GRANT OPTION lets the user pass the privilege on.
  • Error 1142 is a command refused on a table; error 1044 is the database refused entirely.
  • DROP USER removes the account and its privileges, and no data.
  • Least privilege: grant what the work needs and nothing more.

What goes in your journal

Aim, then the statements in the order they were run, and, beside each one, what the second session could do afterwards. Two columns work well: the administrator's statement on the left, the clerk's result on the right.

The write-up must contain at least one refusal with its error number. A journal showing only successful grants has demonstrated that GRANT is a word, not that it does anything, and that is precisely what an examiner will ask you to prove.

munotes.in201

Practical 10: Granting and Revoking Permissions

Test yourself

1. What do GRANT and REVOKE belong to, and what does it control? The Data Control Language, DCL, which controls who may do what to which object.

2. Why is a MySQL user a name and a host together? Because the privilege depends on where the connection comes from. 'u'@'localhost' and 'u'@'%' are different accounts.

3. What can a newly created user do? Connect, and nothing else. It holds USAGE ON ., which is the absence of privileges.

4. Write a statement giving read access to one table. GRANT SELECT ON librarydb.book TO 'libclerk'@'localhost';

5. What is the difference between error 1044 and error 1142? 1142 means one command is refused on one table. 1044 means the user has no access to the database at all.

6. What does WITH GRANT OPTION add? The right to grant that privilege to other users, so the privilege can spread without the administrator.

7. State the principle of least privilege. Give each user exactly the privileges their work requires and no more, so that a mistake or a compromised account can do only limited damage.

Contents This chapter on its own page

munotes.in202

Chapter Fifty-Two

Practical 10: COMMIT and ROLLBACK

Syllabus topic Module 2, Practical 10: "DCL statements: Saving (Commit) and Undoing (rollback)"

Aim

To save work with COMMIT and to undo it with ROLLBACK.

What a transaction is

A transaction is a group of statements treated as one indivisible unit: either all of them happen, or none of them do.

The example everybody uses is a bank transfer, and it is the bank of [Practical 10: Designing the Bank Management System]. Moving 5000 rupees from one account to another is two updates:

UPDATE account SET balance = balance - 5000 WHERE number = 1001;
UPDATE account SET balance = balance + 5000 WHERE number = 1002;

If the machine fails between them, 5000 rupees have left one account and arrived nowhere. No amount of careful programming prevents that, because the failure is not in the program. What prevents it is wrapping both statements in a transaction, so that the server either applies both or neither.

ACID, which is asked by name

The four properties a transaction guarantees:

PropertyMeans
Atomicityall of it happens or none of it does
Consistencyit takes the database from one valid state to another, with every constraint satisfied
Isolationconcurrent transactions do not see each other's unfinished work
Durabilityonce committed, it survives a crash

Atomicity is what ROLLBACK provides and durability is what COMMIT provides, which is why those two statements are this practical.

Autocommit, which is why nothing seems to happen

By default MySQL is in autocommit mode: every single statement is its own transaction and is committed the instant it finishes.

SELECT @@autocommit AS autocommit_is_on;
+------------------+
| autocommit_is_on |
+------------------+
|                1 |
+------------------+

1, so it is on. That is why every UPDATE in this module so far has been permanent immediately, and why the UPDATE with no WHERE in [Practical 2: Inserting, Updating and Deleting Rows] could not be undone.

There are two ways out of it, and both are worth knowing.

START TRANSACTION, and ROLLBACK

SELECT book_id, title, price FROM book WHERE book_id = 1;
+---------+----------------+--------+
| book_id | title          | price  |
+---------+----------------+--------+
|       1 | The C Language | 395.00 |
+---------+----------------+--------+
START TRANSACTION;

UPDATE book SET price = 1.00 WHERE book_id = 1;

SELECT book_id, title, price FROM book WHERE book_id = 1;
+---------+----------------+-------+
| book_id | title          | price |
+---------+----------------+-------+
|       1 | The C Language |  1.00 |
+---------+----------------+-------+

Inside the transaction the price is 1. Now undo it:

ROLLBACK;

SELECT book_id, title, price FROM book WHERE book_id = 1;
+---------+----------------+--------+
| book_id | title          | price  |
+---------+----------------+--------+
|       1 | The C Language | 395.00 |
+---------+----------------+--------+

395 again. The update was real inside the transaction and vanished when the transaction was abandoned. START TRANSACTION suspends autocommit until a COMMIT or a ROLLBACK ends the transaction; BEGIN is another spelling of it.

munotes.in203

Practical 10: COMMIT and ROLLBACK

COMMIT

START TRANSACTION;

UPDATE book SET price = 425.00 WHERE book_id = 1;

COMMIT;

SELECT book_id, title, price FROM book WHERE book_id = 1;
+---------+----------------+--------+
| book_id | title          | price  |
+---------+----------------+--------+
|       1 | The C Language | 425.00 |
+---------+----------------+--------+
ROLLBACK;

SELECT book_id, title, price FROM book WHERE book_id = 1;
+---------+----------------+--------+
| book_id | title          | price  |
+---------+----------------+--------+
|       1 | The C Language | 425.00 |
+---------+----------------+--------+

The second ROLLBACK did nothing, and that is the point. Once a change is committed there is nothing left to roll back. The transaction ended at the COMMIT, and the ROLLBACK after it belongs to a new and empty transaction.

Several statements as one unit

SELECT COUNT(*) AS books_before FROM book;
+--------------+
| books_before |
+--------------+
|            8 |
+--------------+
START TRANSACTION;

INSERT INTO book VALUES (20, 'First Extra',  'Testing', 100.00, '2025-08-01', 1);
INSERT INTO book VALUES (21, 'Second Extra', 'Testing', 200.00, '2025-08-01', 1);
DELETE FROM book WHERE book_id = 8;

SELECT COUNT(*) AS inside_the_transaction FROM book;
+------------------------+
| inside_the_transaction |
+------------------------+
|                      9 |
+------------------------+

Two inserted, one deleted, so nine. And then:

ROLLBACK;

SELECT COUNT(*) AS after_rollback FROM book;
+----------------+
| after_rollback |
+----------------+
|              8 |
+----------------+

All three statements undone together. That is atomicity: the group is the unit, not the statement.

SAVEPOINT: undoing part of a transaction

START TRANSACTION;

INSERT INTO book VALUES (30, 'Keeper',  'Testing', 100.00, '2025-08-01', 1);

SAVEPOINT after_keeper;

INSERT INTO book VALUES (31, 'Mistake', 'Testing', 200.00, '2025-08-01', 1);

SELECT book_id, title FROM book WHERE book_id IN (30, 31) ORDER BY book_id;
+---------+---------+
| book_id | title   |
+---------+---------+
|      30 | Keeper  |
|      31 | Mistake |
+---------+---------+
ROLLBACK TO SAVEPOINT after_keeper;

SELECT book_id, title FROM book WHERE book_id IN (30, 31) ORDER BY book_id;
+---------+--------+
| book_id | title  |
+---------+--------+
|      30 | Keeper |
+---------+--------+

The second insert is gone and the first is still there. The transaction is still open: ROLLBACK TO SAVEPOINT undoes part of it and does not end it, so a COMMIT or a full ROLLBACK is still needed.

ROLLBACK;

SELECT COUNT(*) AS back_to_normal FROM book;
+----------------+
| back_to_normal |
+----------------+
|              8 |
+----------------+

RELEASE SAVEPOINT name; forgets a savepoint without undoing anything.

Switching autocommit off instead

SET autocommit = 0;

SELECT @@autocommit AS autocommit_now;
+----------------+
| autocommit_now |
+----------------+
|              0 |
+----------------+

With autocommit off, a transaction starts by itself at your first statement and lasts until you commit or roll back. Nothing is permanent until you say so.

UPDATE book SET subject = 'Changed' WHERE book_id = 2;

ROLLBACK;

SELECT book_id, subject FROM book WHERE book_id = 2;
munotes.in204

Practical 10: COMMIT and ROLLBACK

+---------+-----------+
| book_id | subject   |
+---------+-----------+
|       2 | Databases |
+---------+-----------+

No START TRANSACTION anywhere, and the update was still undone.

SET autocommit = 1;
START TRANSACTIONSET autocommit = 0
Lasts forone transactionthe rest of the session
Autocommit afterwardsback on by itselfstill off until you set it back
Suitsone careful operation in an ordinary sessiona whole session of careful work

The statement that ends a transaction without being asked

This is the trap from [Practical 3: Altering a Table That Already Holds Data], and it belongs here too.

START TRANSACTION;
UPDATE book SET price = 1;
CREATE TABLE anything (x INT);   -- DDL: commits everything above it
ROLLBACK;                        -- too late

A DDL statement causes an implicit commit. CREATE, ALTER, DROP, TRUNCATE, RENAME and CREATE USER all do it, so a transaction that contains one cannot be rolled back past it. Keep DDL out of transactions.

The engine has to support it

Transactions are a property of the storage engine, not of SQL. MySQL's default engine, InnoDB, supports them. The older MyISAM does not: a START TRANSACTION is accepted, the statements run, and the ROLLBACK silently does nothing.

SELECT table_name AS tbl, engine
FROM information_schema.tables
WHERE table_schema = DATABASE() AND table_name IN ('book', 'member')
ORDER BY table_name;
+--------+--------+
| tbl    | ENGINE |
+--------+--------+
| book   | InnoDB |
| member | InnoDB |
+--------+--------+

InnoDB, so everything in this chapter works. On a table created with ENGINE=MyISAM none of it would, and nothing would tell you. If a rollback in your laboratory appears to do nothing, check the engine first.

What beginners get wrong

Expecting a ROLLBACK to work without a transaction. With autocommit on, each statement has already been committed.

Rolling back after a COMMIT. The transaction is over.

Putting a CREATE or DROP inside a transaction. It commits everything before it.

Assuming every table is transactional. MyISAM tables are not, and fail silently.

Leaving a transaction open. Other sessions wait on the locks it holds, which on a shared server looks to everybody else like the database has stopped.

Thinking ROLLBACK TO SAVEPOINT ends the transaction. It does not.

Forgetting to set autocommit back on after SET autocommit = 0, and then wondering why later work disappears when the connection closes.

Quick revision

  • A transaction is a group of statements that all happen or none do.
  • ACID: atomicity, consistency, isolation, durability.
  • MySQL commits every statement by default: SELECT @@autocommit is 1.
  • START TRANSACTION; or BEGIN; suspends that until COMMIT or ROLLBACK.
  • COMMIT makes the changes permanent; ROLLBACK undoes them all.
  • After a COMMIT there is nothing to roll back.
  • SAVEPOINT name; and ROLLBACK TO SAVEPOINT name; undo part of a transaction without ending it.
  • SET autocommit = 0; makes the whole session work this way.
  • DDL causes an implicit commit; keep it out of transactions.
  • Transactions need InnoDB. MyISAM accepts the statements and ignores them.
munotes.in205

Practical 10: COMMIT and ROLLBACK

What goes in your journal

Aim, then the three runs that are the practical: a change rolled back, a change committed, and a ROLLBACK attempted after the COMMIT and doing nothing.

Print the row before, inside and after each one. Nine short results, and they are the only way to show that the rollback worked: a table that looks the same at the end proves nothing unless the reader has seen it different in the middle.

Add the savepoint run if you have room, and one line saying why a CREATE TABLE must not go inside a transaction.

Test yourself

1. What is a transaction? A group of statements treated as a single unit: either all of them take effect or none of them do.

2. What do the four letters of ACID stand for? Atomicity, consistency, isolation and durability.

3. What is autocommit, and what is it set to by default? The mode in which every statement is committed as soon as it finishes. It is on, @@autocommit is 1.

4. Can you roll back after a COMMIT? No. The commit ended the transaction and made the changes permanent.

5. What does ROLLBACK TO SAVEPOINT s; do? Undoes everything done since the savepoint, and leaves the transaction open so that it still needs a COMMIT or a ROLLBACK.

6. Why must DDL be kept out of a transaction? Because a DDL statement causes an implicit commit, so everything done before it becomes permanent and can no longer be rolled back.

7. A ROLLBACK did nothing and no error appeared. What would you check? The storage engine. MyISAM tables are not transactional and ignore transaction statements silently; transactions need InnoDB.

Contents This chapter on its own page

munotes.in206

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Report or request
Done!