munotes®

Software Testing and Quality Assurance Notes | B.Sc. (Computer Science) Semester 5 | Mumbai University | munotes

Get access to whole semester resourcesSemester Pass

Official Notes munotes.in

Software Testing and Quality Assurance

B.SC. (COMPUTER SCIENCE) · SEMESTER 5

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Computer Science)

For B.Sc. (Computer Science) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in Third Year

Software Testing and Quality Assurance

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I Testing fundamentals and the test process, the SDLC and quality factors, quality, QA, QC, QM and SQA, verification and validation with reviews, inspection and walkthroughs, and the testing strategies from unit to system testing, debugging and test automation

  1. What Software Testing Is 1
  2. Errors, Faults and Failures 8
  3. Why Software Must Be Tested 15
  4. The Seven Principles of Testing 23
  5. The Basic Test Process 29
  6. The Software Testing Life Cycle, Phase by Phase 35
  7. What a Test Case Is, and What Makes a Good One 40
  8. Test Design Techniques: The Three Families 46
  9. Test Execution 51
  10. Test Reporting 57
  11. Writing a Test Plan 62
  12. The Test Documents, From Design to Completion 68
  13. What a Software Development Life Cycle Is 73
  14. The Waterfall Model, and What Royce Actually Said 77
  15. The V-Model: A Test Level for Every Phase 82
  16. Iterative, Incremental and Spiral Models 86
  17. Agile, Scrum and DevOps: Testing in Short Cycles 91
  18. The Role of Testing in Each Phase 97
  19. Software Quality Factors: McCall's Model 102
  20. From McCall to ISO/IEC 25010: The Quality Model Today 106
  21. How Quality Factors Shape Testing 111
  22. What Quality Means 116
  23. Quality in Software Development 120
  24. Quality Control and Quality Assurance 129
  25. Quality Management and Software Quality Assurance 135
  26. Verification and Validation, and Why Both Matter 140
  27. The Kinds of V&V: Static and Dynamic Mechanisms 146
  28. Software Reviews: The Process, the Roles and the Types 152
  29. Inspection 158
  30. Walkthrough, and How It Differs From an Inspection 165
  31. A Strategic Approach to Software Testing 171
  32. Test Levels and Test Types 177
  33. Unit Testing: What It Is For 183
  34. Unit Testing Techniques: Drivers, Stubs and Test Doubles 188
  35. Writing Unit Tests With a Framework 193
  36. Unit Testing Best Practices 198
  37. Integration Testing: Why Units That Work Can Fail Together 204
  38. Top-Down, Bottom-Up and Sandwich Integration 209
  39. Regression Testing, Smoke Testing and Continuous Integration 215
  40. The Challenges of Integration Testing 221
  41. Validation Testing 226
  42. Acceptance Testing: Alpha, Beta and User Acceptance 231
  43. System Testing: The Whole System, End to End 236
  44. Recovery, Security, Stress, Performance and Deployment Testing 241
  45. Load Testing: Users, Ramp-Up, Throughput and Error Rate 247
  46. Cross-Browser and Compatibility Testing 252
  47. Debugging: From a Failure Back to Its Fault 257
  48. Test Automation: What to Automate and What Not To 262
  49. Driving a Browser: WebDriver, Locators and Waits 267
  50. Data-Driven Testing 273
  51. The Page Object Model 278

Module II Black-box, white-box and experience-based test design, software metrics and complexity, defect management, the quality movement, SQA and reliability, ISO 9000, formal technical reviews, quality costs and the seven basic quality tools

  1. Black-Box and White-Box Testing 283
  2. Specification-Based Testing 292
  3. Equivalence Partitioning 300
  4. Boundary Value Analysis 308
  5. Decision Table Testing 315
  6. State Transition Testing 323
  7. Structural Testing: Seeing Inside the Code 330
  8. Statement Testing and Statement Coverage 337
  9. Branch Testing and Branch Coverage 343
  10. Experience-Based Testing 350
  11. Error Guessing 355
  12. Exploratory Testing 360
  13. Checklist-Based Testing 366
  14. What a Software Metric Is, and Why Measure 372
  15. Developing Metrics: Goal, Question, Metric 378
  16. Size Metrics: Lines of Code and Function Points 384
  17. Quality, Process and Test Metrics 391
  18. Object-Oriented Metrics: The CK Suite 396
  19. Cyclomatic Complexity 402
  20. Halstead's Measures and Other Complexity Metrics 408
  21. Why Complexity Matters to Testing: Basis Path Testing 415
  22. What a Defect Is 421
  23. The Defect Life Cycle 426
  24. The Defect Management Process 431
  25. Writing a Defect Report: Severity and Priority 436
  26. Tracking Defects to Closure 441
  27. A Defect Tracker at Work: Bugzilla 446
  28. Defect Metrics 452
  29. Using Defect Data to Improve the Process 458
  30. Quality Concepts: Variation, Design and Conformance 463
  31. The Quality Movement: Shewhart, Deming and Juran 472
  32. The Quality Movement: Feigenbaum, Ishikawa, Crosby and TQM 480
  33. Background Issues in Software Quality Assurance 488
  34. The Challenges in Software Quality Assurance 494
  35. SQA Activities: What the SQA Group Does 500
  36. The Software Quality Assurance Plan 506
  37. Approaches to SQA: Formal Methods, Cleanroom and Process Models 512
  38. Statistical Software Quality Assurance and Six Sigma 520
  39. Software Reliability 526
  40. Statistical Process Control: Control Charts for Software 531
  41. Measuring Software Reliability 538
  42. Improving Software Reliability 543
  43. The ISO 9000 Family and the Seven Principles 548
  44. ISO 9001: The Requirements, Certification and Software 552
  45. Why Reviews Pay: The Cost of a Late Defect 557
  46. Formal Technical Reviews 562
  47. The Benefits of Formal Technical Reviews 568
  48. Quality Improvement Methodologies: PDCA and Kaizen 572
  49. Lean, CMMI and Choosing a Methodology 577
  50. The Cost of Quality 583
  51. Using Quality Costs for Decision Making 587
  52. The Seven Basic Quality Tools 592
  53. Pareto Diagrams 597
  54. Cause-Effect Diagrams 602
  55. Scatter Diagrams 608
  56. Run Charts 612
  57. What the Examination Asks, and How to Answer It 618
munotes.in

Module I

Testing fundamentals and the test process, the SDLC and quality factors, quality, QA, QC, QM and SQA, verification and validation with reviews, inspection and walkthroughs, and the testing strategies from unit to system testing, debugging and test automation

munotes.in

Chapter One

What Software Testing Is

Syllabus topic Module 1, "Software Testing Fundamentals"

In one line

Software testing is checking a piece of software, by running it or by examining it, to find out where it goes wrong and how good it is, before its users find out for you.

In the wording a student can write in an examination: software testing is a set of activities that evaluates the quality of a software product and its work products and discovers defects in them. It may be dynamic, in which the software is executed under specified conditions and its results are observed and evaluated, or static, in which a work product such as a requirement, a design or the code is examined without being executed.

Why testing exists at all

Every program is written by people, and people make mistakes. A programmer misreads a requirement, forgets a case, types < where <= was meant, or builds exactly what was asked for when what was asked for was wrong. None of these mistakes announces itself. The program compiles, it runs, and on most inputs it may even give the right answer.

The mistake shows itself only when the wrong input arrives, and by then the software may be in front of the people who depend on it. Testing exists to move that moment earlier: to find the problem while it is still cheap to fix and nobody has been harmed by it. The next two chapters take the two halves of that sentence separately: what errors, faults and failures actually are (Chapter Two), and why software must be tested, given what a failure costs when it is not found in time (Chapter Three).

What the standards say testing is

Three definitions are worth knowing, because each stresses something different.

The software testing standard, ISO/IEC/IEEE 29119-2:2021, defines testing as a "set of activities conducted to facilitate discovery or evaluation of properties of one or more test items". A test item is simply the thing being tested; the same standard calls it a "work product to be tested", which covers a requirements document as much as a program.

IEEE 730-2014, the standard for software quality assurance processes, defines software testing more narrowly, as an "activity in which a system or component is executed under specified conditions, the results are observed or recorded, and an evaluation is made of some aspect of the system or component". That is dynamic testing: something is run.

The ISTQB Foundation syllabus (version 4.0.1, 2024), the most widely used syllabus for testers, puts it in one line: "Software testing is a set of activities to discover defects and evaluate the quality of software work products."

The definition broken down

Take IEEE 730's sentence apart and each phrase turns out to be doing work.

munotes.in1

What Software Testing Is

  1. A system or component. Testing can be aimed at a single function, a module, a whole application, or a system made of several applications. What is being tested at any moment is the test item, and choosing it is the first decision in any test.
  2. Executed under specified conditions. A test is not "try it and see". The inputs, the starting state of the data, the environment and the order of steps are decided in advance and written down, so that the test can be repeated by someone else and give the same result.
  3. The results are observed or recorded. What the software actually did is captured: the value returned, the page shown, the message printed, the record written to the database.
  4. An evaluation is made. The observed result is compared with the expected result, which was also decided in advance. A test without an expected result is only a demonstration: nobody can say whether it passed. Chapter Seven, on what a test case is, gives the expected result the attention it needs.

The ISO definition adds two things IEEE 730's does not. It speaks of discovery or evaluation, so testing is not only hunting for bugs but also measuring properties such as speed, ease of use or security. And it speaks of test items, not programs, which is what makes a review of a requirements document a form of testing too.

Two kinds of testing: dynamic and static

The standards name the two kinds precisely. ISO/IEC/IEEE 29119-2 defines dynamic testing as "testing in which a test item is evaluated by executing it", and static testing as "testing in which a test item is examined against a set of quality or other criteria without the test item being executed".

Dynamic testingStatic testing
Is the software run?YesNo
What it examinesThe behaviour of running codeRequirements, designs, code, test plans, any document
What it findsFailures, from which defects are tracedDefects directly
Typical methodsUnit, integration, system and acceptance testsReviews, inspections, walkthroughs, static analysis tools
When it can startOnce some code runsAs soon as the first document exists
Covered inModule 1 on strategies, Module 2 on techniquesModule 1 on reviews, inspection and walkthrough

The difference in the "what it finds" row matters. Dynamic testing sees a failure, the software doing the wrong thing, and someone must then work back to the defect that caused it. Static testing looks straight at the work product and sees the defect itself. Chapter Two, on errors, faults and failures, explains those words exactly.

The objectives of testing

Testing is done for more than one reason, and a test plan states which ones apply. The ISTQB syllabus lists the typical objectives as follows; the right-hand column says what each means in practice.

munotes.in2

What Software Testing Is

Objective (ISTQB v4.0.1, 1.1.1)What it means in practice
Evaluating work products such as requirements, user stories, designs, and codeReviewing documents before a line of code exists
Causing failures and finding defectsThe objective most people think of first
Ensuring required coverage of a test objectMaking sure every requirement, or every branch of the code, has been exercised
Reducing the risk level of inadequate software qualityTesting hardest where a failure would hurt most
Verifying whether specified requirements have been fulfilledChecking the product against its specification
Verifying that a test object complies with contractual, legal, and regulatory requirementsFor example, data protection or accessibility rules
Providing information to stakeholders to allow them to make informed decisionsTelling managers whether the product is ready to release
Building confidence in the quality of the test objectEvidence that it works, not only that it fails
Validating whether the test object is complete and works as expected by the stakeholdersChecking it does what users actually need

The syllabus adds that objectives vary with context: the work product, the test level, the risks, the development life cycle in use, and business factors such as time to market. A test of a hospital's dosage calculator and a test of a college canteen's menu page do not have the same objectives.

Worked example: the first test in this book

This book follows one invented system throughout, ExamReg, a college's online examination registration portal. Students log in, fill in the semester examination form, pay the fee and download the hall ticket. Every rule of ExamReg is this book's own example, not a rule of the University.

One of its rules is the late fee, and it is the whole specification the tester is given:

Form submittedLate fee
On or before the last datenone
1 to 7 days lateRs 100
8 to 15 days lateRs 500
More than 15 days lateform not accepted

A programmer writes the function below. A tester, who has not seen the code, writes five test cases from the rule alone, each with an input and an expected result, and runs them.

def late_fee(days_late):
    """Late fee in rupees for a form submitted days_late days after the last date."""
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    if days_late <= 7:
        return 100
    return 500

tests = [            # (days late, expected fee), written from the rule alone
    (0, 0),
    (3, 100),
    (7, 100),
    (8, 500),
    (15, 500),
]
for days, expected in tests:
    actual = late_fee(days)
    verdict = "pass" if actual == expected else "FAIL"
    print(f"{days:>2} days late: expected {expected:>3}, got {actual:>3}  {verdict}")
munotes.in3

What Software Testing Is

 0 days late: expected   0, got 100  FAIL
 3 days late: expected 100, got 100  pass
 7 days late: expected 100, got 100  pass
 8 days late: expected 500, got 500  pass
15 days late: expected 500, got 500  pass

Four tests pass and one fails. Look at what the failure says: a student who submits the form on time is charged Rs 100. The program runs, it gives a sensible-looking number for every input, and on four of the five inputs it is right. Nothing about it looks broken until a test asks the one question it gets wrong.

Notice also where the failing input came from. The tester chose 0 because the rule has a boundary there, between "on time" and "late". Chapter Fifty-Five, on boundary value analysis, turns that instinct into a method.

Testing and debugging are different activities

The failing test has done its job. It has shown that a defect exists. It has not said where the defect is, and finding it is a different activity with a different name. The ISTQB syllabus describes debugging after a dynamic test fails as three steps: reproduction of the failure, diagnosis (finding the defect), and fixing the defect.

Here the three steps are short. Reproduce: call late_fee(0) again and get 100 again. Diagnose: the condition days_late <= 7 is true for 0, so an on-time form falls into the Rs 100 branch; the rule's first case, "on or before the last date", was never written into the code. Fix: handle that case first.

Then two more kinds of testing follow the fix. Confirmation testing reruns the test that failed, to confirm the fix works. Regression testing reruns the tests that passed, to confirm the fix has not broken anything else. Both are run below.

def late_fee(days_late):
    """Late fee in rupees for a form submitted days_late days after the last date."""
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    if days_late <= 0:                  # on or before the last date: the missing case
        return 0
    if days_late <= 7:
        return 100
    return 500

tests = [(0, 0), (3, 100), (7, 100), (8, 500), (15, 500)]
print("confirmation test:", "pass" if late_fee(0) == 0 else "FAIL")
failed = [days for days, expected in tests if late_fee(days) != expected]
print("regression run:", len(tests) - len(failed), "of", len(tests), "pass")
confirmation test: pass
regression run: 5 of 5 pass
TestingDebugging
PurposeTo show that failures occur, or to find defects in a work productTo find the defect behind a failure and remove it
Starts fromA requirement and an expected resultA failure that has been observed
Usually done byTesters, and developers testing their own unitsThe developer who owns the code
Needs knowledge of the code?Not for black-box testingAlways
Ends withA pass or fail verdict, and a defect reportA corrected program, then confirmation and regression tests
munotes.in4

What Software Testing Is

Debugging has a chapter of its own later in Module 1 (Chapter Forty-Seven), with its techniques.

Dijkstra's warning

The fixed function passes all five tests. Is it now correct? The five tests cannot say. They checked five inputs out of every integer the function might receive, and a defect could be sitting on an input nobody tried.

Edsger Dijkstra put this limit in one sentence in his 1972 ACM Turing Lecture, "The Humble Programmer": "program testing can be a very effective way to show the presence of bugs, but is hopelessly inadequate for showing their absence."

The late-fee function shows both halves. The first test run showed the presence of a bug, effectively and cheaply. The second run showed that five tests pass, which is not the same as showing that no bug remains. What happens, for instance, if days_late is negative? The rule does not say, and the tests did not ask. Chapter Four turns this limit into the first of the seven principles of testing.

A static test of the same rule

Go back to the rule itself and read it as a tester, without running anything. It says what happens on time, 1 to 7 days late, 8 to 15 days late and more than 15 days late. It does not say what the portal should do if the date on a form is before the day the form was opened, or if the number of days is missing. Those are gaps in the requirement.

Noticing them is static testing: a work product was examined against a criterion (is every possible input covered?) without executing anything. It found a defect in the requirement, and it found it before a single line of code depended on the missing answer. The same questions asked after release would arrive as complaints.

What testing does not mean

Testing is not only running the program. Reviewing a requirement, a design or the code is testing too: static testing. The ISTQB syllabus calls the belief that testing only consists of executing tests "a common misconception".

A test that passes does not prove the program correct. It proves the program gave the expected result for that input, in that environment, on that day. Dijkstra's sentence is the reason.

Testing is not debugging. Testing shows that something is wrong; debugging finds out why and removes it. They are often done by different people, and they need different information.

munotes.in5

What Software Testing Is

Finding bugs is not the only objective. Measuring quality, checking compliance, building confidence and giving managers the information to decide on a release are objectives too. A test run in which nothing fails is still useful: it is evidence.

A test without an expected result is not a test. Running the program and looking at what comes out, with nothing decided in advance to compare it with, tells nobody whether it passed.

Where testing stops

Testing samples. It runs some inputs out of an enormous number, so it can reduce the risk that a defect reaches users but can never remove it. It also depends on someone knowing the right answer: if the expected result is itself wrong, a correct program fails the test and a wrong one passes. And testing costs time and money, so every project decides how much testing is enough, a question Module 1 returns to in its strategic approach to testing, which asks when testing is complete (Chapter Thirty-One).

Quick revision

  • Software testing: a set of activities that evaluates quality and discovers defects in software and its work products.
  • ISO/IEC/IEEE 29119-2: testing is a "set of activities conducted to facilitate discovery or evaluation of properties of one or more test items".
  • Dynamic testing executes the test item; static testing examines it without executing it (reviews, static analysis).
  • A test needs an input, specified conditions and an expected result; the actual result is compared with it.
  • ISTQB objectives include evaluating work products, finding defects, coverage, reducing risk, verifying requirements, compliance, informing decisions, building confidence and validating.
  • Testing shows failures; debugging reproduces, diagnoses and fixes; then confirmation testing reruns the failed test and regression testing reruns the rest.
  • Dijkstra (1972): testing can show the presence of bugs, never their absence.
  • Worked example: ExamReg's late_fee(0) returned 100 instead of 0; one test in five found it.

Test yourself

1. Define software testing and name its two kinds. Software testing is a set of activities that evaluates the quality of software and its work products and discovers defects in them. Dynamic testing executes the test item under specified conditions and evaluates the results; static testing examines a work product, such as a requirement or the code, without executing it.

2. Why must a test case have an expected result decided in advance? Because a test is an evaluation: the actual result is compared with the expected one. Without an expected result nobody can say whether the test passed, and a wrong output would be accepted simply because it looked reasonable.

3. Differentiate testing from debugging. Testing shows that failures occur, or finds defects directly in a work product; it starts from a requirement and ends with a pass or fail verdict. Debugging starts from an observed failure, reproduces it, diagnoses the defect and fixes it; it is followed by confirmation testing and regression testing.

munotes.in6

What Software Testing Is

4. In the late-fee example, which test failed, and what did the fix require besides changing the code? The test with 0 days late failed: the function returned Rs 100 instead of 0. After the fix, confirmation testing reran that test and regression testing reran the four tests that had passed, to show the fix broke nothing else.

5. Explain Dijkstra's statement about testing, with an example. Testing can show that bugs are present but cannot show that they are absent, because it checks only the inputs actually tried. In the late-fee example, five passing tests say nothing about a negative number of days, which no test tried and the rule does not cover.

6. Is reading a requirements document and finding a gap in it testing? Justify. Yes. It is static testing: a work product is examined against a criterion, here completeness, without anything being executed. It finds the defect directly and earlier than any dynamic test could.

7. State any four objectives of testing. Any four of: evaluating work products; causing failures and finding defects; ensuring required coverage; reducing the risk of inadequate quality; verifying that requirements are fulfilled; verifying compliance with contractual, legal and regulatory requirements; providing information for decisions; building confidence; validating that the product works as stakeholders expect.

Contents This chapter on its own page

munotes.in7

Chapter Two

Errors, Faults and Failures

Syllabus topic Module 1, "Software Testing Fundamentals: Nature of errors"

In one line

A person makes a mistake, the mistake leaves a flaw in the software, and the flaw, when the software runs into it, makes the software do the wrong thing.

In the wording a student can write in an examination: a human error (a mistake) produces a defect (also called a fault or a bug) in a work product, and a defect, when executed, may cause a failure: an observable departure of the system's behaviour from what is required. The fundamental reason behind the error is its root cause.

Why the three words are kept apart

In ordinary speech "error", "bug" and "failure" are used as if they were one thing. In testing they are three different things that happen at three different times, to three different subjects: a person errs, a work product is defective, a running system fails. Keeping them apart is not pedantry. It decides who does what next.

A failure is what a tester or a user sees. A defect is what a developer must find and fix. An error, and the root cause behind it, is what a team must understand if the same kind of defect is not to appear again next month. A report that says only "the fee page has an error" leaves all three questions open.

The definitions

The software engineering vocabulary standard, ISO/IEC/IEEE 24765:2017, defines a mistake as a "human action that produces an incorrect result". The ISTQB syllabus uses error and mistake as the same word: "Human beings make errors (mistakes), which produce defects (faults, bugs), which in turn may result in failures."

A defect is an "imperfection or deficiency in a work product where that work product does not meet its requirements or specifications and needs to be either repaired or replaced" (ISO/IEC 23531:2020). The words fault and bug mean the same thing in everyday testing, and the ISTQB sentence above gives them as alternatives.

A failure is an "event in which any part of a system or system element does not perform as required by its specification" (INCOSE Systems Engineering Handbook, 2023). IEEE 982:2024 says it in six words: "departure of system behavior from system requirements".

A root cause is the "source of a defect such that if it is removed, the defect is decreased or removed" (ISO/IEC/IEEE 24765:2017).

The chain, on the example from Chapter One

Chapter One, on what software testing is, ended with a real example of all four, and it is worth naming each link.

LinkIn the late-fee example
Root causeThe programmer was working from memory of the rule, under a deadline, and never re-read the table
Error (mistake)The programmer assumed the first case of the rule was "up to 7 days late" and forgot "on or before the last date"
Defect (fault, bug)The function has no branch for zero days late, so days_late <= 7 catches it
FailureA student who submits on time is shown a late fee of Rs 100
munotes.in8

Errors, Faults and Failures

Read the table from top to bottom and it is a story in time: the root cause came first, then the mistake, then the defect sat in the code for as long as nobody ran it with zero, and then the failure happened on the first day a student submitted on time.

A defect does not always cause a failure

The ISTQB syllabus is careful here: "Some defects will always result in a failure if executed, while others will only result in a failure in specific circumstances, and some may never result in a failure." The late-fee defect fails every time the input is zero. The next program shows the second kind.

A leap-year check is written with only the first half of the calendar rule: a year divisible by 4 is a leap year. The full Gregorian rule adds that a year divisible by 100 is not, unless it is also divisible by 400. Python's standard library already implements the full rule as calendar.isleap, so it serves as the oracle, the source of the right answer.

import calendar

def is_leap(year):
    return year % 4 == 0          # the defect: the rule for century years is missing

for first, last in [(1901, 2099), (1600, 2400)]:
    wrong = [y for y in range(first, last + 1) if is_leap(y) != calendar.isleap(y)]
    print(f"years {first} to {last}: {last - first + 1} tried, "
          f"{len(wrong)} failure(s) {wrong}")
years 1901 to 2099: 199 tried, 0 failure(s) []
years 1600 to 2400: 801 tried, 6 failure(s) [1700, 1800, 1900, 2100, 2200, 2300]

The defect is present on every run. It causes no failure at all for any year from 1901 to 2099, which includes every year a student's date of birth or an examination could fall in. Over six centuries it fails six times. Tested only on this century's years, the function would pass every test anyone thought to write, and the defect would wait for 2100.

Two lessons follow. A defect's presence and its failure are different facts: testing observes failures, and a defect that cannot be made to fail by the inputs tried stays invisible. And the choice of test inputs decides what can be seen, which is why Module 2 spends a whole row of the syllabus on choosing them (Chapter Fifty-Four onwards, from equivalence partitioning).

Failures without defects, and defects without code

Not every failure comes from a defect in the code. The ISTQB syllabus notes that "Failures can also be caused by environmental conditions, such as when radiation or electromagnetic fields cause defects in firmware." A disk that is full, a network that drops, a clock set to the wrong year or a server configured differently from the test machine can all make correct code fail.

munotes.in9

Errors, Faults and Failures

And not every defect is in the code. Defects "can be found in documentation, such as a requirements specification or a test script, in source code, or in a supporting work product such as a build file." The missing answer for a negative number of days, found while testing the late-fee rule in Chapter One, was a defect in the requirement. The syllabus adds the reason this matters so much: defects in work products produced early, "if undetected, often lead to defective work products later in the lifecycle." One gap in a requirement becomes a gap in the design, then in the code, then in the tests written from the same requirement.

Where errors come from

People do not make mistakes at random. The ISTQB syllabus lists typical reasons: "time pressure, complexity of work products, processes, infrastructure or interactions, or simply because they are tired or lack adequate training." To those every project adds a few of its own:

  • Miscommunication: the person who knows the rule is not the person who codes it, and the rule loses a clause on the way.
  • Changing requirements: the code was right for last month's rule.
  • Unfamiliar technology: a new framework or language whose defaults the programmer does not yet know.
  • Complex interactions: two modules each correct, whose combination was never thought through (Chapter Thirty-Seven, on integration testing, is about exactly this).

These are the root causes that root cause analysis looks for, and Module 2 returns to them in the chapter on using defect data to improve the process (Chapter Eighty).

The word "bug"

Students are often told that Grace Hopper invented the word "bug" when a moth was found in a computer. The object is real, and it is in the Smithsonian's National Museum of American History, which catalogues it as a 1947 log book from the Mark II computer at Harvard University. The museum's own record says the engineers "taped the insect in their logbook and labeled it 'first actual case of bug being found.'"

The same record corrects the legend in three ways. The word was older: "Thomas Edison talked about bugs in electrical circuits in the 1870s." The joke in the label depends on that: it was the first actual bug, a literal insect, in a machine whose faults were already called bugs. And of the log book itself the museum says it "was probably not Hopper's", though Hopper and the Mark II team "helped popularize the use of the term computer bug and the related phrase 'debug.'"

munotes.in10

Errors, Faults and Failures

Other words for the same family

Standards and tools use several more words around a defect. Each has its own shade of meaning, and Module 2's defect management row gives them their full treatment (Chapter Seventy-Three, on what a defect is).

WordStandard definitionShade of meaning
Anomaly"anything observed in the documentation or operation of a system that deviates from expectations based on previously verified system, software, or hardware products or reference documents" (IEEE 1012-2024)Something looks wrong, but it is not yet known whether it is a defect
Incident"anomalous or unexpected event or set of events at any time during the life cycle of a project, product, service, or system" (ISO/IEC/IEEE 12207:2026)The event that was reported; a test incident may turn out to be a defect, a wrong test or an environment problem
Issue"observation that deviates from expectations" (ISO/IEC 20246:2017, the review standard)What a reviewer records during a review
Error, second sense"discrepancy between a computed, observed, or measured value or condition and the true, specified, or theoretically correct value or condition" (ISO/IEC/IEEE 15026-1:2025)A wrong VALUE, not a human mistake

Note the last row. "Error" has two meanings in the standards: the human mistake of the ISTQB sentence, and a wrong value inside a running system. A numerical analyst's "rounding error" is the second meaning. When a textbook says "an error occurred", read the sentence around it to see which is meant.

Classifying defects

A defect that is only counted teaches nothing. Classified, a collection of defects shows where a process is weak. Three classifications are in common use and a defect gets one value in each.

By where it was introduced. Requirements, design, code, documentation, or a supporting item such as a build or configuration file. This tells a team which activity is leaking.

By type. IBM's Orthogonal Defect Classification (Chillarege and six colleagues, IEEE Transactions on Software Engineering, 1992) defines the types by the kind of correction the programmer makes, and distinguishes in each case "between something missing or something incorrect". Its eight types, in the paper's own terms:

ODC defect typeWhat the correction involves
FunctionSignificant capability, end-user or product interfaces, or global data structures; needs a formal design change
AssignmentA few lines of code, such as initialising a variable or a data structure
InterfaceInteraction with other components, modules or drivers through calls, macros, control blocks or parameter lists
CheckingProgram logic that failed to validate data and values before they were used
Timing/serializationManagement of shared and real-time resources
Build/package/mergeLibrary systems, change management or version control
DocumentationPublications and maintenance notes
AlgorithmEfficiency or correctness problems fixed by reimplementing an algorithm or a local data structure, without a design change
munotes.in11

Errors, Faults and Failures

By severity. How bad the failure is for the user, from a crash or data loss down to a cosmetic flaw. Severity, and priority, which is a different thing, are taught with the defect report in Module 2 (Chapter Seventy-Six).

Worked example: classifying three ExamReg defects

Three defects are found in one week on ExamReg. Each is classified below, with the reason, because the reason is what an examiner looks for.

DefectIntroduced inODC typeReason
On-time forms charged Rs 100 (the late-fee test of Chapter One)CodeAlgorithm, missingThe computation lacked a case; the fix adds one branch, with no design change
A negative number of days late is accepted and charged Rs 100Requirements, then codeChecking, missingThe value is never validated before it is used; the rule never said what to do
The fee page calls the payment module with the amount in paise, and it expects rupeesCodeInterface, incorrectTwo modules disagree about a parameter passed between them

The third row is worth a second look. Each module could pass its own tests perfectly, and the pair would still charge a student one hundred times too much. It is a defect of the kind integration testing exists to catch.

When the test is wrong: false positives and false negatives

A test result can itself be wrong, in two ways. If a tester misreads the rule and writes that 7 days late should cost Rs 500, a correct function fails the test: the test reports a defect that does not exist. If the expected result for 0 days had been written as Rs 100, copying the code instead of the rule, the defective function would pass: a real defect goes unreported.

def late_fee(days_late):              # the CORRECTED function from Chapter One
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    if days_late <= 0:
        return 0
    if days_late <= 7:
        return 100
    return 500

expected_from_a_misread_rule = {7: 500}      # the tester's mistake, not the program's
for days, expected in expected_from_a_misread_rule.items():
    actual = late_fee(days)
    print(f"{days} days late: expected {expected}, got {actual}:",
          "pass" if actual == expected else "FAIL, but the program is right")
7 days late: expected 500, got 100: FAIL, but the program is right

The names for these two cases are a trap, because two current standards use them in opposite senses. ISO/IEC 23643:2020 defines a false positive as an "observed defect which does not correspond to a true defect" and a false negative as a "true defect that has not been observed". ISO/IEC TR 29119-11:2020, on testing AI systems, defines a false positive as "incorrect reporting of a pass when in reality it is a failure", which is the reverse. This book uses the first sense, which is also the ISTQB's: a false positive reports a defect that is not there, and a false negative misses one that is. In an answer, define the term before you use it and no examiner can mark you wrong.

munotes.in12

Errors, Faults and Failures

What it does not mean

Every defect does not cause a failure. Some fail every time, some only on particular inputs, and some never. The leap-year defect above failed six times in eight hundred years.

Every failure is not caused by a defect in the code. The environment, a configuration, the data, or a wrong expected result in the test can all produce a failure.

Defects are not only in code. Requirements, designs, test scripts, user manuals and build files can all be defective, and early ones propagate.

"Error" does not always mean a human mistake. In ISO/IEC/IEEE 15026-1 it means a wrong value inside the system.

Grace Hopper did not coin "bug". The Smithsonian's record dates the word to Edison in the 1870s; the 1947 log book is famous because its bug was a real moth.

Quick revision

  • Error (mistake): "human action that produces an incorrect result" (ISO/IEC/IEEE 24765:2017).
  • Defect (fault, bug): an imperfection in a work product that fails its requirements and must be repaired or replaced (ISO/IEC 23531:2020).
  • Failure: the system not performing as required; "departure of system behavior from system requirements" (IEEE 982:2024).
  • Root cause: the source whose removal removes or reduces the defect (ISO/IEC/IEEE 24765:2017).
  • Chain: root cause, error, defect, failure. A defect may fail always, sometimes or never; failures can also come from the environment.
  • Causes (ISTQB): time pressure, complexity, fatigue, lack of training; plus miscommunication and changing requirements.
  • Classify by origin, by ODC type (function, assignment, interface, checking, timing/serialization, build/package/merge, documentation, algorithm; missing or incorrect) and by severity.
  • Leap-year example: no failure from 1901 to 2099, six from 1600 to 2400.
  • False positive (ISTQB sense): a defect reported that is not there. False negative: a real defect not reported.

Test yourself

1. Distinguish error, defect and failure with one example. An error is a human mistake, such as a programmer forgetting the on-time case of a fee rule. It leaves a defect in the work product: the function has no branch for zero days late. When the function runs with zero, the system fails: an on-time student is charged Rs 100.

2. Can a program contain a defect and never fail in testing? Explain. Yes. A defect causes a failure only when the inputs and conditions reach it. The leap-year function missing the century rule gives the right answer for every year from 1901 to 2099, so tests drawn from this century all pass while the defect remains.

munotes.in13

Errors, Faults and Failures

3. Give two causes of failure that are not defects in the code. Environmental conditions, such as radiation or electromagnetic fields affecting firmware, a full disk or a misconfigured server; and a wrong expected result in the test itself, which makes a correct program appear to fail.

4. What is a root cause, and why is finding it worth the effort? It is the source of a defect whose removal removes or reduces the defect, such as time pressure or an ambiguous requirement. Fixing only the defect repairs one instance; removing the root cause prevents the same kind of defect from being made again.

5. Name the eight ODC defect types. Function, assignment, interface, checking, timing/serialization, build/package/merge, documentation and algorithm, each recorded as missing or incorrect.

6. Why must you define "false positive" before using it in an answer? Because standards disagree: ISO/IEC 23643:2020 uses it for a reported defect that does not exist, while ISO/IEC TR 29119-11:2020 uses it for a pass reported when the result was really a failure. Defining it removes the ambiguity.

Contents This chapter on its own page

munotes.in14

Chapter Three

Why Software Must Be Tested

Syllabus topic Module 1, "Software Testing Fundamentals: Nature of errors and the need for testing"

In one line

Software must be tested because software that fails costs money, time and trust, and sometimes lives, and testing is the cheapest way to find a failure before it reaches the people who depend on the software.

In the wording a student can write in an examination: testing is necessary because defects are inevitable in human work, because a defect found late costs far more to fix than one found early, because failures in operation cause financial, legal, reputational and human harm, and because testing provides the evidence on which a release decision, a contract or a regulator relies. Testing reduces the risk of failure in operation; it cannot remove it.

Why the question is worth asking

Every project is under pressure to ship. Testing takes time, needs people, and produces reports that nobody enjoys reading. It is tempting to treat it as the part that can be squeezed. The ISTQB syllabus puts the other side of the argument plainly: software that does not work correctly "can lead to many problems, including loss of money, time or business reputation, and, in extreme cases, even injury or death."

The honest way to answer "why test?" is not with a slogan but with what actually happened when testing was missing. The four failures below are among the best documented in the history of software, and each is told here from the official investigation, not from retellings, because retellings drift.

Ariane 5, 4 June 1996: code reused without the test that mattered

The European Space Agency's inquiry board opens its report with the event: on 4 June 1996 the maiden flight of the Ariane 5 launcher ended in failure, and "Only about 40 seconds after initiation of the flight sequence, at an altitude of about 3700 m, the launcher veered off its flight path, broke up and exploded."

The cause, as the board found it, was in the inertial reference system (the SRI), whose software was reused from Ariane 4. An internal exception "was caused during execution of a data conversion from 64-bit floating point to 16-bit signed integer value." The number being converted, a horizontal bias value called BH, was larger than a 16-bit signed integer can hold. It was larger because "the early part of the trajectory of Ariane 5 differs from that of Ariane 4 and results in considerably higher horizontal velocity values." And the function doing the conversion was not even needed: it performed alignment, and "As soon as the launcher lifts off, this function serves no purpose."

The conversion is easy to reproduce. A 16-bit signed integer holds values from -32768 to 32767, and Python's struct module refuses to pack anything larger into that format, which is exactly the kind of refusal the SRI's software raised.

munotes.in15

Why Software Must Be Tested

import struct

def to_16_bit(value):
    """Pack a number into a 16-bit signed integer, as the SRI's conversion did."""
    return struct.pack(">h", int(value))

for horizontal_bias in [1250.0, 32767.0, 40000.0]:
    try:
        to_16_bit(horizontal_bias)
        print(f"{horizontal_bias:>9}: fits in 16 bits")
    except struct.error as problem:
        print(f"{horizontal_bias:>9}: conversion refused ({problem})")
   1250.0: fits in 16 bits
  32767.0: fits in 16 bits
  40000.0: conversion refused ('h' format requires -32768 <= number <= 32767)

The three values are this book's, chosen to straddle the limit; the report does not publish the flight's value of BH. What the board does say about testing is the lesson. Equipment testing of the SRI was rigorous, "However, no test was performed to verify that the SRI would behave correctly when being subjected to the count-down and flight time sequence and the trajectory of Ariane 5." And then: "Had such a test been performed by the supplier or as part of the acceptance test, the failure mechanism would have been exposed." The test was never written because the SRI's specification did not contain the Ariane 5 trajectory data as a requirement: a defect in a requirement, exactly the kind Chapter Two, on errors, faults and failures, warned propagates.

The Patriot at Dhahran, 25 February 1991: an error too small to see, run for too long

The United States General Accounting Office (GAO) report IMTEC-92-26 begins: "On February 25, 1991, a Patriot missile defense system operating at Dhahran, Saudi Arabia, during Operation Desert Storm failed to track and intercept an incoming Scud. This Scud subsequently hit an Army barracks, killing 28 Americans."

The Patriot's clock counted time "in tenths of seconds", and to predict where a Scud would appear next, the count had to be converted to a real number. Its registers were 24 bits long, so the conversion "cannot be any more precise than 24 bits." One tenth has no exact binary form, so every tick of the clock was stored very slightly short, and the shortfall accumulated for as long as the system ran. At Dhahran the battery "had been operating continuously for over 100 hours."

How short, exactly? GAO says only that the registers were 24 bits long; it does not print the format in which the tenth was held. So the program below tries two readings and lets GAO's own published figures decide. In the first, the tenth is kept to 24 binary places after the point and the rest chopped off. In the second, it is kept to 23, which is what a 24-bit register leaves for the fraction if one bit holds the sign. The figures GAO printed in its Appendix II are typed in beside the program, from the page image of the report.

munotes.in16

Why Software Must Be Tested

from fractions import Fraction

tenth = Fraction(1, 10)

def drift(places, hours):
    """Seconds lost after `hours` if each tenth is chopped to `places` binary places."""
    stored = Fraction(int(tenth * 2**places), 2**places)
    return hours * 3600 * 10 * (tenth - stored)

for places in (24, 23):
    print(f"kept to {places} binary places: drift after 100 h = {float(drift(places, 100)):.4f} s")
print()

# GAO/IMTEC-92-26, Appendix II: hours, calculated time, inaccuracy (seconds)
gao = [(1, "3599.9966", "0.0034"), (8, "28799.9725", "0.0275"),
       (20, "71999.9313", "0.0687"), (48, "172799.8352", "0.1648"),
       (72, "259199.7528", "0.2472"), (100, "359999.6667", "0.3433")]
for hours, gao_time, gao_drift in gao:
    ours = f"{float(drift(23, hours)):.4f}"
    check = hours * 3600 - Fraction(gao_drift)
    note = "" if f"{float(check):.4f}" == gao_time else f"  <- GAO's time should be {float(check):.4f}"
    print(f"{hours:>3} h: drift {ours} s, GAO printed {gao_drift}{note}")
kept to 24 binary places: drift after 100 h = 0.1287 s
kept to 23 binary places: drift after 100 h = 0.3433 s

  1 h: drift 0.0034 s, GAO printed 0.0034
  8 h: drift 0.0275 s, GAO printed 0.0275
 20 h: drift 0.0687 s, GAO printed 0.0687
 48 h: drift 0.1648 s, GAO printed 0.1648
 72 h: drift 0.2472 s, GAO printed 0.2472
100 h: drift 0.3433 s, GAO printed 0.3433  <- GAO's time should be 359999.6567

The first reading gives a drift of about an eighth of a second at 100 hours, which is not what GAO measured, so it is rejected. The second reproduces every figure in GAO's inaccuracy column, to four decimal places: after 100 hours the clock was a third of a second out. In a third of a second a missile flying at "approximately MACH 5" covers hundreds of metres, and GAO's table gives the shift of the radar's range gate as 687 metres, enough for the system to "look in the wrong place for the incoming Scud."

Two small lessons sit inside this one. A model is tested against data, and the first one here failed its test: that is how the 23 was found, not assumed. And the program found something in the report itself. In GAO's last row the calculated time (359999.6667) and the inaccuracy (.3433) cannot both be right, because 360000 minus .3433 is 359999.6567. The inaccuracy column is the one the drift model confirms. It is a small slip in a careful report, and it is why a table of figures is worth recomputing rather than trusting.

The report's account of the timing is the sharpest lesson. Israeli data had shown a loss of accuracy after 8 consecutive hours, and "Army officials modified the software to improve the system's accuracy. However, the modified software did not reach Dhahran until February 26, 1991", the day after the Scud struck. No test had run the system for as long as it was actually used. GAO notes that the Patriot had never before been used against Scud missiles "nor was it expected to operate continuously for long periods of time."

munotes.in17

Why Software Must Be Tested

Therac-25, 1985 to 1987: software trusted to be safe

Nancy Leveson's account "Medical Devices: The Therac-25", drawn from her investigation with Clark S. Turner (IEEE Computer, July 1993), begins: "Between June 1985 and January 1987, a computer-controlled radiation therapy machine, called the Therac-25, massively overdosed six people." Eleven machines had been installed, five in the United States and six in Canada.

The design had shifted safety from hardware to software. Earlier machines in the family had "independent protective circuits" and "mechanical interlocks for policing the machine", but "In the Therac-25, software checks were substituted for many of the traditional hardware interlocks." Code was reused from the earlier machines too, and one of the bugs was later found in the older Therac-20's software, where it had hurt nobody because that machine "includes hardware safety interlocks."

Two software defects stand out, and each was invisible to ordinary testing. At Tyler, Texas, the overdose happened only when an experienced operator edited the prescription quickly: "If the prescription data was edited at a fast pace (as is natural for someone who has repeated the procedure a large number of times), the overdose occurred." The hospital physicist had to practise before he could reproduce it, and the manufacturer's engineer at first "could not reproduce the error" at all. At Yakima, Washington, a one-byte variable called Class3 was "incremented by one in each pass" through a setup routine, and a zero value meant a safety check was skipped: "on every 256th pass through the Set Up Test code, the variable will overflow and have a zero value."

The second defect takes four lines to reproduce.

def skipped_checks(passes, fixed=False):
    """Passes on which the upper collimator check is skipped."""
    class3, skipped = 0, []
    for pass_number in range(1, passes + 1):
        class3 = 1 if fixed else (class3 + 1) % 256   # one byte: 255 + 1 wraps to 0
        if class3 == 0:                                 # zero is read as "nothing to check"
            skipped.append(pass_number)
    return skipped

print("as shipped:", skipped_checks(1000))
print("as fixed  :", skipped_checks(1000, fixed=True))
as shipped: [256, 512, 768]
as fixed  : []

The check is skipped on three passes in a thousand, and only when the operator happens to press the button on exactly one of them. The fix, as the account reports it, was "simple: the program is changed so that the Class3 variable is set to some fixed nonzero value each time through Set Up Test instead of being incremented." A defect with a one-line fix killed and injured people because no test drove the setup routine through hundreds of passes, and none drove the keyboard as fast as a practised operator did.

munotes.in18

Why Software Must Be Tested

Knight Capital, 1 August 2012: one server missed, and dead code woken

The United States Securities and Exchange Commission's order (Release 70694, 2013) records that on August 1, 2012 Knight Capital's automated order router, SMARS, "routed millions of orders into the market over a 45-minute period, and obtained over 4 million executions in 154 stocks for more than 397 million shares." It had been processing just 212 small retail orders. "Knight lost over $460 million from these unwanted positions."

The chain of events is a lesson in two kinds of testing at once. New code was deployed to replace an old, unused function called Power Peg, and the new code "repurposed a flag that was formerly used to activate the Power Peg code." During deployment "one of Knight's technicians did not copy the new code to one of the eight SMARS computer servers," and "Knight did not have a second technician review this deployment". On the eighth server the repurposed flag woke the old code. And that old code had itself been changed years earlier without a retest: in 2005 a function it relied on was moved, and "Knight did not retest the Power Peg code after moving the cumulative quantity function to determine whether Power Peg would still function correctly if called."

Two things were missing: a regression test after a change to old code, and a review (a static check) of the deployment. Chapter Thirty-Nine, on regression testing, smoke testing and continuous integration, and Chapter Twenty-Eight, on software reviews, are about exactly these.

What failures cost a whole economy

Famous disasters are rare; ordinary failures are constant, and their cost adds up. In 2002 the United States National Institute of Standards and Technology (NIST) surveyed software developers and users and estimated that "the national annual cost estimates of an inadequate infrastructure for software testing are estimated to be $59.5 billion." It estimated that "The potential cost reduction from feasible infrastructure improvements is $22.2 billion." Software users bore about 60 per cent of the cost and developers about 40 per cent.

The figures are more than twenty years old and were for one country, so the exact amounts are history. The shape of the finding is what lasts: most of the cost of poor testing is paid by users, not by the people who wrote the software, and a large part of it could be avoided by testing better and earlier.

Why late is expensive

The failures above were found in operation, the most expensive place to find anything. Barry Boehm and Victor Basili, in their 2001 "Software Defect Reduction Top 10 List", put the first item on the list this way: "Finding and fixing a software problem after delivery is often 100 times more expensive than finding and fixing it during the requirements and design phase."

munotes.in19

Why Software Must Be Tested

The reason is easy to see on ExamReg. A missing case in the late-fee rule, caught in a review of the requirement, costs a sentence. Caught in testing, it costs a code change and a retest. Caught after release, it costs refunds to students who were overcharged, a correction to every receipt, a public apology and a patch deployed under pressure, which is exactly the situation in which Knight Capital's technician missed a server. Chapter Ninety-Six returns to this with a defect amplification model and why reviews pay.

What testing contributes

The ISTQB syllabus lists what testing contributes to a project's success, and it is more than finding bugs:

  • It "provides a cost-effective means of detecting defects", which debugging then removes, so testing contributes indirectly to higher quality.
  • It "provides a means of directly evaluating the quality of a test object", and those measures feed decisions such as whether to release.
  • It "provides users with indirect representation on the development project", because testers carry the users' needs into a project users cannot join.
  • It "may also be required to meet contractual or legal requirements, or to comply with regulatory standards."

The syllabus also separates testing from quality assurance: testing is "a product-oriented, corrective approach" and "a major form of quality control", while QA is "a process-oriented, preventive approach". That distinction is Module 1's third row, and Chapter Twenty-Four, on quality control and quality assurance, gives it in full.

The four failures, side by side

Ariane 5 (1996)Patriot (1991)Therac-25 (1985 to 1987)Knight Capital (2012)
What failedLauncher destroyed about 40 seconds into flightMissed an incoming Scud; 28 killedRadiation overdosesMillions of unintended orders in 45 minutes
The defectUnprotected 64-bit to 16-bit conversion in reused codeTime held in 24-bit registers; drift grew with running timeA race on fast editing; a one-byte counter skipping a check every 256th passOld code woken by a repurposed flag on one server
Test that was missingTest with Ariane 5's own trajectoryTest of continuous operation for as long as real useTests at an expert operator's speed, and of hundreds of setup passesRegression test after a change; review of the deployment
Kind of testing, in this bookSystem and acceptance testing with realistic dataStress and endurance testingStress and timing tests; safety reviewsRegression testing; reviews

What it does not mean

Testing does not guarantee a failure-free product. Every one of these systems was tested. Testing reduces risk; the missing test in each case was one that nobody thought to run.

munotes.in20

Why Software Must Be Tested

Reused code is not already tested. Ariane 5 reused Ariane 4's software, and Knight reused an old flag. Code tested in one context is untested in a new one.

Testing is not only for rockets and hospitals. NIST's estimate is dominated by ordinary business software. A portal that overcharges students is a failure in the same sense.

More testing is not automatically better. Testing costs money, and the question is always which tests reduce the most risk for the effort. Chapter Four's principle that testing is context dependent says the same.

Quick revision

  • Failures cost money, time and reputation, and in extreme cases cause injury or death (ISTQB v4.0.1).
  • Ariane 5 (4 June 1996): unprotected 64-bit to 16-bit conversion of BH in reused Ariane 4 code; no test with Ariane 5's trajectory; "Had such a test been performed ... the failure mechanism would have been exposed."
  • Patriot (25 February 1991): time converted in 24-bit registers, each tenth of a second held slightly short; after 100 hours, 0.3433 s of drift and a 687 m range-gate shift; 28 killed; the fix arrived the next day.
  • Therac-25 (June 1985 to January 1987): six people massively overdosed; software checks replaced hardware interlocks; a race on fast editing, and a one-byte counter that skipped a safety check every 256th pass.
  • Knight Capital (1 August 2012): new code missed on one of eight servers, a repurposed flag woke old code never retested; over $460 million lost in 45 minutes.
  • NIST 2002: inadequate testing infrastructure cost the US an estimated $59.5 billion a year; $22.2 billion avoidable.
  • Boehm and Basili 2001: fixing after delivery is often 100 times more expensive than in requirements and design.
  • Testing detects defects, evaluates quality, represents users and meets legal requirements; it is quality control, not quality assurance.

Test yourself

1. Give four reasons why software testing is necessary. Defects are inevitable in human work; defects found late cost far more to fix than those found early; failures in operation cause financial, legal, reputational and human harm; and testing provides the evidence for release decisions and for contractual and regulatory compliance.

2. What caused the Ariane 5 failure, and which test would have exposed it? Software reused from Ariane 4 converted a 64-bit floating-point value, the horizontal bias, into a 16-bit signed integer without protection, and Ariane 5's faster early trajectory made the value too large, raising an exception in both inertial reference units. The inquiry board found that a ground test injecting Ariane 5's own trajectory would have exposed it.

3. Explain how a tiny error made the Patriot miss the Scud at Dhahran. Its clock counted tenths of a second and converted the count to a real number with only 24 bits of precision, so each tick was stored slightly short. The shortfall grew with running time: after about 100 hours of continuous operation it reached about a third of a second, shifting the radar's range gate about 687 metres, so the system looked in the wrong place.

munotes.in21

Why Software Must Be Tested

4. What two missing checks allowed the Knight Capital failure? A review of the deployment, which would have noticed that one of eight servers had not received the new code; and a regression test of the old Power Peg code after it was changed in 2005, which was never done.

5. According to Boehm and Basili, how much more does it cost to fix a problem after delivery? Often 100 times more than finding and fixing it during the requirements and design phase.

6. Is testing the same as quality assurance? No. Testing is product-oriented and corrective, a major form of quality control; quality assurance is process-oriented and preventive, working on the basis that a good process produces a good product.

Contents This chapter on its own page

munotes.in22

Chapter Four

The Seven Principles of Testing

Syllabus topic Module 1, Course Outcome OC 1, "Explain and apply fundamental testing principles"

In one line

Seven short rules, learned the hard way by testers over fifty years, that say what testing can and cannot do and where to spend the effort.

In the wording a student can write in an examination: the seven principles of testing, as the ISTQB Foundation syllabus (v4.0.1) states them, are (1) testing shows the presence, not the absence of defects; (2) exhaustive testing is impossible; (3) early testing saves time and money; (4) defects cluster together; (5) tests wear out; (6) testing is context dependent; and (7) the absence-of-defects fallacy.

Why principles at all

Testing techniques change with every tool and every decade. The principles do not: they describe the nature of the job, not a way of doing it. A student who knows them can reason about a testing situation nobody has written a technique for. The ISTQB syllabus introduces them modestly, as "general guidelines applicable to all testing" that "have been suggested over the years."

MU asks for them in her own words. Course Outcome OC 1 of this paper says a student should be able to "Explain and apply fundamental testing principles". Each principle below is therefore stated, explained, and then applied to ExamReg, the book's invented examination portal.

No.Principle (ISTQB v4.0.1, 2024)Name in the 2018 syllabus (v3.1.1)
1Testing shows the presence, not the absence of defectsTesting shows the presence of defects, not their absence
2Exhaustive testing is impossibleExhaustive testing is impossible
3Early testing saves time and moneyEarly testing saves time and money
4Defects cluster togetherDefects cluster together
5Tests wear outBeware of the pesticide paradox
6Testing is context dependentTesting is context dependent
7Absence-of-defects fallacyAbsence-of-errors is a fallacy

The substance is the same in both versions. Two names changed, and older textbooks and question papers use the older ones, so an answer that gives both is safe.

Principle 1: testing shows the presence, not the absence of defects

The syllabus: "Testing can show that defects are present in the test object, but cannot prove that there are no defects." And: "even if no defects are found, testing cannot prove test object correctness."

This is Dijkstra's point, met in Chapter One on what software testing is, and the leap-year function of Chapter Two on errors, faults and failures is its clearest demonstration: 199 consecutive years tested, no failure, and a defect still present. What testing does achieve is stated in the same place: it "reduces the probability of defects remaining undiscovered".

Applied to ExamReg. When the test lead reports that all 412 test cases passed, the honest reading is that these 412 inputs found no defect, never that ExamReg has none. A release note that promises zero bugs has confused this principle.

munotes.in23

The Seven Principles of Testing

Principle 2: exhaustive testing is impossible

The syllabus: "Testing everything is not feasible except in trivial cases". Exhaustive testing means trying every possible combination of inputs and preconditions, and the numbers grow faster than intuition expects. The program below counts two cases: a function that takes two 32-bit integers, and ExamReg's examination form.

SECONDS_PER_YEAR = 365.25 * 24 * 3600

pairs = 2**32 * 2**32                      # two 32-bit integer inputs
print(f"two 32-bit inputs: {pairs:,} combinations")
print(f"  at a billion tests a second: about {pairs / 10**9 / SECONDS_PER_YEAR:,.0f} years")

form = {                                   # ExamReg's exam form, deliberately simplified
    "semester (1 to 6)": 6,
    "backlog papers (8 tick boxes)": 2**8,
    "days late (0 to 15, or later)": 17,
    "fee concession (yes or no)": 2,
    "payment mode (4 choices)": 4,
}
combos = 1
for field, values in form.items():
    combos *= values
print(f"exam form: {combos:,} combinations")
print(f"  at one manual test a minute: about {combos / 60 / 24:,.0f} days without a break")
two 32-bit inputs: 18,446,744,073,709,551,616 combinations
  at a billion tests a second: about 585 years
exam form: 208,896 combinations
  at one manual test a minute: about 145 days without a break

Even a five-field form outruns a test team, and the form ignores the student's name, the browser, the network and the order in which fields are filled. The syllabus's answer is not to give up but to choose: "test techniques ... test case prioritization ... and risk-based testing ... should be used to focus test efforts." Module 2's techniques exist to choose a few hundred tests that stand in for the millions.

Principle 3: early testing saves time and money

The syllabus: "Defects that are removed early in the process will not cause subsequent defects in derived work products." And therefore "both static testing ... and dynamic testing ... should be started as early as possible."

The mechanism is the one Chapter Two on errors, faults and failures described: a defect in a requirement is copied into the design, the code and the tests written from that requirement. Remove it while it is one sentence and it never multiplies. Boehm and Basili's figure in Chapter Three, on why software must be tested, puts the price of waiting: a fix after delivery is "often 100 times more expensive" than one during requirements and design. This principle is also called shift left, because on a time line drawn left to right it moves testing towards the start.

Applied to ExamReg. The gap in the late-fee rule (what happens with a negative number of days?) was found by reading the rule. Found there, it cost one question to the exam cell. Found after release, it would have cost refunds and a patch.

munotes.in24

The Seven Principles of Testing

Principle 4: defects cluster together

The syllabus: "A small number of system components usually contain most of the defects discovered or are responsible for most of the operational failures." It calls this "an illustration of the Pareto principle", the observation that a few causes account for most of an effect. ExamReg's release 2.0 had 200 defects, and the program sorts them by the module they were found in.

defects = {"Exam form": 74, "Fee payment": 58, "Admin reports": 20, "Hall ticket": 18,
           "Login": 12, "Profile": 10, "Notifications": 8}
total = sum(defects.values())
running = 0
for module, count in sorted(defects.items(), key=lambda kv: -kv[1]):
    running += count
    print(f"{module:<14}{count:>4}  {100 * count / total:5.1f}%   cumulative {100 * running / total:5.1f}%")
Exam form       74   37.0%   cumulative  37.0%
Fee payment     58   29.0%   cumulative  66.0%
Admin reports   20   10.0%   cumulative  76.0%
Hall ticket     18    9.0%   cumulative  85.0%
Login           12    6.0%   cumulative  91.0%
Profile         10    5.0%   cumulative  96.0%
Notifications    8    4.0%   cumulative 100.0%

Two modules out of seven hold two thirds of the defects. The two are not an accident: the exam form holds every eligibility and paper-selection rule, and the fee module every calculation and the link to the payment gateway. Complexity attracts defects. The syllabus draws the practical conclusion: known and predicted clusters "are an important input for risk-based testing". The next release's test effort should go where these defects were found. Module 2 turns this into a method with the Pareto diagram (Chapter One Hundred Four).

Principle 5: tests wear out

The syllabus: "If the same tests are repeated many times, they become increasingly ineffective in detecting new defects." The 2018 syllabus called this the pesticide paradox, and explained the name: tests stop finding defects "just as pesticides are no longer effective at killing insects after a while."

It happens because a suite of tests only asks the questions it was written to ask. Once the defects those questions can find are fixed, the suite goes quiet, and a new defect somewhere it does not look is invisible to it. Below, ExamReg's late-fee function is rewritten in release 2 as a lookup table, and one entry is typed wrongly. The five tests written in Chapter One, on what software testing is, are run again, then one new test.

FEE = {0: 0}
FEE.update({day: 100 for day in range(1, 8)})
FEE.update({day: 500 for day in range(8, 16)})
FEE[12] = 50                                   # release 2: a typing slip, 50 for 500

def late_fee(days_late):
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    return FEE[max(days_late, 0)]

old_suite = [(0, 0), (3, 100), (7, 100), (8, 500), (15, 500)]
passed = sum(late_fee(days) == expected for days, expected in old_suite)
print(f"old suite: {passed} of {len(old_suite)} pass")
print("new test, 12 days late: expected 500, got", late_fee(12))
munotes.in25

The Seven Principles of Testing

old suite: 5 of 5 pass
new test, 12 days late: expected 500, got 50

The old suite passes perfectly and the defect ships unless someone writes a new test. The remedy in the syllabus: "existing tests and test data may need to be modified, and new tests may need to be written." It also notes the one case where repetition is the point: in automated regression testing the same tests are rerun deliberately, to prove that nothing that used to work has broken (Chapter Thirty-Nine, on regression testing).

Principle 6: testing is context dependent

The syllabus: "There is no single universally applicable approach to testing. Testing is done differently in different contexts." What changes with context is how much testing is enough, which techniques are used, how formally it is documented, and who does it.

ContextWhat matters mostWhat testing looks like
A radiation therapy machineSafety of patientsFormal, documented, reviewed by regulators; hardware and software both tested for timing and failure modes (the lesson of Therac-25 in Chapter Three)
A bank's payments systemCorrect money, security, auditHeavy regression suites, security testing, strict change control
ExamReg, a college portalCorrect fees, availability during the form windowRisk-based: deep tests on the exam form and fees, load tests before the deadline rush
A mobile gameFun, speed, many devicesExploratory testing, device and compatibility testing, fast release cycles

Applying one context's testing to another is a mistake in both directions: a game tested like a pacemaker never ships, and a pacemaker tested like a game should never have shipped.

Principle 7: the absence-of-defects fallacy

The syllabus: "It is a fallacy (i.e., a misconception) to expect that software verification will ensure the success of a system." A system can pass every test of its specification, with every defect found fixed, and still "not fulfill the users' needs and expectations", still fail the customer's business goals, and still be worse than its competitors. Hence: "In addition to verification, validation should also be carried out".

Applied to ExamReg. Suppose the specification described the portal for a desktop browser, and the team built and tested exactly that, flawlessly. If most students fill in their forms on phones, a defect-free ExamReg is still a failure. Verification asked "does it meet the specification?" and got yes; validation, which asks "does it meet the users' need?", was never asked. Verification and validation have a chapter of their own (Chapter Twenty-Six).

The seven, and what each tells a tester to do

PrincipleWhat it tells a tester to do
Presence, not absenceReport "no defects found by these tests", never "no defects"
Exhaustive testing is impossibleChoose tests with techniques, priority and risk
Early testing saves time and moneyReview requirements and designs; start testing before code exists
Defects cluster togetherSpend more effort where defects have been found before
Tests wear outReview and add tests; keep repetition for regression
Context dependentFit the testing to the product and its risks
Absence-of-defects fallacyValidate against users' needs, not only the specification
munotes.in26

The Seven Principles of Testing

What it does not mean

"Exhaustive testing is impossible" does not mean testing is pointless. It means testing must be selective, and selection is a skill Module 2 teaches.

"Defects cluster" does not mean the other modules can be skipped. It sets where the extra effort goes; every module still gets tested.

"Tests wear out" does not mean old tests should be thrown away. Regression suites are rerun on purpose. It means old tests alone are not enough to find new defects.

The principles are not a law of any standard. ISTQB calls them general guidelines. They are widely taught because they are true in practice, not because anyone enforces them.

Quick revision

  • ISTQB v4.0.1 section 1.3, seven principles; 2018 names in brackets where they differ.
  • 1. Presence, not absence of defects: passing tests cannot prove correctness.
  • 2. Exhaustive testing is impossible: two 32-bit inputs give about 1.8 × 10^19 combinations, centuries of testing; use techniques, priority and risk.
  • 3. Early testing saves time and money (shift left): static and dynamic testing as early as possible.
  • 4. Defects cluster together (Pareto): ExamReg's exam form and fee modules hold 132 of 200 defects, 66 per cent.
  • 5. Tests wear out (pesticide paradox): repeated tests stop finding new defects; add and revise tests; repetition is right for regression.
  • 6. Testing is context dependent: a therapy machine and a game are not tested alike.
  • 7. Absence-of-defects fallacy (absence-of-errors is a fallacy): verified is not the same as useful; validate too.

Test yourself

1. State the seven principles of testing. Testing shows the presence, not the absence of defects; exhaustive testing is impossible; early testing saves time and money; defects cluster together; tests wear out (the pesticide paradox); testing is context dependent; and the absence-of-defects fallacy.

2. Why is exhaustive testing impossible? Give a number. The combinations of inputs and conditions grow multiplicatively. A function of just two 32-bit integers has 2 to the power 64, about 1.8 × 10^19, input pairs, which would take about 585 years at a billion tests a second. So tests are selected with techniques and risk.

3. Explain the pesticide paradox with an example. Tests repeated unchanged stop finding new defects, because they only ask the questions they were written to ask. In the example, a lookup-table rewrite of the late-fee function stored 50 instead of 500 for 12 days; the old suite of five tests still passed, and only a new test at 12 days found the defect.

munotes.in27

The Seven Principles of Testing

4. What does "defects cluster together" mean, and how should a test manager use it? A small number of components usually contain most of the defects; in ExamReg two of seven modules held 66 per cent of them. The manager uses past and predicted clusters to direct extra test effort in risk-based testing.

5. What is the absence-of-defects fallacy? The misconception that finding and fixing every defect against the specification ensures the system's success. A system can be defect-free against its specification and still fail its users' needs, so validation must accompany verification.

6. Distinguish principles 1 and 7. Principle 1 is about knowledge: testing cannot prove there are no defects. Principle 7 is about value: even a system truly free of defects against its specification can be the wrong system. The first warns against over-trusting a passed test; the second against over-trusting the specification.

Contents This chapter on its own page

munotes.in28

Chapter Five

The Basic Test Process

Syllabus topic Module 1, "Software Testing Fundamentals: Basics of software testing process"

In one line

Testing is not one act of running tests but a sequence of activities: decide what testing is for, work out what to test, design the tests, prepare them, run them, and close the work down properly, while watching progress the whole time.

In the wording a student can write in an examination: the basic software test process consists of test planning, test monitoring and control, test analysis, test design, test implementation, test execution and test completion (ISTQB v4.0.1). Each activity produces its own work products, called testware, and the process is tailored to its context and often carried out iteratively rather than strictly in sequence.

Why a process at all

A student testing their own program runs it a few times and looks at the output. That works for twenty lines. It does not work for ExamReg, where several testers share hundreds of test cases across modules and releases, and where somebody must be able to say, on the day before the examination form opens, which requirements have been tested, which tests failed, and whether the portal is ready.

A process gives testing three things it cannot have otherwise. It makes testing repeatable, so the same tests give the same evidence next month. It makes it visible, so progress can be measured against a plan. And it makes it accountable, so a failure in operation can be traced back to the requirement and the test that should have caught it. The ISTQB syllabus puts it as a warning: there are common sets of activities "without which testing is less likely to achieve test objectives."

The standard's words

The software testing standard ISO/IEC/IEEE 29119-2:2021 defines a test process as a "set of testing activities performed to achieve a test objective". ISO/IEC/IEEE 29119-1:2022 defines testware as "artifacts produced during the test process required to plan, design, and execute tests". And the thing everything else is built from, the test basis, is "information used as the basis for designing and implementing test cases": a requirement, a user story, a design, an interface definition, even a user's reasonable expectation.

The seven groups of activities

The ISTQB syllabus names seven groups. The figure shows them in their logical order, with monitoring and control running beside all the others.

The seven groups of test activities, with monitoring and control alongside

Figure 5.1 The basic test process (ISTQB v4.0.1, section 1.4.1)

1. Test planning. Defining the test objectives and "selecting an approach that best achieves the objectives within the constraints imposed by the overall context." The plan says what will be tested, how, by whom, when, and what counts as finished. Writing a test plan has its own chapter (Chapter Eleven).

2. Test monitoring and test control. Monitoring is "the ongoing checking of all test activities and the comparison of actual progress against the plan." Control is "taking the actions necessary to meet the test objectives": moving testers to a module that is behind, adding tests where defects cluster, or deciding that a module is not ready for the next stage. This is why it is drawn beside the others and not after them.

munotes.in29

The Basic Test Process

3. Test analysis. Analysing the test basis "to identify testable features", then defining and prioritising test conditions. A test condition is a "testable aspect of a component or system, such as a function, transaction, feature, quality attribute, or structural element" (ISO/IEC/IEEE 29119-2). Analysis also reviews the test basis itself for defects and for testability. The syllabus sums it up: analysis answers "what to test?"

4. Test design. "Elaborating the test conditions into test cases and other testware", identifying coverage items that guide the choice of inputs, defining the test data needed and designing the test environment. Design answers "how to test?". The techniques of Module 2 (equivalence partitioning, boundary values, decision tables and the rest) are used here.

5. Test implementation. Creating or acquiring everything execution will need: test data, manual and automated test scripts, test procedures, which order test cases into steps, and test suites, which group procedures. The procedures are arranged into a test execution schedule, and the test environment is built and checked. ISO/IEC/IEEE 29119-2 defines that environment as the "facilities, hardware, software, firmware, and procedures needed to conduct a test".

6. Test execution. "Running the tests in accordance with the test execution schedule", comparing actual results with expected results, logging the results, and analysing anomalies to find their likely cause, so that failures can be reported. Chapter Nine, on test execution, takes this activity in detail.

7. Test completion. At a milestone such as a release: raising change requests or backlog items for defects still open, archiving testware that will be useful again, returning the test environment to an agreed state, recording lessons learned, and writing the test completion report.

Not a waterfall

The order above is logical, not rigid. The syllabus is explicit: "Although many of these activities may appear to follow a logical sequence, they are often implemented iteratively or in parallel." In an agile team, analysis, design, implementation and execution may all happen inside one two-week cycle, for one user story, and then again for the next. Test planning is revisited when the project changes. What matters is that each activity happens, not that they happen once each in a line.

The testware each activity produces

ActivityIts work products (ISTQB v4.0.1, 1.4.3)
Test planningTest plan, test schedule, risk register, entry criteria and exit criteria
Monitoring and controlTest progress reports, control directives, information about risks
Test analysisPrioritised test conditions (such as acceptance criteria); defect reports on the test basis
Test designPrioritised test cases, test charters, coverage items, test data requirements, test environment requirements
Test implementationTest procedures, manual and automated test scripts, test suites, test data, the test execution schedule, test environment items such as stubs, drivers and simulators
Test executionTest logs and defect reports
Test completionTest completion report, action items for improvement, lessons learned, change requests
munotes.in30

The Basic Test Process

Two terms in that table deserve a definition now. Entry criteria are the "states of being that have to be present before an effort can begin successfully", and exit criteria those that "have to be present before an effort can end successfully" (ISO/IEC/IEEE 24765:2017). A test level might enter only when the build installs cleanly, and exit only when every high-priority test has passed.

Traceability: the thread through all of it

The syllabus insists on keeping links "between the test basis elements, testware associated with these elements (e.g., test conditions, risks, test cases), test results, and defects." Those links are traceability, and they are what let a test manager answer the questions that matter. Is every requirement covered by a test? Which requirements are affected by a failed test? If a requirement changes, which tests must change?

The program below is a traceability record for ExamReg's fee rules, in its simplest form: each requirement lists its test cases, and each test case has a result from the latest run.

requirements = {                     # requirement -> the test cases that trace to it
    "FEE-1 no late fee on or before the last date": ["TC-01"],
    "FEE-2 Rs 100 for 1 to 7 days late":             ["TC-02", "TC-03"],
    "FEE-3 Rs 500 for 8 to 15 days late":            ["TC-04", "TC-05"],
    "FEE-4 form refused after 15 days":              ["TC-06"],
    "FEE-5 fee shown before payment":                [],
}
results = {"TC-01": "fail", "TC-02": "pass", "TC-03": "pass",
           "TC-04": "pass", "TC-05": "pass", "TC-06": "pass"}

covered = [r for r, tests in requirements.items() if tests]
print(f"requirements with at least one test: {len(covered)} of {len(requirements)}")
for requirement, tests in requirements.items():
    if not tests:
        status = "NOT TESTED"
    elif all(results[t] == "pass" for t in tests):
        status = "all tests pass"
    else:
        status = "FAILING: " + ", ".join(t for t in tests if results[t] != "pass")
    print(f"  {requirement:<46} {status}")
requirements with at least one test: 4 of 5
  FEE-1 no late fee on or before the last date   FAILING: TC-01
  FEE-2 Rs 100 for 1 to 7 days late              all tests pass
  FEE-3 Rs 500 for 8 to 15 days late             all tests pass
  FEE-4 form refused after 15 days               all tests pass
  FEE-5 fee shown before payment                 NOT TESTED
munotes.in31

The Basic Test Process

Two findings fall straight out of it, and neither could be seen from a list of test results alone. The failing test is linked to a requirement, so the report can say which rule is broken, not only which test failed. And requirement FEE-5 has no test at all: something the portal must do is simply not being checked. Chapter Forty-One, on validation testing, returns to this with the full traceability matrix.

Worked example: the late-fee rule through the whole process

Follow one small piece of ExamReg, its late-fee rule, through all seven activities.

ActivityWhat happens for the late-fee rule
PlanningObjective: show that fees are charged exactly as the rule says, before the form opens. Approach: specification-based tests at every boundary of the rule. Exit criterion: every fee requirement has a passing test
AnalysisTest basis: the fee table. Test conditions: on time, 1 to 7 days, 8 to 15 days, more than 15 days. Reviewing the basis finds a gap: nothing says what a negative number of days means, so a defect report goes to the exam cell
DesignTest cases with inputs and expected results: 0 days pays nothing, 1 and 7 pay Rs 100, 8 and 15 pay Rs 500, 16 is refused (the boundaries, as Chapter Fifty-Five on boundary value analysis will justify)
ImplementationA test procedure (log in as a test student, set the submission date, open the fee page, read the fee), test data (six student accounts), a test suite, and a test environment with the clock set to controlled dates
ExecutionThe suite is run; 0 days shows Rs 100 instead of nothing; the result is logged and a defect report raised with the evidence
Monitoring and controlProgress: six of six tests run, one failed. Control: the fee module stays out of the release build until the fix passes confirmation and regression tests
CompletionCompletion report: all fee requirements tested, one defect found and fixed, one requirement gap resolved by the exam cell; the suite archived for the next release

Test process in context

The same syllabus section adds that the way the process is carried out depends on context: the stakeholders, the team's skills, the business domain and its risks, technical factors, project constraints of time and budget, the organisation, the development life cycle in use, and the tools available. Those factors decide "test strategy, test techniques used, degree of test automation, required level of coverage, level of detail of testware, test reporting". A three-person team building ExamReg for one college and a bank's payments division both follow the seven activities; their test plans differ in length by a factor of fifty.

munotes.in32

The Basic Test Process

Who does what: two roles

The syllabus names two roles. The test management role "takes overall responsibility for the test process, test team and leadership of the test activities", and is mainly concerned with planning, monitoring and control, and completion. The testing role "takes overall responsibility for the engineering (technical) aspect of testing", mainly analysis, design, implementation and execution. They are roles, not job titles: a team leader or a development manager can hold the first, and "one person to take on the roles of testing and test management at the same time" is allowed.

The same process in ISO/IEC/IEEE 29119-2

The software testing standard organises the same work into three layers, which is useful when a question uses its vocabulary.

Layer (ISO/IEC/IEEE 29119-2)Its processesCorresponds to
Organizational test processDevelops and manages organisation-wide test specifications, such as a test policyDecided above any one project
Test management processesTest strategy and planning; test monitoring and control; test completionPlanning, monitoring and control, completion
Dynamic test processesTest design and implementation; test environment set-up and maintenance; test execution; test incident reportingAnalysis, design, implementation, execution

The standard's own definitions make the correspondence clear: the test design and implementation process is the "test process for deriving and specifying test cases and test procedures", and the test execution process executes those procedures "in the prepared test environment" and records the results.

What it does not mean

The process is not strictly sequential. The activities are often iterative and parallel; the order is logical, not a timetable.

Testing does not start at execution. Planning, analysis and design come first, and analysis already finds defects in the test basis before any code runs.

Monitoring and control are not a final step. They run throughout, which is why they are drawn beside the other activities.

Testware is not only test cases. Plans, conditions, data, scripts, environments, logs, reports and lessons learned are all testware.

Quick revision

  • Test process: "set of testing activities performed to achieve a test objective" (ISO/IEC/IEEE 29119-2).
  • Seven groups (ISTQB v4.0.1): planning; monitoring and control; analysis (what to test); design (how to test); implementation; execution; completion.
  • Often iterative or parallel; tailored to context (stakeholders, team, domain, technology, constraints, organisation, SDLC, tools).
  • Testware: artefacts produced to plan, design and execute tests; each activity has its own.
  • Test basis: information used to design test cases. Test condition: a testable aspect identified as a basis for testing.
  • Entry and exit criteria: conditions to begin and to end an effort.
  • Traceability links test basis, conditions, cases, results and defects: coverage, impact of change, audit.
  • Two roles: test management (planning, monitoring, control, completion) and testing (analysis to execution).
munotes.in33

The Basic Test Process

Test yourself

1. List the activities of the basic test process. Test planning, test monitoring and control, test analysis, test design, test implementation, test execution and test completion.

2. Distinguish test analysis from test design. Analysis studies the test basis to decide what to test, producing prioritised test conditions; design turns those conditions into test cases, coverage items, test data requirements and an environment design, deciding how to test.

3. Why are monitoring and control drawn beside the other activities rather than after them? Because they run throughout: progress is checked against the plan continuously, and corrective action, such as moving effort to a module where defects cluster, is taken while the other activities are still going on.

4. Name the testware produced by test implementation. Test procedures, manual and automated test scripts, test suites, test data, the test execution schedule and test environment items such as stubs, drivers and simulators.

5. What does traceability make possible? Give two examples. It shows which requirements are covered by tests and which are not, as with ExamReg's untested requirement FEE-5; and it shows which requirement a failed test affects, so a failure can be reported against the broken rule and the impact of a change can be judged.

6. What happens in test completion? Unresolved defects become change requests or backlog items, reusable testware is archived, the environment is returned to an agreed state, lessons learned are recorded, and a test completion report is written and communicated.

Contents This chapter on its own page

munotes.in34

Chapter Six

The Software Testing Life Cycle, Phase by Phase

Syllabus topic Module 1, "Basics of software testing process"; the paired practical, "as per STLC process"

In one line

The software testing life cycle is the test process of Chapter Five cut into six phases, each with a clear start, a clear finish and a document to show for it, so that testing can be planned and tracked like any other project work.

In the wording a student can write in an examination: the software testing life cycle (STLC) is the sequence of phases through which testing proceeds in a project: requirement analysis, test planning, test case development, test environment set-up, test execution and test cycle closure. Each phase has entry criteria that must hold before it starts, exit criteria that must hold before it ends, and deliverables it produces.

Why the STLC exists, and what it is not

The previous chapter described the test process as the ISTQB syllabus and the ISO/IEC/IEEE 29119-2 standard describe it: groups of activities, often overlapping. A project manager, however, wants phases: something that starts on a date, finishes on a date, and hands over a document. The software testing life cycle is that project view of the same work, and it is how most companies and training courses in India describe testing. The paired practical of this paper uses it by name: a student prepares a test plan, test scenarios, test cases, a test execution report and a defect report "as per STLC process".

It is worth being exact about its status. No international standard defines the STLC. ISO/IEC/IEEE 29119-2 defines test processes and ISTQB defines test activities; the six phases below are the form industry uses, and independent descriptions of it agree closely on the phases and their order. Where they differ, it is in names (test case development is sometimes called test design) or in splitting one phase in two. An examination answer that gives the six phases, their entry and exit criteria and their deliverables, and says how they relate to the standard process, is complete.

The six phases

Phase 1: requirement analysis. The test team studies the requirements from a testing point of view: what is testable, what is ambiguous, what is missing, which requirements are non-functional (performance, security, usability), and where the risks lie. Questions go back to the business analysts and the users. For ExamReg, this is where the question of what a negative number of days late should mean was first asked.

Phase 2: test planning. The test lead defines the objectives, scope and approach, chooses manual or automated testing and the tools, estimates effort, time and cost, assigns roles, sets the entry and exit criteria for each later phase, identifies the deliverables and the risks, and has the test plan reviewed and approved. Chapter Eleven is about writing the test plan.

munotes.in35

The Software Testing Life Cycle, Phase by Phase

Phase 3: test case development. Test scenarios and test cases are written, with test data and expected results, and reviewed. The requirement traceability matrix is updated so every requirement maps to its test cases. Automated test scripts are written here too.

Phase 4: test environment set-up. The hardware, software, network, browsers, databases, test accounts and permissions are prepared, and the environment is checked, usually with a short smoke test that confirms the build installs and its main functions respond. This phase often runs in parallel with phase 3, and it is often run by a separate team.

Phase 5: test execution. The test cases are run as scheduled, results are compared with expected results and logged, failures are reported as defects with severity and priority, fixed defects are retested, and regression tests are run after fixes. Execution usually takes several test cycles, each on a new build.

Phase 6: test cycle closure. The team checks the exit criteria, prepares the test summary (or completion) report, makes sure every defect is closed or knowingly deferred, archives the testware, returns the environment, and holds a meeting on lessons learned.

Entry criteria, exit criteria and deliverables

The phases are made controllable by their criteria. ISO/IEC/IEEE 24765:2017 defines entry criteria as "states of being that have to be present before an effort can begin successfully" and exit criteria as those that "have to be present before an effort can end successfully". The table gives typical ones; a real test plan states them for its own project.

PhaseEntry criteria (typical)Exit criteria (typical)Deliverables
Requirement analysisRequirements document available; access to people who can answer questionsTestable requirements listed; questions answered or loggedRequirement review comments; list of testable requirements; first risk list
Test planningRequirements analysed; scope and timeline knownTest plan reviewed and signed offTest plan; effort and cost estimate; schedule
Test case developmentApproved test plan; stable requirementsTest cases written, reviewed and approved; traceability matrix completeTest scenarios; test cases; test data; automation scripts; traceability matrix
Test environment set-upEnvironment requirements known; build availableEnvironment ready and smoke test passedReady environment; smoke test result
Test executionApproved test cases; ready environment; build handed overPlanned tests run; exit thresholds met, for example no open critical defectTest logs; defect reports; execution reports each cycle
Test cycle closureExecution finished; exit criteria evaluatedReport signed off; testware archivedTest summary (completion) report; lessons learned

The six phases against the standard process

STLC phaseISTQB v4.0.1 activity (the basic test process, Chapter Five)ISO/IEC/IEEE 29119-2 process
Requirement analysisTest analysis (reviewing the test basis)Test design and implementation (its first part)
Test planningTest planningTest strategy and planning
Test case developmentTest design and test implementationTest design and implementation
Test environment set-upTest implementation (building the environment)Test environment set-up and maintenance
Test executionTest executionTest execution; test incident reporting
Test cycle closureTest completionTest completion
(throughout)Test monitoring and test controlTest monitoring and control
munotes.in36

The Software Testing Life Cycle, Phase by Phase

Read the table and one difference stands out. The STLC puts requirement analysis before planning, while ISTQB puts planning first and analysis after. In practice both happen together: a plan cannot be made without knowing what the requirements ask for, and the analysis cannot be complete without knowing the plan's scope. The last row matters too. Monitoring and control are not a phase of the STLC, but they run through all six, exactly as in the basic test process of Chapter Five.

Worked example: does execution exit?

ExamReg release 2.0 has 240 planned test cases, 40 of them critical, 80 high priority and 120 medium. The test plan's exit criteria for the execution phase are: every critical test has been run; at least 90 per cent of all planned tests have been run; at least 95 per cent of the tests run have passed; and no critical defect is open. After the first execution cycle the figures are as below, and the program checks each criterion.

planned = {"critical": 40, "high": 80, "medium": 120}
run     = {"critical": 40, "high": 78, "medium": 101}
passed  = {"critical": 38, "high": 72, "medium": 96}
open_defects = {"critical": 1, "major": 3, "minor": 7}

total_planned = sum(planned.values())
total_run = sum(run.values())
total_passed = sum(passed.values())
criteria = [
    ("every critical test run", run["critical"] == planned["critical"]),
    (f"at least 90% of planned tests run ({total_run} of {total_planned}, "
     f"{100 * total_run / total_planned:.1f}%)", total_run >= 0.90 * total_planned),
    (f"at least 95% of tests run passed ({total_passed} of {total_run}, "
     f"{100 * total_passed / total_run:.1f}%)", total_passed >= 0.95 * total_run),
    ("no critical defect open", open_defects["critical"] == 0),
]
for text, met in criteria:
    print(("met     " if met else "NOT MET ") + text)
print("execution phase may close:", all(met for _, met in criteria))
met     every critical test run
met     at least 90% of planned tests run (219 of 240, 91.2%)
NOT MET at least 95% of tests run passed (206 of 219, 94.1%)
NOT MET no critical defect open
execution phase may close: False

The phase does not close. Two criteria are met and two are not: the pass rate is just under the threshold, and one critical defect is still open. Test control now acts: the critical defect is fixed first, the failed tests are rerun on the next build, and a second cycle begins. Nobody argues about whether testing is finished, because the plan said in advance what finished means. That is the whole value of exit criteria.

munotes.in37

The Software Testing Life Cycle, Phase by Phase

The STLC inside the SDLC

The STLC is not a separate life cycle running beside development; it runs inside the software development life cycle (SDLC), which Module 1's second row covers from Chapter Thirteen. Its first two phases can begin as soon as requirements exist, long before code, which is principle 3 (early testing) in practice. In the V-model (Chapter Fifteen), each development phase on the left of the V has its test level on the right, and a complete STLC, from analysis to closure, can be run for each test level.

In agile development the six phases compress into each short iteration. Requirement analysis becomes refining a user story and its acceptance criteria, planning becomes the sprint's test tasks, case development and execution happen within the sprint, and closure becomes the sprint review and retrospective. The phases are still there; they are just two weeks long.

SDLCSTLC
PurposeTo build the softwareTo evaluate the software and find its defects
PhasesRequirements, design, implementation, testing, deployment, maintenanceRequirement analysis, test planning, test case development, environment set-up, execution, closure
Main outputThe working product and its documentsTest plan, test cases, logs, defect reports, completion report
RelationshipThe whole life cycleRuns inside it, starting as early as requirements

What it does not mean

The STLC is not an international standard. It is industry's phase view of the standard test process; ISO/IEC/IEEE 29119-2 and ISTQB define the process it describes.

The STLC does not make a product defect-free. Some descriptions give its objective as a defect-free product; no process can deliver that (principle 1, Chapter Four). Its objective is to find defects and give evidence about quality in a planned, controlled way.

The phases are not strictly sequential. Environment set-up usually overlaps case development, and execution runs in cycles with fixes between them.

The STLC phases are not test levels. Unit, integration, system and acceptance testing are levels (Chapter Thirty-Two); the STLC can be followed within any of them.

Quick revision

  • STLC: the phases testing goes through in a project; industry's view of the test process, not an international standard.
  • Six phases: requirement analysis, test planning, test case development, test environment set-up, test execution, test cycle closure.
  • Each has entry criteria (to begin), exit criteria (to end) and deliverables.
  • Deliverables: review comments; test plan; test cases, data, traceability matrix; ready environment and smoke test result; logs, defect reports, execution reports; summary report and lessons learned.
  • Maps onto ISTQB's activities and ISO/IEC/IEEE 29119-2's processes; monitoring and control run throughout.
  • Worked example: 219 of 240 tests run (91.25 per cent), 206 of them passed (about 94.1 per cent), one critical defect open, so execution does not exit.
  • In agile, all six phases happen inside each iteration.
munotes.in38

The Software Testing Life Cycle, Phase by Phase

Test yourself

1. What is the software testing life cycle? Name its phases. It is the sequence of phases through which testing proceeds in a project: requirement analysis, test planning, test case development, test environment set-up, test execution and test cycle closure, each with entry criteria, exit criteria and deliverables.

2. Define entry and exit criteria, with one example of each for test execution. Entry criteria are the conditions that must hold before a phase begins, such as approved test cases and an environment that has passed its smoke test. Exit criteria are those that must hold before it ends, such as every critical test run and no critical defect open.

3. What are the deliverables of test cycle closure? The test summary or completion report, confirmation that defects are closed or knowingly deferred, archived testware, the environment returned, and documented lessons learned.

4. How does the STLC relate to the SDLC? It runs inside the SDLC, as its testing thread. Its first phases can begin as soon as requirements exist, before any code, and in the V-model a full testing cycle can be run for each test level.

5. In the worked example, why could the execution phase not close? Two of its four exit criteria were not met: only 94.1 per cent of the tests run had passed against a threshold of 95 per cent, and one critical defect was still open.

6. Is it correct to say the objective of the STLC is a defect-free product? Justify. No. Testing can show the presence of defects but never their absence (principle 1), so no process can guarantee a defect-free product. The STLC's objective is to find defects and give evidence of quality in a planned and controlled way.

Contents This chapter on its own page

munotes.in39

Chapter Seven

What a Test Case Is, and What Makes a Good One

Syllabus topic Module 1, "Software Testing Fundamentals: Test case design principles"; the paired practical, "Test Scenarios"

In one line

A test case is one precisely written check: from this starting state, give the software this input, and it should do exactly this.

In the wording a student can write in an examination: a test case is a "set of preconditions, inputs and expected results, developed to drive the execution of a test item to meet test objectives" (ISO/IEC/IEEE 29119-2:2021). A test scenario is a higher-level situation to be tested, from which several test cases are derived. A good test case is traceable to a requirement, has one clear objective, states its preconditions and its expected result in advance, and can be repeated by anyone.

Why a test case is written down

A tester who clicks around ExamReg and notices that something looks wrong has done useful work, but nobody can repeat it, count it, or show which requirement it covered. A written test case fixes all three. It can be run again tomorrow, on the next build, by a different person, and give comparable evidence. It can be counted, so progress can be reported. And it can be traced to the requirement it checks, so coverage can be measured.

The written test case also forces the one decision that separates a test from a demonstration: the expected result, decided before the software is run. Chapter One made the point with the late-fee function: without an expected result, the Rs 100 charged to an on-time student would have looked perfectly reasonable.

The standard definitions

Three definitions, from newest to oldest wording, all saying the same thing:

TermDefinitionSource
Test case"set of preconditions, inputs and expected results, developed to drive the execution of a test item to meet test objectives"ISO/IEC/IEEE 29119-2:2021
Test case"set of test inputs, execution conditions, and expected results developed for a particular objective, such as to exercise a particular program path or to verify compliance with a specific requirement"IEEE 1012-2024
Test case specification"documentation of a set of one or more test cases"ISO/IEC/IEEE 29119-2:2021

Take the first apart. Preconditions are what must be true before the test starts: a user logged in, a form half filled, a date on the server's clock. Inputs are what the tester supplies. Expected results are what the software must do in response, decided in advance. And the whole is developed to meet test objectives: a test case exists for a reason, which is why it is traced to a requirement or a risk.

The parts of a written test case

Test management tools and templates vary, but the same fields appear in nearly all of them. Here is one ExamReg test case written in full.

FieldContent
Test case IDTC-FEE-01
TitleNo late fee for a form submitted on the last date
Requirement tracedFEE-1: no late fee on or before the last date
PriorityHigh (every student who submits on time is affected)
PreconditionsTest student SC23001 is logged in; the exam form is complete; the portal's clock is set to the last date
Test dataStudent SC23001, regular papers only, no backlog papers
Steps1. Open the fee page. 2. Read the late fee line. 3. Read the total
Expected resultLate fee shows Rs 0; total equals the regular form fee alone
Actual result(filled in during execution)
Status(pass, fail or blocked, filled in during execution)
PostconditionsForm remains unpaid, so the account can be reused
munotes.in40

What a Test Case Is, and What Makes a Good One

Two fields are filled in only when the test runs: the actual result and the status. Everything else is written before, which is what makes the test repeatable. The postcondition is easily forgotten and quietly important: it says what state the test leaves behind, so the next test does not start from a surprise.

The expected result, and where it comes from

The part of a test case that decides pass or fail has a name. ISO/IEC TR 29119-11:2020 defines a test oracle as a "source of information for determining whether a test has passed or failed". For ExamReg's fee test the oracle is the fee table in the requirements. Oracles come from several places, and a tester should know which one each test relies on:

  • The specification, as with the fee table.
  • An independent calculation, as when Chapter Two, on errors, faults and failures, used Python's own calendar.isleap to judge a leap-year function.
  • A previous version of the software, trusted for the features that did not change.
  • A comparable product, such as a published calculator for the same rule.
  • A person's judgement, for things like readability or layout, where no document can say what "right" is.

The same technical report names the difficulty that sits under all of this, the test oracle problem: the "challenge of determining whether a test has passed or failed for a given set of test inputs and state". When nobody can say what the right answer is, the test cannot be judged, however carefully it was run.

From test condition to test suite

A test case sits in a hierarchy, and the practical's documents use every level of it.

LevelDefinitionExamReg example
Test condition"testable aspect of a component or system ... identified as a basis for testing" (ISO/IEC/IEEE 29119-2)The late fee for 1 to 7 days
Test scenario"situation or setting for a test item used as the basis for generating test cases" (ISO/IEC TR 29119-11)A student submits the form a few days late and pays the fee
Test casePreconditions, inputs and expected results (ISO/IEC/IEEE 29119-2)Submitted 7 days late: late fee Rs 100
Test procedure"sequence of test cases in execution order, with associated actions required to set up preconditions and perform wrap-up activities post execution" (ISO/IEC/IEEE 29119-2)Log in, set the date, run the four late-fee cases in order, log out
Test script"document specifying one or more test procedures" (ISO/IEC/IEEE 29119-1)The procedure written for a person, or as an automated program
Test suite"set of test cases or test procedures" (ISO/IEC/IEEE 29119-1)Every fee test, run before each release
munotes.in41

What a Test Case Is, and What Makes a Good One

Test scenarios and test cases

In industry, and in the practical of this paper, a test scenario is usually written as one line saying what situation will be tested, and several test cases are then written saying exactly how, each with its own data and expected result. A scenario is quick to write and quick to review; it lets a business analyst check the testers have thought of the right situations before anyone writes a single step.

Test scenarioTest cases derived from it
TS-01: A student submits the exam form late and pays the late feeTC-FEE-02: 1 day late, Rs 100. TC-FEE-03: 7 days late, Rs 100. TC-FEE-04: 8 days late, Rs 500. TC-FEE-05: 15 days late, Rs 500. TC-FEE-06: 16 days late, form refused
TS-02: A student with backlog papers fills the exam formTC-FRM-11: one backlog paper added. TC-FRM-12: the maximum number of backlog papers. TC-FRM-13: a backlog paper from a semester not yet attempted is refused
TS-03: Payment fails midwayTC-PAY-21: gateway timeout, no fee recorded. TC-PAY-22: payment succeeds but the receipt page fails to load; fee recorded once, not twice
Test scenarioTest case
SaysWhat situation to testExactly how: preconditions, inputs, steps, expected result
Level of detailOne lineSeveral fields
Derived fromRequirements, user stories, use casesA scenario, or a test condition
NumberFewSeveral per scenario
Can be executed as it stands?NoYes
Best forReviewing coverage with stakeholders earlyExecution, repetition and pass or fail evidence

What makes a good test case

These are principles of practice, each followed because of what goes wrong without it.

  1. It is traceable. It names the requirement or risk it checks. Without that link, a failure cannot be reported against a rule and coverage cannot be measured (Chapter Five, on the basic test process).
  2. It has one objective. A case that checks the fee, the receipt and the email at once gives one "fail" for three possible causes, and the defect report cannot say which.
  3. Its preconditions are explicit. Most "cannot reproduce" arguments between testers and developers are really two different starting states.
  4. Its expected result is exact and decided in advance. "The fee is calculated correctly" is not an expected result; "Rs 100" is.
  5. It is repeatable and independent. It gives the same result every time on the same build, and does not depend on another test having run first unless its preconditions say so.
  6. It includes the unwelcome inputs. Invalid values, empty fields, the maximum, the moment just after a deadline. Most defects live in the cases a developer did not think of; a suite of only valid, typical inputs finds few of them.
  7. It is not redundant. Each case should be able to find something the others cannot. Ten cases that all enter 3 days late add effort and no evidence. Module 2's techniques are, in effect, methods for choosing non-redundant cases.
  8. Another tester could run it. Clear steps and data, no private knowledge.
  9. It is prioritised. ISTQB v4.0.1 section 5.1.5 describes prioritising test cases, most commonly by risk, so that if time runs out the most important ones have already run.
munotes.in42

What a Test Case Is, and What Makes a Good One

Worked example: a table of test cases, executed

The six late-fee cases of scenario TS-01 are written below as data, one record per case, with the same fields as the table above. A short program then does what a tester does: runs each case, records the actual result, and sets the status. The function under test is the corrected version from Chapter One.

def late_fee(days_late):
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    if days_late <= 0:
        return 0
    if days_late <= 7:
        return 100
    return 500

test_cases = [   # id, requirement, input (days late), expected result
    ("TC-FEE-01", "FEE-1", 0,  0),
    ("TC-FEE-02", "FEE-2", 1,  100),
    ("TC-FEE-03", "FEE-2", 7,  100),
    ("TC-FEE-04", "FEE-3", 8,  500),
    ("TC-FEE-05", "FEE-3", 15, 500),
    ("TC-FEE-06", "FEE-4", 16, "refused"),
]
print(f"{'ID':<10}{'req':<7}{'input':>6}  {'expected':<9}{'actual':<9}status")
for case_id, requirement, days, expected in test_cases:
    try:
        actual = late_fee(days)
    except ValueError:
        actual = "refused"
    status = "pass" if actual == expected else "FAIL"
    print(f"{case_id:<10}{requirement:<7}{days:>6}  {str(expected):<9}{str(actual):<9}{status}")
ID        req     input  expected actual   status
TC-FEE-01 FEE-1       0  0        0        pass
TC-FEE-02 FEE-2       1  100      100      pass
TC-FEE-03 FEE-2       7  100      100      pass
TC-FEE-04 FEE-3       8  500      500      pass
TC-FEE-05 FEE-3      15  500      500      pass
TC-FEE-06 FEE-4      16  refused  refused  pass

Notice what the records contain and what the program adds. Each record was written before the run: identifier, requirement, input and expected result. The program supplies only the actual result and the status, exactly the two fields left blank in the written test case. Notice also that the test for 16 days has an expected result that is not a number: a refusal is a result too, and a test case that forgot to say what should happen after 15 days would have been incomplete.

munotes.in43

What a Test Case Is, and What Makes a Good One

What it does not mean

A test scenario is not a test case. A scenario says what situation to test; it cannot be executed or judged until test cases with data and expected results are derived from it.

A test case is not a test script or a procedure. The case says what to check; the procedure says in what order and with what setup to run cases; the script is the document or program that carries the procedure.

More test cases do not mean better testing. Ten redundant cases are worth less than three well-chosen ones; what counts is what the cases can find.

Expected results do not come from running the software. An expected result copied from the program's own output tests nothing: it certifies whatever the program already does, including its defects.

Quick revision

  • Test case: "set of preconditions, inputs and expected results, developed to drive the execution of a test item to meet test objectives" (ISO/IEC/IEEE 29119-2:2021).
  • Fields: ID, title, requirement traced, priority, preconditions, test data, steps, expected result, actual result, status, postconditions.
  • Test oracle: the source that decides pass or fail: specification, independent calculation, previous version, comparable product, human judgement. The oracle problem is not knowing the right answer.
  • Hierarchy: test condition, scenario, case, procedure, script, suite.
  • Scenario says what situation; case says exactly how, with data and expected result; several cases per scenario.
  • A good test case: traceable, one objective, explicit preconditions, exact expected result in advance, repeatable, includes invalid inputs, non-redundant, runnable by others, prioritised.

Test yourself

1. Define a test case and name its essential parts. A set of preconditions, inputs and expected results developed to drive the execution of a test item to meet test objectives. In practice it also carries an identifier, the requirement it traces to, steps, a priority, postconditions, and, after execution, the actual result and a status.

2. Differentiate a test scenario from a test case, with an example. A scenario is a one-line situation to be tested, such as a student submitting the form late and paying the late fee; it cannot be executed as it stands. A test case derived from it is executable and exact, such as a form submitted 7 days late with an expected late fee of Rs 100.

3. What is a test oracle? Give three kinds. The source of information that decides whether a test passed or failed. Examples: the specification, an independent calculation such as a trusted library function, a previous version of the software, a comparable product, or a person's judgement.

4. Why must the expected result be decided before the test is run? Because the test is a comparison, and a result decided after seeing the output simply accepts whatever the program did, defects included.

munotes.in44

What a Test Case Is, and What Makes a Good One

5. State any five principles of good test case design. Any five of: traceable to a requirement; one objective per case; explicit preconditions; exact expected result decided in advance; repeatable and independent; includes invalid and unexpected inputs; not redundant; clear enough for another tester; prioritised by risk.

6. In what order are these produced during testing: test case, test suite, test condition, test scenario, test procedure? Test condition first (in test analysis), then the test scenario and the test cases derived from it (in test design), then the test procedure that puts cases in execution order with their setup, and the test suite that groups them (both in test implementation).

Contents This chapter on its own page

munotes.in45

Chapter Eight

Test Design Techniques: The Three Families

Syllabus topic Module 1, "Software Testing Fundamentals: Test case design principles and techniques"

In one line

A test design technique is a systematic way of choosing a few good tests out of the millions possible, and the techniques come in three families: those that read the specification, those that read the code, and those that draw on the tester's experience.

In the wording a student can write in an examination: a test design technique is a "procedure used to create or select a test model, identify test coverage items, and derive corresponding test cases" (ISO/IEC/IEEE 29119-2:2021). Techniques are classified as black-box (specification-based), which derive tests from the specified behaviour without reference to the internal structure; white-box (structure-based), which derive tests from the internal structure of the code; and experience-based, which derive tests from the tester's knowledge and experience (ISTQB v4.0.1).

Why techniques exist

Chapter Four's second principle, that exhaustive testing is impossible, leaves a tester with a question: out of every possible input, which few hundred should be tried? Picking them by instinct gives tests that cluster around typical values and miss the edges, and two testers picking by instinct give two different suites with no way to say which is better.

A technique answers the question systematically. The ISTQB syllabus puts it in one line: test techniques "help to develop a relatively small, but sufficient, set of test cases in a systematic way." Systematic means two testers applying the same technique to the same specification arrive at essentially the same tests, and it means the result can be measured: the technique says what has to be covered, so it can say how much has been.

Two words the techniques rely on

A test coverage item is a "measurable attribute of a test item that is the focus of testing" (ISO/IEC/IEEE 29119-2). Each technique names its own: for equivalence partitioning the coverage items are the partitions; for branch testing they are the branches of the code.

Test coverage is then the "degree, expressed as a percentage, to which specified test coverage items have been exercised by a test case or test cases" (ISO/IEC/IEEE 29119-2). If the fee rule has five partitions and the tests touch four of them, partition coverage is 80 per cent. Coverage is how a technique turns "we tested it" into a number.

Family one: black-box, or specification-based

ISO/IEC/IEEE 29119-1:2022 defines specification-based testing as testing "in which the principal test basis is the external inputs and outputs of the test item, commonly based on a specification, rather than its implementation in source code or executable software". The tester treats the software as a box whose inside cannot be seen: only what goes in and what comes out are known.

The ISTQB syllabus notes the great advantage of this: "the test cases are independent of how the software is implemented. Consequently, if the implementation changes, but the required behavior stays the same, then the test cases are still useful." The black-box techniques MU names, each with its own chapter in Module 2, are equivalence partitioning (Chapter Fifty-Four), boundary value analysis (Chapter Fifty-Five), decision table testing (Chapter Fifty-Six) and state transition testing (Chapter Fifty-Seven).

munotes.in46

Test Design Techniques: The Three Families

Family two: white-box, or structure-based

ISO/IEC/IEEE 29119-1 defines structure-based testing as "dynamic testing in which the tests are derived from an examination of the structure of the test item". The box is now transparent: the tester reads the code, finds its statements, decisions and paths, and designs tests to exercise them.

This family has a mirror-image limitation. The syllabus: "As the test cases are dependent on how the software is designed, they can only be created after the design or implementation of the test object." And if the code changes, the tests may need to change with it. The white-box techniques MU names are statement testing (Chapter Fifty-Nine) and branch testing (Chapter Sixty), with structural testing introduced as a whole in Chapter Fifty-Eight.

Family three: experience-based

ISO/IEC/IEEE 29119-4:2021 defines experience-based testing as a "class of test case design techniques based on using the experience of testers to generate test cases". A tester who has seen date fields fail on 29 February, and forms fail when a user double-clicks Submit, tries those things first.

The syllabus is candid about both sides: the effectiveness of these techniques "depends heavily on the tester's skills", and yet they "can detect defects that may be missed using the black-box test techniques and white-box test techniques. Hence, experience-based test techniques are complementary". MU names three: error guessing (Chapter Sixty-Two), exploratory testing (Chapter Sixty-Three) and checklist-based testing (Chapter Sixty-Four).

The three families compared

Black-box (specification-based)White-box (structure-based)Experience-based
Tests derived fromSpecification: requirements, rules, interfacesThe code or design: statements, decisions, pathsThe tester's knowledge of where software fails
Needs the code?NoYesNo
Can startAs soon as a specification existsOnly after design or code existsAny time, often once something runs
Tests survive a rewrite of the code?Yes, if behaviour is unchangedOften notUsually
Finds wellMissing or wrong behaviour against the specificationUntested code, including code that should not be thereDefects of kinds nobody specified
MissesCode the specification does not mentionBehaviour the code does not implement at allWhatever the tester has never met
Coverage measured asPartitions, boundaries, rules, transitionsStatements, branchesHard to measure; checklists and session notes help
MU's techniques (Module 2)Equivalence partitioning, boundary value analysis, decision table, state transitionStatement testing, branch testingError guessing, exploratory, checklist-based

The "misses" row is the reason all three are used. A black-box tester cannot see a stray line of code that does something nobody asked for. A white-box tester cannot see a requirement the programmer forgot, because there is no code for it to cover. And both follow rules, which is exactly where an experienced tester's instinct for the unexpected earns its place.

munotes.in47

Test Design Techniques: The Three Families

Worked example: three families, three defects

Below, ExamReg's late-fee function contains three planted defects of three different kinds. The oracle is the fee table, with one clarification the exam cell gave after the review of the requirement in Chapter One, on what software testing is: a negative number of days, or anything that is not a whole number of days, must be refused cleanly. Each family's tests are chosen the way that family chooses them, and the program reports what each finds.

def late_fee(days_late):
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    if days_late <= 0:                  # defect C: the input is never validated
        return 0
    if days_late == 13:                 # defect B: a forgotten debugging shortcut
        return 0
    if days_late <= 6:                  # defect A: the rule says 7, not 6
        return 100
    return 500

def expected(days):                     # the oracle: the fee table and the exam cell's ruling
    if not isinstance(days, int) or days < 0 or days > 15:
        return "refused"
    return 0 if days == 0 else (100 if days <= 7 else 500)

def actual(days):
    try:
        return late_fee(days)
    except ValueError:
        return "refused"
    except Exception as problem:        # anything else is a crash, not a clean refusal
        return "crash: " + type(problem).__name__

families = {
    "black-box, from the fee table":    [0, 1, 7, 8, 15, 16],
    "white-box, one test per branch":   [16, 0, 13, 3, 10],
    "experience-based, likely mistakes": [-1, "7", 7.5],
}
for family, inputs in families.items():
    found = [f"{d!r}: expected {expected(d)}, got {actual(d)}"
             for d in inputs if actual(d) != expected(d)]
    print(f"{family}: {len(inputs)} tests, {len(found)} failure(s)")
    for line in found:
        print("   ", line)
black-box, from the fee table: 6 tests, 1 failure(s)
    7: expected 100, got 500
white-box, one test per branch: 5 tests, 1 failure(s)
    13: expected 500, got 0
experience-based, likely mistakes: 3 tests, 3 failure(s)
    -1: expected refused, got 0
    '7': expected refused, got crash: TypeError
    7.5: expected refused, got 500

Each family found what the others could not. The black-box tests were chosen at the edges of the fee table, so 7 days, the boundary the programmer got wrong, was among them; but 13 is an ordinary value in the middle of a band, and no specification-based technique would single it out. The white-box tests were chosen to make every branch of the code run at least once, so the strange == 13 branch was forced to execute and gave itself away; but no branch exists for "exactly 7", so branch testing had no reason to try it. And the experience-based tests tried what users actually do wrong: a negative number, a number typed as text, a fraction. All three failed, and all three are the same defect showing three faces: the function never validates its input, so it charges nothing for minus one day, crashes on text, and charges Rs 500 for seven and a half days. The fee table never mentioned any of these inputs, so no specification-based test would have tried them.

munotes.in48

Test Design Techniques: The Three Families

The lesson is the one the ISTQB syllabus gives: the families are complementary. A test plan that uses only one of them has chosen, in advance, which kinds of defect it will not find.

The wider catalogue

MU's syllabus names nine techniques. The software testing standard ISO/IEC/IEEE 29119-4:2021 defines more, and a student meeting their names elsewhere should know which family each belongs to.

FamilyTechniques named in ISO/IEC/IEEE 29119-4:2021On MU's syllabus
Specification-basedEquivalence partitioning, boundary value analysis, decision table testing, state transition testing, classification tree method, cause-effect graphing, syntax testing, scenario testing, requirements-based testing, random testing, metamorphic testingThe first four
Structure-basedStatement testing, branch testing, branch condition testing, branch condition combination testing, modified condition/decision coverage (MC/DC) testing, data flow testingThe first two
Experience-basedError guessingError guessing, and ISTQB's exploratory and checklist-based testing

Choosing a technique

No technique is best everywhere, which is Chapter Four's sixth principle (testing is context dependent) applied to design. A few rules of thumb from practice:

  • Input fields with ranges (marks, days, amounts): equivalence partitioning with boundary value analysis.
  • Business rules combining several conditions (fee waivers, eligibility): decision tables.
  • Behaviour that depends on history (login attempts, an order's status): state transition testing.
  • Safety-critical or complex code: white-box coverage targets, set in the test plan.
  • New, poorly specified or rushed features: exploratory testing, guided by checklists.
  • Every time: some error guessing, because it costs little and finds what rules do not.

What it does not mean

Black-box does not mean careless or blind. It means the tests are derived from the specification without looking at the code; the design is as systematic as any white-box technique.

White-box testing is not only for developers. Anyone who can read the code can use it, though in practice developers apply it most, at unit level.

Experience-based testing is not random clicking. It is disciplined by charters, checklists and notes, as Module 2 shows; and random testing is a separate, specification-based technique in ISO/IEC/IEEE 29119-4.

100 per cent coverage does not mean no defects remain. It means every coverage item of that technique was exercised. In the worked example, 100 per cent branch coverage missed defect A.

munotes.in49

Test Design Techniques: The Three Families

Quick revision

  • Test design technique: a "procedure used to create or select a test model, identify test coverage items, and derive corresponding test cases" (ISO/IEC/IEEE 29119-2).
  • Test coverage: the percentage of specified coverage items exercised by the tests.
  • Black-box (specification-based): from the specification; tests survive code changes; misses code nobody specified. MU: equivalence partitioning, boundary value analysis, decision tables, state transitions.
  • White-box (structure-based): from the code; needs the code; misses behaviour never coded. MU: statement and branch testing.
  • Experience-based: from the tester's experience; complementary to both. MU: error guessing, exploratory, checklist-based.
  • Worked example: each family found exactly one kind of planted defect the others missed.

Test yourself

1. What is a test design technique, and why are techniques used? A procedure for identifying test coverage items and deriving test cases from a test basis. They are used because exhaustive testing is impossible, and they produce a small but sufficient set of tests systematically, so the result is repeatable and its coverage can be measured.

2. Name the three families of techniques and give MU's examples of each. Black-box or specification-based: equivalence partitioning, boundary value analysis, decision table testing and state transition testing. White-box or structure-based: statement testing and branch testing. Experience-based: error guessing, exploratory testing and checklist-based testing.

3. Why can white-box tests not be written at the start of a project? Because they are derived from the internal structure of the code or detailed design, which does not exist until design or implementation has been done.

4. Give one kind of defect each family is good at finding and one it misses. Black-box finds wrong or missing behaviour against the specification but misses code nobody specified; white-box finds untested or unexpected code but misses required behaviour that was never coded; experience-based finds unusual real-world inputs but misses whatever the tester has never encountered.

5. In the worked example, which family found the forgotten debugging shortcut, and why could the others not? White-box branch testing, because it had to execute every branch, including the one for 13 days. Black-box tests had no reason to choose 13, an ordinary value in the middle of a band, and experience-based tests looked at invalid inputs instead.

6. Does 100 per cent branch coverage guarantee a correct program? No. It guarantees each branch ran at least once. In the worked example the branch tests achieved full branch coverage and still missed the wrong boundary at 7 days, because no branch exists for a value the code handles wrongly within a range.

Contents This chapter on its own page

munotes.in50

Chapter Nine

Test Execution

Syllabus topic Module 1, "Software Testing Fundamentals: Test execution"; the paired practical, "Screenshot Capture and Logging Mechanism"

In one line

Test execution is actually running the tests: following each procedure, writing down exactly what the software did, comparing it with what it should have done, and keeping the evidence.

In the wording a student can write in an examination: test execution is the "process of running a test on the test item, producing actual results" (ISO/IEC/IEEE 29119-2:2021). It includes running the test procedures according to the test execution schedule, comparing actual results with expected results, recording the outcome in a test log, analysing anomalies and reporting failures as incidents, and then retesting fixes and running regression tests.

Why execution needs discipline

Execution is the part of testing everyone imagines, and it is easy to do badly. A tester who runs a test, sees something odd, shrugs and moves on has spent the time and kept none of the value. A tester who runs the right test on the wrong build, or with yesterday's data, produces a result nobody can use. And a tester who reports "the fee page is broken" without saying which test, which build, which data and what exactly was seen, gives the developer a puzzle instead of a defect.

Everything in the earlier activities, the plan, the conditions, the cases and procedures, was prepared so that execution could be quick, repeatable and evidential. This chapter is about making it so.

Before the first test runs

Three things are checked before execution starts, and each has its own name in the standards.

The entry criteria are met. The build has been delivered and installs; the test cases are approved; the people are available (Chapter Six, on the software testing life cycle).

The environment and data are ready. ISO/IEC/IEEE 29119-2 even names the documents that say so: a test environment readiness report, "document that describes the status of the test environment", and a test data readiness report, "document describing the status of each test data requirement". For ExamReg that means the test server has the release 2.0 build, its clock can be set, the payment gateway's test sandbox is reachable, and the six test student accounts exist with the right papers.

A smoke test passes. A short, broad check that the build's main functions respond at all (can a student log in, open the form, reach the fee page?) before hours are spent on detailed tests. A build that fails its smoke test goes back unused. Smoke testing has more to it and returns in Chapter Thirty-Nine, on regression testing, smoke testing and continuous integration.

Running a test and judging it

The tester follows the test procedure exactly as written, supplies the specified inputs, and records the actual results: the "set of behaviors or conditions of a test item, or set of conditions of associated data or the test environment, observed as a result of test execution" (ISO/IEC/IEEE 29119-2). Note that the definition includes the data and the environment: a fee shown correctly on screen but stored wrongly in the database is an actual result too, and a tester who checks only the screen misses it.

munotes.in51

Test Execution

Then comes the verdict. ISO/IEC/IEEE 29119-2 defines a test result as the "indication of whether a specific test case has passed or failed, i.e. if the actual results correspond to the expected results or if deviations were observed". In practice four states are recorded:

StatusMeaningExamReg example
PassActual results match the expected results7 days late, fee shown Rs 100
FailActual results differ from the expected results0 days late, fee shown Rs 100, expected nothing
BlockedThe test cannot be run, because a precondition cannot be metThe payment test cannot run because the gateway sandbox is down
Not runNot yet executed in this cycleHall ticket tests, scheduled for tomorrow

A blocked test is neither a pass nor a fail, and it must never be quietly counted as either. A report that shows 95 per cent passing because the blocked tests were left out of the total is misleading.

When the result is unexpected: anomalies and incidents

A failed comparison is not yet a defect in the software. The ISTQB syllabus describes what happens next: "Anomalies are analyzed to identify their likely causes. This analysis allows us to report the anomalies based on the failures observed." A failure can come from the software, from the test itself (a wrong expected result, the false positive of Chapter Two on errors, faults and failures), from the environment (a misconfigured server) or from the test data.

ISO/IEC/IEEE 29119-3:2021 calls the event a test incident: an "event occurring during the execution of a test that requires investigation". The tester records it in an incident report, "documentation of the occurrence, nature, and status of an incident" (ISO/IEC/IEEE 29119-2), which becomes a defect report once the cause is confirmed to be a defect. Writing a defect report has a chapter of its own in Module 2 (Chapter Seventy-Six).

The test log

Everything that happens during execution goes into the test log, which ISO/IEC/IEEE 24765:2017 defines as a "chronological record of relevant details about the execution of tests". ISO/IEC/IEEE 29119-3 calls the document a test execution log, one that "records details of the execution of one or more test procedures". A useful log entry answers: which test, on which build, in which environment, with which data, when, by whom, with what result, and where the evidence is.

munotes.in52

Test Execution

Evidence: logs and screenshots

The paired practical sets an exercise called "Screenshot Capture and Logging Mechanism", in which an automated test takes a screenshot when it fails and writes its progress to a log. The idea behind it is older than any tool. A failure that cannot be shown did not, for practical purposes, happen: the developer who receives the report will try to reproduce it, and if they cannot, the report is closed. A screenshot shows what the user saw at the moment of failure. A log shows what the program did on the way there. Together they turn "it went wrong" into evidence.

Logs are written at levels, so that a quiet run and a detailed investigation can use the same code. Apache Log4j 2's manual describes levels "from debug to fatal"; Python's standard logging module has DEBUG, INFO, WARNING, ERROR and CRITICAL. A test run typically logs each test's start at INFO, each failure at ERROR with its evidence, and detailed steps at DEBUG, switched on only when chasing a problem.

One currency point for the practical. Apache's own page for Log4j 1 records that "On August 5, 2015 the Logging Services Project Management Committee announced that Log4j 1.x had reached end of life", and it lists vulnerabilities in Log4j 1 of which it says "none of the issues listed will be fixed". A practical written today should use Log4j 2, as Apache recommends.

Worked example: an execution run with its log and evidence

The program below runs four ExamReg fee tests against release 2.0's function, which still has the on-time defect found in Chapter One, on what software testing is. It logs every step with Python's logging module, marks a test blocked when its precondition fails (the payment sandbox is down), and on a failure saves a text snapshot of what the fee page showed, the stand-in here for a screenshot. The log format leaves out the time so that the output is the same on every run; a real log starts each line with a timestamp.

import logging, os, sys

logging.basicConfig(stream=sys.stdout, level=logging.INFO,
                    format="%(levelname)-7s %(message)s")
log = logging.getLogger("examreg-tests")

def late_fee(days_late):                     # release 2.0, with the on-time defect
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    return 100 if days_late <= 7 else 500

def fee_page(days_late):                     # what the student would see
    return f"ExamReg fee page\nLate fee: Rs {late_fee(days_late)}\n"

sandbox_up = False                           # the payment gateway's test sandbox
tests = [("TC-FEE-01", 0, 0, None), ("TC-FEE-03", 7, 100, None),
         ("TC-FEE-04", 8, 500, None), ("TC-PAY-21", 3, 100, "sandbox")]
os.makedirs("evidence", exist_ok=True)
tally = {"pass": 0, "fail": 0, "blocked": 0}
log.info("build 2.0.1, environment TEST-2, cycle 1")
for case_id, days, expected, needs in tests:
    if needs == "sandbox" and not sandbox_up:
        log.warning("%s blocked: payment sandbox unreachable", case_id)
        tally["blocked"] += 1
        continue
    actual = late_fee(days)
    if actual == expected:
        log.info("%s pass: %d days late, fee Rs %d", case_id, days, actual)
        tally["pass"] += 1
    else:
        evidence = os.path.join("evidence", case_id + ".txt")
        with open(evidence, "w") as page:
            page.write(fee_page(days))
        log.error("%s FAIL: %d days late, expected Rs %d, got Rs %d; evidence %s",
                  case_id, days, expected, actual, evidence)
        tally["fail"] += 1
log.info("cycle 1 totals: %s", tally)
print(open("evidence/TC-FEE-01.txt").read(), end="")
munotes.in53

Test Execution

INFO    build 2.0.1, environment TEST-2, cycle 1
ERROR   TC-FEE-01 FAIL: 0 days late, expected Rs 0, got Rs 100; evidence evidence/TC-FEE-01.txt
INFO    TC-FEE-03 pass: 7 days late, fee Rs 100
INFO    TC-FEE-04 pass: 8 days late, fee Rs 500
WARNING TC-PAY-21 blocked: payment sandbox unreachable
INFO    cycle 1 totals: {'pass': 2, 'fail': 1, 'blocked': 1}
ExamReg fee page
Late fee: Rs 100

Read the log as a developer would. The build and environment come first, so there is no argument about which version was tested. The failure names the test, the input, the expected and actual results and where the evidence is, and the evidence file shows exactly what a student would have seen. The blocked test is reported as blocked, with its reason, and counted separately: it is a problem to fix in the environment, not a defect in the fee rule.

After the fix: retesting and regression testing

When the developer fixes the defect and delivers a new build, two different things are tested, and they have different names.

Retesting, also called confirmation testing, is "testing performed to check that modifications made to correct a fault have successfully removed the fault" (ISO/IEC/IEEE 29119-2). Here: run TC-FEE-01 again and see Rs 0.

Regression testing is "testing performed following modifications to a test item or to its operational environment, to identify whether failures in unmodified parts of the test item occur" (ISO/IEC/IEEE 29119-1). Here: rerun the other fee tests, and the form and payment tests near the changed code, to show the fix broke nothing else.

Retesting (confirmation testing)Regression testing
QuestionIs this defect really fixed?Did the change break anything that used to work?
Tests runThe test that failedTests that previously passed
ScopeNarrow: the fixed defectBroad: the areas around, or the whole product
WhenAfter a fix is deliveredAfter any change: a fix, a new feature, an environment change
Automated?Often manualThe first candidate for automation, because it is repeated every build

Stopping and restarting: suspension and resumption

Sometimes testing must stop. ISO/IEC/IEEE 29119-1 defines suspension criteria as "criteria used to (temporarily) stop all or a portion of the testing activities". A test plan states them in advance, together with the resumption requirements that must be met to restart. For ExamReg, for example: suspend if the build fails its smoke test, or if more than a quarter of the scheduled tests are blocked; resume when a build passes the smoke test and the blocking problem is fixed. Deciding this in advance stops testers wasting days on a build that was never going to be usable, and stops a manager pushing testing on regardless.

munotes.in54

Test Execution

What it does not mean

A failed test is not automatically a defect in the software. The test, the data or the environment may be wrong; anomalies are analysed first.

Blocked is not failed, and not passed. It means the test could not run, and it is counted separately.

Retesting and regression testing are not the same. One checks the fix; the other checks everything else around it.

Execution is not only for automated tests. Manual execution follows the same rules: exact procedures, recorded actual results, a log and evidence.

Quick revision

  • Test execution: "process of running a test on the test item, producing actual results" (ISO/IEC/IEEE 29119-2).
  • Before it: entry criteria, environment and data readiness, a passing smoke test.
  • Actual results include data and environment, not only the screen. Statuses: pass, fail, blocked, not run.
  • Anomalies are analysed: the cause may be the software, the test, the environment or the data. Test incident: an event during execution that needs investigation; recorded in an incident report.
  • Test log: "chronological record of relevant details about the execution of tests" (ISO/IEC/IEEE 24765).
  • Evidence: logs (by level) and screenshots on failure. Log4j 1 reached end of life on 5 August 2015; use Log4j 2.
  • Retesting (confirmation): checks a fix. Regression testing: checks unmodified parts after a change.
  • Suspension criteria stop testing; resumption requirements restart it.

Test yourself

1. What is test execution? List its main tasks. Running tests on the test item to produce actual results: executing procedures as scheduled, comparing actual with expected results, logging outcomes, analysing anomalies and reporting incidents, then retesting fixes and running regression tests.

2. Distinguish retesting from regression testing. Retesting reruns the test that failed, to confirm a fix removed the fault. Regression testing reruns tests that previously passed, after any change, to check that failures have not appeared in unmodified parts.

3. What does "blocked" mean as a test status, and why must it be reported separately? The test could not be run because a precondition could not be met, such as a gateway sandbox being down. It is neither a pass nor a fail, and hiding blocked tests in either count distorts the picture of quality.

4. Why is evidence such as a log or a screenshot attached to a failure? So the failure can be shown and reproduced: it records what the user saw and what the program did, with the build, environment and data, which lets the developer find the defect instead of arguing about whether it exists.

munotes.in55

Test Execution

5. What are suspension criteria? Give one example. Criteria, decided in advance, for temporarily stopping all or part of testing; for example, suspending when a build fails its smoke test, and resuming when a build passes it.

6. Which Log4j should a practical use today, and why? Log4j 2. Apache announced on 5 August 2015 that Log4j 1.x had reached end of life, and the vulnerabilities found in it since will not be fixed.

Contents This chapter on its own page

munotes.in56

Chapter Ten

Test Reporting

Syllabus topic Module 1, "Software Testing Fundamentals: Test execution, reporting"; the paired practical, "Test Execution Report"

In one line

Test reporting is telling the people who make decisions what the testing has found and how far it has got, clearly enough that they can act on it.

In the wording a student can write in an examination: test reporting summarises and communicates test information during and after testing (ISTQB v4.0.1). A test progress report, which ISO/IEC/IEEE 29119 calls a test status report, is produced regularly during testing to support test control; a test completion report, formerly called a test summary report, is produced once, when a test level, cycle or project ends, and evaluates the testing and the product against the test plan.

Why reporting matters

Testing produces information, and information that stays in a tester's notebook changes nothing. The release manager deciding whether ExamReg can open the examination form on Monday does not read four hundred test logs. The developer deciding what to fix first needs to know which failures are critical. The college principal wants to know whether students' fees will be charged correctly. Reporting is how testing's findings reach each of them in a form they can use.

It is also where testing is most easily misrepresented. A pass rate quoted without saying how many tests were not run, or a "green" dashboard that leaves out blocked tests, can send a system into production that should not go. A good report is honest before it is reassuring.

The two kinds of test report

ISO/IEC/IEEE 29119-2:2021 defines both. A test status report is a "report that provides information about the status of the testing that is being performed in a specified reporting period". A test completion report is a "report that provides a summary of the testing that was performed".

The ISTQB syllabus gives their purposes. Progress reports "support the ongoing test control and must provide enough information to make modifications to the test schedule, resources, or test plan". Completion reports "summarize a specific test activity (e.g., test level, test cycle, iteration) and can give information for subsequent testing".

Test progress (status) reportTest completion report
WhenRegularly during testing: daily, weekly, each cycleOnce, when a level, cycle, iteration or project ends
PurposeSupport test control: change the schedule, the resources, the planSummarise the testing and evaluate the product against the plan
Typical contents (ISTQB v4.0.1)Testing period; progress, ahead or behind, with deviations; impediments and workarounds; test metrics; new and changed risks; testing planned for next periodTest summary; evaluation against objectives and exit criteria; deviations from the plan; impediments and workarounds; metrics; unmitigated risks and defects not fixed; lessons learned
FormalityOften informal within the teamFollows a set template
Older nameTest status reportTest summary report (IEEE 829)
munotes.in57

Test Reporting

What the reports measure

The ISTQB syllabus lists the common test metrics, and a report draws on several of them:

  • Project progress: task completion, resource usage, test effort.
  • Test progress: test cases implemented; environment readiness; test cases run and not run, passed and failed; execution time.
  • Product quality: availability, response time, mean time to failure.
  • Defects: number and priorities found and fixed, defect density, defect detection percentage.
  • Risk: the residual risk level.
  • Coverage: requirements coverage, code coverage.
  • Cost: the cost of testing, the organisation's cost of quality.

Most of these have chapters of their own in Module 2, where metrics, defect metrics and quality costs are rows of the syllabus. A report uses them; it does not invent new ones.

The test execution report

The paired practical asks students to prepare a test execution report. It is the name industry uses for a test progress report focused on execution: after a day or a cycle of running tests, it says what was run, what passed, what failed, what was blocked, which defects are open, and what that means. Its usual sections are:

  1. Identification: project, release and build, test environment, cycle, reporting period, author.
  2. Execution summary: planned, executed, passed, failed, blocked, not run, as numbers and percentages.
  3. By module or feature: the same figures broken down, so problems can be located.
  4. Defects: open defects by severity, new and closed in the period.
  5. Impediments: what blocked testing and what is being done about it.
  6. Risks and recommendation: what the figures mean for the release, measured against the exit criteria.

Worked example: ExamReg, cycle 1, day 3

ExamReg release 2.0 has 240 planned test cases across five modules. At the end of day 3 of the first execution cycle, the counts per module are as follows, and the program produces the execution summary from them. Computing it, rather than typing it, is not a formality: percentages typed by hand are one of the commonest errors in real reports.

modules = {        # planned, passed, failed, blocked (not run is what remains)
    "Login":         (30, 29, 1, 0),
    "Exam form":     (80, 61, 9, 0),
    "Fee payment":   (60, 42, 5, 8),
    "Hall ticket":   (40, 12, 1, 0),
    "Admin reports": (30, 0, 0, 0),
}
open_defects = {"critical": 1, "major": 5, "minor": 6, "cosmetic": 3}

def pct(part, whole):
    return f"{100 * part / whole:5.1f}%" if whole else "    -"

print("ExamReg release 2.0, build 2.0.1, environment TEST-2, cycle 1, day 3")
print(f"{'module':<14}{'planned':>8}{'run':>5}{'pass':>6}{'fail':>6}{'blocked':>9}{'not run':>9}{'pass rate':>11}")
totals = [0, 0, 0, 0, 0, 0]
for name, (planned, passed, failed, blocked) in modules.items():
    run = passed + failed
    not_run = planned - run - blocked
    row = [planned, run, passed, failed, blocked, not_run]
    totals = [t + r for t, r in zip(totals, row)]
    print(f"{name:<14}{planned:>8}{run:>5}{passed:>6}{failed:>6}{blocked:>9}{not_run:>9}{pct(passed, run):>11}")
planned, run, passed, failed, blocked, not_run = totals
print(f"{'TOTAL':<14}{planned:>8}{run:>5}{passed:>6}{failed:>6}{blocked:>9}{not_run:>9}{pct(passed, run):>11}")
print()
print("executed", run, "of", planned, "planned:", pct(run, planned).strip())
print("passed", passed, "of", planned, "planned:", pct(passed, planned).strip())
print("open defects:", ", ".join(f"{n} {s}" for s, n in open_defects.items()))
munotes.in58

Test Reporting

ExamReg release 2.0, build 2.0.1, environment TEST-2, cycle 1, day 3
module         planned  run  pass  fail  blocked  not run  pass rate
Login               30   30    29     1        0        0      96.7%
Exam form           80   70    61     9        0       10      87.1%
Fee payment         60   47    42     5        8        5      89.4%
Hall ticket         40   13    12     1        0       27      92.3%
Admin reports       30    0     0     0        0       30          -
TOTAL              240  160   144    16        8       72      90.0%

executed 160 of 240 planned: 66.7%
passed 144 of 240 planned: 60.0%
open defects: 1 critical, 5 major, 6 minor, 3 cosmetic

The report then says, in words, what the numbers mean:

  • Progress. Two thirds of the planned tests have run in three of the five days scheduled. Admin reports have not started; the team is behind schedule there and ahead on login.
  • Quality. Nine in ten executed tests pass, but the failures are concentrated: the exam form has 9 of the 16, which matches the defect clustering seen in Chapter Four's principles of testing.
  • Impediments. Eight fee-payment tests are blocked because the gateway's test sandbox is down; the vendor has promised it for tomorrow. They are not counted as passed.
  • Risk and recommendation. One critical defect is open (on-time students charged a late fee). The exit criteria set in Chapter Six, on the software testing life cycle, cannot be met while it is open, so the recommendation is: fix it first, and move one tester from login, which is finished, to admin reports.

Notice the two different percentages at the end. Passed of executed is 90 per cent and sounds good. Passed of planned is 60 per cent, because 80 tests have not yet run or cannot run. Both are true, and a report that gave only the first would mislead the release manager into thinking the product was nearly ready. The honest report gives both.

Audience, formality and channel

The ISTQB syllabus notes that "Different audiences require different information in the reports and influence the degree of formality and the frequency of test reporting." The developers want the list of failures with evidence; the project manager wants progress against schedule and the blockers; senior management wants the risk and the recommendation, in a sentence.

The channel changes with the audience too. The syllabus lists verbal communication, dashboards (for example continuous integration dashboards, task boards and burn-down charts), email and chat, online documentation, and formal test reports. A team in one room talks; a team spread across offices and time zones writes more down.

munotes.in59

Test Reporting

Principles of an honest report

  1. Numbers come from the data, not from memory. Counts and percentages are computed from the test management records.
  2. Blocked and not-run tests are shown. Leaving them out inflates the pass rate.
  3. Every percentage names its base. "90 per cent" of what: of the tests run, or of the tests planned?
  4. Facts are separated from judgement. The figures first, then the tester's interpretation, clearly labelled as a recommendation.
  5. The report is measured against the plan. Exit criteria and the schedule are the yardstick; without them "good" and "bad" have no meaning.
  6. Risks are stated, not softened. An open critical defect is written as what it is.

What it does not mean

A test report is not a list of every test. It summarises; the logs and the test management tool hold the detail.

A high pass rate does not mean a product is ready. It depends on how many tests were run, which ones failed, and whether the exit criteria are met.

The completion report is not the last progress report. It evaluates the whole activity against the plan, records deviations and lessons learned, and is written once.

Reporting is not the tester's opinion. It is evidence first; the recommendation that follows must be traceable to that evidence.

Quick revision

  • Test reporting: summarising and communicating test information during and after testing.
  • Test status (progress) report: regular, during testing, supports control. Test completion report: once, at the end of a level, cycle or project; evaluates against the plan (formerly the IEEE 829 test summary report).
  • Progress report contents: period; progress and deviations; impediments; metrics; risks; next period's plan.
  • Completion report contents: summary; evaluation against objectives and exit criteria; deviations; impediments; metrics; unmitigated risks and unfixed defects; lessons learned.
  • Metric families (ISTQB): project progress, test progress, product quality, defects, risk, coverage, cost.
  • Test execution report (the practical): identification, execution summary, by module, defects, impediments, recommendation.
  • Worked example: 160 of 240 run, 144 passed: 90 per cent of executed, 60 per cent of planned; 8 blocked; one critical defect open.

Test yourself

1. Distinguish a test progress report from a test completion report. A progress (status) report is produced regularly during testing to support test control, covering progress, deviations, impediments, metrics, risks and the next period's plan. A completion report is produced once when a level, cycle or project ends, and evaluates the testing and the product against the plan's objectives and exit criteria, with deviations and lessons learned.

2. List the contents of a test execution report. Identification (release, build, environment, cycle, period); an execution summary (planned, executed, passed, failed, blocked, not run); a breakdown by module; open, new and closed defects by severity; impediments; and risks with a recommendation against the exit criteria.

munotes.in60

Test Reporting

3. Why must a report give both "passed of executed" and "passed of planned"? Because they answer different questions. In the worked example 90 per cent of executed tests passed, but only 60 per cent of planned tests have passed so far; quoting only the first hides the 80 tests not yet run or blocked.

4. Name four families of test metrics. Any four of: project progress, test progress, product quality, defect, risk, coverage and cost metrics.

5. How does audience change a test report? Developers need the failures and their evidence; project managers need progress, blockers and schedule; senior management needs the risk and the recommendation. The audience also sets the formality and frequency, from a daily conversation to a formal completion report.

Contents This chapter on its own page

munotes.in61

Chapter Eleven

Writing a Test Plan

Syllabus topic Module 1, "Software Testing Fundamentals: Test execution, reporting, and documentation"; the paired practical, "Prepare a Test Plan"

In one line

A test plan is the document that says, before testing starts, what will be tested and why, how, by whom, with what, by when, and what "finished" will mean.

In the wording a student can write in an examination: a test plan is a "detailed description of test objectives to be achieved and the means and schedule for achieving them, organized to coordinate testing activities for some test item or set of test items" (ISO/IEC/IEEE 29119-2:2021); in the older wording, a "document describing the scope, approach, resources, and schedule of intended test activities" (IEEE 1012-2024). The classic outline, from IEEE 829-1998, has sixteen headings, from the test plan identifier to approvals.

Why a plan is written at all

A plan's first value is not the document but the thinking it forces. The ISTQB syllabus says it well: test planning "guides the testers' thinking and forces the testers to confront the future challenges related to risks, schedules, people, tools, costs, effort, etc." A tester who has to write down what will not be tested, what will stop testing, and what happens if the test environment arrives late, has thought about all three before they happen.

The syllabus lists four things a finished plan does. It documents the means and schedule for achieving the test objectives. It helps ensure the test activities meet the established criteria. It serves as communication with the team and other stakeholders. And it shows that testing follows the organisation's test policy and strategy, or explains why it does not.

Policy, strategy, plan

The plan sits at the bottom of a three-level hierarchy, and the words are easy to confuse.

DocumentDefinitionScope
Test policy"executive-level document that describes the purpose, goals, principles, and scope of testing within an organization" (ISO/IEC/IEEE 29119-3:2021)The whole organisation, for years
Test strategy"part of the test plan that describes the approach to testing for a specific project, test level, or test type" (ISO/IEC/IEEE 29119-2:2021)One project, level or type
Test planObjectives, means and schedule for testing a test item or set of items (ISO/IEC/IEEE 29119-2:2021)One project or one test level

Inside the strategy sits the test approach, a "high-level test implementation choice, typically made as part of the test strategy design activity" (ISO/IEC/IEEE 29119-1:2022): for example, risk-based, with specification-based techniques at system level and automated regression at every build. A project may have one master test plan covering all levels and a level test plan for each level, which is how IEEE 829-2008 organised them.

The classic outline: IEEE 829-1998's sixteen headings

IEEE 829, the Standard for Software Test Documentation, was first published in 1983, revised in 1998 and again in 2008, and has since been superseded by ISO/IEC/IEEE 29119-3. Its 1998 outline for a test plan is still the one most textbooks and many companies use, and it is a sound checklist. The sixteen headings, each with what goes under it:

munotes.in62

Writing a Test Plan

No.HeadingWhat goes under it
1Test plan identifierA unique name and version for this plan
2IntroductionPurpose, scope, objectives, and the documents this plan relies on
3Test itemsThe software items to be tested, with their versions
4Features to be testedEach feature or requirement that will be tested
5Features not to be testedWhat is deliberately left out, and why
6ApproachHow testing will be done: levels, types, techniques, tools, automation
7Item pass/fail criteriaHow to decide whether each test item has passed
8Suspension criteria and resumption requirementsWhen to stop testing, and what must be true to restart
9Test deliverablesEvery document and artefact testing will produce
10Testing tasksThe tasks, their dependencies and the skills they need
11Environmental needsHardware, software, network, tools, test data, facilities
12ResponsibilitiesWho does what: test, fix, provide environments, approve
13Staffing and training needsThe people needed and any training they require
14ScheduleMilestones and dates, tied to the development schedule
15Risks and contingenciesWhat could go wrong with testing, and the fallback for each
16ApprovalsWho must sign the plan, with names and dates

Heading 5 surprises students most. Writing down what will not be tested is not an admission of failure; it is the point at which a manager can object. If the plan says load testing of the payment gateway is not in scope because the vendor certifies it, the principal can either accept that risk or insist otherwise, before the deadline rush rather than during it.

The same content, as ISTQB describes a plan today

The ISTQB syllabus (2024) lists the typical content of a test plan more compactly: the context of testing (scope, objectives, test basis); assumptions and constraints; stakeholders (roles, responsibilities, hiring and training needs); communication (forms, frequency, templates); a risk register of product and project risks; the test approach (levels, types, techniques, deliverables, entry and exit criteria, independence of testing, metrics, test data and environment requirements, deviations from policy and strategy); and budget and schedule. The two lists cover the same ground:

ISTQB v4.0.1 contentIEEE 829-1998 headings that carry it
Context of testing2 Introduction, 3 Test items, 4 and 5 Features to be tested and not to be tested
Assumptions and constraints2 Introduction, 15 Risks and contingencies
Stakeholders12 Responsibilities, 13 Staffing and training needs, 16 Approvals
Communication(not a heading of its own in 1998)
Risk register15 Risks and contingencies
Test approach6 Approach, 7 Item pass/fail criteria, 8 Suspension and resumption, 9 Deliverables, 11 Environmental needs
Budget and schedule10 Testing tasks, 14 Schedule
munotes.in63

Writing a Test Plan

Entry and exit criteria in the plan

The plan is where entry and exit criteria are set, for each test level. The ISTQB syllabus gives typical ones. Entry criteria concern the availability of resources (people, tools, environments, test data, budget, time), of testware (test basis, testable requirements, test cases), and the initial quality of the test object, such as "all smoke tests have passed". Exit criteria are either measures of thoroughness (coverage achieved, unresolved defects, defect density, failed test cases) or yes/no criteria (planned tests executed, static testing performed, all defects reported, regression tests automated).

The syllabus adds a realistic note: "Running out of time or budget can also be viewed as valid exit criteria", provided "the stakeholders have reviewed and accepted the risk to go live without further testing." And in agile teams the same ideas have other names: exit criteria are the Definition of Done, and the entry criteria a user story must meet before work starts are the Definition of Ready.

How much effort? Estimation in the plan

The schedule and budget depend on an estimate of test effort. The ISTQB syllabus describes four techniques. Estimation based on ratios uses figures from past projects, such as a test-to-development effort ratio. Extrapolation measures the current project early and projects forward, for example averaging the last three iterations. Wideband Delphi has experts estimate separately, discuss the outliers and re-estimate until they agree; Planning Poker is its agile variant. Three-point estimation asks for an optimistic estimate a, a most likely m and a pessimistic b, and combines them as E = (a + 4m + b) / 6, with a spread SD = (b - a) / 6.

The program applies three-point estimation to four of ExamReg's test tasks, in person-hours.

tasks = {                                   # optimistic, most likely, pessimistic
    "design fee and form tests":    (16, 24, 44),
    "set up test environment":      (6, 8, 16),
    "execute cycle 1":              (40, 56, 90),
    "execute regression cycle":     (12, 16, 26),
}
total = 0
for task, (a, m, b) in tasks.items():
    estimate = (a + 4 * m + b) / 6
    spread = (b - a) / 6
    total += estimate
    print(f"{task:<28} E = {estimate:5.1f}  SD = {spread:4.1f}  "
          f"(likely {estimate - spread:.1f} to {estimate + spread:.1f})")
print(f"{'total expected effort':<28} E = {total:5.1f} person-hours")
design fee and form tests    E =  26.0  SD =  4.7  (likely 21.3 to 30.7)
set up test environment      E =   9.0  SD =  1.7  (likely 7.3 to 10.7)
execute cycle 1              E =  59.0  SD =  8.3  (likely 50.7 to 67.3)
execute regression cycle     E =  17.0  SD =  2.3  (likely 14.7 to 19.3)
total expected effort        E = 111.0 person-hours
munotes.in64

Writing a Test Plan

Notice that each estimate E is larger than the most likely value m. The pessimistic figure pulls it up, because in testing, as in most work, things tend to take longer rather than shorter than hoped: a blocked environment or a defect-heavy build costs far more time than a smooth one saves. A plan that used m alone would be optimistic by construction.

Worked example: a test plan for ExamReg release 2.0

Below is a complete plan in the IEEE 829-1998 outline, the form the paired practical asks for. It is short because ExamReg is small; a bank's plan under the same headings runs to fifty pages.

1. Test plan identifier. STP-EXAMREG-2.0, version 1.1.

2. Introduction. This plan covers system testing and user acceptance testing of ExamReg release 2.0, the college's online examination registration portal, before the examination form opens for the November session. Objectives: show that fees are charged exactly as the fee rules say, that eligible students can register and ineligible ones cannot, and that the portal handles the last-day rush. References: ExamReg requirements v2.0, fee rules circular, test policy TP-01.

3. Test items. ExamReg web application build 2.0.x; the fee calculation service; the hall ticket generator; the payment gateway adapter (the gateway itself is the vendor's).

4. Features to be tested. Login and lockout; exam form, including backlog papers and eligibility; fee calculation, including late fees and concessions; payment and receipts; hall ticket generation; admin reports.

5. Features not to be tested. The payment gateway's internal processing (certified by the vendor; only our adapter is tested); the college's existing student database (read only, unchanged in this release); translations (English only in 2.0).

6. Approach. Risk-based. Specification-based techniques (equivalence partitioning, boundary values, decision tables for fee rules, state transitions for login); automated regression of fee and form tests on every build; a load test of 500 simultaneous users before release; exploratory sessions on the exam form. Independent testing by the college IT cell's test team.

7. Item pass/fail criteria. An item passes when all its high-priority tests pass and it has no open critical or major defect.

8. Suspension criteria and resumption requirements. Suspend if a build fails the smoke test or more than 25 per cent of scheduled tests are blocked. Resume on a build that passes the smoke test with the blocking cause fixed.

9. Test deliverables. This plan; test scenarios and cases; traceability matrix; test data; execution reports for each cycle; defect reports; load test report; test completion report.

10. Testing tasks. Requirement review; test design; environment set-up; cycle 1; defect fixing and retesting; regression cycle; load test; user acceptance testing; completion report. Test design depends on the approved requirements; execution depends on build delivery.

munotes.in65

Writing a Test Plan

11. Environmental needs. Test server TEST-2 matching production; a controllable server clock; the payment gateway's sandbox; six test student accounts with prepared papers; Chrome, Firefox and Edge on desktop and on Android.

12. Responsibilities. Test lead: plan, monitor, report. Two testers: design and execution. Developers: fixes and unit tests. IT cell: environments. Exam cell: acceptance testing and answers to requirement questions.

13. Staffing and training needs. Two testers and a test lead for four weeks; one tester trained on the load testing tool.

14. Schedule. Test design weeks 1 and 2; environment ready end of week 1; cycle 1 week 3; regression and load test week 4; acceptance testing and sign-off at the end of week 4, one week before the form opens.

15. Risks and contingencies. Late build: reduce cycle 1 scope to high-priority tests. Gateway sandbox unavailable: test the adapter against a simulator and retest when the sandbox returns. Requirement changes from the exam cell: frozen after week 1 except for defects.

16. Approvals. Test lead; development lead; head of the exam cell; principal. Names, signatures and dates.

What it does not mean

A test plan is not a list of test cases. It plans the testing; test cases are designed later, in test design, and live in their own documents.

IEEE 829-1998 is not the current standard. It was superseded by IEEE 829-2008 and then by ISO/IEC/IEEE 29119-3; its sixteen headings survive because they make a good checklist.

"Features not to be tested" is not a weakness. It records a decision and its reason, so the people who carry the risk can accept or reject it.

A plan is not written once and frozen. It is revised when the project changes, and test control acts on it throughout.

Quick revision

  • Test plan: objectives, means and schedule for testing a test item or items (ISO/IEC/IEEE 29119-2); scope, approach, resources and schedule (IEEE 1012-2024).
  • Hierarchy: test policy (organisation) above test strategy (project, level or type) above the test plan; the test approach is chosen within the strategy.
  • IEEE 829-1998's 16 headings: identifier; introduction; test items; features to be tested; features not to be tested; approach; item pass/fail criteria; suspension and resumption; deliverables; testing tasks; environmental needs; responsibilities; staffing and training; schedule; risks and contingencies; approvals.
  • ISTQB v4.0.1 content: context, assumptions and constraints, stakeholders, communication, risk register, test approach, budget and schedule.
  • Entry and exit criteria per level; in agile, Definition of Ready and Definition of Done.
  • Estimation: ratios, extrapolation, Wideband Delphi, three-point (E = (a + 4m + b) / 6, SD = (b - a) / 6).
munotes.in66

Writing a Test Plan

Test yourself

1. Define a test plan and state its purposes. A document describing the test objectives and the means, resources and schedule for achieving them. It documents how the objectives will be met, helps ensure the test activities meet their criteria, communicates with the team and stakeholders, and shows that testing follows the test policy and strategy.

2. List the sixteen headings of the IEEE 829-1998 test plan. Test plan identifier; introduction; test items; features to be tested; features not to be tested; approach; item pass/fail criteria; suspension criteria and resumption requirements; test deliverables; testing tasks; environmental needs; responsibilities; staffing and training needs; schedule; risks and contingencies; approvals.

3. Distinguish test policy, test strategy and test plan. The policy is an executive document stating the purpose, goals and principles of testing across the organisation. The strategy describes the approach to testing for a particular project, level or type. The plan sets the objectives, means, resources and schedule for testing specific items.

4. Give typical entry and exit criteria for system testing. Entry: the build has passed its smoke test, the environment and test data are ready, and the test cases are approved. Exit: all high-priority tests executed and passed, requirements coverage complete, and no open critical defect.

5. Using three-point estimation, estimate a task with a = 6, m = 9 and b = 18 person-hours. E = (6 + 4 × 9 + 18) / 6 = 10 person-hours, with SD = (18 - 6) / 6 = 2, so the task is likely to take between 8 and 12 person-hours. This is the worked example in the ISTQB syllabus.

6. Why should a test plan say which features will not be tested? Because excluding a feature is a risk decision, and recording it with its reason lets the stakeholders who carry the risk accept or reject it before testing, not discover it after release.

Contents This chapter on its own page

munotes.in67

Chapter Twelve

The Test Documents, From Design to Completion

Syllabus topic Module 1, "Software Testing Fundamentals: Test execution, reporting, and documentation"

In one line

Testing produces a small set of standard documents, one for each job: planning the testing, specifying the tests, and recording what happened when they ran.

In the wording a student can write in an examination: IEEE 829-1998, the Standard for Software Test Documentation, defines eight documents in three groups: test planning (the test plan); test specification (the test design specification, the test case specification and the test procedure specification); and test reporting (the test item transmittal report, the test log, the test incident report and the test summary report). IEEE 829 has since been superseded by ISO/IEC/IEEE 29119-3, which defines the equivalent documents under names such as test status report, test completion report, test execution log and incident report.

Why testing is documented

Test documents are not paperwork for its own sake. They carry four things that cannot be carried any other way.

Evidence. When ExamReg goes live, the principal who approved it, an auditor, or a regulator in a safety-critical field can ask what was tested and what was found. Only documents answer.

Repeatability. A test that exists only in one tester's head cannot be rerun by someone else, or by the same person next year.

Handover. Test design, execution and reporting are often done by different people; the documents are how one hands over to the next.

Improvement. The logs, incident reports and completion reports of this release are the data that make the next release's testing better.

ISO/IEC/IEEE 24765:2017 defines test documentation simply as "documentation describing plans for, or results of, the testing of a system or component".

IEEE 829-1998: the eight documents

The 1998 standard organises its documents by the three things testing needs to write down.

GroupDocumentWhat it is for
PlanningTest planThe scope, approach, resources and schedule of the testing: what will be tested, the tasks, who does them, and the risks (Chapter Eleven, on writing a test plan)
SpecificationTest design specificationRefines the approach for a group of features: which features the tests cover, which test cases and procedures are needed, and the pass/fail criteria for each feature
SpecificationTest case specificationThe actual input values and the expected outputs of each test case, and any constraints the case places on how it is run; kept separate from the design so a case can be reused in more than one design
SpecificationTest procedure specificationEvery step needed to run the specified test cases, in order, written to be followed step by step without extra detail
ReportingTest item transmittal reportIdentifies the items being handed over for testing, when a separate development group delivers builds to a test group, or when a formal start of execution is wanted
ReportingTest logThe test team's record of what happened during execution
ReportingTest incident reportDescribes any event during execution that needs further investigation
ReportingTest summary reportSummarises the testing done against one or more test designs, and evaluates the results
munotes.in68

The Test Documents, From Design to Completion

Two points about the set are worth making in an answer. First, the specification is split three ways on purpose: the design says what to cover, the case says with what values, the procedure says in what steps. Keeping them apart lets one test case be reused in several designs and one procedure run many cases. Second, the reporting documents follow execution in time: the transmittal report comes before it, the log during it, the incident reports as problems appear, and the summary after.

IEEE 829-2008 and ISO/IEC/IEEE 29119-3

The 2008 revision renamed and extended the set around test levels: a Master Test Plan over all levels and a Level Test Plan for each; Level Test Design, Level Test Case and Level Test Procedure documents; a Level Test Log; an Anomaly Report; a Level Interim Test Status Report; a Level Test Report; and a Master Test Report. The name anomaly report was chosen deliberately, as Wikipedia's account of the standard explains, because a discrepancy between expected and actual results can have causes other than a fault in the system, such as a wrong expected result.

IEEE 829-2008 has in turn been superseded by ISO/IEC/IEEE 29119-3, whose current edition, the 2021 one, is the standard SEVOCAB draws its test documentation terms from. The names map as follows.

IEEE 829-1998ISO/IEC/IEEE 29119 (2021 and 2022 editions)29119 definition, where SEVOCAB gives one
Test planTest plan"detailed description of test objectives to be achieved and the means and schedule for achieving them" (29119-2)
Test design specificationTest design, within the test specificationTest specification: "complete documentation of the test design, test cases, and test procedures for a specific test item" (29119-2)
Test case specificationTest case specification"documentation of a set of one or more test cases" (29119-2)
Test procedure specificationTest procedure, within the test specification"sequence of test cases in execution order, with associated actions required to set up preconditions and perform wrap-up activities post execution" (29119-2)
Test item transmittal report(no single equivalent; readiness is reported instead)Test environment readiness report: "document that describes the status of the test environment" (29119-2)
Test logTest execution log"record of the execution of one or more test procedures" (29119-2)
Test incident reportIncident report"documentation of the occurrence, nature, and status of an incident" (29119-2)
Test summary reportTest completion report"report that provides a summary of the testing that was performed" (29119-2)
(none)Test status report"report that provides information about the status of the testing that is being performed in a specified reporting period" (29119-2)
(none)Test data requirements, test environment requirementsTest environment requirements: "description of the necessary properties of the test environment" (29119-2)
(none)Test traceability matrix"document, spreadsheet, or other automated tool used to identify related items in documentation and software, such as requirements with associated tests" (29119-3)
munotes.in69

The Test Documents, From Design to Completion

The newer standard adds documents for things 1998 left implicit: the status report during testing (Chapter Ten, on test reporting), the readiness of the environment and the data (Chapter Nine, on test execution), and the traceability matrix.

Must every document be produced?

No. Wikipedia's account of IEEE 829-2008 notes that the standard "specified the format of these documents, but did not stipulate whether they must all be produced". Which documents a project writes is decided in its test plan, and the answer depends on context: a medical device project may produce every one, formally reviewed and signed; a three-person team building ExamReg may keep test cases and logs in a test management tool and write only a plan and a completion report as documents.

Agile teams go further. The Agile Manifesto values "Working software over comprehensive documentation", and it adds that "while there is value in the items on the right, we value the items on the left more". In agile testing the same information is still kept, but lighter: acceptance criteria on user stories instead of design specifications, automated tests that are themselves the test cases, and the continuous integration server's history instead of a hand-written log.

Traceability across the documents

The documents are only as useful as the links between them. A requirement should lead forward to its test cases, their procedures, their log entries and any incidents; and an incident should lead back to the requirement it affects. The program below records those links for one ExamReg requirement across two execution cycles, then follows them in both directions.

requirement = {"FEE-1": "no late fee on or before the last date"}
test_cases = {"TC-FEE-01": {"requirement": "FEE-1", "procedure": "TP-FEE"}}
log = [   # test log entries: cycle, test case, build, result, incident raised
    {"cycle": 1, "case": "TC-FEE-01", "build": "2.0.1", "result": "fail", "incident": "IR-014"},
    {"cycle": 2, "case": "TC-FEE-01", "build": "2.0.2", "result": "pass", "incident": None},
]
incidents = {"IR-014": {"case": "TC-FEE-01", "status": "closed after retest on 2.0.2"}}

print("forward, from the requirement:")
for rid, text in requirement.items():
    print(f"  {rid}: {text}")
    for case, info in test_cases.items():
        if info["requirement"] == rid:
            print(f"    test case {case}, run by procedure {info['procedure']}")
            for entry in log:
                if entry["case"] == case:
                    note = f", incident {entry['incident']}" if entry["incident"] else ""
                    print(f"      cycle {entry['cycle']} on build {entry['build']}: {entry['result']}{note}")

print("backward, from the incident:")
case = incidents["IR-014"]["case"]
rid = test_cases[case]["requirement"]
print(f"  IR-014 ({incidents['IR-014']['status']}) <- {case} <- {rid}: {requirement[rid]}")
munotes.in70

The Test Documents, From Design to Completion

forward, from the requirement:
  FEE-1: no late fee on or before the last date
    test case TC-FEE-01, run by procedure TP-FEE
      cycle 1 on build 2.0.1: fail, incident IR-014
      cycle 2 on build 2.0.2: pass
backward, from the incident:
  IR-014 (closed after retest on 2.0.2) <- TC-FEE-01 <- FEE-1: no late fee on or before the last date

Forward, the chain answers "was this requirement tested, and what happened?": it failed on build 2.0.1, an incident was raised, and it passed on 2.0.2. Backward, it answers "which rule does this incident break?" in one step. Without the links in the documents, both questions need a person's memory.

What it does not mean

IEEE 829 is not the current standard. The 1998 version was superseded in 2008, and 829-2008 by ISO/IEC/IEEE 29119-3. Its eight documents are still taught because they divide the work cleanly.

Not every document must be written. The standards define formats; the test plan decides which documents a project produces.

A test case specification is not a test design specification. The design says which features are covered and how they pass; the case gives the actual values and expected outputs.

Less documentation in agile does not mean less information. The same facts are kept in lighter forms: acceptance criteria, automated tests, tool histories.

Quick revision

  • IEEE 829-1998: eight documents in three groups.
  • Planning: test plan.
  • Specification: test design specification (what to cover, feature pass/fail criteria); test case specification (input values and expected outputs); test procedure specification (steps to run cases).
  • Reporting: test item transmittal report (items handed over); test log (what happened); test incident report (events to investigate); test summary report (summary and evaluation).
  • IEEE 829-2008: master and level test plans, level design, case, procedure, log, anomaly report, interim status report, level and master test reports.
  • ISO/IEC/IEEE 29119-3 supersedes IEEE 829: test status report, test completion report, test execution log, incident report, readiness reports, test traceability matrix.
  • Not every document is mandatory; agile keeps the same information more lightly.

Test yourself

1. Name the eight documents of IEEE 829-1998 in their groups. Planning: the test plan. Specification: the test design specification, the test case specification and the test procedure specification. Reporting: the test item transmittal report, the test log, the test incident report and the test summary report.

2. Distinguish the test design, test case and test procedure specifications. The design specification says which features are covered, which cases and procedures are needed, and the pass/fail criteria per feature. The case specification gives the actual input values and expected outputs of each case. The procedure specification lists the steps to run the cases, in order.

munotes.in71

The Test Documents, From Design to Completion

3. What is a test item transmittal report, and when is it used? A document identifying the items being handed over for testing, used when a separate development group delivers builds to a test group, or when a formal start of test execution is wanted.

4. Which standard replaced IEEE 829, and what do the log and summary report become there? ISO/IEC/IEEE 29119-3, now in its 2021 edition. The test log becomes the test execution log and the test summary report becomes the test completion report; the incident report keeps its name, and a test status report is added for progress during testing.

5. Why did IEEE 829-2008 call it an anomaly report rather than a fault report? Because a difference between expected and actual results can arise from causes other than a fault in the system, such as a wrong expected result or a test run incorrectly.

6. Are all the documents mandatory for every project? No. The standards specify the format of each document, not whether it must be produced; the test plan decides, according to context, and agile teams keep the same information in lighter forms.

Contents This chapter on its own page

munotes.in72

Chapter Thirteen

What a Software Development Life Cycle Is

Syllabus topic Module 1, "Software Development Life Cycle (SDLC): Overview of SDLC"

In one line

A software development life cycle is the sequence of stages a piece of software passes through, from the first idea to the day it is switched off, and an SDLC model is a particular way of arranging those stages.

In the wording a student can write in an examination: a life cycle is the "evolution of a system, product, service, project or other human-made entity from conception through retirement" (ISO/IEC/IEEE 12207:2026). A life cycle model is a "framework of processes and activities concerned with the life cycle that can be organized into stages, acting as a common reference for communication and understanding" (the same standard). The typical phases of software development are feasibility, requirements, design, implementation, testing, deployment and maintenance; models such as waterfall, the V-model, spiral and agile arrange these phases differently.

Why life cycle models exist

Building ExamReg involves a dozen people over months: the exam cell that knows the rules, a business analyst who writes them down, designers, programmers, testers, the IT cell that runs the servers. Without an agreed shape for the work, each of them decides privately when their part starts and what they hand over, and the handovers fail. A life cycle model is that agreed shape: which activities happen, in what order or in what cycles, and what each produces for the next.

The ISTQB syllabus puts it precisely: an SDLC model "defines how different development phases and types of activities performed within this process relate to each other, both logically and chronologically." The model is not the work itself; it is the map everyone reads from.

The phases every life cycle contains

Whatever the model, the same kinds of work have to be done. Models differ in how they order and repeat them, not in whether they exist.

The generic phases of software development and the work product of each

Figure 13.1 The phases of a software development life cycle, with maintenance feeding back into requirements

PhaseWhat happensStandard definition, where one existsWork product
FeasibilityDecide whether the system is worth buildingA feasibility study is a "study to identify and analyze a problem and its potential solutions in order to determine their viability, costs, and benefits" (ISO/IEC 2382:2015)Feasibility report
RequirementsFind out and write down what the system must doRequirements analysis is the "process of studying user needs to arrive at a definition of system, hardware, or software requirements" (ISO/IEC/IEEE 24765:2017)Requirements specification
DesignDecide how the system will do it: architecture, modules, data, interfacesSoftware design is the "use of scientific principles, technical information, and imagination in the definition of a software system to perform pre-specified functions with maximum economy and efficiency" (ISO/IEC/IEEE 24765:2017)Design documents
ImplementationWrite the code, and the developers' own unit tests(coding)Source code, unit tests
TestingIntegrate and test the system against its requirements(the testing fundamentals and test documents of Chapters One to Twelve)Test reports, fixed defects
DeploymentPut the system into useDeployment is the "stage of a project in which a system is put into operation and transition issues are resolved" (IEEE 2675-2021)The released system, user guides
MaintenanceCorrect, improve and adapt it after releaseMaintenance is the "process of modifying a software system or component after delivery to correct faults, improve performance or other attributes, or adapt to a changed environment" (ISO/IEC 25051:2014)Fixes, changes, new releases
munotes.in73

What a Software Development Life Cycle Is

The arrow from maintenance back to requirements in the figure is the part students forget. A system that is used gets changed, and every change runs through requirements, design, code and test again. For most successful software, far more of its life is spent in that loop than in the first build.

The families of models

The ISTQB syllabus groups the models into three families, and adds the agile methods that sit alongside them.

FamilyHow it arranges the phasesExamples (ISTQB v4.0.1)Chapter
SequentialOne phase after another, each finished before the next beginsWaterfall model, V-modelFourteen and Fifteen
IterativeThe phases are repeated in cycles, each refining the productSpiral model, prototypingSixteen
IncrementalThe product is built and delivered in pieces, each adding capabilityUnified ProcessSixteen
Agile methods and practicesShort iterations, incremental delivery, change welcomedScrum, Kanban, extreme programming, test-driven, behaviour-driven and acceptance test-driven developmentSeventeen

The standard vocabulary defines the sequential model's defining property precisely. ISO/IEC/IEEE 24765:2017 describes the waterfall model as one in which the phases "are performed in that order, possibly with overlap but with little or no iteration", and incremental development as a technique in which requirements, design, implementation and testing "occur in an overlapping, iterative (rather than sequential) manner, resulting in incremental completion of the overall software product."

Why the choice of model matters to a tester

This is a testing paper, and the SDLC appears on its syllabus because the model decides how testing is done. The ISTQB syllabus lists what the choice of SDLC affects: the "Scope and timing of test activities", the "Level of detail of test documentation", the "Choice of test techniques and test approach", the "Extent of test automation" and the "Role and responsibilities of a tester".

In practice the differences are large:

  • In a sequential model, testers review requirements and design tests early, but "dynamic testing cannot be performed early in the SDLC", because the code arrives late (ISTQB).
  • In iterative and incremental models each iteration delivers something that runs, so "both static testing and dynamic testing may be performed at all test levels" in every iteration, and frequent delivery "requires fast feedback and extensive regression testing".
  • In agile development, change is expected throughout, so "lightweight work product documentation and extensive test automation to make regression testing easier are favored", and much manual testing uses experience-based techniques.
munotes.in74

What a Software Development Life Cycle Is

Chapter Eighteen, on the role of testing in each phase, takes the phases one by one from the tester's side.

Worked example: ExamReg release 2.0, phase by phase

PhaseWhat ExamReg release 2.0 neededWhere testing was already at work
FeasibilityCan the college move the paper exam form online before the November session, within the IT cell's budget?Testers estimated the effort to test fee rules and the last-day load
RequirementsThe fee rules, eligibility rules, backlog papers, hall ticket format, 500 simultaneous users on the last dayReviewing the rules found the missing answer for negative days late (Chapter One, on what software testing is)
DesignFive modules, a fee service, an adapter to the payment gateway, the hall ticket generatorTest levels planned against the design; integration order chosen
ImplementationCode for each module, with developers' unit testsStatic analysis and code reviews
TestingIntegration, system and acceptance testingThe test plan of Chapter Eleven
DeploymentInstall on the production server one week before the form opensSmoke test in production; the load test report
MaintenanceRelease 2.1 in March for the next session's rule changeRegression testing of everything the change touches

Choosing a model

No model is right for every project, which is the SDLC form of the principle that testing is context dependent. The usual considerations, stated as practice:

If the project has ...... a sensible choice isBecause
Stable, well-understood requirements and a fixed contractSequential (waterfall or V-model)Planning once is efficient, and the V-model pairs every phase with its test
High technical or business riskSpiralEach cycle begins by resolving the biggest risk
Requirements users cannot state until they see somethingPrototyping or incrementalEarly working versions draw out the real requirements
Changing requirements and an available customerAgile (for example Scrum)Short cycles absorb change and deliver value early
Safety or regulatory evidence to produceV-model, often inside a larger iterative planEvery requirement is traced to a test that verifies it

What it does not mean

The SDLC is not a single fixed sequence. It is a set of kinds of work; the model decides the order and the repetition.

The life cycle does not end at deployment. Maintenance is usually the longest part of a product's life, and every change passes through the phases again.

Testing is not only one phase of the SDLC. It appears as a phase in the table, but test activities run alongside every phase (principle 3, early testing, and Chapter Eighteen).

munotes.in75

What a Software Development Life Cycle Is

"Agile" does not mean without a life cycle. Agile teams still do requirements, design, coding, testing and deployment, in short cycles instead of long phases.

Quick revision

  • Life cycle: evolution of a system from conception through retirement (ISO/IEC/IEEE 12207:2026).
  • Life cycle model: a framework of processes and activities organised into stages, a common reference for everyone (ISO/IEC/IEEE 12207:2026).
  • Generic phases: feasibility, requirements, design, implementation, testing, deployment, maintenance; maintenance loops back.
  • Families (ISTQB): sequential (waterfall, V-model), iterative (spiral, prototyping), incremental (Unified Process), plus agile methods (Scrum, Kanban, XP, TDD, BDD, ATDD).
  • The SDLC decides the scope and timing of testing, documentation detail, techniques, automation and the tester's role.
  • Choose by requirement stability, risk, customer availability and the evidence required.

Test yourself

1. Define the software development life cycle and a life cycle model. The life cycle is the evolution of a software system from conception through retirement. A life cycle model is a framework of processes and activities organised into stages, acting as a common reference for communication and understanding, such as the waterfall or spiral model.

2. List the phases of the SDLC with the work product of each. Feasibility (feasibility report), requirements (requirements specification), design (design documents), implementation (source code and unit tests), testing (test reports and fixed defects), deployment (the released system and user guides), maintenance (fixes, changes and new releases).

3. Name the three families of SDLC models with an example of each. Sequential, such as the waterfall model and the V-model; iterative, such as the spiral model and prototyping; incremental, such as the Unified Process. Agile methods such as Scrum combine iteration and increments.

4. How does the choice of SDLC model affect testing? Give three effects. It sets the scope and timing of test activities, the level of detail of test documentation, the choice of techniques and approach, the extent of automation, and the tester's role. For example, in a sequential model dynamic testing comes late, while in agile development automated regression testing runs every iteration.

5. Why is maintenance drawn as a loop back to requirements? Because a system in use keeps changing, and each change must again be specified, designed, coded and tested; for most software this loop makes up most of its life.

Contents This chapter on its own page

munotes.in76

Chapter Fourteen

The Waterfall Model, and What Royce Actually Said

Syllabus topic Module 1, "Software Development Life Cycle (SDLC): Overview of SDLC"

In one line

The waterfall model builds software in a fixed sequence of phases, requirements, then design, then code, then test, then operation, each finished before the next begins, like water falling from one step to the next.

In the wording a student can write in an examination: the waterfall model is a "model of the software development process in which the constituent activities, typically a concept phase, requirements phase, design phase, implementation phase, test phase, and installation and checkout phase, are performed in that order, possibly with overlap but with little or no iteration" (ISO/IEC/IEEE 24765:2017). It is the classic sequential model, usually traced to Winston Royce's 1970 paper, which in fact argued that the simple sequence is risky and must be supplemented.

Why a sequence was attractive

In the 1960s large software was new, and it was being built for customers like governments and space programmes who paid by contract. A fixed sequence of phases, each ending in an approved document, suited them. Each phase had a clear deliverable to sign off, progress could be measured by which phase was finished, and a customer could be held to requirements agreed at the start. The same logic still suits some projects today: where requirements are stable and well understood, planning the whole job once is efficient.

What Royce drew

Royce's paper, "Managing the Development of Large Software Systems", was given at the IEEE WESCON conference in August 1970. He begins with the simplest possible development, just analysis followed by coding, which is enough for a small program used by the people who wrote it. For a large system delivered to a customer he draws seven steps:

  1. System requirements
  2. Software requirements
  3. Analysis
  4. Program design
  5. Coding
  6. Testing
  7. Operations

That staircase, his Figure 2, is the diagram every textbook reproduces. What the textbooks usually drop is what he said next.

Royce's seven steps, with the iteration he drew between steps and the return from testing to design

Figure 14.1 Royce 1970: the seven steps (Figure 2), iteration between successive steps (Figure 3), and testing sending the work back to design (Figure 4)

What Royce actually said about it

Straight after the staircase, Royce writes: "I believe in this concept, but the implementation described above is risky and invites failure." The problem, he explains, is where testing sits. "The testing phase which occurs at the end of the development cycle is the first event for which timing, storage, input/output transfers, etc., are experienced as distinguished from analyzed." If those real behaviours fail to meet the constraints, the fix is not a small patch but a redesign, and the redesign can violate the requirements on which the whole design rested. The project is back at the start, with, in his estimate, up to a 100 per cent overrun in schedule or cost.

munotes.in77

The Waterfall Model, and What Royce Actually Said

He does not throw the sequence away: "However, I believe the illustrated approach to be fundamentally sound." Instead he adds five features to reduce the risk, and every one of them is still good advice.

Royce's stepWhat he meantWhere it lives in this book
1. Program design comes firstDo a preliminary program design before analysis, so that storage, timing and interfaces are considered earlyDesign reviews (Chapter Twenty-Eight, on software reviews)
2. Document the design"quite a lot" of documentation; he calls "ruthless enforcement of documentation requirements" the first rule of managing software developmentThe test documents (Chapter Twelve)
3. Do it twiceBuild a pilot version first, a simulation of the whole process in miniature, so the delivered version is really the secondPrototyping (Chapter Sixteen, on iterative, incremental and spiral models)
4. Plan, control and monitor testingTesting is the biggest user of resources and the phase of greatest risk; plan it, use specialists, inspect every line, test every pathThe test process, plan and levels (Chapters Five, Eleven and Thirty-Two)
5. Involve the customerCommit the customer formally at points after the requirements are defined, not only at the endValidation and acceptance testing (Chapters Forty-One and Forty-Two)

Royce on testing, fifty years early

Step 4 is worth reading in full, because it anticipates half of this syllabus. Royce writes: "Without question the biggest user of project resources, whether it be manpower, computer time, or management judgment, is the test phase. It is the phase of greatest risk in terms of dollars and schedule." He then recommends four things:

  1. Independent testers. Many parts of testing are best done by specialists who did not contribute to the design; if only the designer can test the design, the documentation has failed. This is the independent test group of Chapter Thirty-One, on the strategic approach to testing.
  2. Visual inspection by a second person. Most errors are obvious and can be spotted by a second party scanning the analysis and the code, "dropped minus signs, missing factors of two, jumps to wrong addresses". This is the review and inspection of Chapters Twenty-Eight and Twenty-Nine, six years before Fagan published inspections.
  3. Every logic path, at least once. "Test every logic path in the computer program at least once with some kind of numerical check." This is path and branch coverage (Chapters Fifty-Nine to Sixty and Seventy-Two).
  4. The computer last. Only after the simple errors are removed should the software go to formal checkout on the machine.

The paper that gave the world the sequential model also argued for early reviews, independent testing, coverage and a pilot version. The model that took its name kept the staircase and dropped the advice.

munotes.in78

The Waterfall Model, and What Royce Actually Said

The waterfall model as it is taught

The textbook waterfall keeps Royce's staircase in its simplest form. Its properties follow directly from the sequence.

StrengthsWeaknesses
Simple to understand and manage; progress is visible by phaseWorking software appears only at the end, so users see nothing until late
Each phase ends in a reviewed document, useful for contracts and auditsRequirements are frozen early, but users often discover what they need only when they see the system
Suits stable, well-understood requirements and fixed-price contractsDefects in requirements and design are found late, in testing, when they are most expensive (Chapter Three, on why software must be tested)
Easy to assign people to phasesTesting is squeezed when earlier phases overrun, because the deadline does not move
Works where regulation demands documented phase gatesPoor at absorbing change; going back a phase is costly

Where testing sits in the waterfall

In the pure waterfall, testing is a phase near the end: the code is complete, then it is tested. That is exactly the placement Royce warned about. The ISTQB syllabus describes the consequence for any sequential model: "in the initial phases testers typically participate in requirement reviews, test analysis, and test design. The executable code is usually created in the later phases, so typically dynamic testing cannot be performed early in the SDLC."

The sentence contains the remedy. Even in a strict waterfall, testers need not wait: they can review the requirements and designs as they are written (static testing) and design their tests while the code is being built. The V-model, the next chapter, is essentially the waterfall with that remedy drawn into it.

Worked example: ExamReg in a pure waterfall

Suppose ExamReg release 2.0 were run as a pure waterfall over sixteen weeks: requirements weeks 1 to 3, design 4 to 6, coding 7 to 12, testing 13 to 15, deployment week 16. Coding overruns by two weeks, as coding often does. The deadline cannot move, because the examination form opens on a fixed date. Testing shrinks from three weeks to one.

In that one week, testing finds that the portal slows to a crawl above 300 simultaneous users, far short of the 500 expected on the last day. The cause is the design of the fee service, fixed in week 5. This is Royce's Figure 4 exactly: a real behaviour, "experienced as distinguished from analyzed", met for the first time in the test phase, and the fix is a redesign, not a patch. The team now chooses between releasing a portal that will fail on the last day and missing the date.

Every remedy in this book is a way of moving that discovery earlier: a design review in week 5 asking how the fee service handles 500 users; a small load test of a pilot version, Royce's "do it twice"; or an iterative model in which a working fee service exists by week 6.

munotes.in79

The Waterfall Model, and What Royce Actually Said

What it does not mean

Royce did not recommend the simple waterfall. He called it "risky" and said it "invites failure", then added five steps to fix it.

The word "waterfall" is not in Royce's paper. It does not occur anywhere in the text; the name was attached to his staircase later.

Waterfall does not forbid going back. The standard definition allows "overlap" and "little" iteration, and Royce drew iteration between successive steps; what the model resists is going back more than a step.

Sequential does not mean testers wait until coding ends. Reviews and test design run alongside the early phases even in a waterfall.

Quick revision

  • Waterfall model: phases performed in order, possibly with overlap, "with little or no iteration" (ISO/IEC/IEEE 24765:2017); the classic sequential model.
  • Royce 1970 (IEEE WESCON): seven steps: system requirements, software requirements, analysis, program design, coding, testing, operations.
  • Royce: "risky and invites failure", because testing is the first time real behaviour is met; up to 100 per cent overrun.
  • His five fixes: program design comes first; document the design; do it twice; plan, control and monitor testing; involve the customer.
  • His testing advice: independent specialists, visual inspection by a second person, every logic path tested at least once, the computer last.
  • Strengths: simple, documented, good for stable requirements and contracts. Weaknesses: late working software, late defects, poor with change, testing squeezed.

Test yourself

1. Explain the waterfall model with its phases. A sequential model in which each phase is completed before the next begins: in Royce's form, system requirements, software requirements, analysis, program design, coding, testing and operations. Each phase ends in a document that becomes the input to the next, with little or no iteration.

2. What did Royce say was wrong with the simple sequence? That it is risky and invites failure, because testing at the end is the first time timing, storage and input/output are experienced rather than analysed. A failure then usually demands a redesign that can break the requirements, sending the project back to the start with up to a 100 per cent overrun.

3. State Royce's five recommended steps. Program design comes first; document the design; do it twice (build a pilot version); plan, control and monitor testing; involve the customer.

4. Give three advantages and three disadvantages of the waterfall model. Advantages: simple to manage, with progress visible by phase; documented phase outputs useful for contracts and audits; efficient when requirements are stable. Disadvantages: working software only at the end; requirement and design defects found late and expensively; poor at absorbing change, with testing squeezed when earlier phases overrun.

munotes.in80

The Waterfall Model, and What Royce Actually Said

5. Where does testing sit in the waterfall, and how can testers still start early? Dynamic testing is a late phase, after coding. Testers can still start early by reviewing the requirements and designs as they are written and by designing their tests during the earlier phases.

Contents This chapter on its own page

munotes.in81

Chapter Fifteen

The V-Model: A Test Level for Every Phase

Syllabus topic Module 1, "Software Development Life Cycle (SDLC): phases and their relationship to testing"

In one line

The V-model is the waterfall bent into a V: the development phases run down the left side, the test levels run up the right side, and each test level checks the work product of the phase opposite it.

In the wording a student can write in an examination: the V-model is a sequential SDLC model in which each development phase is paired with a corresponding test level: user requirements with acceptance testing, the system specification with system testing, the architectural design with integration testing, and the detailed design with component (unit) testing. Unlike the waterfall model, it "integrates the test process throughout the development process, implementing the principle of early testing" (ISTQB, 2018 syllabus), because the tests for each level are designed while the matching development phase is under way.

Why bend the waterfall

The last chapter ended on Royce's warning: when testing is a single phase at the end, real problems are met too late. The V-model answers it without giving up the sequence. It keeps the phases of the waterfall, but it recognises that "testing" is not one activity: checking a component, checking that components work together, checking the whole system and checking that users can accept it are four different jobs, each with its own test basis. And each of them can be prepared as soon as its test basis exists.

The ISTQB Foundation syllabus of 2018 describes it in exactly those terms: "Unlike the Waterfall model, the V-model integrates the test process throughout the development process, implementing the principle of early testing. Further, the V-model includes test levels associated with each corresponding development phase, which further supports early testing". The current syllabus (2024) still lists it among the sequential models, with the waterfall.

The shape of the V

The V-model: development phases down the left, test levels up the right, and the tests of each level designed from the phase opposite

Figure 15.1 The V-model: each test level checks the work product of the development phase opposite it

Read the figure in three directions.

Down the left arm is development, as in the waterfall: user requirements, then the system specification, then the architectural design, then the detailed design, and coding at the point of the V.

Up the right arm is test execution, in the order the pieces exist: components are tested first, then their integration, then the whole system, then acceptance.

Across the V, the dashed arrows, is what makes the model useful. The test level on the right is designed from the work product on the left, and it is designed when that work product is written, not when the code arrives. The acceptance tests are written while the user requirements are being agreed; the system tests while the specification is written; and so on down.

munotes.in82

The V-Model: A Test Level for Every Phase

Each phase and its test level

Development phase (left)Its work product, which is the test basisTest level (right)What that level checks
User requirementsBusiness needs, user requirements, acceptance criteriaAcceptance testingThat the system fulfils the users' business needs and is ready to deploy
System specificationSystem requirements: functional and non-functionalSystem testingThe behaviour and capabilities of the whole system, end to end
Architectural designThe architecture: components and their interfacesIntegration testingThe interfaces and interactions between components
Detailed designThe design of each componentComponent (unit) testingEach component in isolation

The right-hand column comes from the current ISTQB syllabus's own descriptions of the levels: component testing "focuses on testing components in isolation"; component integration testing "focuses on testing the interfaces and interactions between components"; system testing "focuses on the overall behavior and capabilities of an entire system or product"; and acceptance testing "focuses on validation and on demonstrating readiness for deployment, which means that the system fulfills the user's business needs." The syllabus adds a fifth level, system integration testing, which tests "the interfaces of the system under test and other systems and external services"; in a V it sits beside system testing. Test levels have their own chapter (Chapter Thirty-Two).

The relationship between phases and testing

This is the part of the syllabus MU words as "phases and their relationship to testing", and the V states the relationship in four rules, all of which the current ISTQB syllabus lists as good practice whatever model is used:

  1. Every development activity has a matching test activity. "For every software development activity, there is a corresponding test activity, so that all development activities are subject to quality control."
  2. Each test level has its own objectives. Different levels "have specific and different test objectives, which allows for testing to be appropriately comprehensive while avoiding redundancy". Unit tests do not re-test what system tests cover, and system tests do not repeat unit tests.
  3. Test design starts in the matching phase. "Test analysis and design for a given test level begins during the corresponding development phase of the SDLC, so that testing can adhere to the principle of early testing".
  4. Testers review drafts early. "Testers are involved in reviewing work products as soon as drafts of these work products are available".

The third rule is the one that pays. Writing an acceptance test forces the question how will we know this requirement is met? while the requirement can still be changed at the cost of a sentence. A requirement nobody can write a test for is usually a requirement nobody can build to either.

Worked example: ExamReg release 2.0 on a V

Left arm: when it is writtenTests designed at that momentRight arm: when they run
Week 2, user requirements: a student who submits on or before the last date pays no late feeAcceptance test: an exam cell clerk submits a form for a real student on the last date and checks the receiptWeek 15, acceptance testing by the exam cell
Week 3, system specification: the fee table, the 500-user load target, the three browsersSystem tests: every row of the fee table through the web interface; a 500-user load test; each page on each browserWeek 13 and 14, system testing by the test team
Week 5, architectural design: the fee service is called by the form module and calls the gateway adapterIntegration tests: the form passes days late and backlog papers to the fee service; the fee service passes the amount in rupees, not paise, to the adapterWeek 12, integration testing
Week 6, detailed design: the late-fee function's signature and rulesUnit tests: the late-fee cases of Chapter Seven, on test casesWeeks 8 to 11, as each function is coded
munotes.in83

The V-Model: A Test Level for Every Phase

Look at the integration row. The paise-for-rupees interface defect of Chapter Two, on errors, faults and failures, could be caught by an integration test designed in week 5, from the architecture document, before either module was coded. In a pure waterfall nobody would have written that test until week 13.

Strengths and weaknesses

StrengthsWeaknesses
Testing is planned and designed from the start, so defects in requirements and designs are found earlyStill sequential: working software appears late, and changes to requirements are expensive
Each test level has a clear test basis and objective, so nothing is tested twice and nothing is missedAssumes requirements can be fixed early, which suits some projects and not others
Traceability from each requirement to its test is built in, which regulators and auditors valueTest execution is still late; the early work is design, not running code
Simple to explain and manageHeavy documentation if applied rigidly

The V-model is common in safety-critical and regulated industries, such as vehicles, medical devices and aviation, where every requirement must be shown to have been verified. There it often sits inside a larger iterative plan: a V for each increment.

The W-model, briefly

Some writers add a second V inside the first, drawn as a W: alongside each development phase on the left is a review of that phase's work product. The idea is already in rule 4 above: static testing of the requirements, the specification and the designs, as soon as drafts exist, before any dynamic test runs. The W is the V with static testing drawn in.

What it does not mean

The V-model is not the waterfall with testing at the end. Its test execution is at the end, but its test design is at the start, phase by phase.

munotes.in84

The V-Model: A Test Level for Every Phase

The V does not mean each level is tested only once. Defects found at a level are fixed and retested, and regression testing follows every change.

Unit testing does not check the requirements. Each level checks the work product opposite it; only acceptance testing checks the users' needs directly.

The V-model is not obsolete. The current ISTQB syllabus still lists it, and regulated industries use it, often within iterative development.

Quick revision

  • V-model: a sequential model pairing each development phase with a test level; test design starts in the matching phase, so it "integrates the test process throughout the development process" (ISTQB 2018).
  • Pairs: user requirements with acceptance testing; system specification with system testing; architectural design with integration testing; detailed design with component (unit) testing; coding at the point.
  • Left arm down, right arm up; the arrows across are tests designed early from each work product.
  • Good practice for any model (ISTQB v4.0.1): a test activity for every development activity; distinct objectives per level; test design during the matching phase; early review of drafts.
  • Strengths: early test design, clear levels, traceability. Weaknesses: still sequential, late working software, costly change.
  • W-model: a review beside every development phase.

Test yourself

1. Explain the V-model with a diagram. Draw development phases down the left (user requirements, system specification, architectural design, detailed design), coding at the point, and test levels up the right (component, integration, system, acceptance testing), with each level opposite the phase whose work product is its test basis. Tests for each level are designed during the matching phase and executed on the way up.

2. Which test level is paired with the architectural design, and why? Integration testing, because the architecture defines the components and their interfaces, and integration testing checks exactly those interfaces and interactions.

3. How does the V-model improve on the waterfall model? It integrates testing throughout development: tests for each level are designed as soon as the matching work product exists, so defects in requirements and designs are found early instead of in a single late test phase.

4. State two strengths and two weaknesses of the V-model. Strengths: testing planned and designed from the start; clear objectives and traceability for each level. Weaknesses: still sequential, so working software comes late; assumes stable early requirements, so changes are costly.

5. What does "phases and their relationship to testing" mean in the V-model? That every development phase has a corresponding test level whose test basis is that phase's work product, and whose test design begins during that phase.

Contents This chapter on its own page

munotes.in85

Chapter Sixteen

Iterative, Incremental and Spiral Models

Syllabus topic Module 1, "Software Development Life Cycle (SDLC): Overview of SDLC"

In one line

Instead of doing each phase once, iterative and incremental models do them many times: each round refines the product, adds a piece to it, or, in the spiral model, attacks the biggest remaining risk.

In the wording a student can write in an examination: iterative development is the "repeated use of concurrent planning, developing, and testing activities" (ISO/IEC/IEEE 26515:2018); incremental development is a technique in which requirements, design, implementation and testing "occur in an overlapping, iterative (rather than sequential) manner, resulting in incremental completion of the overall software product" (ISO/IEC/IEEE 24765:2017). The spiral model (Boehm, 1988) organises the iterations around risk: each cycle determines objectives, alternatives and constraints, evaluates the alternatives and resolves risks, develops and verifies the next level of the product, and plans the next cycle.

Why go round more than once

The waterfall and the V-model assume the requirements can be settled before design begins. For many systems that assumption fails. Users cannot describe what they want until they see something; the technology is new and nobody knows whether the design will be fast enough; the business changes while the software is being built. A sequential model meets all of these problems at the end, in testing, which is Royce's warning from Chapter Fourteen on the waterfall model. The iterative family meets them early and repeatedly, by producing something that runs early and learning from it.

Iterative and incremental: two different ideas

The two words are often used together, and the models usually combine them, but they mean different things.

IterativeIncremental
The ideaRepeat the cycle to improve the same productBuild the product in pieces, each adding capability
Standard definitionIteration: "repeating the application of the same process or set of processes at the same level of abstraction on the same system or system element" (ISO/IEC/IEEE 12207:2026)Increment: "tested, deliverable version of a product that provides new or modified capabilities" (ISO/IEC/IEEE 24748-8:2026)
PictureSketch the whole portrait roughly, then refine itPaint the portrait one finished section at a time
ExamReg exampleA rough version of the whole portal, then a better one, then the release versionRelease 1: login and exam form. Release 2: fees and payment. Release 3: hall tickets and reports

The ISTQB syllabus lists the spiral model and prototyping as examples of iterative models, and the Unified Process as an incremental one. In practice most modern development is both: each iteration delivers an increment, and each increment is refined in later iterations.

Prototyping

A prototype is a "preliminary type, form, or instance of a system that serves as a model for later stages or for the final, complete version of the system" (ISO/IEC/IEEE 24765). Prototyping is the "hardware and software development technique in which a preliminary version of part or all of the hardware or software is developed to permit user feedback, determine feasibility, or investigate timing or other issues" (the same vocabulary).

munotes.in86

Iterative, Incremental and Spiral Models

There are two kinds. A throwaway prototype is built quickly to learn something, usually what users want from a screen, and then discarded. An evolutionary prototype is built carefully and grown into the product. For ExamReg, a clickable mock-up of the exam form shown to twenty students before any code is written is a throwaway prototype: it costs a day and settles which fields confuse them. Royce's step 3, "do it twice", from the waterfall chapter, is a prototype in all but name.

The spiral model

Barry Boehm, then chief scientist of TRW's Defense Systems Group, published "A Spiral Model of Software Development and Enhancement" in the magazine Computer in May 1988. It opens on the debate of its day ("The waterfall model is dead." "No, it isn't, but it should be.") and states its own central idea plainly: "The major distinguishing feature of the spiral model is that it creates a risk-driven approach to the software process rather than a primarily document-driven or code-driven process."

The spiral model: four quadrants, the spiral growing outward from the centre

Figure 16.1 Boehm's spiral model, simplified: each cycle passes through the four quadrants, and cost grows with the radius

Boehm reads his own figure this way: "The radial dimension in Figure 2 represents the cumulative cost incurred in accomplishing the steps to date; the angular dimension represents the progress made in completing each cycle of the spiral."

Each cycle goes through four stages, one per quadrant:

  1. Determine objectives, alternatives and constraints. For the part of the product being worked on: its objectives (performance, functionality, ability to change), the alternative ways of building it (design A, design B, reuse, buy), and the constraints (cost, schedule, interfaces).
  2. Evaluate alternatives; identify and resolve risks. Evaluating the alternatives against the objectives exposes uncertainties, and those are the risks. They are resolved by whatever is cheapest and convincing: prototyping, simulation, benchmarking, reference checking, user questionnaires, analytic modelling.
  3. Develop and verify the next-level product. What is built next depends on the risk that remains. If performance or user-interface risks dominate, the next step is another, more detailed prototype. If those are settled and the risks are in program development, the next step follows the familiar waterfall stages, each "followed by a validation step".
  4. Plan the next phases. Each cycle ends, in Boehm's account, with a review by the people concerned with the product, covering everything produced in the cycle and the plans for the next, so that all parties commit to the next step.

The cycle then repeats, further out on the spiral, until the product is complete. Boehm adds a small Round 0 in his worked example, a feasibility study before the first full cycle.

munotes.in87

Iterative, Incremental and Spiral Models

Why "risk-driven" matters to a tester

The spiral's defining idea is that risk decides what to do next. Boehm extends that to testing directly: risk considerations "can determine the amount of time and effort that should be devoted to such other project activities as" quality assurance, formal verification and testing. That is risk-based testing, which Chapter Eleven's test plan used as its approach, stated as a property of the whole process.

Testing also appears in every quadrant. Quadrant 2 uses prototypes, benchmarks and simulations, which are experiments, often run by testers. Quadrant 3 ends each level of the product with verification and validation. And quadrant 4's review is a static test of everything the cycle produced.

Testing in iterative and incremental models generally

The ISTQB syllabus states the consequence for any model of this family: "it is assumed that each iteration delivers a working prototype or product increment. This implies that in each iteration both static testing and dynamic testing may be performed at all test levels. Frequent delivery of increments requires fast feedback and extensive regression testing."

The last sentence is the one to remember. Every new increment is added to software that already works, and each one can break what came before. By the third ExamReg release the regression suite includes the tests of releases one and two, and it runs on every build. This is why iterative development and test automation grew up together, and why Chapter Thirty-Nine, on regression testing and continuous integration, belongs to the same story.

Worked example: ExamReg release 2.0 as a spiral

CycleQuadrant 1: objectives and alternativesQuadrant 2: the biggest risk, and how it was resolvedQuadrant 3: what was built and verifiedQuadrant 4: review and plan
Round 0Can the form go online this session?Budget and the payment gateway contractA feasibility reportThe principal commits to the project
1A form students can fill without helpStudents misreading backlog papersA throwaway prototype, tried with twenty studentsForm layout agreed; prototype discarded
2Handle 500 students on the last dayThe fee service may not scale; two designs comparedA load test of both designs on a thin working versionDesign B chosen; plan the full build
3The complete releaseRemaining risks are ordinary development risksWaterfall-like stages with their test levelsExam cell acceptance; release

Compare it with the pure waterfall story of Chapter Fourteen. There, the load problem was discovered in the last week, with no time left. Here it was the first risk attacked, in cycle 2, while choosing between two designs cost nothing but a benchmark.

munotes.in88

Iterative, Incremental and Spiral Models

Strengths and weaknesses of the spiral model

StrengthsWeaknesses
Risk is confronted early and explicitlyDepends on people skilled at identifying and judging risk
Accommodates prototyping, reuse, specification and waterfall stages, chosen by riskMore complex to manage than a sequential model
Each cycle ends with a review and a commitmentCan be costly for small, low-risk projects
Suits large, complex, novel systemsBoehm himself called it "not yet as fully elaborated" as established models

The last row is Boehm's own honesty: his third conclusion was that the model "is not yet as fully elaborated as the more established models", needing more work on contracting, milestones, reviews, scheduling and status monitoring before it would suit every situation.

What it does not mean

Iterative does not mean unplanned. Each iteration is planned; the spiral plans every cycle explicitly.

Incremental is not the same as iterative. Incremental adds pieces; iterative refines the whole. Most real projects do both.

The spiral is not a fixed number of loops. The number of cycles, and what each builds, is decided by the risks that remain.

A prototype is not the product. A throwaway prototype is discarded once it has answered its question; building on it by accident is a common and costly mistake.

Quick revision

  • Iterative development: repeated planning, developing and testing to refine the same product. Incremental development: overlapping stages delivering the product in increments.
  • Increment: a tested, deliverable version providing new or modified capabilities (ISO/IEC/IEEE 24748-8:2026).
  • Prototyping: a preliminary version for user feedback, feasibility or timing; throwaway or evolutionary.
  • Spiral model (Boehm, Computer, May 1988): risk-driven; radius is cumulative cost, angle is progress.
  • Four quadrants: determine objectives, alternatives and constraints; evaluate alternatives and resolve risks; develop and verify the next-level product; plan the next phases, with a review and commitment.
  • Testing: static and dynamic testing in every iteration; fast feedback and extensive regression testing (ISTQB).

Test yourself

1. Distinguish iterative from incremental development, with an example. Iterative development repeats the cycle to refine the same product, such as a rough version of the whole portal improved in later rounds. Incremental development delivers the product in pieces, each adding capability, such as login and the exam form first, then fees and payment, then hall tickets.

2. Explain the spiral model with a diagram. Draw two axes and a spiral growing outward from the centre, clockwise, through four quadrants: determine objectives, alternatives and constraints; evaluate alternatives and identify and resolve risks; develop and verify the next-level product; plan the next phases. The radius shows cumulative cost and the angle progress within a cycle; each cycle ends with a review and commitment.

3. Why is the spiral model called risk-driven? Because at each cycle the remaining risks decide what to do next: a prototype, a simulation, a specification or waterfall-style development, and even how much effort goes into testing and verification.

munotes.in89

Iterative, Incremental and Spiral Models

4. What is a prototype? Distinguish throwaway from evolutionary prototypes. A preliminary version of a system used for feedback, feasibility or investigation. A throwaway prototype is built quickly to answer a question and then discarded; an evolutionary prototype is built carefully and grown into the product.

5. How does testing change in iterative and incremental models? Every iteration delivers something that runs, so static and dynamic testing happen in every iteration and at all levels, and frequent delivery demands fast feedback and extensive, usually automated, regression testing.

Contents This chapter on its own page

munotes.in90

Chapter Seventeen

Agile, Scrum and DevOps: Testing in Short Cycles

Syllabus topic Module 1, "Software Development Life Cycle (SDLC): Overview of SDLC"

In one line

Agile development builds software in short cycles of a few weeks, each ending in working software, and testing moves from a phase at the end into every day of every cycle.

In the wording a student can write in an examination: agile development is a "development approach based on iterative development, frequent inspection and adaptation, and incremental deliveries, in which requirements and solutions evolve through collaboration in cross-functional teams" (SEVOCAB). Its values are stated in the Agile Manifesto (2001). Scrum is the most widely used agile framework: a Scrum Team works in Sprints of one month or less toward a Product Goal, and each Sprint produces an Increment that meets the Definition of Done. DevOps extends the same thinking to operations, with continuous integration and delivery. In all of them testing is continuous, largely automated, and shared by the whole team.

Why agile arose

By the late 1990s many teams found the sequential models too slow for software whose requirements changed every few months: by the time a year-long waterfall delivered, the business had moved on. The spiral had shown that iteration could manage risk; agile methods pushed iteration to its limit, with cycles of weeks rather than months, and working software at the end of every one. In February 2001 seventeen practitioners, among them Kent Beck, Martin Fowler, Ken Schwaber and Jeff Sutherland, put their shared values into a one-page manifesto.

The Agile Manifesto

The manifesto's four values are its whole argument:

  • "Individuals and interactions over processes and tools"
  • "Working software over comprehensive documentation"
  • "Customer collaboration over contract negotiation"
  • "Responding to change over following a plan"

Its next sentence is the one most often forgotten: "That is, while there is value in the items on the right, we value the items on the left more." The manifesto does not reject documentation, plans, contracts or processes; it ranks them below the things on the left.

Twelve principles follow. Four of them bear most directly on testing:

  • "Our highest priority is to satisfy the customer through early and continuous delivery of valuable software."
  • "Welcome changing requirements, even late in development."
  • "Deliver working software frequently, from a couple of weeks to a couple of months, with a preference to the shorter timescale."
  • "Working software is the primary measure of progress."

Each one raises the demand on testing. Continuous delivery means continuous testing; welcoming late change means every change must be retested quickly; and if working software is the measure of progress, somebody must be able to show, every few weeks, that the software works.

Scrum

The Scrum Guide, maintained by its creators Ken Schwaber and Jeff Sutherland and last revised in November 2020, defines a small framework with three parts.

munotes.in91

Agile, Scrum and DevOps: Testing in Short Cycles

The Scrum Team is "one Scrum Master, one Product Owner, and Developers", with no sub-teams or hierarchies, "focused on one objective at a time, the Product Goal." The guide speaks of three accountabilities rather than roles. The Product Owner is accountable for the value of the product and orders the Product Backlog. The Developers create the Increment, and "Developers" includes everyone who builds it: programmers, testers, designers. The Scrum Master is accountable for the team's effectiveness with Scrum. There is no separate tester accountability: testing is part of creating a usable Increment, and the Developers are accountable for, among other things, "Instilling quality by adhering to a Definition of Done".

The events. The Sprint is the container for all the others: Sprints are "fixed length events of one month or less". Sprint Planning decides the Sprint Goal and the work to reach it. The Daily Scrum is "a 15-minute event for the Developers" to inspect progress toward the Sprint Goal. The Sprint Review inspects the Increment with stakeholders. The Sprint Retrospective plans improvements to how the team works.

The artifacts and their commitments. The Product Backlog (committed to the Product Goal), the Sprint Backlog (committed to the Sprint Goal) and the Increment, committed to the Definition of Done.

The Definition of Done: where testing lives in Scrum

The Scrum Guide defines it precisely: "The Definition of Done is a formal description of the state of the Increment when it meets the quality measures required for the product." It is also absolute: "If a Product Backlog item does not meet the Definition of Done, it cannot be released or even presented at the Sprint Review." The ISTQB syllabus draws the connection to testing: in agile development, exit criteria are called the Definition of Done, and the entry criteria a user story must meet before work begins are the Definition of Ready.

A Definition of Done for ExamReg's team might read: code reviewed; unit tests written and passing; acceptance tests for the story automated and passing; no open defect of major severity or above; regression suite green on the build server; exam cell has seen it working. Every line is a test or a review. In Scrum, "done" means "tested".

Testing in agile development

The ISTQB syllabus summarises how agile changes testing: agile "assumes that change may occur throughout the project. Therefore, lightweight work product documentation and extensive test automation to make regression testing easier are favored in agile projects. Also, most of the manual testing tends to be done using experience-based test techniques ... that do not require extensive prior test analysis and design."

Three practices, which the syllabus groups as approaches in which tests direct development, make the tests come first:

munotes.in92

Agile, Scrum and DevOps: Testing in Short Cycles

ApproachWhat it does (ISTQB v4.0.1, 2.1.3)
Test-driven development (TDD)"Directs the coding through test cases"; tests are written first, then code to satisfy them, then both are refactored
Acceptance test-driven development (ATDD)"Derives tests from acceptance criteria as part of the system design process", before the feature is built
Behaviour-driven development (BDD)Expresses behaviour in simple natural language, usually "Given/When/Then", which is then turned into executable tests

A BDD scenario for ExamReg's late fee reads like this, and a tool can run it as a test once each step is bound to code:

Given a student has completed the examination form
And today is 3 days after the last date
When the student opens the fee page
Then the late fee shown is Rs 100

TDD's cycle is worked through in full in Chapter Thirty-Six, on unit testing best practices.

DevOps and continuous integration

The ISTQB syllabus describes DevOps as "an organizational approach aiming to create synergy by getting development (including testing) and operations to work together to achieve a set of common goals." Its technical core is continuous integration (CI), a "technique that continually merges artifacts, including source code updates from all developers on a team, into a shared mainline to build and test the developed system" (IEEE 2675-2021), and continuous delivery (CD), which keeps the system always ready to release.

For testing, the syllabus lists the benefits: fast feedback on code quality; CI "promotes shift left in testing ... by encouraging developers to submit high quality code accompanied by component tests and static analysis"; stable automated test environments; more visibility of non-functional qualities such as performance; less repetitive manual testing; and a smaller regression risk because automated regression tests run at scale. It lists the costs too: the pipeline must be built, the CI and CD tools maintained, and "Test automation requires additional resources and may be difficult to establish and maintain." And it adds a warning worth repeating: even with this much automation, manual testing from the user's perspective will still be needed. Continuous integration, and Jenkins as the tool the practical uses for it, has its own chapter (Chapter Thirty-Nine).

Shift left, and retrospectives

Shift left is the syllabus's name for the principle of early testing applied deliberately: reviewing specifications from a tester's point of view, writing test cases before code, using CI with automated component tests, running static analysis before dynamic testing, and starting non-functional testing at component level where possible. The syllabus adds the balancing sentence: shift left "does not mean that testing later in the SDLC should be neglected."

Retrospectives, held at the end of each iteration, ask what went well, what did not, and how to improve. The syllabus lists their benefits for testing, including more effective testing, better testware, better requirements and better cooperation between development and testing. They are the agile form of the lessons learned in test completion (Chapter Five, on the basic test process).

munotes.in93

Agile, Scrum and DevOps: Testing in Short Cycles

The test pyramid

When tests run every day, their cost and speed matter. The test pyramid is the model most teams use to balance them. Martin Fowler's account of it records that most people know it "due to Mike Cohn, when he described it in his 2009 book Succeeding with Agile", and that Cohn had drawn it in conversation with Lisa Crispin in 2003 and 2004.

The test pyramid: many unit tests, fewer service tests, few UI tests

Figure 17.1 The test pyramid (Cohn): many small, fast tests at the base and few broad, slow ones at the top

The ISTQB syllabus explains the shape: "The higher the layer, the lower the test granularity, the lower the test isolation ... and the higher the test execution time." Tests at the base are small, isolated and fast, so many are needed; tests at the top are end-to-end and slow, so few are used. Cohn's original three layers are "unit tests", "service tests" and "UI tests". Fowler adds the reason not to invert it: tests that run end to end through the user interface are "brittle, expensive to write, and time consuming to run."

The testing quadrants

A second agile model, the testing quadrants, defined by Brian Marick, sorts tests along two lines: whether they face the business or the technology, and whether they support the team (guide development) or critique the product. The ISTQB syllabus fills in the four quadrants:

Support the teamCritique the product
Business facingQ2: functional tests, examples, user story tests, prototypes, API tests, simulationsQ3: exploratory testing, usability testing, user acceptance testing
Technology facingQ1: component and component integration tests, automated in CIQ4: smoke tests and non-functional tests (except usability)

The quadrants are a checklist: a team whose tests all sit in Q1 has fast unit tests and no idea whether users can use the product.

Worked example: one ExamReg Sprint

The team runs two-week Sprints. In Sprint 4 the Sprint Goal is: students with a fee concession can submit the form and pay only the backlog and late fees.

DayWhat happens, with the testing in it
1Sprint Planning: the story is refined with the exam cell; its acceptance criteria are written as three BDD scenarios (Definition of Ready met)
2 to 9Developers write unit tests first (TDD); CI builds and runs the whole regression suite on every commit; a tester automates the BDD scenarios and runs an exploratory session on the concession screens
10Daily Scrum reports one failing scenario: a concession student is charged the form fee on late forms. Fixed the same day; the scenario passes
11 to 13Performance check of the fee page (Q4); exam cell tries the feature on the test server (Q3)
14Sprint Review shows the working increment to the exam cell; Retrospective notes that the BDD scenarios caught the only serious defect and agrees to write them for every story
munotes.in94

Agile, Scrum and DevOps: Testing in Short Cycles

Compare the waterfall of Chapter Fourteen. There, the first working fee page existed in week 13 of 16. Here, a tested increment exists every two weeks.

What it does not mean

Agile does not mean no documentation or no planning. The manifesto values the items on the right; it values those on the left more. Scrum plans every Sprint.

Agile does not mean no testers. Scrum folds testing into the Developers' accountability; people who specialise in testing are Developers.

Automation does not replace all manual testing. The syllabus says user-perspective manual testing will still be needed; exploratory and usability testing sit in Q3.

Shift left does not mean skipping later testing. System and acceptance testing still happen; they just find fewer defects.

Quick revision

  • Agile development: iterative, incremental, frequent inspection and adaptation, cross-functional collaboration.
  • Manifesto values: individuals and interactions, working software, customer collaboration, responding to change, over processes and tools, comprehensive documentation, contract negotiation, following a plan.
  • Scrum (Guide, November 2020): Scrum Team of Product Owner, Scrum Master and Developers; Sprint (one month or less), Sprint Planning, Daily Scrum (15 minutes), Sprint Review, Sprint Retrospective; Product Backlog, Sprint Backlog, Increment; Definition of Done = exit criteria; Definition of Ready = entry criteria.
  • TDD, ATDD, BDD: tests first; BDD in Given/When/Then.
  • DevOps: development and operations together; CI and CD; fast feedback; manual user-perspective testing still needed.
  • Shift left: test earlier, without neglecting later testing. Retrospectives: continuous improvement.
  • Test pyramid (Cohn): many unit, some service, few UI tests. Quadrants (Marick): business or technology facing, supporting the team or critiquing the product.

Test yourself

1. State the four values of the Agile Manifesto. Individuals and interactions over processes and tools; working software over comprehensive documentation; customer collaboration over contract negotiation; responding to change over following a plan. The items on the right still have value; those on the left are valued more.

2. Describe the Scrum framework. A Scrum Team of one Product Owner, one Scrum Master and Developers works in Sprints of one month or less. Each Sprint has Sprint Planning, Daily Scrums of 15 minutes, a Sprint Review and a Sprint Retrospective, and produces an Increment that meets the Definition of Done, drawn from the Product Backlog through the Sprint Backlog.

3. What is the Definition of Done, and how does it relate to testing? A formal description of the state of the Increment when it meets the quality measures required for the product. It works as the exit criteria for a backlog item: an item that does not meet it cannot be released or even shown at the Sprint Review, so it normally lists the reviews and tests that must pass.

munotes.in95

Agile, Scrum and DevOps: Testing in Short Cycles

4. Distinguish TDD, ATDD and BDD. TDD writes unit tests before the code and directs coding through them; ATDD derives tests from acceptance criteria before the feature is developed; BDD writes the expected behaviour in simple Given/When/Then language that is turned into executable tests.

5. Explain the test pyramid. A model with many small, fast, isolated unit tests at the base, fewer service or API tests in the middle, and few slow, broad UI end-to-end tests at the top; higher layers have lower granularity and isolation and longer execution time.

6. Give two benefits and two challenges of DevOps for testing. Benefits: fast feedback on code quality, and a smaller regression risk because automated regression tests run on every change. Challenges: the delivery pipeline and its CI and CD tools must be built and maintained, and test automation needs resources and can be hard to maintain.

Contents This chapter on its own page

munotes.in96

Chapter Eighteen

The Role of Testing in Each Phase

Syllabus topic Module 1, "Software Development Life Cycle (SDLC): Role of testing in each phase"

In one line

Testing is not one phase near the end: in every phase of development there is something a tester should be doing, from questioning the requirements on the first day to checking each change years after release.

In the wording a student can write in an examination: the role of testing in each SDLC phase is: in requirements, reviewing requirements for ambiguity, completeness and testability and deriving acceptance criteria; in design, reviewing designs and designing integration and system tests; in implementation, unit testing, static analysis and code reviews; in the testing phase, integration, system and acceptance testing; at deployment, installation, smoke and operational acceptance checks; and in maintenance, confirmation and regression testing of every change, guided by impact analysis.

Why every phase needs a tester

Chapter Four's third principle, early testing saves time and money, and Boehm and Basili's figure in Chapter Three, on why software must be tested, that a fix after delivery "is often 100 times more expensive than finding and fixing it during the requirements and design phase", lead to one conclusion: the cheapest place to find a defect is the phase that made it. A misunderstood rule found while the requirement is being written costs a conversation. The same rule found by a student in the examination hall costs refunds, apologies and an emergency release.

The ISTQB syllabus turns this into four good practices that hold for any SDLC model: "For every software development activity, there is a corresponding test activity"; each test level has its own objectives; "Test analysis and design for a given test level begins during the corresponding development phase"; and "Testers are involved in reviewing work products as soon as drafts of these work products are available". This chapter goes through the phases one at a time and says what those practices mean in each.

Requirements: the most valuable testing there is

ISO/IEC/IEEE 12207:2026 defines a requirement as a "statement that translates or expresses a need and its associated constraints and conditions". Testing's job in this phase is to check that statements are right before anything is built on them.

What a tester does:

  • Reviews each requirement for ambiguity, contradiction and gaps, asking what happens at every edge: on the last date, one day after it, with no backlog papers, with the maximum.
  • Checks testability. ISO/IEC/IEEE 24765 defines testability as the "extent to which an objective and feasible test can be designed to determine whether a requirement is met". The portal should be fast fails that test; the fee page appears within 2 seconds for 500 simultaneous users passes it.
  • Writes acceptance criteria with the users, which become the acceptance tests of the V-model (Chapter Fifteen) and the Definition of Ready in agile work (Chapter Seventeen).
  • Starts traceability, giving every requirement an identifier that tests will point back to.
  • Identifies risks, which will decide where testing effort goes.
munotes.in97

The Role of Testing in Each Phase

Design: testing the plan before the building

What a tester does:

  • Reviews the architecture and detailed design against the requirements: does every requirement have a place in the design? Can the fee service really take 500 requests at once?
  • Designs integration tests from the interfaces between components, as the V-model pairs them.
  • Designs system tests and the test environment, including test data, from the specification.
  • Reviews the design for testability: can the clock be set for late-fee tests? Can the payment gateway be replaced by a simulator in the test environment?

Implementation: the developer's testing

What happens:

  • Unit (component) testing by developers, often written first under test-driven development (Chapter Thirty-Three onwards covers unit testing).
  • Static analysis by tools that read code without running it, finding unused variables, unreachable code and suspicious patterns (Chapter Twenty-Seven, on the kinds of V&V).
  • Code reviews and inspections by other developers (Chapters Twenty-Eight to Thirty).
  • Continuous integration, which builds and runs the tests on every change (Chapter Thirty-Nine).

The testing phase: integration, system and acceptance

Here the test levels of the V-model execute: integration testing of the interfaces, system testing of the whole against its specification, system integration testing against external systems such as the payment gateway, and acceptance testing by the users. Module 1's last row takes each of these in turn, from Chapter Thirty-One, on the strategic approach to testing.

Deployment: the last checks before users arrive

What a tester does:

  • Installation or deployment testing: the system installs correctly in production, with the right configuration.
  • A smoke test in production: the main functions respond after deployment.
  • Operational acceptance testing: backups, recovery, monitoring and the other things the IT team will need to run the system (a form of acceptance testing, Chapter Forty-Two).

Maintenance: testing never stops

A system in use changes, and the ISTQB syllabus names the kinds of maintenance: "corrective, adaptive to changes in the environment or improve performance or maintainability". It also names the triggers for maintenance testing: modifications (planned enhancements, corrective changes, hot fixes); upgrades or migrations of the operational environment, including data conversion; and retirement, which may need testing of data archiving and of restoring archived data.

Testing a change "includes both evaluating the success of the implementation of the change and the checking for possible regressions in parts of the system that remain unchanged (which is usually most of the system)". The size of that effort depends, in the syllabus's words, on "The degree of risk of the change", "The size of the existing system" and "The size of the change". The tool for judging it is impact analysis: the "identification of all system and software products that a change request affects" (ISO/IEC/IEEE 24765).

munotes.in98

The Role of Testing in Each Phase

The phases side by side

SDLC phaseTesting's roleStatic or dynamicMain work products
RequirementsReview requirements; check testability; write acceptance criteria; start traceability; identify risksStaticReview comments; acceptance criteria; risk list
DesignReview designs; design integration and system tests and the environmentStatic, plus test designTest designs; environment requirements
ImplementationUnit testing; static analysis; code reviews; continuous integrationBothUnit tests; review records; analysis reports
TestingIntegration, system, system integration and acceptance testingDynamicTest logs; defect reports; completion report
DeploymentInstallation, smoke and operational acceptance testingDynamicDeployment checklist; smoke test result
MaintenanceImpact analysis; confirmation and regression testing of each changeBothUpdated regression suite; test reports

Worked example: testing the requirements before anything is built

The cheapest test in this whole book runs on sentences. A reviewer reading ExamReg's draft requirements looks first for words that no test can check, because a requirement nobody can test is a requirement nobody can prove was met. The program below does a crude version of that first pass: it flags any requirement containing a word from a list of vague terms.

VAGUE = ["fast", "quickly", "user-friendly", "easy", "appropriate", "as needed",
         "etc", "approximately", "reasonable", "adequate", "efficient", "normally"]

requirements = {
    "R1": "A form submitted 1 to 7 days after the last date pays a late fee of Rs 100.",
    "R2": "The fee page should load quickly even when many students use the portal.",
    "R3": "The hall ticket shows the seat number, subjects, dates etc.",
    "R4": "The portal locks an account for 30 minutes after 3 wrong passwords in a row.",
    "R5": "Error messages must be user-friendly and appropriate.",
    "R6": "The fee page appears within 2 seconds for 500 simultaneous users.",
}
for rid, text in requirements.items():
    words = text.lower().replace(".", " ").replace(",", " ").split()
    found = [v for v in VAGUE if (v in words if " " not in v else v in text.lower())]
    verdict = "testable as written" if not found else "VAGUE: " + ", ".join(found)
    print(f"{rid}: {verdict}")
R1: testable as written
R2: VAGUE: quickly
R3: VAGUE: etc
R4: testable as written
R5: VAGUE: user-friendly, appropriate
R6: testable as written

Three of six requirements come back vague, and each flag is a question for the exam cell before design starts. How quickly, for how many students? Which fields does a hall ticket show, exactly, since "etc." cannot be tested? What does a user-friendly error message say? R6 is what R2 becomes once someone answers the first question, and it can be tested with a load test.

munotes.in99

The Role of Testing in Each Phase

The program is deliberately crude, and its limits are the lesson. A word list catches vague words; it cannot catch a requirement that is precise and wrong, like the fee table that forgot negative days in Chapter One, on what software testing is. That needs a person reading the requirement and asking what happens at every edge, which is why requirement reviews are done by people, with tools only to help.

What it does not mean

The testing phase is not the only place testing happens. It is where dynamic testing of the integrated system happens; testing activities run in every phase.

Early testing does not replace late testing. Shift left moves effort earlier, and system and acceptance testing still run; they simply find fewer defects.

Maintenance testing is not only retesting the fix. It is also regression testing of the unchanged system, usually the larger part, sized by impact analysis.

Testability is not only a property of code. A requirement is testable or not before a line of code exists.

Quick revision

  • ISTQB good practices: a test activity for every development activity; distinct objectives per level; test design during the matching phase; testers review drafts early.
  • Requirements: review for ambiguity, gaps, testability ("extent to which an objective and feasible test can be designed"); acceptance criteria; traceability; risks.
  • Design: review designs; design integration and system tests; design the environment; check design testability.
  • Implementation: unit testing, static analysis, code reviews, continuous integration.
  • Testing: integration, system, system integration, acceptance testing.
  • Deployment: installation, smoke, operational acceptance testing.
  • Maintenance: triggers are modifications, upgrades or migrations, retirement; test the change and regressions; scope by risk and size; impact analysis.
  • Fixing after delivery is often 100 times dearer than in requirements and design (Boehm and Basili 2001).

Test yourself

1. Explain the role of testing in each phase of the SDLC. Requirements: review for ambiguity, gaps and testability, write acceptance criteria. Design: review designs and design integration and system tests. Implementation: unit tests, static analysis and code reviews. Testing: integration, system and acceptance testing. Deployment: installation and smoke testing. Maintenance: confirmation and regression testing guided by impact analysis.

2. What is testability of a requirement? Give a testable and an untestable example. The extent to which an objective and feasible test can be designed to show the requirement is met. Untestable: the fee page should load quickly. Testable: the fee page appears within 2 seconds for 500 simultaneous users.

3. What triggers maintenance testing? Modifications such as planned enhancements, corrective changes and hot fixes; upgrades or migrations of the operational environment, including data conversion; and retirement, which may require testing of data archiving and restoration.

munotes.in100

The Role of Testing in Each Phase

4. What decides how much maintenance testing is needed? The degree of risk of the change, the size of the existing system and the size of the change, judged with impact analysis of what the change affects.

5. Why is testing in the requirements phase the most valuable? Because a defect in a requirement is copied into the design, code and tests built on it, and fixing it after delivery can cost around a hundred times more than fixing it while the requirement is being written.

Contents This chapter on its own page

munotes.in101

Chapter Nineteen

Software Quality Factors: McCall's Model

Syllabus topic Module 1, "Software Development Life Cycle (SDLC): Software quality factors"

In one line

McCall's model breaks the vague word "quality" into eleven specific factors, grouped by what a user does with the software: operate it, revise it, or move it to a new setting.

In the wording a student can write in an examination: McCall's quality model (McCall, Richards and Walters, 1977) defines eleven software quality factors in three groups: product operation (correctness, reliability, efficiency, integrity, usability), product revision (maintainability, flexibility, testability) and product transition (portability, reusability, interoperability). Each factor is refined into measurable criteria, such as traceability, error tolerance and modularity, and each criterion into metrics.

Why a quality model was needed

By the mid-1970s the United States Air Force was buying a great deal of software, and it had no way to say, in a contract, what quality it wanted. Specifications described what a program must do; they said nothing about whether it would be easy to fix, change or move. The report's own diagnosis was that the emphasis had always been on the product's first operation: "Specifications and testing only stress factors such as the correctness or reliability." Yet life-cycle costs showed that maintenance and redesign cost more than initial development.

So the Air Force's Rome Air Development Center asked General Electric to define software quality in parts that could be specified and measured. The result, published as RADC-TR-77-369 in November 1977, is the report students know as McCall's model.

Three ways of looking at a product

McCall's team found that everything a customer does with delivered software falls into three activities: operating it, revising it, and transitioning it to a new environment. The eleven factors hang from those three, and the report's Figure 3.1-1 gives each factor the question a user would ask. The questions below are McCall's, read from the figure.

ActivityFactorMcCall's question for it
Product operationCorrectnessDoes it do what I want?
ReliabilityDoes it do it accurately all of the time?
EfficiencyWill it run on my hardware as well as it can?
IntegrityIs it secure?
UsabilityCan I run it?
Product revisionMaintainabilityCan I fix it?
FlexibilityCan I change it?
TestabilityCan I test it?
Product transitionPortabilityWill I be able to use it on another machine?
ReusabilityWill I be able to reuse some of the software?
InteroperabilityWill I be able to interface it with another system?

The grouping is the most examined part of the model, and the questions make it easy to remember: operation is about using the software today; revision is about changing it; transition is about moving it or its parts somewhere else.

The eleven factors, as McCall defined them

The report's Table 3.1-1 defines each factor in one line.

munotes.in102

Software Quality Factors: McCall's Model

FactorMcCall's definition (1977)
Correctness"Extent to which a program satisfies its specifications and fulfills the user's mission objectives."
Reliability"Extent to which a program can be expected to perform its intended function with required precision."
Efficiency"The amount of computing resources and code required by a program to perform a function."
Integrity"Extent to which access to software or data by unauthorized persons can be controlled."
Usability"Effort required to learn, operate, prepare input, and interpret output of a program."
Maintainability"Effort required to locate and fix an error in an operational program."
Testability"Effort required to test a program to insure it performs its intended function."
Flexibility"Effort required to modify an operational program."
Portability"Effort required to transfer a program from one hardware configuration and/or software system environment to another."
Reusability"Extent to which a program can be used in other applications"
Interoperability"Effort required to couple one system with another."

Notice how many definitions begin with effort. For the revision and transition factors, quality is measured by how much work a change costs, which is exactly the cost the Air Force was trying to control.

From factors to criteria to metrics

A factor is how a user sees quality, and it cannot be measured directly. McCall's model therefore works in three layers.

  1. Factors: the user's view (reliability).
  2. Criteria: attributes of the software that produce the factor. The report gives reliability's as "error tolerance, consistency, accuracy, and simplicity", and it defines each criterion, for example error tolerance as "Those attributes of the software that provide continuity of operation under nonnominal conditions."
  3. Metrics: measurements of each criterion, taken from the code, the documents or the test results.

The report defines twenty-three criteria, and many serve more than one factor. The program below records, from the report's Table 4.1-1, which factors each criterion supports, then counts them both ways.

related = {   # McCall 1977, Table 4.1-1: criterion -> the factors it supports
    "traceability": ["correctness"],
    "completeness": ["correctness"],
    "consistency": ["correctness", "reliability", "maintainability"],
    "accuracy": ["reliability"],
    "error tolerance": ["reliability"],
    "simplicity": ["reliability", "maintainability", "testability"],
    "modularity": ["maintainability", "flexibility", "testability",
                   "portability", "reusability", "interoperability"],
    "generality": ["flexibility", "reusability"],
    "expandability": ["flexibility"],
    "instrumentation": ["testability"],
    "self-descriptiveness": ["flexibility", "maintainability", "testability",
                             "portability", "reusability"],
    "execution efficiency": ["efficiency"],
    "storage efficiency": ["efficiency"],
    "access control": ["integrity"],
    "access audit": ["integrity"],
    "operability": ["usability"],
    "training": ["usability"],
    "communicativeness": ["usability"],
    "software system independence": ["portability", "reusability"],
    "machine independence": ["portability", "reusability"],
    "communications commonality": ["interoperability"],
    "data commonality": ["interoperability"],
    "conciseness": ["maintainability"],
}
print(len(related), "criteria")
factors = {}
for criterion, fs in related.items():
    for f in fs:
        factors.setdefault(f, []).append(criterion)
for f in sorted(factors, key=lambda f: -len(factors[f])):
    print(f"{f:<17}{len(factors[f])} criteria: {', '.join(factors[f])}")
print()
widest = sorted(related, key=lambda c: -len(related[c]))[:3]
for c in widest:
    print(f"'{c}' supports {len(related[c])} factors")
23 criteria
maintainability  5 criteria: consistency, simplicity, modularity, self-descriptiveness, conciseness
reusability      5 criteria: modularity, generality, self-descriptiveness, software system independence, machine independence
reliability      4 criteria: consistency, accuracy, error tolerance, simplicity
testability      4 criteria: simplicity, modularity, instrumentation, self-descriptiveness
flexibility      4 criteria: modularity, generality, expandability, self-descriptiveness
portability      4 criteria: modularity, self-descriptiveness, software system independence, machine independence
correctness      3 criteria: traceability, completeness, consistency
interoperability 3 criteria: modularity, communications commonality, data commonality
usability        3 criteria: operability, training, communicativeness
efficiency       2 criteria: execution efficiency, storage efficiency
integrity        2 criteria: access control, access audit

'modularity' supports 6 factors
'self-descriptiveness' supports 5 factors
'consistency' supports 3 factors
munotes.in103

Software Quality Factors: McCall's Model

Two facts stand out from the counts. Modularity supports six of the eleven factors and self-descriptiveness five: a program built from independent, well-explained modules is easier to fix, change, test, move and reuse, all at once. That is the strongest argument in the report for good design. And testability appears as a factor in its own right, built from simplicity, modularity, instrumentation and self-descriptiveness: a design decision, made long before testing, decides how hard testing will be.

Factors pull against each other

The report is honest that the factors cannot all be maximised together. Its section 4.2 lists the trade-offs generally found, and almost all of them involve efficiency:

Trade-off (McCall, section 4.2)Why
Integrity against efficiencyAccess control adds code and processing, lengthening run time and using storage
Usability against efficiencyEasing the operator's task costs code and processing
Maintainability and testability against efficiencyModular, instrumented, commented code carries overhead; highly optimised code is hard to maintain and test
Portability against efficiencyDirect, optimised code ties the program to one system
Flexibility, reusability and interoperability against efficiencyGenerality and standard interfaces add overhead
Flexibility, reusability and interoperability against integrityGeneral data structures, reusable code and coupled systems open more ways in

The practical consequence for a project is that its quality requirements must rank the factors. ExamReg must be correct and secure above all, because it handles fees and student records; it can afford to be a little less efficient than a hand-optimised program in exchange for being maintainable, because the fee rules change every session.

Worked example: ranking McCall's factors for ExamReg

FactorImportance for ExamRegReason
CorrectnessVery highA wrong fee or a wrong eligibility decision harms a student directly
IntegrityVery highMarks, fees and personal data must be protected
ReliabilityHighThe portal must work every time during the two-week form window
UsabilityHighFirst-year students use it once a semester, without training
Maintainability, flexibilityHighFee rules and paper lists change every session
TestabilityHighEvery rule change must be retested quickly
EfficiencyMediumMatters on the last day's peak, less otherwise
InteroperabilityMediumMust couple to the payment gateway and the student database
Portability, reusabilityLowOne college, one server, one release line
munotes.in104

Software Quality Factors: McCall's Model

The ranking is not a table for its own sake: it decides what the testing concentrates on, which is the subject of Chapter Twenty-One, on how quality factors shape testing.

What it does not mean

McCall's model is not the current standard. It dates from 1977; the international product quality model today is ISO/IEC 25010:2023, the next chapter. Its factors survive there under other names.

Factors are not measured directly. They are refined into criteria and then metrics, which are what is actually measured.

More of every factor is not always better. The factors trade against each other, and a project chooses its balance.

Testability is not a property of the test team. In McCall's model it is a property of the software, built from simplicity, modularity, instrumentation and self-descriptiveness.

Quick revision

  • McCall, Richards and Walters, "Factors in Software Quality", RADC-TR-77-369, US Air Force, November 1977.
  • Three activities, eleven factors: operation (correctness, reliability, efficiency, integrity, usability); revision (maintainability, flexibility, testability); transition (portability, reusability, interoperability).
  • Questions: Does it do what I want? Can I fix it? Will I be able to use it on another machine?
  • Layers: factors (user view), criteria (23 software attributes), metrics (measures).
  • Modularity supports six factors and self-descriptiveness five.
  • Trade-offs: efficiency against most factors; integrity against flexibility, reusability and interoperability.

Test yourself

1. List McCall's quality factors under their three product activities. Product operation: correctness, reliability, efficiency, integrity, usability. Product revision: maintainability, flexibility, testability. Product transition: portability, reusability, interoperability.

2. Define reliability and maintainability as McCall did. Reliability is the extent to which a program can be expected to perform its intended function with required precision. Maintainability is the effort required to locate and fix an error in an operational program.

3. What are criteria in McCall's model? Give the criteria of reliability. Criteria are attributes of the software that produce a factor and can be measured through metrics. Reliability's are error tolerance, consistency, accuracy and simplicity.

4. Why can the factors not all be maximised together? Give two trade-offs. Because improving one often costs another. Integrity against efficiency: access control adds code and processing time. Portability against efficiency: optimised, direct code ties a program to one system.

5. Which criterion supports the most factors, and what does that suggest? Modularity, which supports six factors: maintainability, flexibility, testability, portability, reusability and interoperability. It suggests that a modular design raises many qualities at once.

Contents This chapter on its own page

munotes.in105

Chapter Twenty

From McCall to ISO/IEC 25010: The Quality Model Today

Syllabus topic Module 1, "Software Development Life Cycle (SDLC): Software quality factors"

In one line

McCall's eleven factors of 1977 became the international standard's product quality model, which has changed twice since: six characteristics in ISO/IEC 9126, eight in ISO/IEC 25010:2011, and nine in the current ISO/IEC 25010:2023.

In the wording a student can write in an examination: ISO/IEC 25010:2023 defines a product quality model of nine characteristics, each divided into subcharacteristics: functional suitability, performance efficiency, compatibility, interaction capability, reliability, security, maintainability, flexibility and safety. It replaced ISO/IEC 25010:2011 (eight characteristics), which had replaced ISO/IEC 9126 (six: functionality, reliability, usability, efficiency, maintainability and portability).

Why an international model replaced McCall's

McCall's model, and others like it from the same years, were written for one customer, the United States Air Force, and each research group used its own names. A company buying software from a supplier in another country needed both sides to mean the same thing by "reliable" or "maintainable". The International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC) therefore standardised a quality model, and have revised it as software changed: security was an afterthought in 1991 and a headline in 2011, and safety became a characteristic of its own in 2023 as software moved into cars, medical devices and machines.

ISO/IEC 9126: six characteristics

ISO/IEC 9126 was first issued in December 1991 and revised in 2001. Its product quality model had six characteristics, and this is the model most textbooks, including those MU names, still teach:

ISO/IEC 9126 characteristicIts subcharacteristics
FunctionalitySuitability, accuracy, interoperability, security, functionality compliance
ReliabilityMaturity, fault tolerance, recoverability, reliability compliance
UsabilityUnderstandability, learnability, operability, attractiveness, usability compliance
EfficiencyTime behaviour, resource utilisation, efficiency compliance
MaintainabilityAnalysability, changeability, stability, testability, maintainability compliance
PortabilityAdaptability, installability, co-existence, replaceability, portability compliance

ISO/IEC 25010:2011: eight characteristics

In March 2011, ISO/IEC 25010:2011 replaced ISO/IEC 9126-1 as part of a new series called SQuaRE (Systems and software Quality Requirements and Evaluation). It had eight characteristics. Functionality became functional suitability; efficiency became performance efficiency; and two new characteristics appeared: compatibility, taking co-existence from portability and interoperability from functionality, and security, which had been a mere subcharacteristic of functionality. Usability, reliability, maintainability and portability kept their names, with changes inside them.

ISO/IEC 25010:2023: nine characteristics

The second edition, published in November 2023, is the current standard. ISO's own summary of it: "The product quality model is composed of nine characteristics (which are further subdivided into subcharacteristics) that relate to quality properties of the products." The ISTQB syllabus's release notes record the headline changes: the 2023 edition "renames 'usability' to 'interaction capability', 'portability' to flexibility, and adds a new characteristic 'safety'".

Characteristic (2023)Definition (ISO/IEC 25010:2023, through SEVOCAB)Subcharacteristics
Functional suitability"capability of a product to provide functions that meet stated and implied needs of intended users when it is used under specified conditions"Functional completeness, functional correctness, functional appropriateness
Performance efficiency"capability of a product to perform its functions within specified time and throughput parameters and be efficient in the use of resources"Time behaviour, resource utilisation, capacity
Compatibility"capability of a product to exchange information with other products, or to perform its required functions while sharing the same common environment and resources"Co-existence, interoperability
Interaction capability"capability of a product to be interacted with by specified users to exchange information between a user and a system via the user interface"Appropriateness recognisability, learnability, operability, user error protection, user engagement, inclusivity, user assistance, self-descriptiveness
Reliability"capability of a product to perform specified functions under specified conditions for a specified period of time without interruptions and failures"Faultlessness, availability, fault tolerance, recoverability
Security"capability of a product to protect information and data so that persons or other products have the degree of data access appropriate to their types and levels of authorization"Confidentiality, integrity, non-repudiation, accountability, authenticity, resistance
Maintainability"capability of a product to be modified by the intended maintainers with effectiveness and efficiency"Modularity, reusability, analysability, modifiability, testability
Flexibility"capability of a product to be adapted to changes in its requirements, contexts of use, or system environment"Adaptability, scalability, installability, replaceability
Safety"capability of a product under defined conditions to avoid a state in which human life, health, property, or the environment is endangered"Operational constraint, risk identification, fail safe, hazard warning, safe integration
munotes.in106

From McCall to ISO/IEC 25010: The Quality Model Today

A few of the newer subcharacteristics deserve a sentence each, because they are the ones older books lack:

  • Faultlessness (reliability) replaces the old "maturity": the "capability of a product to perform specified functions without fault under normal operation".
  • Resistance (security) is new: the "capability of a product to sustain operations while under attack from a malicious actor".
  • Inclusivity (interaction capability) is new: the "capability of a product to be utilized by people of various backgrounds".
  • Scalability (flexibility) is new: the "capability of a product to handle growing or shrinking workloads or to adapt its capacity to handle variability".
  • Fail safe (safety) is the "capability of a product to automatically place itself in a safe operating mode, or to revert to a safe condition in the event of a failure". Therac-25, in Chapter Three on why software must be tested, is what its absence looks like.

McCall's factors in the 2023 model

Nearly every McCall factor survives, sometimes as a characteristic and sometimes as a subcharacteristic. The program below records where each lands, and then asks the question in reverse: which 2023 characteristics have no McCall ancestor at all?

mccall_to_2023 = {        # McCall 1977 factor -> ISO/IEC 25010:2023 characteristic (subcharacteristic)
    "correctness":      ("functional suitability", "functional correctness"),
    "reliability":      ("reliability", "faultlessness"),
    "efficiency":       ("performance efficiency", "time behaviour, resource utilization"),
    "integrity":        ("security", "confidentiality, integrity"),
    "usability":        ("interaction capability", "operability, learnability"),
    "maintainability":  ("maintainability", "analysability, modifiability"),
    "flexibility":      ("maintainability", "modifiability"),
    "testability":      ("maintainability", "testability"),
    "portability":      ("flexibility", "adaptability, installability"),
    "reusability":      ("maintainability", "reusability"),
    "interoperability": ("compatibility", "interoperability"),
}
iso_2023 = ["functional suitability", "performance efficiency", "compatibility",
            "interaction capability", "reliability", "security", "maintainability",
            "flexibility", "safety"]

for factor, (characteristic, sub) in mccall_to_2023.items():
    print(f"{factor:<17} -> {characteristic} ({sub})")
used = {c for c, _ in mccall_to_2023.values()}
print("2023 characteristics with no McCall factor:", [c for c in iso_2023 if c not in used])
landing = {}
for factor, (c, _) in mccall_to_2023.items():
    landing.setdefault(c, []).append(factor)
print("maintainability now holds:", landing["maintainability"])
munotes.in107

From McCall to ISO/IEC 25010: The Quality Model Today

correctness       -> functional suitability (functional correctness)
reliability       -> reliability (faultlessness)
efficiency        -> performance efficiency (time behaviour, resource utilization)
integrity         -> security (confidentiality, integrity)
usability         -> interaction capability (operability, learnability)
maintainability   -> maintainability (analysability, modifiability)
flexibility       -> maintainability (modifiability)
testability       -> maintainability (testability)
portability       -> flexibility (adaptability, installability)
reusability       -> maintainability (reusability)
interoperability  -> compatibility (interoperability)
2023 characteristics with no McCall factor: ['safety']
maintainability now holds: ['maintainability', 'flexibility', 'testability', 'reusability']

Three things come out of the mapping. Safety is the one characteristic with no ancestor in McCall, which is the 2023 edition's most important addition. Maintainability absorbs four of McCall's factors: his maintainability, flexibility, testability and reusability are all, in 2023, about how effectively a product can be modified. And there is a trap in the names: McCall's flexibility (the effort to modify a program) now sits under ISO's maintainability, while ISO's flexibility is what used to be called portability. The same word, two meanings, fifty years apart.

Quality in use: the other model

ISO/IEC 25010:2011 contained a second model, quality in use, which describes the effect of the product on its users in a real context: effectiveness, efficiency, satisfaction, freedom from risk and context coverage. The current SQuaRE series places the quality-in-use model in its own standard, ISO/IEC 25019:2023, whose vocabulary SEVOCAB also carries. The distinction is worth knowing for an examination: product quality is a property of the software, measurable in testing; quality in use is the result of using it, measured with real users in real contexts.

Worked example: ExamReg's quality requirements in 2023 terms

2023 characteristicExamReg requirement that expresses it, written to be testable
Functional suitabilityEvery fee in the fee table is charged exactly for every combination of days late, backlog papers and concession
Performance efficiencyThe fee page appears within 2 seconds for 500 simultaneous users
CompatibilityPayments pass to the gateway and receipts return, with no loss
Interaction capabilityNine of ten first-year students complete the form unaided in a usability test
ReliabilityThe portal is available 99.5 per cent of the time during the two-week form window
SecurityA student can see only their own form, fees and hall ticket
MaintainabilityA change to the late-fee table can be made and retested within one day
FlexibilityThe portal runs unchanged when the college moves it to a new server
SafetyNot applicable: nothing ExamReg does can endanger life, health, property or the environment
munotes.in108

From McCall to ISO/IEC 25010: The Quality Model Today

The last row is as important as the others. A quality model is a checklist of what to consider, not a list of what every product must have; "not applicable, because" is a valid and useful entry.

What it does not mean

ISO/IEC 9126 is not current. It was replaced in 2011; a modern answer names ISO/IEC 25010:2023 and may mention 9126 as its predecessor.

The 2023 model does not drop usability and portability. It renames them, as interaction capability and flexibility, and adds safety.

McCall's flexibility is not ISO's flexibility. McCall's is about modifying a program; ISO's is about adapting it to new requirements, contexts or environments.

Every product need not score on every characteristic. The model is used to decide which characteristics matter, and "not applicable" is a legitimate answer.

Quick revision

  • ISO/IEC 9126 (1991, 2001): six characteristics: functionality, reliability, usability, efficiency, maintainability, portability.
  • ISO/IEC 25010:2011: eight; adds compatibility and security; functionality became functional suitability, efficiency became performance efficiency.
  • ISO/IEC 25010:2023 (current, November 2023): nine: functional suitability, performance efficiency, compatibility, interaction capability (was usability), reliability, security, maintainability, flexibility (was portability), safety (new).
  • New subcharacteristics include faultlessness, resistance, inclusivity, user engagement, self-descriptiveness, scalability and the five of safety.
  • McCall's factors map into the 2023 model; only safety has no McCall ancestor; maintainability absorbs four McCall factors.
  • Quality in use (effect on users) is a separate model, now ISO/IEC 25019:2023.

Test yourself

1. Name the nine quality characteristics of ISO/IEC 25010:2023. Functional suitability, performance efficiency, compatibility, interaction capability, reliability, security, maintainability, flexibility and safety.

2. What changed between ISO/IEC 25010:2011 and 2023? Usability was renamed interaction capability, portability was renamed flexibility, and safety was added as a ninth characteristic, with new subcharacteristics such as faultlessness, resistance, inclusivity and scalability.

3. List the six characteristics of ISO/IEC 9126 and say what replaced it. Functionality, reliability, usability, efficiency, maintainability and portability. It was replaced by ISO/IEC 25010:2011, which has since been revised as ISO/IEC 25010:2023.

4. Define security and give its subcharacteristics in the 2023 model. The capability of a product to protect information and data so that persons or other products have the degree of data access appropriate to their authorisation. Its subcharacteristics are confidentiality, integrity, non-repudiation, accountability, authenticity and resistance.

5. Which McCall factor has a namesake with a different meaning in ISO/IEC 25010:2023? Flexibility. McCall's flexibility is the effort to modify an operational program, which now falls under maintainability; ISO's flexibility is the capability to be adapted to new requirements, contexts or environments, which was formerly portability.

munotes.in109

From McCall to ISO/IEC 25010: The Quality Model Today

6. Distinguish product quality from quality in use. Product quality is a property of the software itself, described by the nine characteristics and measured largely in testing. Quality in use is the effect of the product on its users in a real context, such as their effectiveness, satisfaction and freedom from risk.

Contents This chapter on its own page

munotes.in110

Chapter Twenty-One

How Quality Factors Shape Testing

Syllabus topic Module 1, "Software Development Life Cycle (SDLC): Software quality factors and their impact on testing"

In one line

Each quality a product must have calls for its own kind of test, so the quality factors a project cares about decide which tests it runs, how deep they go, and where the effort is spent.

In the wording a student can write in an examination: software quality factors affect testing in four ways. They decide the test types (functional testing for functional suitability; performance, usability, security, reliability, compatibility, maintainability and portability testing for the non-functional characteristics); they require quality requirements to be testable, which is itself a factor (testability); they decide the priority of testing through product risk analysis, each characteristic's risk level being its likelihood combined with its impact; and they decide when testing starts, since non-functional defects found late are the most dangerous.

Why quality factors change the testing

A test suite that checks only whether each function gives the right answer tests one characteristic out of nine. ExamReg could pass every one of its fee tests and still collapse under 500 users on the last day, show one student's marks to another, or baffle a first-year student with its form. Each of those is a quality failure, and none of them would be found by a functional test. The quality model of the last chapter is, for a tester, a list of the kinds of testing that might be needed; the project's priorities among the factors decide which are.

Functional and non-functional testing

The ISTQB syllabus divides testing by what it evaluates. Functional testing "evaluates the functions that a component or system should perform"; its objective is "checking the functional completeness, functional correctness and functional appropriateness", which are the three subcharacteristics of functional suitability. Non-functional testing "evaluates attributes other than functional characteristics", which the syllabus calls testing "how well the system behaves". The other eight characteristics of ISO/IEC 25010 are its subject.

The syllabus makes two observations that change how non-functional testing is planned. First, many non-functional tests are derived from functional ones: they run the same function but check that "a non-functional constraint is satisfied (e.g., checking that a function performs within a specified time". Second, timing matters: "The late discovery of non-functional defects can pose a serious threat to the success of a project", so some non-functional testing should start early, in reviews and at component level.

From each characteristic to its tests

Characteristic (ISO/IEC 25010:2023)Test typeStandard definition of the test type (through SEVOCAB)ExamReg example
Functional suitabilityFunctional testing"testing conducted to evaluate the compliance of a system or component with specified functional requirements" (ISO/IEC/IEEE 24765)Every fee rule, eligibility rule and paper choice
Performance efficiencyPerformance, load, stress and volume testingPerformance testing evaluates "the degree to which a test item accomplishes its designated functions within given constraints of time and other resources" (ISO/IEC/IEEE 29119-2)500 users on the last day; the fee page within 2 seconds
CompatibilityInteroperability and co-existence testing; cross-browser testingInteroperability testing helps ensure a system "retains the capability of exchanging information with systems of different types" (ISO/IEC/IEEE 24765)Payments to the gateway; Chrome, Firefox, Edge and Android
Interaction capabilityUsability and accessibility testingUsability testing is an "evaluation that involves representative users performing specific tasks with the system" (ISO TR 25060:2023)Ten first-year students fill the form while observed
ReliabilityReliability, recovery and failover testingReliability testing includes "evaluating the frequency with which failures occur" (ISO/IEC/IEEE 29119-1)The portal restarts cleanly after a server crash, with no half-paid forms
SecuritySecurity testing, vulnerability scanning, penetration testingSecurity testing evaluates whether a test item and its data are protected "so that unauthorized persons or systems cannot use, read, or modify them" (ISO/IEC/IEEE 29119-1)A student cannot open another student's form by changing a number in the address
MaintainabilityMaintainability testing; static analysis; reviewsMaintainability testing evaluates "the degree of effectiveness and efficiency with which a test item can be modified" (ISO/IEC/IEEE 29119-1)Time to change the late-fee table and retest it; complexity of the fee code
FlexibilityPortability, installability and adaptability testingPortability testing evaluates "the ease with which a test item can be transferred from one hardware or software environment to another" (ISO/IEC/IEEE 29119-1)Install on the new college server; run on the next database version
SafetySafety analysis and fail-safe testing(safety-critical products only)Not applicable to ExamReg
munotes.in111

How Quality Factors Shape Testing

Chapter Forty-Four takes the system-level types among these (recovery, security, stress, performance and deployment testing) in detail, and Chapters Forty-Five and Forty-Six the load and cross-browser testing that the practical sets.

Testability: the factor that governs all the others

One characteristic is different from the rest, because it is about testing itself. ISO/IEC 25010:2023 defines testability as the "capability of a product to enable an objective and feasible test to be designed and performed to determine whether a requirement is met". It sits under maintainability, and McCall made it a factor of its own in 1977.

A product with poor testability makes every other test expensive or impossible. If ExamReg's late-fee code reads the real system clock and nothing else, no test can check the fee for "16 days late" without waiting sixteen days. If the fee service can only talk to the real payment gateway, every test costs real money. Testability is designed in: a clock that tests can set, a gateway that tests can replace with a simulator, logs that show what happened. Testers who review designs (Chapter Eighteen, on the role of testing in each phase) ask for these things, because they cannot be added cheaply later.

munotes.in112

How Quality Factors Shape Testing

Quality requirements must be measurable

A quality factor becomes testable only when its requirement is stated with a number and a condition. The same rule met in the last two chapters bears repeating here, because non-functional requirements break it most often:

Untestable as writtenTestable
The portal should be fastThe fee page appears within 2 seconds for 500 simultaneous users
The portal should be reliableAvailable 99.5 per cent of the time during the form window
The form should be easy to useNine of ten first-year students complete it unaided
The code should be maintainableA change to the fee table is made and retested within one working day

Deciding where the effort goes: product risk

A project cannot test every characteristic deeply, so it tests hardest where failure would hurt most. The ISTQB syllabus calls the approach product risk analysis, and connects it directly to the quality model: product risks "are related to the product quality characteristics (e.g., described in the ISO 25010 quality model)". A risk has two attributes, likelihood and impact, and together they express the risk level. In the quantitative approach, the syllabus says, "the risk level is calculated as the multiplication of risk likelihood and risk impact."

The results decide, in the syllabus's list, the test scope, the test levels and types, the techniques and coverage, the effort for each task, and the order of testing, "to find the critical defects as early as possible".

Worked example: sharing ExamReg's test effort by risk

The ExamReg test lead scores each relevant characteristic from 1 to 5 for likelihood of failure and for impact if it fails. The scores are judgements, written down so they can be argued with. The program multiplies them into risk levels, as the syllabus describes, and shares 400 hours of test effort in proportion.

EFFORT_HOURS = 400
risks = {                          # characteristic: (likelihood 1-5, impact 1-5)
    "functional suitability": (4, 5),   # many fee and eligibility rules; wrong fee harms students
    "performance efficiency": (4, 4),   # a last-day rush never tested before
    "security":               (3, 5),   # student data and payments
    "interaction capability": (3, 3),   # first-year students, once a semester
    "reliability":            (2, 4),   # a crash during the window is serious
    "compatibility":          (3, 2),   # gateway adapter, several browsers
    "maintainability":        (2, 2),   # fee table changes each session
    "flexibility":            (1, 2),   # one server
}
levels = {c: l * i for c, (l, i) in risks.items()}
total = sum(levels.values())
print(f"{'characteristic':<24}{'L':>3}{'I':>3}{'level':>7}{'share':>8}{'hours':>7}")
for c in sorted(levels, key=lambda c: -levels[c]):
    l, i = risks[c]
    share = levels[c] / total
    print(f"{c:<24}{l:>3}{i:>3}{levels[c]:>7}{share:>8.1%}{round(EFFORT_HOURS * share):>7}")
print(f"{'total':<24}{'':>6}{total:>7}")
munotes.in113

How Quality Factors Shape Testing

characteristic            L  I  level   share  hours
functional suitability    4  5     20   25.0%    100
performance efficiency    4  4     16   20.0%     80
security                  3  5     15   18.8%     75
interaction capability    3  3      9   11.2%     45
reliability               2  4      8   10.0%     40
compatibility             3  2      6    7.5%     30
maintainability           2  2      4    5.0%     20
flexibility               1  2      2    2.5%     10
total                              80

The table turns a quality model into a test plan. Functional suitability gets the most effort, but performance efficiency and security together get more than a third of the total, far more than a team that thought only of functional testing would have given them. Flexibility gets a few hours of installation testing. And because the scores are written down, the exam cell can challenge them: if they believe a crash during the window is catastrophic, they can raise reliability's impact to 5 and rerun the arithmetic.

What it does not mean

Non-functional testing is not optional extra testing. It is testing of eight of the nine characteristics, and its failures are often the most visible.

A risk score is not an objective measurement. It is a structured judgement; its value is that it is written down, argued over and revisited.

Testability is not the testers' responsibility alone. It is designed into the product by architects and developers, and testers ask for it in reviews.

Each characteristic does not need its own separate test phase. Many non-functional tests reuse functional tests with an extra check, and several run at every level.

Quick revision

  • Quality factors decide the test types, the testability needed, the priorities and the timing of testing.
  • Functional testing checks functional suitability (completeness, correctness, appropriateness); non-functional testing checks "how well the system behaves" against the other characteristics.
  • Pairs: performance efficiency with performance, load, stress and volume testing; compatibility with interoperability and cross-browser testing; interaction capability with usability and accessibility testing; reliability with reliability and recovery testing; security with security testing; maintainability with static analysis and maintainability testing; flexibility with portability and installation testing; safety with safety analysis.
  • Testability: the capability of a product to enable an objective and feasible test to be designed and performed (ISO/IEC 25010:2023); designed in, not added later.
  • Product risk: likelihood and impact; quantitatively, risk level = likelihood × impact (ISTQB); used to set scope, types, techniques, effort and order.
  • Worked example: 400 hours shared by risk level; functional suitability first, performance efficiency and security next.

Test yourself

1. How do software quality factors affect testing? They decide which test types are needed, one or more for each characteristic that matters; they require quality requirements to be measurable and the product to be testable; they set priorities through product risk analysis; and they push some non-functional testing early, because late non-functional defects are the most dangerous.

munotes.in114

How Quality Factors Shape Testing

2. Distinguish functional from non-functional testing, with one example of each. Functional testing checks what the system does against its functional requirements, for example that a form 3 days late is charged Rs 100. Non-functional testing checks how well it behaves, for example that the fee page appears within 2 seconds for 500 simultaneous users.

3. Which test types check performance efficiency and security? Performance efficiency: performance, load, stress and volume testing. Security: security testing, including vulnerability scanning and penetration testing.

4. What is testability, and how is it achieved? The capability of a product to enable an objective and feasible test to be designed and performed to show a requirement is met. It is achieved in design, for example by letting tests set the clock, replace external services with simulators, and read logs of what happened.

5. How is a product risk level calculated, and what is it used for? In the quantitative approach, as risk likelihood multiplied by risk impact. It is used to decide the test scope, the levels, types and techniques, the effort for each area, and the order of testing.

Contents This chapter on its own page

munotes.in115

Chapter Twenty-Two

What Quality Means

Syllabus topic Module 1, "Definition of Quality and Quality Assurance: Understanding quality"

In one line

Quality is how well something does what it is meant to do for the people who depend on it; the standards make that precise as the degree to which a product's characteristics fulfil its requirements.

In the wording a student can write in an examination: ISO 9000 defines quality as the "degree to which a set of inherent characteristics ... of an object ... fulfils requirements", where a requirement is a "need or expectation that is stated, generally implied, or obligatory". Two classic short definitions sit behind it: Joseph Juran's fitness for use and Philip Crosby's conformance to requirements. David Garvin (1984) showed that people mean five different things by quality: transcendent, product-based, user-based, manufacturing-based and value-based.

Why "quality" needs defining at all

Everybody uses the word and almost nobody means the same thing by it. A student says ExamReg is good quality because it looks modern; the exam cell says so because it charges the right fees; the IT cell because it never crashes; the principal because it cost less than expected. All four are talking about quality, and they could each be satisfied while the others are not.

David Garvin, writing in MIT Sloan Management Review in the autumn of 1984, found the same confusion in the research literature: scholars in philosophy, economics, marketing and operations management had studied quality, "but each group has viewed it from a different vantage point", producing "a host of competing perspectives". A testing and quality course cannot proceed on a word with four meanings. It needs one definition to measure against, and an understanding of the others so that no stakeholder's meaning is forgotten.

ISO's definition

The international quality management standard's vocabulary, ISO 9000, now in its 2026 edition, defines quality in one sentence. ISO's own summary on its website gives it: quality is the "degree to which a set of inherent characteristics [or distinguishing features] of an object", where an object is "anything perceivable or conceivable, such as a product, service, process, person, organization, system or resource", "fulfils requirements."

Take the definition apart; every word is doing work.

  1. Degree. Quality is a scale, not a yes or no. A product has more or less of it.
  2. Inherent characteristics. Characteristics the object has in itself, such as correctness, speed or security, not ones assigned to it from outside, such as its price or its owner.
  3. Of an object. Anything: a program, a service, a process, even an organisation. The definition applies to the testing process as much as to the software.
  4. Fulfils requirements. Quality is measured against requirements, and the standards define a requirement broadly: a "need or expectation that is stated, generally implied, or obligatory" (SEVOCAB, from ISO/IEC 19770-1:2017).
munotes.in116

What Quality Means

The word "implied" in the last line matters as much as any. A student never writes down the portal must not show my marks to someone else, but it is an expectation that is generally implied, and a portal that breaks it has poor quality however well it meets its written specification. The standards' word for fulfilling a requirement is conformity, and its opposite, a nonconformity, is the "non-fulfillment of a requirement".

Fitness for use, and conformance to requirements

ASQ, the American Society for Quality, records in its glossary that quality is "A subjective term for which each person or sector has its own definition", and gives two technical meanings: "the characteristics of a product or service that bear on its ability to satisfy stated or implied needs" and "a product or service free of deficiencies". It then names the two classic short definitions: "According to Joseph Juran, quality means 'fitness for use'; according to Philip Crosby, it means 'conformance to requirements.'"

Fitness for use (Juran)Conformance to requirements (Crosby)
AsksDoes it serve the user's actual purpose?Does it meet what was specified?
Judged byThe user, in useMeasurement against the specification
StrengthKeeps the user's real need in viewObjective, measurable, testable
WeaknessHard to measure; users differA product can conform to a wrong specification
In testing termsValidationVerification

The last row connects this chapter to the rest of the syllabus. Checking conformance to requirements is verification; checking fitness for use is validation. Chapter Four's seventh principle, the absence-of-defects fallacy, is exactly the gap between the two: software can conform perfectly and still not be fit for use. Verification and validation have their own chapter (Chapter Twenty-Six).

Garvin's five approaches to quality

Garvin's article identifies "five major approaches to the definition of quality": "(1) the transcendent approach of philosophy; (2) the product-based approach of economics; (3) the user-based approach of economics, marketing, and operations management; and (4) the manufacturing-based and (5) value-based approaches of operations management." Each is a way a real person judges software.

ApproachQuality is ...How a person using it judges ExamReg
TranscendentSomething recognised on sight but hard to defineIt just feels solid and well made
Product-basedA measurable amount of desirable attributes the product hasCounts the features: backlog papers, concessions, online payment, hall ticket download
User-basedWhatever satisfies the particular userA first-year student: I filled it in five minutes without help
Manufacturing-basedConformance to the specification: getting it right the first timeThe test lead: every requirement verified, no open defect
Value-basedPerformance at an acceptable priceThe principal: it does what the paper form did, faster, for the budget
munotes.in117

What Quality Means

The practical lesson is that a quality plan has to satisfy several approaches at once, because a project has stakeholders who use each of them. ExamReg's tests serve the manufacturing-based view; its usability sessions serve the user-based view; its budget serves the value-based view. A project that pleases only one of its judges has not delivered quality in the eyes of the others.

Quality is not grade

One distinction clears up a common confusion. SEVOCAB defines grade as a "scalar, qualitative measure or ranking to distinguish the relative fitness for use of similar items". A budget smartphone and a premium one are different grades; either can be of high or low quality. A low-grade product is not a failure; a low-quality product is.

For software, grade is the set of features and capabilities chosen (a basic ExamReg without online payment is a lower grade than one with it); quality is how well whichever grade was chosen fulfils its requirements. A basic ExamReg that charges every fee correctly and never crashes has higher quality than a feature-rich one that overcharges students.

Satisfaction: quality as the customer experiences it

The standards also name the outcome that quality is for. Customer satisfaction is the "state of fulfillment in which the needs of a customer are met or exceeded for the customer's expected experiences as assessed by the customer at the moment of evaluation" (SEVOCAB, from ISO/IEC/IEEE 24765). Three phrases in it are worth an examiner's attention: the needs are met or exceeded; it is assessed by the customer, not by the supplier; and it is judged at the moment of evaluation, so it can change. A portal that satisfied students in May can disappoint them in November, when their expectations have risen or a competitor has shown them something better.

Worked example: is ExamReg release 2.0 of good quality?

QuestionEvidenceJudgement
Does it conform to its stated requirements?412 test cases passed; one defect open, of minor severityHigh conformance (manufacturing-based view)
Does it meet implied needs?Security testing found no way to see another student's data; no requirement had said so explicitlyImplied need met
Is it fit for use by first-year students?Nine of ten completed the form unaided in usability sessionsFit for use, with one improvement to make (user-based view)
What grade is it?Online payment, backlog papers, hall ticket download; no mobile appA middle grade, chosen deliberately
Is it good value?Delivered within budget; replaced three weeks of manual form checkingGood value (value-based view)

A single-word answer, yes, would hide everything useful in the table. The honest answer is that release 2.0 has high conformance, meets the implied needs that were checked, is fit for use with one improvement, and is good value at its grade.

munotes.in118

What Quality Means

What it does not mean

Quality does not mean luxury or expense. It is the degree to which requirements are fulfilled; a simple, cheap product can have excellent quality.

Quality is not only conformance to the written specification. Implied needs count, and a product can conform to a wrong specification.

Quality is not only the user's opinion either. Satisfaction is assessed by the customer, but conformance is measured objectively; both matter.

Quality is not a yes or no. ISO's definition begins with "degree".

Quick revision

  • Quality (ISO 9000:2026, via ISO): "degree to which a set of inherent characteristics ... of an object ... fulfils requirements".
  • Requirement: a need or expectation that is stated, generally implied, or obligatory. Conformity: fulfilment of a requirement; nonconformity: non-fulfilment.
  • Juran: fitness for use (validation). Crosby: conformance to requirements (verification).
  • Garvin (MIT Sloan Management Review, Fall 1984): five approaches: transcendent, product-based, user-based, manufacturing-based, value-based.
  • Grade ranks similar items by capability; it is not quality.
  • Customer satisfaction: needs met or exceeded, as assessed by the customer at the moment of evaluation.

Test yourself

1. Define quality as ISO 9000 does, and explain its key terms. Quality is the degree to which a set of inherent characteristics of an object fulfils requirements. Degree makes it a scale; inherent characteristics belong to the object itself, not assigned ones like price; an object can be a product, service, process or organisation; and requirements are needs or expectations that are stated, generally implied or obligatory.

2. Distinguish Juran's and Crosby's definitions of quality. Juran defined quality as fitness for use: whether the product serves the user's actual purpose. Crosby defined it as conformance to requirements: whether the product meets what was specified. The first corresponds to validation, the second to verification.

3. List Garvin's five approaches to quality with an example of each. Transcendent (recognised on sight), product-based (the amount of desirable attributes, such as features), user-based (whatever satisfies the user), manufacturing-based (conformance to specification) and value-based (performance at an acceptable price).

4. Distinguish quality from grade. Grade ranks similar items by their capabilities or features; quality is how well an item of any grade fulfils its requirements. A low-grade product can have high quality, and a high-grade one low quality.

5. Why does the word "implied" in the definition of requirement matter to a tester? Because users expect things they never write down, such as privacy of their data, and a product that breaks an implied need has poor quality even if it meets its written specification; testers must therefore test implied needs too.

Contents This chapter on its own page

munotes.in119

Chapter Twenty-Three

Quality in Software Development

Syllabus topic Module 1, "Definition of Quality and Quality Assurance: Understanding quality, in software development"

In one line

Software quality is quality as the last chapter defined it, applied to a product that is designed but never manufactured, that does not wear out but changes constantly, and that can be judged from three places: its code, its behaviour under test, and its use by real people.

In the wording a student can write in an examination: software quality is the "capability of a software product to satisfy stated and implied needs when used under specified conditions" (ISO/IEC 25000:2014), or, more narrowly, the "degree to which a software product meets established requirements" (IEEE 730-2014, since replaced by a 2026 edition). Understanding quality in software development means understanding two things. First, why software is different: its faults are design faults, not physical ones; it does not wear out, but it declines when its use or its code changes; and, in Frederick Brooks's analysis, it is complex, must conform to interfaces other people designed, is constantly changed, and is invisible. Second, that its quality has three views: internal quality (the code and documents), external quality (the software running) and quality in use (the software in its users' hands).

Why software strains the ordinary idea of quality

Much of quality management grew up in manufacturing. ASQ's history of quality records that factory quality was long kept by inspecting products, and that during the Second World War the United States armed forces moved from inspecting every unit to sampling inspection. Both ideas assume many copies of one design, any of which can come out slightly wrong.

Software has no such copies. Every copy of ExamReg is the same program; if the fee rule is wrong in one, it is wrong in all of them, and inspecting a sample of copies would find nothing that inspecting one would not. The quality of software is decided almost entirely before the first copy exists, in its requirements, design and code. Michael Lyu's Handbook of Software Reliability Engineering (1996) puts the difference in one sentence: "Unlike hardware faults which are mostly physical faults, software faults are design faults", which are "harder to visualize, classify, detect, and correct."

The consequence runs through this whole course. Quality in software development cannot be inspected into the product at the end; it has to be built in while the product is designed and written. That is why so much of the syllabus is about reviews, process and prevention, the subject of quality assurance in Chapter Twenty-Four, and not only about testing the finished program.

Software does not wear out, but it does decline

The second difference is time. A pump, a tyre or a hard disk wears out: its parts age, and it fails more often until it is replaced. Lyu states the contrast plainly: software reliability differs from hardware reliability "in the sense that software does not wear out, burn out, or deteriorate, i.e., its reliability does not decrease with time." An unchanged program, given the same input in the same conditions, does on its thousandth run what it did on its first.

munotes.in120

Quality in Software Development

That does not mean software quality stays fixed. Lyu names the two ways it falls: "software may experience reliability decrease due to abrupt changes of its operational usage or incorrect modifications to the software." Both happen to ExamReg every session.

  • A change in use. The code that handled a few dozen forms a day in the first week meets 300 on the last date. Nothing in the program has changed, but it is being used as it never was before, and a fault that no ordinary day reached is reached.
  • A change in the code. The exam cell revises the fee table, a developer edits the fee function, and the edit breaks a case that used to work. Every correction is itself a change to a design, and can bring a new fault with it. Guarding against that is the job of regression testing (Chapter Thirty-Nine).

The same passage names the good news. Software "generally enjoys reliability growth during testing and operation", because each fault found and removed is gone for good from every copy.

HardwareSoftware
Where its faults come fromMostly physical: wear, material, manufactureDesign: requirements, design and code
CopiesEach unit can come out differentlyEvery copy is identical
Left unchanged over timeWears out, and fails more oftenDoes not wear out
What lowers its reliabilityAge and wearA change in how it is used, or a faulty change to it
Finding and fixing a faultRestores the unit to its designChanges the design: reliability can grow, or a new fault can enter

Chapter Ninety, on software reliability, returns to this contrast with the failure curves of hardware and software.

Brooks: four properties that make software quality hard

In "No Silver Bullet", a paper first given at the IFIP World Computing Conference in 1986 and later reprinted in his book The Mythical Man-Month, Frederick Brooks separated the difficulties of software into accidental ones, which better tools can remove, and essential ones, which belong to software itself. He named "the inherent properties of this irreducible essence of modern software systems: complexity, conformity, changeability, and invisibility." Each is a reason quality is hard to achieve in software development, and each makes a demand on testing.

Complexity. "Software entities are more complex for their size than perhaps any other human construct, because no two parts are alike (at least above the statement level)." Brooks traced unreliability straight to it: "From the complexity comes the difficulty of enumerating, much less understanding, all the possible states of the program, and from that comes the unreliability." For a tester this is why exhaustive testing is impossible, the second of the seven principles in Chapter Four, and why test design techniques exist: to choose, from far more states than anyone could try, the few tests most likely to find faults.

munotes.in121

Quality in Software Development

Conformity. A physicist can hope that nature obeys a few simple laws. Software must instead fit interfaces that other people designed for their own reasons, and Brooks observed that "much complexity comes from conformation to other interfaces; this cannot be simplified out by any redesign of the software alone." ExamReg must fit the paper codes it is given, the exam cell's fee rules, the payment gateway's protocol and the browsers students use. Its developers can simplify none of them, and each is a place where the program can be right by its own logic and wrong against the thing it must fit. Integration testing and compatibility testing exist for exactly these seams.

Changeability. "All successful software gets changed." Software, Brooks wrote, "is pure thought-stuff, infinitely malleable", and "The pressures for extended function come chiefly from users who like the basic function and invent new uses for it." A program that will certainly change has to be built to be changed, which is why maintainability and flexibility are quality characteristics in their own right in the quality model today (Chapter Twenty), and why every change needs its regression tests.

Invisibility. "Software is invisible and unvisualizable." A building has a floor plan, on which "Contradictions become obvious, omissions can be caught"; a program's structure, when anyone tries to draw it, turns out to be "not one, but several, general directed graphs, superimposed one upon another." Nobody can look at a program and see whether it is good, the way an inspector can look at a weld. Its quality has to be made visible by other means: documents that can be reviewed, models, the measures of Module 2, and tests that turn behaviour into evidence.

Brooks's propertyWhat it meansWhat it asks of quality work
ComplexityMore states than anyone can list; no two parts alikeTest design techniques, risk-based choice of tests, reviews
ConformityMust fit interfaces designed by othersClear interface requirements; integration and compatibility testing
ChangeabilityAll successful software gets changedMaintainable design, regression testing, control of changes
InvisibilityNo single drawing shows the whole of itReviews of documents and models, measures, tests as evidence

Three views of quality: internal, external and in use

The last difference is where quality is seen from. The standards distinguish three views, and each is judged by different people with different evidence.

munotes.in122

Quality in Software Development

Internal quality is the "totality of attributes of a product that determine its ability to satisfy stated and implied needs when used under specified conditions" (ISO/IEC/IEEE 24765). It is the quality of the code, design and documents themselves, judged without running the program: by reviews, by static analysis, and by internal measures, each a "measure of the product itself, either direct or indirect", such as the size of a module or its complexity.

External quality is the "extent to which a product satisfies stated and implied needs when used under specified conditions". It is judged by running the software, as testers do. An external measure is an "indirect measure of a product derived from measures of the behavior of the system of which it is a part", such as the number of failures in a test cycle or the response time under load.

Quality in use is, in ISO/IEC 25019:2023, the "extent to which the system or product, when it is used in a specified context of use, satisfies or exceeds" what its stakeholders need "to achieve specified beneficial goals or outcomes". The context of use is the "combination of users, goals and tasks, resources, and environment" (ISO TR 25060:2023). Quality in use is judged by the people who use the software, for their own purposes, in their own conditions. ISO's own page describes the 2023 standard as "a quality-in-use model composed of three characteristics". They are beneficialness, the "extent of benefit resulting from the use of a product, system, or service"; freedom from risk, the "extent to which a product or system mitigates the potential risk to economic status, human life, health, society, financial values, enterprise activities, or the environment"; and acceptability.

Internal quality influences external quality, which influences quality in use; each view depends on the one before it

Figure 23.1 Three views of software quality: internal, external and in use

The older product quality standard, ISO/IEC 9126, linked the three views in the chain the figure shows. Internal quality influences external quality: well-structured, well-reviewed code is more likely to behave well when it runs. External quality influences quality in use: software that behaves well under test is more likely to serve its users. Read the other way, each depends on the one before it: quality in use cannot be had without external quality, nor external quality without internal quality. The influence is real, but it is not a guarantee, and the gaps between the views are where some of the most instructive failures live. The worked example below is one of them.

Who judges software quality

WhoWhat they seeView of qualityTheir evidence
Developer and reviewerCode, design and documentsInternalReview findings, static analysis, internal measures
TesterThe running system, in a test environmentExternalTest results, failures found, external measures
User (a student filling in the form)The system in real useIn useWhether the task got done, and without harm
Customer (the college that paid for it)The system's cost and benefitIn use, and valueBenefit for the money, complaints, risks avoided
Operator (the IT cell)The system in productionExternal and in useAvailability, incidents, how hard it is to run
munotes.in123

Quality in Software Development

Each of these people can be satisfied while another is not. The developers can be proud of clean code that the students find confusing; the students can be happy with a form that the IT cell must restart every night. This is Garvin's point from Chapter Twenty-Two, on what quality means, seen from inside a software project: a quality plan has to ask every one of them.

Worked example: one test result, two very different qualities in use

Two builds of ExamReg's fee function each contain one defect. Build A has lost the on-time case, Chapter One's defect from the chapter on what software testing is: it charges a form submitted on or before the last date as if it were late. Build B refuses a form submitted on the fifteenth day late, which the rule still accepts. The portal records any on-time form as 0 days late, and the test team runs one test for each value from 0 to 20 days late; both builds fail exactly one test. By that external measure, they are equally good.

The exam cell also has last session's record of how many forms arrived on each day. The counts below are this book's illustration, not a real college's data. The program weighs each build's failures by how many students would actually have met them.

def rule(days_late):                     # the exam cell's rule; 0 means on or before the last date
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    if days_late == 0:
        return 0
    if days_late <= 7:
        return 100
    return 500

def build_a(days_late):                  # defect A: Chapter One's, the on-time case lost
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    if days_late <= 7:
        return 100
    return 500

def build_b(days_late):                  # defect B: the fifteenth day refused
    if days_late < 0 or days_late >= 15:
        raise ValueError("form not accepted")
    if days_late == 0:
        return 0
    if days_late <= 7:
        return 100
    return 500

def outcome(fee, days_late):
    try:
        return fee(days_late)
    except ValueError:
        return "refused"

days = list(range(0, 21))                # one test for each value of days late, 0 to 20
forms = [720,                            # last session's forms on or before the last date
         50, 30, 25, 20, 15, 12, 28,     # 1 to 7 days late
         18, 10, 8, 7, 6, 5, 9, 14,      # 8 to 15 days late
         8, 6, 4, 3, 2]                  # 16 to 20 days late (refused, as the rule says)
total = sum(forms)
print("tests:", len(days), " forms in the session:", total)
for name, build in [("build A", build_a), ("build B", build_b)]:
    wrong = [d for d in days if outcome(build, d) != outcome(rule, d)]
    hit = sum(n for d, n in zip(days, forms) if d in wrong)
    print(f"{name}: wrong on day {', '.join(map(str, wrong))};"
          f" tests failed {len(wrong)} of {len(days)} ({len(wrong) / len(days):.1%});"
          f" students hit {hit} of {total} ({hit / total:.1%})")
munotes.in124

Quality in Software Development

tests: 21  forms in the session: 1000
build A: wrong on day 0; tests failed 1 of 21 (4.8%); students hit 720 of 1000 (72.0%)
build B: wrong on day 15; tests failed 1 of 21 (4.8%); students hit 14 of 1000 (1.4%)

By the test results the two builds are identical: each fails 1 of its 21 tests, about 4.8 per cent. In use they are nothing alike. Build A's defect meets every student who submitted on time, which is most of them: it would have charged 720 of the 1,000 students a late fee they did not owe, 72 per cent of all users. Build B's defect would have met 14 students, 1.4 per cent.

But a count of students is not the whole of quality in use either. Build A's harm is Rs 100 taken wrongly, which can be refunded. Build B's harm is a refused form: under the rule, those 14 students cannot sit the examination unless somebody notices in time. Quality in use includes freedom from risk, and a defect that meets few users can still do the most damage. A tester who reports only 1 of 21 tests failed for each build has told the exam cell almost nothing it needs to know. The defect report has to say who is hit and how badly, which is what severity and priority in a defect report (Chapter Seventy-Six) exist to record.

The general lesson is about test design. A test count treats every input as equally important; users do not. The on-time case deserves more tests, and more careful ones, than the eleventh day late, because that is where the students are.

When is software good enough?

No release is free of defects; testing can show that defects are present but never that they are absent, the first of the seven principles in Chapter Four. Every release is therefore a decision that the software is good enough. The NIST study of 2002 called this the central difficulty: "The major problem for the software industry is deciding when a firm should stop testing". It reported that commercial developers decide with a combination of rules of thumb: a sufficient percentage of test cases passing, statistics from a code coverage tool, counts and trends of defects by severity, beta testing by real users, and the number of new problem reports falling below a threshold. NIST described these as nonanalytical: none of them calculates the risk that remains.

munotes.in125

Quality in Software Development

The three views make the decision more honest. Internal quality asks whether any serious review finding is still open. External quality asks whether the exit criteria are met: the planned tests run, the pass rate reached, no critical defect open. Quality in use asks whether the software has been tried in its real context of use, with its real users and its busiest day. The worked example shows why the third question cannot be skipped. A release rule of at least 95 per cent of tests passed would let build A through, with 20 of 21 tests passed, about 95.2 per cent, and it would still overcharge 720 students.

The cost of getting it wrong

Chapter Three, on why software must be tested, showed what poor quality has cost when software fails in the field, from a lost rocket to a national estimate of $59.5 billion a year. ExamReg's scale is smaller but the pattern is the same. Released, build A's defect would have taken 720 × 100 = 72,000 rupees from students who owed nothing, and the college would then pay again: in refunds, corrected receipts, apologies and an emergency release made under pressure. Found in a review of the code, the same defect would have cost one changed line and a retest. Chapter One Hundred One, on the cost of quality, turns this into the four kinds of quality cost and shows how to count them.

What it does not mean

Software not wearing out does not mean its quality is permanent. Its reliability falls when its use changes or when a change to the code brings in a new fault.

Internal quality is not a private concern of developers. External quality and quality in use are built on it, and it decides how costly every later change will be.

Passing tests is not the same as quality in use. Tests measure external quality over the inputs someone chose; users meet the inputs their context produces.

Good enough is not an excuse for poor quality. It is a decision, made with evidence, that the remaining risk is acceptable to the people who will carry it.

Quick revision

  • Software quality: "capability of a software product to satisfy stated and implied needs when used under specified conditions" (ISO/IEC 25000:2014); "degree to which a software product meets established requirements" (IEEE 730-2014).
  • Software faults are design faults, and every copy is identical, so quality must be built in during development, not inspected in at the end.
  • Software does not wear out (Lyu 1996), but its reliability falls with a change in use or a faulty change to the code, and grows as faults are found and removed.
  • Brooks (1986), the essential properties of software: complexity, conformity, changeability, invisibility.
  • Internal quality: the product itself, judged without running it. External quality: the software running, judged by tests. Quality in use: in a context of use (users, goals and tasks, resources, environment); ISO/IEC 25019:2023 has three characteristics: beneficialness, freedom from risk, acceptability.
  • Internal quality influences external quality, which influences quality in use (the chain of ISO/IEC 9126).
  • Worked example: two builds each failed 1 of 21 tests; in use one would have hit 720 of 1,000 students, the other 14 students with a worse harm.
  • Deciding when software is good enough is, in NIST's words of 2002, the major problem; decide with all three views.
munotes.in126

Quality in Software Development

Test yourself

1. What is software quality? Give two standard definitions. ISO/IEC 25000:2014 defines it as the capability of a software product to satisfy stated and implied needs when used under specified conditions. IEEE 730-2014 defines it more narrowly as the degree to which a software product meets established requirements.

2. Why is the quality of software different from the quality of a manufactured product? Software's faults are design faults, not physical ones, and every copy is identical, so quality cannot be controlled by inspecting copies; it must be built in during development. Software does not wear out, but its reliability falls when its use changes or when a change to the code introduces a fault.

3. Explain Brooks's four essential properties of software and their effect on testing. Complexity: software has more states than can be listed, so exhaustive testing is impossible and tests must be designed. Conformity: it must fit interfaces designed by others, which calls for integration and compatibility testing. Changeability: all successful software gets changed, which calls for maintainable design and regression testing. Invisibility: its structure cannot be seen, so quality must be made visible through reviews, measures and tests.

4. Distinguish internal quality, external quality and quality in use, with an example of each for an online registration portal. Internal quality is the quality of the product itself, judged without running it, for example the complexity of the fee code found in a review. External quality is the quality of the running software, judged by testing, for example the number of failures in a test cycle. Quality in use is the quality experienced by real users in their context of use, for example whether students register correctly and without harm on the last date.

munotes.in127

Quality in Software Development

5. Two builds each fail one of 21 tests. Why might one be far worse than the other? Because the test count weighs every input equally, while real users do not arrive equally. A defect in the commonest case, such as an on-time form, can hit hundreds of users while a defect on a rare day hits a few, and a defect that hits few can still do more harm, such as refusing a registration. Quality in use depends on the context of use and on the severity of the harm.

Contents This chapter on its own page

munotes.in128

Chapter Twenty-Four

Quality Control and Quality Assurance

Syllabus topic Module 1, "Definition of Quality and Quality Assurance: Distinction between Quality Assurance (QA), Quality Control (QC)"

In one line

Quality control checks the product to find what is wrong with it; quality assurance works on the process, so that there is less to find, and gives everyone justified confidence that the product will meet its requirements.

In the wording a student can write in an examination: quality assurance (QA) is a "process that is focused on providing confidence that quality requirements are fulfilled" (ISO/IEC/IEEE 12207:2026); quality control (QC) is the "tasks to evaluate the quality of services as delivered or the quality of developed or manufactured products" (the same standard). QA is process-oriented and preventive: it defines, audits and improves the way work is done. QC is product-oriented and corrective: it examines what has been built and finds its defects. Testing is, in the ISTQB syllabus's words, "a major form of quality control", and reviews of work products are another.

Two terms that are often confused

In everyday use the two terms blur. ASQ's glossary warns that they have "many interpretations" and that in practice they are often used interchangeably, for anything done to ensure quality. The ISTQB syllabus notes a second confusion, between quality assurance and testing: people often use the two terms as if they meant one thing, but "testing and QA are not the same." By the standards' definitions, testing is quality control.

The difference matters because the two do different jobs, need different skills, and fail in different ways. A project with excellent testing and no quality assurance finds its defects expensively, late, and over and over again. A project with fine process documents and no quality control has no evidence that the product works.

Quality control: checking the product

The standard definition is short: QC is the "tasks to evaluate the quality of services as delivered or the quality of developed or manufactured products" (ISO/IEC/IEEE 12207:2026). An older definition in the same vocabulary adds what follows the evaluation: "monitoring service performance or product quality, recording results, and recommending necessary changes" (ISO/IEC/IEEE 24765).

The object of quality control is always a product: code, a design, a requirements document, a test plan, a build, a released system. Its methods examine that product:

  • Testing, which runs the software and compares what it does with what it should do.
  • Reviews and inspections of work products, which read them for defects without running anything (Chapters Twenty-Eight and Twenty-Nine, on reviews and inspection).
  • Other checks that the syllabus lists beside testing: formal methods, simulation and prototyping.

The ISTQB syllabus describes testing, and with it this whole side of quality work, in one sentence: "Testing is a product-oriented, corrective approach that focuses on those activities supporting the achievement of appropriate levels of quality." It then places testing inside quality control: "Testing is a major form of quality control, while others include formal methods (model checking and proof of correctness), simulation and prototyping."

munotes.in129

Quality Control and Quality Assurance

A note for readers who meet older copies. The release notes of version 4.0.1 record that, in this section, the word QC was replaced with testing, "because the section compares QA with testing, not with QC". An older copy may therefore call QC product-oriented and corrective where version 4.0.1 says it of testing. The two versions agree on the substance: the product-oriented, corrective side of quality work is quality control, and testing is its largest part.

Quality assurance: confidence from the process

Quality assurance is defined by what it gives: confidence. ISO/IEC/IEEE 12207:2026 calls it a "process that is focused on providing confidence that quality requirements are fulfilled"; ISO/IEC/IEEE 15288:2023 places it as the "part of quality management focused on providing confidence that quality requirements are fulfilled". The word itself has a standard meaning: assurance is "grounds for justified confidence that a claim has been or will be achieved" (ISO/IEC/IEEE 15026-1:2025). QA's product, in other words, is evidence that the organisation's way of working can be trusted to produce what was promised.

The ISTQB syllabus gives its method: "QA is a process-oriented, preventive approach that focuses on the implementation and improvement of processes. It works on the basis that if a good process is followed correctly, then it will generate a good product." And its scope: "QA applies to both the development and testing processes, and is the responsibility of everyone on a project."

Its typical activities act on the way work is done:

  • Defining processes and standards: how requirements are written, how code is reviewed, what a test plan must contain, when a change may be released.
  • Training people to follow them.
  • Audits: an audit is an "independent examination of a work product or set of work products to assess compliance with specifications, standards, contractual agreements, or other criteria" (ISO/IEC/IEEE 12207:2026). A QA audit asks whether the process was followed, and produces findings about the process.
  • Measuring and improving the process, using the data that quality control produces, which is the subject of the worked example below.

Prevention and detection

The simplest way to hold the two apart is by what they do to a defect. Quality assurance tries to stop a defect from being made; quality control tries to find the defects that were made anyway.

ExamReg shows the difference on a single kind of defect. Suppose the fee form keeps accepting impossible input, such as a negative number of backlog papers. Quality control finds each instance: a tester enters -1, the form accepts it, a defect report is written, the developer fixes that field, and the test is added to the regression suite. Next release, a different form has the same kind of defect, and the cycle repeats. Quality assurance asks why the same kind keeps appearing, and changes the process: a shared validation component that every form must use, a rule in the coding standard, and an item on the code review checklist. The defect stops being made.

munotes.in130

Quality Control and Quality Assurance

Neither replaces the other. Prevention is never perfect, so detection is always needed; and detection alone is expensive, because each defect is found only after it has been built, and a process that keeps making the same mistake keeps paying for it.

The same results, used twice

The ISTQB syllabus names the point where the two meet: "Test results are used by QA and testing. In testing they are used to fix defects, while in QA they provide feedback on how well the development and test processes are performing." The program below reads the ExamReg release 2.0 defect data, fixed once for this whole book, first as quality control reads it and then as quality assurance does.

found_by = {"requirements review": 16, "design review": 24, "code review": 34,
            "unit testing": 44, "integration testing": 32, "system testing": 28,
            "acceptance testing": 10, "after release": 12}
by_type = {"input validation": 58, "logic and computation": 44, "interface": 30,
           "user interface": 24, "data and database": 18, "documentation": 12,
           "performance": 8, "security": 6}
total = sum(found_by.values())
assert total == sum(by_type.values()) == 200          # ExamReg release 2.0

# Quality control reads the data for the PRODUCT: what is wrong, and was it caught in time?
escaped = found_by["after release"]
print(f"QC: {total} defects in release 2.0; {total - escaped} found before release;"
      f" {escaped} found by students after it")

# Quality assurance reads the same data for the PROCESS: where do defects come from,
# and which activities catch them?
reviews = sum(n for activity, n in found_by.items() if activity.endswith("review"))
tests = sum(n for activity, n in found_by.items() if activity.endswith("testing"))
print(f"QA: caught by reviews {reviews} ({reviews / total:.0%}), by tests {tests}"
      f" ({tests / total:.0%}), by students {escaped} ({escaped / total:.0%})")
commonest = max(by_type, key=by_type.get)
print(f"QA: the commonest kind is {commonest}: {by_type[commonest]} of {total}"
      f" ({by_type[commonest] / total:.0%})")
QC: 200 defects in release 2.0; 188 found before release; 12 found by students after it
QA: caught by reviews 74 (37%), by tests 114 (57%), by students 12 (6%)
QA: the commonest kind is input validation: 58 of 200 (29%)

Quality control's reading is about this release. There were 200 defects; 188 were caught before release and 12 reached students. Every one must be fixed and its fix confirmed, and the 12 in production come first, because students are meeting them now.

munotes.in131

Quality Control and Quality Assurance

Quality assurance's reading is about the next release. Reviews caught 74 defects, 37 per cent, before any code ran, which says the review process is earning its time. But 12 escaped, 6 per cent of the total, and the commonest kind of defect, input validation, is 58 of 200, 29 per cent: nearly three in every ten defects are the same kind of mistake. That is a process signal, not a product one. The response is the one in the last section: a shared validation component, a coding rule and a review checklist item, followed by a check in release 2.1 that the count has fallen. Chapter Eighty, on using defect data to improve the process, takes this analysis much further.

A review is quality control; checking that reviews happen is quality assurance

One pair of activities causes more confusion than any other, because both involve reading documents.

  • A code review of the fee function reads the product and finds its defects. It is quality control, of a static kind.
  • An audit at the end of a sprint, which checks whether every change to the fee rules actually went through a code review as the process requires, reads records of the process. It is quality assurance.

The rule that settles every such case is to ask what the activity examines. If it examines the product, it is quality control; if it examines or changes the way the product is made, it is quality assurance.

QA and QC side by side

Quality assurance (QA)Quality control (QC)
Standard definition"process that is focused on providing confidence that quality requirements are fulfilled""tasks to evaluate the quality of services as delivered or the quality of developed or manufactured products"
FocusThe processThe product
AimPrevent defects; give justified confidenceFind defects so they can be corrected
ApproachPreventive, proactiveCorrective, reactive
WhenThroughout, starting before the product existsOnce a work product exists to examine
Typical activitiesProcess definition, standards, training, process audits, process improvementTesting, reviews and inspections of work products
WhoEveryone on the project, often led by an SQA groupTesters, reviewers and inspectors
OutputDefined processes, audit findings, improvementsDefect reports, test results, pass or fail decisions
Question it asksIs the work being done in a way that will produce quality?Does this product meet its requirements?

Worked contrast: the ExamReg project's quality activities

Activity on the ExamReg projectQA or QCWhy
Writing the coding standard, with its rules for validating inputQADefines how code is written, before the code exists
Training the developers on the payment gateway's interfaceQAPrevents defects by building skill
Deciding that every change to a fee rule must be code reviewedQASets the process
Code reviewing the fee function against the checklistQCExamines the product and finds its defects
Running the late-fee test cases on build 2.0.1QCExamines the product by running it
Checking the hall ticket PDF against its specificationQCExamines a product against its requirements
Auditing whether every fee change in the sprint was reviewedQAExamines whether the process was followed
Analysing release 2.0's defects by type and changing the standardQAImproves the process from the product's data
munotes.in132

Quality Control and Quality Assurance

What it does not mean

Quality assurance is not another name for testing. Testing is quality control; QA is about the process, and a tester's job title does not change that.

Quality control is not only testing. Reviews and inspections of documents and code examine the product too.

Quality assurance does not make quality control unnecessary. A good process lowers the number of defects made; it never lowers it to zero.

Quality assurance is not one department's job alone. In the syllabus's words, it "is the responsibility of everyone on a project", even when a separate SQA group leads it.

Quick revision

  • QA: "process that is focused on providing confidence that quality requirements are fulfilled" (ISO/IEC/IEEE 12207:2026); "part of quality management" (ISO/IEC/IEEE 15288:2023). Process-oriented, preventive.
  • QC: "tasks to evaluate the quality of services as delivered or the quality of developed or manufactured products" (ISO/IEC/IEEE 12207:2026). Product-oriented, corrective.
  • Assurance: grounds for justified confidence that a claim has been or will be achieved.
  • Testing is "a major form of quality control"; so are reviews and inspections of work products (ISTQB v4.0.1, section 1.2.2).
  • QA prevents defects; QC detects them. Both are needed.
  • The same test results serve both: QC uses them to fix defects, QA as feedback on the process.
  • A review of the product is QC; an audit of whether reviews happen is QA.
  • ExamReg release 2.0: 188 of 200 defects caught before release, 12 after; reviews caught 74 (37 per cent); input validation is 58 of 200 (29 per cent), a process signal.

Test yourself

1. Define quality assurance and quality control as the standards do. Quality assurance is a process focused on providing confidence that quality requirements are fulfilled (ISO/IEC/IEEE 12207:2026), and part of quality management (ISO/IEC/IEEE 15288:2023). Quality control is the tasks that evaluate the quality of delivered services or of developed or manufactured products (ISO/IEC/IEEE 12207:2026).

2. Distinguish QA from QC under five headings. Focus: QA the process, QC the product. Aim: QA prevents defects and gives confidence, QC finds defects to be corrected. Approach: QA preventive, QC corrective. Timing: QA throughout, from before the product exists; QC once a work product exists. Activities: QA defines, audits and improves processes; QC tests, reviews and inspects work products.

munotes.in133

Quality Control and Quality Assurance

3. Is testing quality assurance or quality control? Explain. Quality control. It examines the product by running it and finds defects so that they can be corrected; the ISTQB syllabus calls it product-oriented and corrective, and a major form of quality control, while QA is process-oriented and preventive.

4. How can the same test results be used for both QA and QC? QC uses them to fix defects in the product. QA uses them as feedback on the process: for example, finding that 58 of ExamReg's 200 defects were input validation defects shows a weakness in how code is written and reviewed, which QA corrects with a shared validation component, a coding rule and a checklist item.

5. Classify with reasons: a code review of a module, and an audit of whether code reviews were held. The code review is quality control, because it examines the product for defects. The audit is quality assurance, because it examines whether the process was followed.

Contents This chapter on its own page

munotes.in134

Chapter Twenty-Five

Quality Management and Software Quality Assurance

Syllabus topic Module 1, "Definition of Quality and Quality Assurance: Distinction between ... Quality Management (QM), and Software Quality Assurance (SQA)"

In one line

Quality management is everything an organisation does to direct and control quality, from its policy and objectives to its improvements; quality assurance and quality control are two parts of it; and software quality assurance is quality assurance applied to the processes that produce software.

In the wording a student can write in an examination: quality management (QM) is the "coordinated activities to direct and control an organization with regard to quality" (ISO/IEC/IEEE 12207:2026). It sets a quality policy and quality objectives, and achieves them through quality planning, quality control, quality assurance and quality improvement. Software quality assurance (SQA) is the "set of activities that define and assess the adequacy of software processes to provide evidence that establishes confidence that the software processes are appropriate for and produce software products of suitable quality for their intended purposes" (IEEE 730-2014, whose 2026 edition now replaces it). QA and QC are both parts of QM; SQA is QA for software processes, ideally carried out with independence from the people who develop the software.

Quality management: the whole

Chapter Twenty-Four separated quality control, which checks the product, from quality assurance, which gives confidence in the process. Neither says who decides what quality the organisation is aiming for, who pays for the checks, or who acts when the results are poor. That is quality management.

The standard definition is deliberately broad: the "coordinated activities to direct and control an organization with regard to quality". The organisation does not direct and control quality by accident; it builds a management system, which the standards define as a "set of interrelated or interacting elements to establish policy and objectives and to achieve those objectives" (ISO/IEC 19770-1:2017). For quality, those elements are the following.

  • A quality policy: the organisation's stated intentions on quality, set by its leadership.
  • Quality objectives: measurable targets that turn the policy into something that can be checked. ISO's page for ISO 9001:2026 says the planning in a quality management system "must include measures designed to achieve an organization's quality objectives and continuously improve the system's effectiveness."
  • Quality planning, quality control and quality improvement, the three processes of Joseph Juran's trilogy. The Juran Institute describes the idea as one in which "organizations must use three universal processes"; ASQ's glossary gives each in a phrase: "quality planning (developing the products and processes required to meet customer needs), quality control (meeting product and process goals) and quality improvement (achieving unprecedented levels of performance)."
  • Quality assurance, which ISO/IEC/IEEE 15288:2023 places explicitly inside the whole: it is the "part of quality management focused on providing confidence that quality requirements are fulfilled".
Quality management holds policy and objectives, planning, quality assurance with SQA inside it, quality control and improvement

Figure 25.1 Quality management is the whole; QA and QC are parts of it, and SQA is QA applied to software processes

munotes.in135

Quality Management and Software Quality Assurance

The standard an organisation's quality management system can be certified against is ISO 9001, whose requirements, in ISO's own words, "define how to establish, implement, maintain, and continually improve a quality management system (QMS)." Its current edition, ISO 9001:2026, and the family around it are the subject of Chapters Ninety-Four and Ninety-Five.

Software quality assurance

For software, IEEE 730 applies the same ideas. It defines software quality management in words that mirror the general definition: "coordinated activities to direct and control an organization with regard to software quality". Within it sits SQA, defined as the "set of activities that define and assess the adequacy of software processes to provide evidence that establishes confidence that the software processes are appropriate for and produce software products of suitable quality for their intended purposes."

Every phrase of that definition does work, and together they make a complete short answer.

  1. Define and assess. SQA both sets the processes (what the project will do, and to what standard) and checks them.
  2. The adequacy of software processes. Its object is the process, which places SQA on the quality assurance side of Chapter Twenty-Four's distinction, not the quality control side.
  3. Evidence that establishes confidence. Its output is evidence: audit findings, records, measurements, that let the customer and management trust the work.
  4. Appropriate for, and produce, software products of suitable quality. It asks two questions: is the process fit for this project, and does it actually yield good products? So SQA looks at products too, as evidence about the process.
  5. For their intended purposes. Quality is judged against use, which is fitness for use from Chapter Twenty-Two, on what quality means.

A note on currency. The definitions above are from IEEE 730-2014, the edition whose wording the public software engineering vocabulary (SEVOCAB) still carries. IEEE's own page lists that edition as inactive since 27 March 2025 and superseded by IEEE 730-2026, published on 21 August 2026, which establishes "Requirements for initiating, planning, controlling, and executing the software quality assurance (SQA) processes of a software development or maintenance project". The 2026 text is not freely available; a student quoting the definitions above should name them as IEEE 730-2014's.

Assure, not ensure

IEEE 730-2014 draws a distinction that explains what an SQA group can and cannot do. To assure is "to promise or state with certainty by one person to another person or group"; to ensure is "to make certain that things occur or events take place". An SQA group assures: it gives the customer and management evidence they can rely on. It does not, by itself, ensure quality; the developers who write the code and the testers who check it do that. An SQA group that finds a process being skipped cannot fix the product. What it can do is make the gap visible to the people who can.

munotes.in136

Quality Management and Software Quality Assurance

Independence: why SQA stands apart from development

Evidence is only worth something if it is objective. A reviewer who reports to the project manager, days before a deadline, is under pressure to find that everything is fine. IEEE 730-2014 therefore defines independence of SQA as the "situation in which SQA is free from technical, managerial, and financial influences, intentional or unintentional", and names three kinds:

Kind of independenceIEEE 730-2014's definitionFor the ExamReg project
TechnicalSQA "uses personnel who are not involved in the development of the system or its elements"The SQA auditor has written none of ExamReg's code
Managerial"the responsibility of the SQA effort is vested in an organization separate from the development and project management organizations"The auditor reports to the head of quality, not to ExamReg's project manager
Financial"control of the SQA budget is vested in an organization independent of the development organization"The project manager cannot cut the audits to save money

A small organisation may not manage all three, and the definitions show what is lost when one is missing: an auditor who wrote the code, reports to the project manager or depends on the project's budget is judging their own side. The ISTQB syllabus's reminder from the last chapter still holds as well: quality assurance "is the responsibility of everyone on a project", even where an independent group leads it.

The four terms compared

Quality management (QM)Quality assurance (QA)Quality control (QC)Software quality assurance (SQA)
Standard definition"coordinated activities to direct and control an organization with regard to quality""part of quality management focused on providing confidence that quality requirements are fulfilled""tasks to evaluate the quality of services as delivered or the quality of developed or manufactured products""set of activities that define and assess the adequacy of software processes to provide evidence that establishes confidence" (IEEE 730-2014)
ScopeThe whole organisation's qualityConfidence in the processesThe productsThe processes of software projects
FocusDirection and controlProcessProductSoftware process
NatureManagerialPreventiveCorrectivePreventive, with independent evidence
Question it asksWhat quality are we aiming for, and are we getting there?Can we trust the way we work?Is this product right?Can we trust the way this software is built?
Led byTop managementEveryone, often led by a QA functionTesters, reviewers, inspectorsAn SQA group, ideally independent of development
ExamplesQuality policy, objectives, management reviewProcess standards, process auditsTesting, reviews, inspectionsThe SQA plan, process audits, product assessments
How it relates to the othersContains QA and QCPart of QMPart of QMQA applied to software
munotes.in137

Quality Management and Software Quality Assurance

Worked example: does ExamReg's software house meet its quality objectives?

Quality management becomes real only when its objectives are measurable and someone checks them. Suppose the software house that builds ExamReg set three quality objectives for the year. The targets are this book's illustration; the measured values come from ExamReg release 2.0's defect data, used throughout the book.

release = {"defects": 200, "critical": 8, "found by reviews": 74, "found after release": 12}

objectives = [            # (objective, measured value in per cent, "at most" or "at least", target)
    ("defects found by reviews", 100 * release["found by reviews"] / release["defects"], "at least", 35),
    ("defects found after release", 100 * release["found after release"] / release["defects"], "at most", 5),
    ("critical defects among all defects", 100 * release["critical"] / release["defects"], "at most", 5),
]
for name, actual, sense, target in objectives:
    met = actual <= target if sense == "at most" else actual >= target
    goal = f"{sense} {target}%"
    print(f"{name:<36}{actual:>5.1f}%   target {goal:<14}{'met' if met else 'NOT MET'}")
defects found by reviews             37.0%   target at least 35%  met
defects found after release           6.0%   target at most 5%    NOT MET
critical defects among all defects    4.0%   target at most 5%    met

Two objectives are met and one is not: 12 of 200 defects, 6 per cent, reached students, against a target of at most 5 per cent. Each of the four terms has a different part to play in what happens next.

  • Quality control fixes the 12 defects in the product and confirms each fix.
  • Software quality assurance audits the test process for release 2.0 and reports, with evidence, where it let the 12 through: which were missed by a planned test, and which had no test planned at all.
  • Quality management takes the missed objective to its management review, decides the improvement (a change to the test process, more time for it, or both), funds it, and sets the objective again for release 2.1.
  • Quality assurance, in the broad sense, is the confidence the whole cycle produces: evidence that when an objective is missed, the organisation notices and acts.

Chapter Eighty, on using defect data to improve the process, shows how the analysis of the 12 would be done.

What it does not mean

Quality management is not a department. It is the whole organisation's direction and control of quality, and it starts with top management.

SQA is not testing. It defines and assesses the processes; testing is quality control of the product.

SQA assures; it does not ensure. It supplies objective evidence. The quality itself is made by those who build and check the software.

munotes.in138

Quality Management and Software Quality Assurance

An SQA group is not the only one responsible for quality. Its independence makes its evidence trustworthy; it does not take the responsibility away from anyone else.

Quick revision

  • QM: "coordinated activities to direct and control an organization with regard to quality" (ISO/IEC/IEEE 12207:2026): policy and objectives, then planning, control, assurance and improvement.
  • Juran's trilogy: quality planning, quality control, quality improvement.
  • QA: "part of quality management" giving confidence (ISO/IEC/IEEE 15288:2023). QC: evaluates the products.
  • SQA (IEEE 730-2014): activities that define and assess the adequacy of software processes to provide evidence that establishes confidence that they produce software of suitable quality. IEEE 730-2026, published 21 August 2026, now replaces the 2014 edition.
  • Assure (give evidence to others) is not ensure (make it happen).
  • SQA independence: technical, managerial, financial.
  • Nesting: QM contains QA and QC; SQA is QA for software.
  • Worked example: reviews 37 per cent (target at least 35, met); after release 6 per cent (target at most 5, not met); critical 4 per cent (target at most 5, met).

Test yourself

1. Define quality management and name its parts. Quality management is the coordinated activities to direct and control an organization with regard to quality. It sets a quality policy and quality objectives, and achieves them through quality planning, quality control, quality assurance and quality improvement.

2. Define software quality assurance and explain the definition. SQA is the set of activities that define and assess the adequacy of software processes to provide evidence that establishes confidence that the processes are appropriate for, and produce, software of suitable quality for its intended purposes (IEEE 730-2014). It sets and checks processes, its object is the process, its output is evidence, and it judges the process both by its fitness and by the products it yields.

3. Distinguish QM, QA, QC and SQA. QM is the whole: direction and control of quality, from policy to improvement, led by top management. QA is the part of QM that gives confidence that processes will meet quality requirements; it is preventive. QC is the part that evaluates products and finds their defects; it is corrective. SQA is QA applied to the processes of software projects, ideally by a group independent of development.

4. What are the three kinds of SQA independence, and why do they matter? Technical independence (SQA staff are not involved in developing the system), managerial independence (SQA reports to an organisation separate from development and project management) and financial independence (the SQA budget is controlled outside the development organisation). They matter because SQA's value is objective evidence, which pressure from the project would weaken.

5. What is the difference between assuring and ensuring quality? To assure is to state with certainty to others, backed by evidence; to ensure is to make certain that things happen. SQA assures quality by supplying evidence; the developers and testers who build and check the software ensure it.

Contents This chapter on its own page

munotes.in139

Chapter Twenty-Six

Verification and Validation, and Why Both Matter

Syllabus topic Module 1, "Verification and Validation (V&V): Definition of V&V and its significance in software development"

In one line

Verification checks that each product of development matches what was specified for it; validation checks that the finished product does what its users actually need; and a product can pass either check while failing the other.

In the wording a student can write in an examination: verification is "confirmation, through the provision of objective evidence, that specified requirements have been fulfilled", and validation is "confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled" (both ISO/IEC/IEEE 12207:2026). In the plain words of CMMI, verification ensures that "you built it right" and validation that "you built the right thing". Both matter because each catches what the other cannot: verification finds where the product departs from its specification, validation finds where the specification itself departs from what is needed. When the stakes are high, the work is given to independent verification and validation (IV&V), performed by an organisation independent of the developers.

Two questions about every product

Put the standard definitions side by side and they differ in only one phrase. Both are "confirmation, through the provision of objective evidence, that" something "has been fulfilled". Verification confirms specified requirements: what the specification, the design or the previous phase's document said. Validation confirms the requirements for a specific intended use: what the product is actually for.

CMMI for Development states the same pair in its glossary. Verification is "Confirmation that work products properly reflect the requirements specified for them"; validation is "Confirmation that the product or service, as provided (or as it will be provided), will fulfill its intended use." And it gives them their memorable form: "In other words, verification ensures that 'you built it right'; whereas, validation ensures that 'you built the right thing.'" Put as two questions a student can carry into the examination hall, verification asks are we building the product right? and validation asks are we building the right product?

IEEE 1012-2024, the standard for verification and validation, defines both as processes and makes each richer:

  • Verification is the "process of evaluating a system or component to determine whether the products of a given development phase satisfy the conditions imposed at the start of that phase". It works phase by phase: the design is verified against the requirements, the code against the design.
  • Validation is the "process of providing evidence that the system, software, or hardware and its associated products satisfy requirements allocated to it at the end of each life cycle activity, solve the right problem (e.g., correctly model physical laws, implement business rules, and use the proper system assumptions), and satisfy intended use and user needs". The middle phrase is the heart of it: validation asks whether the product solves the right problem.
munotes.in140

Verification and Validation, and Why Both Matter

The general vocabulary combines the two into one definition of V&V: the "process of determining whether the requirements for a system or component are complete and correct, the products of each development phase fulfill the requirements or conditions imposed by the previous phase, and the final system or component complies with specified requirements" (ISO/IEC/IEEE 24765).

Verification and validation side by side

VerificationValidation
Standard definitionConfirmation by objective evidence that specified requirements have been fulfilledConfirmation by objective evidence that the requirements for a specific intended use have been fulfilled
In CMMI's plain words"you built it right""you built the right thing"
The questionAre we building the product right?Are we building the right product?
Checks againstThe specification: the output of the previous phaseThe users' needs and the product's intended use
WhenAt the end of each phase, on each work productEarly and throughout, and finally on the product in its real setting
Typical methodsReviews, inspections, walkthroughs, static analysis, unit, integration and system testingPrototypes tried by users, requirements reviews with users, acceptance, alpha and beta testing
Mostly done byDevelopers, testers and reviewersUsers and customers, with testers
What it findsPlaces where a product departs from its specificationWrong or missing requirements; a product that does not fit its use

Two cautions about the table. First, validation is not only an end-of-project activity. CMMI says validation "is performed early (concept/exploration phases) and incrementally throughout the product lifecycle", and the cheapest validation of all is showing users the requirements, or a prototype, before anything is built. Second, the two use similar methods. In CMMI's words, "Validation activities use approaches similar to verification (e.g., test, analysis, inspection, demonstration, simulation). Often, the end users and other relevant stakeholders are involved in the validation activities." What makes an activity verification or validation is what it checks against, not the technique.

The V-model (Chapter Fifteen) draws the same line: its lower test levels verify each phase's product against the phase opposite, and acceptance testing at the top validates the system against the users' needs.

Worked example: built right, but not the right thing

ExamReg gives some students a fee concession. As the Scrum sprint in Chapter Seventeen set it out, a concession waives the form fee only: a concession student who submits late still pays the late fee. Suppose, for this example, that the analyst who wrote the specification misread the exam cell's note and wrote that concession students "pay no fees", late fees included. That is the book's invented mistake, but it is an ordinary kind of requirements defect: a rule written down slightly wrong.

munotes.in141

Verification and Validation, and Why Both Matter

The program checks three builds on 42 cases: every value of days late from 0 (the portal's value for any form on or before the last date) to 20, with and without a concession. Verification compares each build with the specification it was coded from; validation compares it with what the exam cell actually needs.

def late_fee(days_late):                 # the late-fee rule; 0 means on or before the last date
    if days_late < 0 or days_late > 15:
        return "refused"
    if days_late == 0:
        return 0
    return 100 if days_late <= 7 else 500

def need(days_late, concession):         # what the exam cell actually applies: a concession
    return late_fee(days_late)           # waives the form fee only; late fees apply to all

def spec_v1(days_late, concession):      # the specification as the analyst wrote it:
    fee = late_fee(days_late)            # concession students pay no fees, late ones included
    return 0 if concession and fee != "refused" else fee

spec_v2 = need                           # the specification after validation corrected it

def build_1(days_late, concession):      # coded faithfully from spec v1
    if days_late < 0 or days_late > 15:
        return "refused"
    if concession or days_late == 0:
        return 0
    return 100 if days_late <= 7 else 500

def build_2(days_late, concession):      # coded from spec v1, with Chapter One's on-time defect
    if days_late < 0 or days_late > 15:
        return "refused"
    if concession:
        return 0
    return 100 if days_late <= 7 else 500

def build_3(days_late, concession):      # coded from spec v2
    if days_late < 0 or days_late > 15:
        return "refused"
    if days_late == 0:
        return 0
    return 100 if days_late <= 7 else 500

cases = [(d, c) for c in (False, True) for d in range(0, 21)]
print(len(cases), "cases: days late from 0 to 20, with and without a concession")
for name, build, label, spec in [("build 1", build_1, "spec v1", spec_v1),
                                 ("build 2", build_2, "spec v1", spec_v1),
                                 ("build 3", build_3, "spec v2", spec_v2)]:
    verified = sum(build(d, c) == spec(d, c) for d, c in cases)
    validated = sum(build(d, c) == need(d, c) for d, c in cases)
    print(f"{name}, coded from {label}: verification {verified} of {len(cases)},"
          f" validation {validated} of {len(cases)}")
42 cases: days late from 0 to 20, with and without a concession
build 1, coded from spec v1: verification 42 of 42, validation 27 of 42
build 2, coded from spec v1: verification 41 of 42, validation 26 of 42
build 3, coded from spec v2: verification 42 of 42, validation 42 of 42

Build 1 is the case to remember. It passes verification perfectly, 42 of 42: every test written from the specification agrees with it, because the developer did exactly what the specification said. It fails validation on 15 cases, every concession student from 1 to 15 days late, whom it lets off a late fee the exam cell requires. No amount of verification could have found this, because verification's yardstick, the specification, was itself wrong. This is Chapter Four's seventh principle, the absence-of-defects fallacy, on a small scale: a product that conforms perfectly can still be the wrong product.

munotes.in142

Verification and Validation, and Why Both Matter

Build 2 shows that verification still earns its keep. It departs from its own specification in one case, an on-time form charged Rs 100 (the defect of Chapter One, on what software testing is), and verification finds that one case, cheaply and early, in a unit test. Validation fails 16 cases: the same one, and the 15 inherited from the specification.

Build 3 is what both together produce. Validation, ideally a review of the specification with the exam cell before any code was written, corrected the specification; verification then confirmed that the code matches the corrected specification. It passes 42 of 42 on both.

Why both matter

The significance of V&V in software development follows from the example.

  1. Verification alone builds the wrong thing well. Every error in the requirements passes straight through it, as build 1 did.
  2. Validation alone finds defects late and vaguely. Users see that an outcome is wrong, not which phase went wrong, and much of what matters (the security of a password store, the accuracy of an audit log) is invisible in ordinary use.
  3. Both produce objective evidence. The definitions insist on "objective evidence", which is what a customer, an auditor or a regulator can rely on, where an assurance of it works on my machine is not.
  4. Checking each phase before the next builds on it keeps defects cheap. Verification at every phase stops a design defect from being coded; early validation stops a requirements defect from being designed.
  5. The effort can match the risk. IEEE 1012-2024 specifies V&V requirements "for different integrity levels": the more harm a failure could do, the more V&V the software receives.

Independent verification and validation (IV&V)

For software whose failure would be serious, V&V is often handed to people who did not build it. IEEE 1012-2024 defines independent verification and validation (IV&V) as "verification and validation performed by an organization that is technically, managerially, and financially independent of the development organization." The three kinds of independence are the same three that Chapter Twenty-Five gave for software quality assurance: IV&V staff did not build the system; they report outside the development and project management chain; and their budget cannot be cut by the project they are checking.

Independence brings two things that an internal team struggles to supply. The first is freedom from pressure: an IV&V team does not have to meet the development schedule it is judging. The second is a different set of assumptions. The analyst who misread the concession note in the worked example would read the specification the same way on every review; a reader from outside, who has never seen the note before, has a better chance of asking why concession students should pay no late fee at all.

munotes.in143

Verification and Validation, and Why Both Matter

IV&V costs money, so it is kept for software whose failure would cost far more than the checking, such as software that controls a medical device or an aircraft: in IEEE 1012's terms, software at a high integrity level. A college portal would not usually have a full IV&V effort, but ExamReg's college could reasonably ask an outside auditor to validate its fee calculations before the first session goes live, which is IV&V in miniature.

What it does not mean

Verification is not only testing. Reviews, inspections, walkthroughs and static analysis verify too, without running anything; the next chapter, Chapter Twenty-Seven, sorts the static and dynamic kinds of V&V.

Validation is not only acceptance testing at the end. It starts with the requirements and continues throughout.

Passing verification does not prove the product is right. It proves the product matches its specification, which may itself be wrong.

Independent does not mean detached. An IV&V team works throughout the project and reports its findings promptly to the developers as well as to management.

Quick revision

  • Verification: "confirmation, through the provision of objective evidence, that specified requirements have been fulfilled" (ISO/IEC/IEEE 12207:2026). CMMI: "you built it right".
  • Validation: "confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled". CMMI: "you built the right thing".
  • IEEE 1012-2024: verification checks each phase's products against the conditions set at its start; validation shows the product solves the right problem and meets intended use and user needs.
  • Verification checks against the specification; validation against the users' needs. The difference is the yardstick, not the technique.
  • Validation starts early (requirements, prototypes) and continues to acceptance testing.
  • Worked example: build 1 verified 42 of 42 but validated 27 of 42, because the specification was wrong.
  • IV&V: V&V by an organisation technically, managerially and financially independent of development (IEEE 1012-2024); used at high integrity levels.

Test yourself

1. Define verification and validation, and state the difference in one sentence. Verification is confirmation, through objective evidence, that specified requirements have been fulfilled; validation is confirmation, through objective evidence, that the requirements for a specific intended use have been fulfilled. Verification asks whether we built the product right, validation whether we built the right product.

2. Can software pass verification and fail validation? Give an example. Yes. If the specification is wrong, code that matches it perfectly passes verification and fails validation: an ExamReg build coded from a specification that wrongly excused concession students from late fees agreed with its specification in every case but charged the wrong fee in 15 cases the exam cell cared about.

munotes.in144

Verification and Validation, and Why Both Matter

3. Give three methods of verification and three of validation. Verification: reviews and inspections of work products, static analysis, and unit, integration and system testing against the specification. Validation: requirements reviews and prototypes with users, acceptance testing, and alpha and beta testing in the users' own setting.

4. Explain the significance of V&V in software development. Verification stops defects passing from one phase to the next while they are cheap to fix; validation makes sure the requirements, and so the product, fit the real need. Both produce objective evidence for the customer and regulators, and the effort can be scaled to the software's integrity level.

5. What is IV&V and when is it used? Independent verification and validation: V&V performed by an organisation that is technically, managerially and financially independent of the developers (IEEE 1012-2024). It is used where failure would be serious, for example in software that controls medical devices or aircraft, because independence removes schedule pressure and brings fresh assumptions.

Contents This chapter on its own page

munotes.in145

Chapter Twenty-Seven

The Kinds of V&V: Static and Dynamic Mechanisms

Syllabus topic Module 1, "Verification and Validation (V&V): Different types of V&V mechanisms"

In one line

V&V mechanisms come in two kinds: static ones examine a work product without running it, and dynamic ones run it; a project needs both, because each finds defects the other cannot.

In the wording a student can write in an examination: static V&V mechanisms evaluate a work product "without the test item being executed" (ISO/IEC/IEEE 29119-2): reviews (informal reviews, walkthroughs, technical reviews and inspections), audits, desk checking, static analysis by tools, and formal proof of correctness. Dynamic V&V mechanisms evaluate a test item "by executing it": testing at every level, dynamic analysis, simulation and prototyping. Static mechanisms can be used from the first requirement onwards and find defects directly; dynamic mechanisms need something that runs, and find failures from which defects are then traced.

The dividing line: is the product run?

IEEE 1012-2024 lists what V&V processes do: they "include the analysis, evaluation, review, inspection, assessment, and testing of products." Those mechanisms fall on either side of one line, drawn by two definitions in the software testing standard, ISO/IEC/IEEE 29119-2:

  • Static testing is "testing in which a test item is examined against a set of quality or other criteria without the test item being executed".
  • Dynamic testing is "testing in which a test item is evaluated by executing it".

The ISTQB syllabus draws the same line, and adds how static work is done: "In contrast to dynamic testing, in static testing the software under test does not need to be executed." Work products "are evaluated through manual examination (e.g., reviews) or with the help of a tool (e.g., static analysis)." It also answers a question students often get wrong: "Static testing can be applied for both verification and validation." Static and dynamic describe how a mechanism works; verification and validation, from the last chapter, describe what it checks against. The two classifications cross.

The static mechanisms

MechanismWhat it isExaminesTypical defects found
ReviewsPeople read a work product and comment on it; the types run from informal review to inspectionAnything that can be read: requirements, designs, code, test plansAmbiguities, omissions, contradictions, design weaknesses, logic errors
AuditAn "independent examination of a work product or set of work products to assess compliance with specifications, standards, contractual agreements, or other criteria" (ISO/IEC/IEEE 12207:2026)Products and records, against standards and agreementsNon-compliance: a skipped review, a missing sign-off, a departure from the standard
Desk checkingA "manual simulation of program execution to detect faults through step-by-step examination of the source program" (ISO/IEC 2382), usually by the code's own authorSource code, listingsLogic and arithmetic slips the author can trace by hand
Static analysisThe "process of evaluating a system or component based on its form, structure, content, or documentation" (ISO/IEC/IEEE 24765), in practice by a toolAnything with a formal structure: code, modelsUnreachable code, undefined variables, standards violations, some security weaknesses, complexity
Formal proofProof of correctness is a "formal technique used to prove mathematically that a computer program satisfies its specified requirements" (ISO/IEC/IEEE 24765); formal verification includes model checkingA program or design against a formal specificationAny departure from the specification, for the properties proved
munotes.in146

The Kinds of V&V: Static and Dynamic Mechanisms

Reviews have the next three chapters to themselves: the review process and its types (Chapter Twenty-Eight), inspection (Chapter Twenty-Nine) and the walkthrough (Chapter Thirty).

Static analysis deserves a word more here, because it is the static mechanism a machine does. The ISTQB syllabus notes that it "can identify problems prior to dynamic testing while often requiring less effort, since no test cases are required", and that it "is often incorporated into CI frameworks", running on every change. It has one requirement: "for static analysis, work products need a structure against which they can be checked (e.g., models, code or text with a formal syntax)." A tool can analyse code; it cannot yet read a requirements document the way a reviewer does.

The dynamic mechanisms

MechanismWhat it isTypical use
TestingEvaluating a test item "by executing it", at the unit, integration, system and acceptance levelsChecking behaviour against expected results
Dynamic analysis"evaluating a system or component based on its behavior during execution" (ISO/IEC/IEEE 24765)Measuring coverage, memory use, timing while the program runs
SimulationA "model that behaves or operates like a given system when provided a set of controlled inputs" (ISO/IEC/IEEE 24765)Trying a design or an environment that is too costly or risky to use for real, such as 500 students on the last date
PrototypingBuilding "a preliminary version of part or all of the hardware or software" to "permit user feedback, determine feasibility, or investigate timing or other issues" (ISO/IEC/IEEE 24765)Validating requirements and risky design choices early

The ISTQB syllabus names the same set when it places testing among the forms of quality control: "Testing is a major form of quality control, while others include formal methods (model checking and proof of correctness), simulation and prototyping."

Worked example: a static analyser finds two defects without running anything

A static analysis tool is a program that reads other programs. The one below is small, about forty lines, but it does what real analysers do. It reads the source of three versions of ExamReg's fee function, never calls any of them, and follows each function's if-statements in order, keeping track of which values of days_late can still reach each line. When a branch can never be reached, or some values run off the end of the function with no return, it reports it.

munotes.in147

The Kinds of V&V: Static and Dynamic Mechanisms

import ast

SOURCE = '''
def fee_a(days_late):
    if days_late < 0:
        raise ValueError("days late cannot be negative")
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    if days_late <= 7:
        return 100
    if days_late <= 0:
        return 0
    return 500

def fee_b(days_late):
    if days_late < 0:
        raise ValueError("days late cannot be negative")
    if days_late <= 0:
        return 0
    if days_late <= 7:
        return 100
    if days_late < 15:
        return 500

def fee_c(days_late):
    if days_late < 0:
        raise ValueError("days late cannot be negative")
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    if days_late <= 0:
        return 0
    if days_late <= 7:
        return 100
    return 500
'''

INF = float("inf")
TRUE_FOR = {ast.Lt: lambda k: (-INF, k - 1), ast.LtE: lambda k: (-INF, k),
            ast.Gt: lambda k: (k + 1, INF), ast.GtE: lambda k: (k, INF)}
FALSE_FOR = {ast.Lt: lambda k: (k, INF), ast.LtE: lambda k: (k + 1, INF),
             ast.Gt: lambda k: (-INF, k), ast.GtE: lambda k: (-INF, k - 1)}

def meet(a, b):                          # the whole numbers in both ranges
    return (max(a[0], b[0]), min(a[1], b[1]))

def empty(r):
    return r[0] > r[1]

def shown(r):
    return f"{r[0]} upward" if r[1] == INF else f"{r[0]} to {r[1]}"

def analyse(fn):
    """Walk fn's statements in order, tracking which values of its argument can still reach
    each one. Nothing is run: the program reads the source's structure."""
    arg = fn.args.args[0].arg
    live = (-INF, INF)
    findings = []
    for stmt in fn.body:
        if isinstance(stmt, ast.If) and isinstance(stmt.test, ast.Compare):
            test = stmt.test
            op, k = type(test.ops[0]), test.comparators[0]
            if isinstance(test.left, ast.Name) and test.left.id == arg and isinstance(k, ast.Constant):
                if empty(meet(live, TRUE_FOR[op](k.value))):
                    findings.append(f"line {stmt.lineno}: 'if {ast.unparse(test)}' can never be true here")
                if isinstance(stmt.body[-1], (ast.Return, ast.Raise)):
                    live = meet(live, FALSE_FOR[op](k.value))
                continue
        if isinstance(stmt, (ast.Return, ast.Raise)):
            live = (1, 0)                # nothing gets past a return or a raise
    if not empty(live):
        findings.append(f"{arg} from {shown(live)} reaches the end and returns None")
    return findings or ["no finding"]

for fn in ast.parse(SOURCE).body:
    for finding in analyse(fn):
        print(f"{fn.name}: {finding}")
fee_a: line 9: 'if days_late <= 0' can never be true here
fee_b: days_late from 15 upward reaches the end and returns None
fee_c: no finding

Each finding points straight at a real defect.

  • fee_a checks days_late <= 7 before days_late <= 0. Every on-time value is at most 7, so the on-time branch can never run, and an on-time form is charged Rs 100. It is Chapter One's defect in another shape, from the chapter on what software testing is, and here no test case found it: the analyser worked it out from the structure alone.
  • fee_b has no final return and no check for forms more than 15 days late. For 15 days and above it falls off the end and returns None: a form 15 days late gets no fee at all instead of Rs 500, and a form 30 days late is never refused.
  • fee_c, the corrected rule, gets no finding.
munotes.in148

The Kinds of V&V: Static and Dynamic Mechanisms

The limits of the tool are just as instructive. It knows nothing about the fee rule: if fee_c said return 50 instead of return 100, the analyser would report nothing, because the structure would be perfect and only the value wrong. Finding that needs a test with an expected result, which is dynamic testing, or a reviewer who knows the rule, which is a different static mechanism. This is the ISTQB syllabus's point exactly: static and dynamic testing "complement each other", and "there are some defect types that can only be found by either static or dynamic testing."

Static against dynamic

The ISTQB syllabus lists the differences, and they make a complete answer to a common examination question.

Static mechanismsDynamic mechanisms
Is the product run?NoYes
What they findDefects, directlyFailures, from which defects are found "through subsequent analysis"
What they can examineAny work product, including documentsOnly executable products
When they can startWith the first requirementOnce something runs
Hard-to-reach codeExamined as easily as any otherReached only if a test drives execution there
Quality characteristicsThose that do not need execution, such as maintainabilityThose that do, such as performance efficiency
Typical tools and methodsReviews, audits, desk checking, static analysis tools, formal proofTesting, dynamic analysis, simulation, prototyping

Some defects are found far more easily by one side. The syllabus lists, among those easier or cheaper to find statically, defects in requirements such as "inconsistencies, ambiguities, contradictions, omissions, inaccuracies, duplications", design defects, coding defects such as "variables with undefined values, undeclared variables, unreachable or duplicated code, excessive code complexity", deviations from standards, and "Incorrect interface specifications (e.g., mismatched number, type or order of parameters)". Dynamic mechanisms own everything that only shows when the software runs in its environment: timing, load, memory, and the behaviour of the whole system with its real data.

Where each mechanism sits in the life cycle

PhaseStatic V&VDynamic V&V
RequirementsReviews and inspections of the requirements; audits against standardsPrototypes tried by users; simulation of the process being automated
DesignDesign reviews and inspections; static analysis of models; formal proof of critical partsSimulation of the design under load; prototypes of risky parts
ImplementationDesk checking, code reviews, static analysis on every buildUnit testing; dynamic analysis such as coverage measurement
Integration and systemReviews of test plans and test cases; audits of the test processIntegration testing and system testing, including performance testing
Acceptance and releaseAudits of the release against the processAcceptance testing, alpha and beta testing
munotes.in149

The Kinds of V&V: Static and Dynamic Mechanisms

The pattern is the one the ISTQB syllabus gives as the value of static testing: it "can detect defects in the earliest phases of the SDLC, fulfilling the principle of early testing". Static mechanisms dominate the left of the table, where nothing runs yet; dynamic mechanisms take over as the product becomes executable; and both continue to the end.

What it does not mean

Static does not mean verification, nor dynamic validation. A review of the requirements with the users is static validation; a unit test against the design is dynamic verification.

Static analysis does not replace testing. It finds structural problems without knowing what the right answer is; only a test or a knowledgeable reviewer can say that a fee is wrong.

A static finding is not always a defect. Tools report suspicious patterns, and some are harmless; a person judges each one.

Formal proof is not proof that the software is right for its users. It proves the program meets its formal specification, which can itself be wrong, as the last chapter showed.

Quick revision

  • Static testing: a test item "examined against a set of quality or other criteria without the test item being executed"; dynamic testing: evaluated "by executing it" (ISO/IEC/IEEE 29119-2).
  • Static mechanisms: reviews (informal, walkthrough, technical, inspection), audits, desk checking, static analysis, formal proof.
  • Dynamic mechanisms: testing, dynamic analysis, simulation, prototyping.
  • Static finds defects directly, works on any work product, starts earliest; dynamic finds failures, needs something executable, and measures run-time qualities such as performance.
  • Static analysis needs a formal structure (code, models); it "can identify problems prior to dynamic testing while often requiring less effort".
  • Static and dynamic both serve verification and validation.
  • Worked example: an ast-based analyser found an unreachable on-time branch and a fall-off-the-end without running the code, and could not have found a wrong fee value.

Test yourself

1. Distinguish static from dynamic V&V mechanisms, with examples of each. Static mechanisms examine a work product without executing it: reviews, inspections, walkthroughs, audits, desk checking, static analysis and formal proof. Dynamic mechanisms evaluate it by executing it: testing at every level, dynamic analysis, simulation and prototyping.

2. What is static analysis, and what can it find? The process of evaluating a system or component based on its form, structure, content or documentation, usually by a tool and without running it. It finds structural defects such as unreachable code, undefined variables, standards violations, some security weaknesses and excessive complexity, and it needs a work product with a formal structure, such as code.

munotes.in150

The Kinds of V&V: Static and Dynamic Mechanisms

3. Give three differences between static and dynamic testing. Static testing finds defects directly, while dynamic testing causes failures that must then be analysed; static testing can examine non-executable work products such as requirements, while dynamic testing needs executable software; and static testing reaches rarely executed code as easily as any other, while dynamic testing reaches only what its tests drive.

4. Which V&V mechanisms can be used in the requirements phase? Static ones such as reviews and inspections of the requirements and audits against standards, and dynamic ones that do not need the final product, such as prototypes shown to users and simulations.

5. Why is static analysis not enough on its own? Because it checks structure, not meaning: it cannot know what the right output is. A function that returns the wrong fee through perfectly structured code passes static analysis and is caught only by a test with an expected result or a reviewer who knows the rule.

Contents This chapter on its own page

munotes.in151

Chapter Twenty-Eight

Software Reviews: The Process, the Roles and the Types

Syllabus topic Module 1, "Verification and Validation (V&V): Concepts of Software Reviews"

In one line

A software review is a planned reading of a work product by people who can judge it, to find its anomalies and agree its quality; it follows a defined process, gives each participant a role, and ranges in formality from an informal review to an inspection.

In the wording a student can write in an examination: a review is a "process or meeting during which a work product or a set of work products, is presented to project personnel, managers, users, customers, or other stakeholders for comment or approval" (ISO/IEC 29110-1-2:2024). The generic review process of ISO/IEC 20246, as the ISTQB syllabus gives it, has five activities: planning, review initiation, individual review, communication and analysis, and fixing and reporting. Its principal roles are manager, author, moderator (facilitator), scribe (recorder), reviewer and review leader. Its commonly used types, from least to most formal, are the informal review, walkthrough, technical review and inspection.

Why review at all

A review is the cheapest defect-finding activity in software development, for a reason Chapter Twenty-Seven gave when it sorted the static and dynamic mechanisms: it needs nothing that runs, so it can start with the first page of requirements. The ISTQB syllabus puts the economics plainly: "Even though reviews can be costly to implement, the overall project costs are usually much lower than when no reviews are performed because less time and effort needs to be spent on fixing defects later in the project."

A review also does what no test can. It can examine a requirement, a design or a test plan; it can ask whether a document is clear, complete and consistent; and it brings people together. In the syllabus's words, since reviews can happen early, "a shared understanding can be created among the involved stakeholders." In ExamReg's release 2.0, the requirements, design and code reviews together found 74 of the 200 defects, before a single test had run.

The review process

The ISTQB syllabus takes its process from ISO/IEC 20246, which "defines a generic review process that provides a structured but flexible framework from which a specific review process may be tailored to a particular situation. If the required review is more formal, then more of the tasks described for the different activities will be needed."

The five activities of the review process, from planning to fixing and reporting, with a follow-up review when needed

Figure 28.1 The generic review process (ISTQB v4.0.1, section 3.2.2, after ISO/IEC 20246)

  1. Planning. The scope of the review is defined: "the purpose, the work product to be reviewed, quality characteristics to be evaluated, areas to focus on, exit criteria, supporting information such as standards, effort and the timeframes for the review".
  2. Review initiation. Everyone and everything is made ready: each participant "has access to the work product under review, understands their role and responsibilities and receives everything needed to perform the review."
  3. Individual review. Each reviewer works alone, applying "one or more review techniques (e.g., checklist-based reviewing, scenario-based reviewing)", and logs every anomaly, recommendation and question they find.
  4. Communication and analysis. The logged items are discussed, usually in a review meeting, because "the anomalies identified during a review are not necessarily defects". For each one, "the decision should be made on its status, ownership and required actions", and the participants decide the quality level of the work product and the follow-up needed.
  5. Fixing and reporting. "For every defect, a defect report should be created so that corrective actions can be followed up. Once the exit criteria are reached, the work product can be accepted."
munotes.in152

Software Reviews: The Process, the Roles and the Types

One word in the process needs care. An anomaly is "anything observed in the documentation or operation of a system that deviates from expectations based on previously verified system, software, or hardware products or reference documents" (IEEE 1012-2024). A reviewer logs anomalies; only the analysis decides which of them are defects. A question that turns out to have a good answer is an anomaly, but not a defect.

Review techniques: how a reviewer reads

ISO/IEC 20246 names several techniques for the individual review. A reviewer can use more than one.

TechniqueISO/IEC 20246's definitionOn ExamReg
Ad hocAn "unstructured independent review technique"Read the fee requirement and note whatever seems wrong
Checklist-basedA "review technique guided by a list of questions or required attributes"A checklist: is every boundary stated? is every term defined?
Scenario-basedThe review "is guided by determining the ability of the work product to address specific scenarios"Walk a student who submits exactly on the last date through the requirement
Role-basedReviewers "review a work product from the perspective of different stakeholder roles"Read it as the exam cell clerk, as a concession student, as the accounts office
Perspective-basedA "form of role-based reviewing that uses checklists and involves the creation of prototype deliverables"The tester drafts test cases from the requirement while reading it

The roles

The ISTQB syllabus names six principal roles, and one person may hold several.

RoleResponsibility (ISTQB v4.0.1)In the review of ExamReg's fee requirement
Manager"decides what is to be reviewed and provides resources, such as staff and time for the review"The project manager, who books two hours for four people
Author"creates and fixes the work product under review"The analyst who wrote the fee requirement
Moderator (facilitator)"ensures the effective running of review meetings, including mediation, time management, and a safe review environment in which everyone can speak freely"A senior developer from another team
Scribe (recorder)"collates anomalies from reviewers and records review information, such as decisions and new anomalies found during the review meeting"A junior tester
Reviewer"performs reviews", and "may be someone working on the project, a subject matter expert, or any other stakeholder"A tester, a developer and a clerk from the exam cell
Review leader"takes overall responsibility for the review such as deciding who will be involved, and organizing when and where the review will take place"The test lead
munotes.in153

Software Reviews: The Process, the Roles and the Types

The four review types

The syllabus observes that review types range "from informal reviews to formal reviews", and ISO/IEC 20246 defines the two ends: an informal review is a "form of review that does not follow a defined process and has no formal documented output", and a formal review is a "form of review that follows a defined process with formal documented output". Four types are in common use.

  • Informal review. "Informal reviews do not follow a defined process and do not require a formal documented output. The main objective is detecting anomalies." A colleague reading a draft and scribbling in the margin is an informal review, and so is much of pair programming.
  • Walkthrough. "A walkthrough, which is led by the author, can serve many objectives", from "evaluating quality and building confidence in the work product" to "educating reviewers, gaining consensus, generating new ideas" and "detecting anomalies". ISO/IEC 20246 defines it as a "formal review in which an author leads members of the review through a work product, and the participants ask questions and make comments about possible issues." Chapter Thirty is about the walkthrough.
  • Technical review. "A technical review is performed by technically qualified reviewers and led by a moderator." Its objectives are "to gain consensus and make decisions regarding a technical problem", and also to detect anomalies and evaluate quality. ISO/IEC 20246 defines it as a "formal peer review of a work product by a team of technically qualified personnel that examines the suitability of the work product for its intended use and identifies discrepancies from specifications and standards".
  • Inspection. "As inspections are the most formal type of review, they follow the complete generic process". "The main objective is to find the maximum number of anomalies", and "Metrics are collected and used to improve the SDLC, including the inspection process. In inspections, the author cannot act as the review leader or scribe." Chapter Twenty-Nine is about the inspection.
Informal reviewWalkthroughTechnical reviewInspection
FormalityNone: no defined processFormal in ISO/IEC 20246; flexible in practiceFormalThe most formal: the complete generic process
Led byAnyone, often the authorThe authorA moderatorA trained leader, never the author
Main objectiveDetect anomaliesMany: understanding, consensus, ideas, education, anomaliesConsensus and decisions on a technical problem; anomaliesThe maximum number of anomalies
Individual preparationOptionalOptionalExpectedRequired
Documented outputNot requiredYes, as a formal reviewYesYes, with metrics
Metrics collectedNoNoSeldomYes, and used to improve the process
Typical useEveryday drafts, pair workPresenting a design or code to the teamChoosing between technical designsCritical requirements, designs and code
munotes.in154

Software Reviews: The Process, the Roles and the Types

The level of formality is a choice. The syllabus says it depends on "the SDLC being followed, the maturity of the development process, the criticality and complexity of the work product being reviewed, legal or regulatory requirements, and the need for an audit trail", and that "The same work product can be reviewed with different review types, e.g., first an informal one and later a more formal one."

The formal technical review of Module 2 (Chapter Ninety-Seven) returns to reviews from the quality assurance side, as an SQA activity with its own guidelines; this chapter's process and types are the foundation it builds on.

Worked example: reviewing ExamReg's fee requirement

Here is the requirement as the analyst first wrote it.

IdRequirement (draft)
R-FEE-3A form submitted after the last date pays a late fee: Rs 100 if up to a week late, Rs 500 if up to 15 days late. Later forms are not accepted. Concession students do not pay fees.

Planning and initiation. The review leader chooses a technical review, because the fee rules decide money. The scope is R-FEE-3 and the exam cell's note it came from; the focus is boundaries and exceptions; the exit criterion is that no defect of major severity remains open. Each reviewer receives the requirement, the note and the team's requirements checklist.

Individual review. Each reviewer reads with a different technique. The tester uses the checklist and drafts test cases as they read; the exam cell clerk reads in role, as a student would meet the rule; the developer reads for anything that cannot be coded unambiguously. Their logs, merged by the scribe:

No.Anomaly loggedLogged by
1What is charged on the last date itself? The requirement says nothing about day 0Tester
2Is a form exactly 7 days late charged Rs 100 or Rs 500? "Up to a week" does not sayTester, developer
3Is a form exactly 15 days late accepted? "Up to 15 days" does not sayTester
4Our note says a concession waives the form fee only; this says concession students pay no fees at allExam cell clerk
5Are the days calendar days or working days?Developer
6The fee amounts change most sessions; can they be set without a code change?Developer

Communication and analysis. The meeting, run by the moderator, decides each item's status.

munotes.in155

Software Reviews: The Process, the Roles and the Types

No.StatusAction
1Defect (omission)Add: on or before the last date, no late fee
2Defect (ambiguity)State the band as 1 to 7 days late, inclusive
3Defect (ambiguity)State the band as 8 to 15 days late, inclusive; more than 15 days, refused
4Defect (contradicts the source note), majorCorrect: a concession waives the form fee only; late fees apply
5Question, answered: calendar daysNo defect; add the word "calendar" for clarity
6RecommendationPass to the design as a requirement for a configurable fee table

Fixing and reporting. The analyst rewrites R-FEE-3; the scribe's record goes to the project; the exit criterion is met once the major defect, item 4, is corrected and checked.

Four defects in one short requirement, found in a two-hour meeting. Item 1 is the on-time defect that Chapter One, on what software testing is, found with a failing test after the code was written; item 4 is the wrong specification that Chapter Twenty-Six showed no amount of verification could catch; and items 1 to 3 are the boundaries that the static analyser of Chapter Twenty-Seven found in code. Found here, each costs a sentence to fix.

What makes reviews succeed

The ISTQB syllabus lists the success factors, and every one of them can be seen in the example.

  • "Defining clear objectives and measurable exit criteria. Evaluation of participants should never be an objective"
  • "Choosing the appropriate review type to achieve the given objectives"
  • "Performing reviews on small chunks, so that reviewers do not lose concentration"
  • Feedback to stakeholders and authors, adequate time to prepare, and support from management
  • "Making reviews part of the organization's culture, to promote learning and process improvement"
  • Adequate training for all participants, and facilitated meetings

The first deserves emphasis. A review examines the work, never the person. If authors fear that their anomalies will be counted against them, they stop bringing work to review early, and the cheapest defect-finding activity in the project is lost.

What it does not mean

A review is not an inspection by another name. Inspection is one type, the most formal; the informal review, the walkthrough and the technical review are reviews too.

Every anomaly is not a defect. Anomalies are analysed; some are questions with good answers, and some are recommendations.

A review is not an assessment of the author. Evaluating participants "should never be an objective".

Reviews are not only for code. Any work product that can be read, from a requirement to a test plan, can be reviewed.

Quick revision

  • Review: a process or meeting in which work products are presented to stakeholders for comment or approval.
  • Process (ISO/IEC 20246, ISTQB v4.0.1): planning, review initiation, individual review, communication and analysis, fixing and reporting.
  • Anomaly: a deviation from expectations; analysis decides whether it is a defect.
  • Techniques: ad hoc, checklist-based, scenario-based, role-based, perspective-based.
  • Roles: manager, author, moderator (facilitator), scribe (recorder), reviewer, review leader.
  • Types: informal review (no defined process), walkthrough (led by the author), technical review (led by a moderator, technical consensus), inspection (most formal, maximum anomalies, metrics, author never leader or scribe).
  • Formality depends on the SDLC, process maturity, criticality, regulation and the need for an audit trail.
  • Success: clear objectives and exit criteria, never evaluating participants, small chunks, time, training, management support.
munotes.in156

Software Reviews: The Process, the Roles and the Types

Test yourself

1. Describe the activities of the review process. Planning defines the scope, focus, exit criteria, effort and timeframes. Review initiation makes sure everyone has the work product and understands their role. In individual review each reviewer applies review techniques and logs anomalies, recommendations and questions. In communication and analysis the anomalies are discussed, usually in a meeting, and each gets a status, owner and action. In fixing and reporting, defects are reported and fixed, the exit criteria are checked, and the results are reported.

2. Name the roles in a review and their responsibilities. The manager decides what is reviewed and provides resources; the author creates and fixes the work product; the moderator runs the meeting and keeps it safe and on time; the scribe records anomalies and decisions; the reviewers perform the review; and the review leader takes overall responsibility and organises it.

3. Compare the four types of review. An informal review follows no defined process and needs no documented output, and aims to detect anomalies. A walkthrough is led by the author and serves many objectives, including education and consensus. A technical review is led by a moderator with technically qualified reviewers, to reach decisions on technical problems. An inspection is the most formal, follows the complete process, aims to find the most anomalies, collects metrics, and never lets the author lead or record.

4. Why is every anomaly found in a review not a defect? Because an anomaly is only a deviation from what a reviewer expected; on analysis it may turn out to be a question with a good answer, a recommendation, or a misunderstanding by the reviewer. The meeting decides which anomalies are defects.

5. Give four factors for a successful review. Clear objectives and measurable exit criteria, with evaluation of participants never an objective; the right review type for the objectives and work product; reviewing in small chunks with adequate preparation time; and management support with training for all participants.

Contents This chapter on its own page

munotes.in157

Chapter Twenty-Nine

Inspection

Syllabus topic Module 1, "Verification and Validation (V&V): Inspection"

In one line

An inspection is the most formal kind of review: a trained moderator leads a small team with defined roles through five planned operations, the meeting's only aim is to find errors, and the errors found are classified, counted and used to improve the process that made them.

In the wording a student can write in an examination: an inspection is a "formal review of a work product to identify issues, which uses defined team roles and measurement to improve the review process" (ISO/IEC 20246:2017). The method is M. E. Fagan's, published in the IBM Systems Journal in 1976. Its five operations are overview, preparation, inspection, rework and follow-up; its four roles are moderator, designer, coder/implementor and tester, with one participant acting as reader. Fagan reported that inspections found 82 per cent of the errors in one application before testing, raised coding productivity by 23 per cent in one systems programming study, and produced 38 per cent fewer errors than a comparable walk-through sample.

Where inspection came from

Fagan's paper opens from the cost of late errors: "The cost of reworking errors in programs becomes higher the later they are reworked in the process, so every attempt should be made to find and fix errors as early in the process as possible." His answer was a process with inspections at fixed points, each with exit criteria, and he described inspections in three adjectives: "Inspections are a formal, efficient, and economical method of finding errors in design and code."

Fagan's programming process placed inspections after each major work product:

  • I0, after the internal (module) specifications;
  • I1, the design-complete inspection, after the logic specifications;
  • I2, the code inspection, once the code reaches its first clean compilation;
  • IT1 and IT2, inspections of the test plan and of the test cases;
  • PI0, PI1 and PI2, inspections of the publications, because poor documentation "can mislead the user, causing him to make errors quite as important as errors in the program."

The same method serves all of them. The paper describes it through I1 and I2, and says the others "retain the same essential properties" but differ "in materials inspected, number of participants, and some other minor points."

The five operations

Fagan's Table 3 lists the operations, the objective of each, and the rate at which each goes for systems programming.

OperationWhoObjective (Fagan's Table 3)Design I1 rateCode I2 rate
1. OverviewWhole teamCommunication, education500 lines per hourNot necessary
2. PreparationEach person aloneEducation100 lines per hour125 lines per hour
3. InspectionWhole teamFind errors130 lines per hour150 lines per hour
4. ReworkDesigner or coderRework and resolve errors found by inspection20 hours per K.NCSS16 hours per K.NCSS
5. Follow-upModeratorSee that all errors, problems and concerns have been resolvedNone givenNone given
munotes.in158

Inspection

K.NCSS is a thousand non-commentary source statements, roughly a thousand lines of code without comments. Each operation has its own rules.

  1. Overview. "The designer first describes the overall area being addressed and then the specific area he has designed in detail", and the design documents are handed out. A code inspection needs no overview, because the same people inspected the design.
  2. Preparation. Participants "literally do their homework to try to understand the design, its intent and logic." They also study the ranked distributions of error types found by recent inspections, and checklists of clues, so that they look where errors are most likely.
  3. Inspection. A reader chosen by the moderator, usually the coder, paraphrases the design or code. "Every piece of logic is covered at least once, and every branch is taken at least once." Each error found is noted, its type classified and its severity (major or minor) recorded, and the reading moves on. No one designs solutions at the meeting: "The inspection is not intended to redesign, evaluate alternate design solutions, or to find solutions to errors; it is intended just to find errors!" Within one day the moderator writes the inspection report.
  4. Rework. "All errors or problems noted in the inspection report are resolved by the designer or coder/implementor."
  5. Follow-up. The moderator checks that every issue is resolved. "If more than five percent of the material has been reworked, the team should reconvene and carry out a 100 percent reinspection."

Two practical rules come from Fagan's experience. The meeting's error detection falls off after two hours, so "it is advisable to schedule inspection sessions of no more than two hours at a time. Two two-hour sessions per day are acceptable." And the time for inspections and rework "must be scheduled and managed with the same attention as other important project activities", because under pressure inspections are the first thing a project drops.

Fagan also defined what the meeting is looking for: "an error is defined as any condition that causes malfunction or that precludes the attainment of expected or previously specified results."

The roles

"The inspection team is best served when its members play their particular roles", Fagan wrote, and he named four.

RoleFagan's description
Moderator"The key person in a successful inspection." A competent programmer, but not necessarily an expert on the program; best from an unrelated project, to preserve objectivity. Manages the team; in Fagan's words, "he is the coach". Schedules the meetings, reports within one day, follows up the rework, and should be specially trained
Designer"The programmer responsible for producing the program design."
Coder/implementor"The programmer responsible for translating the design into code."
Tester"The programmer responsible for writing and/or executing test cases or otherwise testing the product of the designer and coder."
munotes.in159

Inspection

The reader is not a fifth person: it is a job the moderator gives, usually to the coder. When one person has done two jobs, the roles are refilled from outside: if the same person designed and coded the work, they take the designer's role and "a coder from some related or similar program will perform the role of the coder." On size, "Four people constitute a good-sized inspection team", and the team "should not be artificially increased over four" unless the code touches several interfaces whose owners should be present.

The modern syllabus keeps the essential rule in a sentence: "In inspections, the author cannot act as the review leader or scribe." Fagan's moderator is today's review leader and moderator; his hand-written notes are today's scribe's log.

Looking where the errors are: checklists and error types

Fagan observed that finding errors has to be taught: "it is one thing to direct people to find errors in design or code. It is quite another problem for them to find errors." His answer had two parts. Inspectors study the ranked distribution of error types from recent inspections, so they concentrate on the most common and costly kinds; and they use checklists of clues for each type. His Figure 5 shows part of the design checklist for logic, with questions such as "Are All Constants Defined?" and "Are All Increment Counts Properly Initialized (0 or 1)?" Every error found is also classified as missing, wrong or extra.

The program below recomputes the error distributions from Fagan's Figures 3 and 4, read row by row off the page, and adds the arithmetic of his Tables 1 and 3.

import math

# Fagan 1976, Table 1: the Aetna application, errors found per K.NCSS
by_inspection, by_test, after = 38, 8, 0
total = by_inspection + by_test + after
print(f"Table 1: inspections found {by_inspection} of {total} errors per K.NCSS,"
      f" {by_inspection / total:.1%}")

# Figures 3 and 4: errors by type as (missing, wrong, extra), read off the page
design = {"logic": (126, 57, 24), "prologue/prose": (44, 38, 7), "CB usage": (18, 17, 1),
          "other": (15, 10, 10), "more detail": (24, 6, 2), "interconnect calls": (18, 9, 0),
          "test and branch": (12, 7, 2), "CB definition": (16, 2, 0),
          "maintainability": (8, 5, 3), "return code/msg": (5, 7, 2),
          "interconnect reqts": (4, 5, 2), "performance": (1, 2, 3),
          "register usage": (1, 2, 0), "higher level docu": (1, 0, 1), "FPFS": (1, 0, 0),
          "mod attributes": (1, 0, 0), "pass data areas": (0, 1, 0)}
code = {"logic": (33, 49, 10), "design error": (31, 32, 14), "prologue/prose": (25, 24, 3),
        "CB usage": (3, 21, 1), "code comments": (5, 17, 1), "interconnect calls": (7, 9, 3),
        "maintainability": (5, 7, 2), "PL/S or BAL use": (4, 9, 1), "performance": (3, 2, 5),
        "F1": (0, 8, 0), "test and branch": (2, 5, 0), "register usage": (4, 2, 0),
        "storage usage": (1, 0, 0)}
for name, table in [("Design (Figure 3)", design), ("Code (Figure 4)", code)]:
    errors = sum(sum(row) for row in table.values())
    missing, wrong, extra = (sum(row[i] for row in table.values()) for i in range(3))
    print(f"{name}: {errors} errors; missing {missing / errors:.0%},"
          f" wrong {wrong / errors:.0%}, extra {extra / errors:.0%}")
    ranked = sorted(table, key=lambda t: -sum(table[t]))[:3]
    print("  most frequent:", ", ".join(f"{t} {sum(table[t])} ({sum(table[t]) / errors:.1%})"
                                         for t in ranked))

# Table 3: rates of progress for systems programming, per person, for 1,000 lines
people, lines = 4, 1000
for name, overview, preparation, meeting, rework in [("design I1", 500, 100, 130, 20),
                                                     ("code I2", None, 125, 150, 16)]:
    hours = {"overview": people * lines / overview if overview else 0,
             "preparation": people * lines / preparation,
             "inspection": people * lines / meeting,
             "rework": rework}
    sessions = math.ceil(lines / meeting / 2)            # no session longer than two hours
    print(f"{name}: " + ", ".join(f"{k} {v:.1f}" for k, v in hours.items() if v)
          + f"; total {sum(hours.values()):.1f} people-hours; {sessions} two-hour meetings")
munotes.in160

Inspection

Table 1: inspections found 38 of 46 errors per K.NCSS, 82.6%
Design (Figure 3): 520 errors; missing 57%, wrong 32%, extra 11%
  most frequent: logic 207 (39.8%), prologue/prose 89 (17.1%), CB usage 36 (6.9%)
Code (Figure 4): 348 errors; missing 35%, wrong 53%, extra 11%
  most frequent: logic 92 (26.4%), design error 77 (22.1%), prologue/prose 52 (14.9%)
design I1: overview 8.0, preparation 40.0, inspection 30.8, rework 20.0; total 98.8 people-hours; 4 two-hour meetings
code I2: preparation 32.0, inspection 26.7, rework 16.0; total 74.7 people-hours; 4 two-hour meetings

The recomputed totals match what Fagan printed: 520 design errors split 57, 32 and 11 per cent, logic at 39.8 per cent, and 348 code errors with logic at 26.4 per cent. Three lessons come out of the numbers.

  • Design inspections find mostly what is missing; code inspections mostly what is wrong. More than half the design errors were missing items, while more than half the code errors were wrong ones. A design checklist should therefore ask what has been left out?, and a code checklist what is incorrect?
  • Logic dominates both, and documentation is close behind. Prologue and prose errors, in the comments and descriptions, are the second commonest design error and the third commonest code error; Fagan inspected them because the next programmer relies on them.
  • The rates turn into a plan. Inspecting 1,000 lines of design with four people costs about 99 people-hours by Table 3's rates, and the meeting alone needs four sessions of at most two hours. Fagan's own estimate for the whole process, "overview through follow-up", was "about 90 to 100 people-hours for systems programming". Code, needing no overview and read faster, costs about 75.
munotes.in161

Inspection

On ExamReg, the same arithmetic tells the test lead what a code inspection of the fee module will cost before the first meeting is booked, which is exactly what Fagan meant by managing inspections "with the same attention as other important project activities".

What Fagan reported

Fagan gave results from two settings, and warned in the paper that they "cannot be considered representative of every situation".

A systems programming study at IBM. A piece of an operating system component, designed by three programmers and coded by 13, went through I1 and I2 inspections for the first time. The net saving "translated into a 23 percent increase in the productivity of the coding operation alone." A control sample, taken once the inspections were routine, differed by only 0.9 per cent, so the gain was not a novelty effect. The net savings were 94 programmer hours per K.NCSS from I1 and 51 from I2, while a third inspection after unit test, I3, cost 20 hours more than it saved, and "As a consequence, I3 is no longer in effect." In testing after unit test, the inspected sample had "38 percent less errors" than a comparable piece built with walk-throughs.

An application at Aetna Life and Casualty. A COBOL program of 4,439 non-commentary statements in eight modules, written by two programmers with inspections as the only change to their process, was estimated to need 62 programmer days and took 46.5, including inspection meetings: "The resulting saving in programmer resources was 25 percent." Table 1 records where its errors were found: 38 per K.NCSS by the design and code inspections, 8 by unit and preparation for acceptance testing, and none in acceptance testing or in six months of use. That is Fagan's error detection efficiency of 82 per cent (38 ÷ 46 is 0.826, and the table prints 82), where

error detection efficiency = errors found by an inspection ÷ total errors in the product before inspection.

The reason inspections pay, in Fagan's words, is where they find errors: rework at the early levels "is 10 to 100 times less expensive than if it is done in the last half of the process." Chapter Ninety-Six, on why reviews pay, returns to that cost.

munotes.in162

Inspection

Inspection results are not for appraising people

Fagan was emphatic that inspection data belongs to the programmer: the results "should not under any circumstances be used for programmer performance appraisal." The reason is practical. An inspection works only if authors bring work early and errors are reported freely; the moment error counts are used against the people who made them, both stop. It is the same rule as the ISTQB success factor in Chapter Twenty-Eight, on software reviews: evaluation of participants should never be an objective.

What makes an inspection formal

FeatureIn an inspection
Entry and exit criteriaEach inspection has a defined point, such as first clean compilation for code, and cannot be claimed complete until its rework is done
A trained moderatorLeads, schedules, reports within a day and follows up; never the author
Defined rolesModerator, designer, coder/implementor, tester, and a reader
PreparationEvery participant studies the material, the error-type distributions and the checklists beforehand
A single objective in the meetingFind errors; no design, no solutions
ClassificationEvery error by type, as missing, wrong or extra, and as major or minor
Written recordsThe error list, the module summary and the inspection summary report
Verified reworkFollow-up by the moderator; full reinspection when more than 5 per cent is reworked
MeasurementError rates and types analysed to improve both the product and the process

What it does not mean

An inspection is not a meeting to fix the code. Solutions are noted if obvious and worked out afterwards; the meeting only finds errors.

An inspection does not replace testing. In Fagan's Aetna data inspections found 82 per cent of the errors; testing found the rest.

The moderator is not the author's manager or the author. The moderator should come from an unrelated project, and the author cannot lead or record.

Fagan's rates are not universal. They were measured for systems programming and are, by his note, conservative; application code went four to six times faster.

Quick revision

  • Inspection (ISO/IEC 20246): a formal review that uses defined team roles and measurement to improve the review process.
  • Fagan (IBM Systems Journal, 1976): inspection points I0, I1 (design complete), I2 (code), IT1, IT2 (test plan, test cases), PI (publications).
  • Operations: overview (communication, education), preparation (education), inspection (find errors), rework, follow-up.
  • Roles: moderator ("the key person", trained, from an unrelated project), designer, coder/implementor, tester; a reader paraphrases; four is a good size.
  • Rules: at most two hours per session; report within one day; no solution hunting; reinspect if more than 5 per cent reworked; never use results to appraise programmers.
  • Error detection efficiency = errors found by the inspection ÷ total errors before inspection; 82 per cent at Aetna.
  • Results: coding productivity up 23 per cent; 38 per cent fewer errors than walk-throughs; Aetna 25 per cent fewer programmer days.
  • Design errors are mostly missing (57 per cent); code errors mostly wrong (53 per cent); logic tops both.
munotes.in163

Inspection

Test yourself

1. Describe the five operations of a Fagan inspection. Overview: the designer presents the design to the whole team and hands out the documents. Preparation: each participant studies the material, the common error types and the checklists. Inspection: a reader paraphrases the work, every piece of logic and every branch is covered, and errors are noted and classified without seeking solutions; the moderator reports within a day. Rework: the designer or coder resolves every error. Follow-up: the moderator verifies the rework and calls a full reinspection if more than 5 per cent was reworked.

2. What are the roles in an inspection, and why must the moderator be trained and independent? Moderator, designer, coder/implementor and tester, with a reader chosen by the moderator. The moderator runs the whole process and the meeting, so needs training in leading it; coming from an unrelated project preserves objectivity, and the author may never lead or record.

3. Define error detection efficiency and compute it for Fagan's Aetna data. It is the errors found by an inspection divided by the total errors in the product before inspection. At Aetna, inspections found 38 errors per K.NCSS out of 46 found in total, about 82 per cent.

4. What results did Fagan report for inspections? In a systems programming study, a 23 per cent increase in coding productivity and 38 per cent fewer errors than a walk-through sample; in an application at Aetna, 25 per cent fewer programmer days than estimated, 82 per cent of errors found by inspection, and no errors in acceptance testing or six months of use.

5. What makes an inspection more formal than other reviews? Entry and exit criteria, a trained moderator who is not the author, defined roles, required preparation with checklists, a meeting with the single objective of finding errors, classification and written records of every error, verified rework, and measurement used to improve the process.

Contents This chapter on its own page

munotes.in164

Chapter Thirty

Walkthrough, and How It Differs From an Inspection

Syllabus topic Module 1, "Verification and Validation (V&V): Walkthrough"

In one line

A walkthrough is a review led by the author, who takes colleagues step by step through a work product so that they understand it, discuss it and point out problems; an inspection is led by a trained moderator, has one aim in its meeting (finding errors), and measures and follows up what it finds.

In the wording a student can write in an examination: a walkthrough is a "formal review in which an author leads members of the review through a work product, and the participants ask questions and make comments about possible issues" (ISO/IEC 20246:2017). The ISTQB syllabus lists its objectives as "evaluating quality and building confidence in the work product, educating reviewers, gaining consensus, generating new ideas, motivating and enabling authors to improve and detecting anomalies", and says individual preparation is not required. It differs from an inspection in who leads it (the author against a trained moderator), its objective (understanding and consensus as well as errors, against errors alone), its process (a meeting with optional preparation, against five operations with rework and follow-up), its roles, checklists and data (informal, against defined and measured) and its repeatability.

What a walkthrough is

Three definitions in the software engineering vocabulary describe the same activity from slightly different angles.

  • ISO/IEC/IEEE 24765 calls it a "static analysis technique in which a designer or programmer leads members of the development team and other interested parties through a segment of documentation or code, and the participants ask questions and make comments about possible errors, violation of development standards, and other problems".
  • ISO/IEC 20246 calls it a "formal review in which an author leads members of the review through a work product, and the participants ask questions and make comments about possible issues".
  • ISO/IEC 2382 defines a structured walkthrough as a "systematic examination of the requirements, design, or implementation of a system, or any part of it, by qualified personnel".

Two features are common to all three. The author leads: the person who produced the work presents it, in the order they choose. And the participants respond: they ask questions and comment as the author goes. Everything else, how much preparation, how many people, what is recorded, varies from one organisation to the next.

Why teams hold walkthroughs

The ISTQB syllabus gives a walkthrough more objectives than any other review type: "A walkthrough, which is led by the author, can serve many objectives, such as evaluating quality and building confidence in the work product, educating reviewers, gaining consensus, generating new ideas, motivating and enabling authors to improve and detecting anomalies. Reviewers might perform an individual review before the walkthrough, but this is not required."

That breadth is the reason to choose one. A walkthrough suits:

munotes.in165

Walkthrough, and How It Differs From an Inspection

  • teaching: a new team member learns the fee module fastest by having its author walk through it;
  • early drafts: an author who wants reactions before finishing a design;
  • consensus: a team choosing between two approaches, where discussing alternatives is the point;
  • stakeholders who are not programmers: an analyst walking the exam cell through the requirements, which is validation in the sense of Chapter Twenty-Six;
  • work of moderate risk, where the cost of a full inspection is not justified.

Roles in a walkthrough

A walkthrough needs fewer roles than an inspection, and they are less fixed.

RoleIn a walkthrough
Author (presenter)Leads the session and presents the work product, in the order they choose
Participants (reviewers)Colleagues, and sometimes users or other stakeholders; they ask questions and comment
ScribeOften present, recording the issues raised; in informal practice sometimes the author
ModeratorOptional; where there is one, they keep the meeting on time and on topic

Leading a code walkthrough: tracing test cases by hand

A code walkthrough can drift into a line-by-line reading that nobody remembers. A more effective way for the author to lead is to trace test cases through the code by hand: participants bring a few test cases, each with its expected result, and the author "executes" the code on paper for each one, saying aloud what every line does to every variable. The participants follow, and a wrong value shows up at the line where it happens.

Here is the author of ExamReg's fee function leading such a walkthrough. The fee rules are the book's fixed ones: a form fee of Rs 800, Rs 150 for each backlog paper, the late fee of Chapter One, on what software testing is, forms more than 15 days late refused, a negative number of days refused as invalid input, and a concession that waives the form fee only, as the Scrum team of Chapter Seventeen established.

FORM_FEE, BACKLOG_FEE = 800, 150

def total_fee(days_late, backlog_papers, concession):
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    fee = FORM_FEE + BACKLOG_FEE * backlog_papers
    if concession:
        fee = 0                                  # the concession
    if days_late <= 0:
        late = 0
    elif days_late <= 7:
        late = 100
    else:
        late = 500
    return fee + late

print("author's own check before the walkthrough, TC1:", total_fee(0, 0, False))

The author ran one case before the meeting, and it passed:

author's own check before the walkthrough, TC1: 800

The participants bring three test cases:

Test caseDays lateBacklog papersConcessionExpected total
TC100NoRs 800
TC2101NoRs 1,450
TC332YesRs 400

The author traces each one aloud, and the scribe writes the trace on the board.

munotes.in166

Walkthrough, and How It Differs From an Inspection

StepTC1 (0, 0, No)TC2 (10, 1, No)TC3 (3, 2, Yes)
days_late < 0 or days_late > 15?NoNoNo
fee = 800 + 150 × backlog8009501,100
concession?NoNoYes: fee = 0
late0 (on time)500 (8 to 15 days)100 (1 to 7 days)
Returned8001,450100
Expected8001,450400

TC3 stops the room. At the concession line the trace sets fee to 0, wiping out the Rs 300 for two backlog papers along with the form fee. A participant asks what the concession is supposed to waive; the answer, the form fee only, is the fix: subtract the form fee instead of zeroing the total. The failing case was never run: the defect was found on paper, at the line that caused it. After the walkthrough the author makes the change, and the three test cases are run on both versions to confirm what the trace showed.

FORM_FEE, BACKLOG_FEE = 800, 150

def total_fee(days_late, backlog_papers, concession):         # as walked through
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    fee = FORM_FEE + BACKLOG_FEE * backlog_papers
    if concession:
        fee = 0                                                 # the concession
    if days_late <= 0:
        late = 0
    elif days_late <= 7:
        late = 100
    else:
        late = 500
    return fee + late

def total_fee_fixed(days_late, backlog_papers, concession):   # after the walkthrough
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    fee = FORM_FEE + BACKLOG_FEE * backlog_papers
    if concession:
        fee -= FORM_FEE                                         # waives the form fee only
    if days_late <= 0:
        late = 0
    elif days_late <= 7:
        late = 100
    else:
        late = 500
    return fee + late

cases = [("TC1", (0, 0, False), 800), ("TC2", (10, 1, False), 1450), ("TC3", (3, 2, True), 400)]
for version in (total_fee, total_fee_fixed):
    print(version.__name__)
    for name, args, expected in cases:
        actual = version(*args)
        mark = "" if actual == expected else "  <- the walkthrough's finding"
        print(f"  {name} {args}: expected {expected}, got {actual}{mark}")
total_fee
  TC1 (0, 0, False): expected 800, got 800
  TC2 (10, 1, False): expected 1450, got 1450
  TC3 (3, 2, True): expected 400, got 100  <- the walkthrough's finding
total_fee_fixed
  TC1 (0, 0, False): expected 800, got 800
  TC2 (10, 1, False): expected 1450, got 1450
  TC3 (3, 2, True): expected 400, got 400

The run agrees with the hand trace line for line: the version walked through returns 100 for TC3, and the fixed one 400. Notice also what this walkthrough did and did not do. It found a real defect and it taught the room how the concession works. But it did not use a checklist, did not classify the defect, did not measure anything, and nobody is assigned to verify the fix except the author. Those are exactly the things an inspection adds.

munotes.in167

Walkthrough, and How It Differs From an Inspection

Walkthrough against inspection

Fagan, who defined the inspection, compared the two in his 1976 paper. His starting point was that walk-throughs "are practiced in many different ways in different places, with varying regularity and thoroughness. This inconsistency causes the results of walk-throughs to vary widely and to be nonrepeatable. Inspections, however, having an established process and a formal procedure, tend to vary less and produce more repeatable results."

His Table 4 sets the two processes side by side:

Inspection operationObjectiveWalk-through operationObjective
1. OverviewEducation (group)NoneNone
2. PreparationEducation (individual)1. PreparationEducation (individual)
3. InspectionFind errors! (group)2. Walk-throughEducation (group); discuss design alternatives; find errors
4. ReworkFix problemsNoneNone
5. Follow-upEnsure all fixes correctly installedNoneNone

His note under the table makes the key point: "Note the separation of objectives in the inspection process." An inspection meeting does one thing; a walk-through meeting does three at once.

His Table 5 compares their key properties:

Property (Fagan, Table 5)InspectionWalk-through
Formal moderator trainingYesNo
Definite participant rolesYesNo
Who drives itModeratorOwner of the material (designer or coder)
Uses how-to-find-errors checklistsYesNo
Uses the distribution of error types to look forYesNo
Follow-up to reduce bad fixesYesNo
Fewer future errors from detailed error feedback to the programmerYesIncidental
Improves inspection efficiency from analysis of resultsYesNo
Analysis of data leads to process problems and improvementsYesNo

And in one comparison of equivalent pieces of an operating system, the inspected sample had "38 percent less errors" in later testing than the walk-through sample, as Chapter Twenty-Nine, on inspection, reported.

Putting Fagan's comparison together with the modern syllabus gives the table a 5-mark answer needs.

WalkthroughInspection
Led byThe authorA trained moderator; never the author
Main objectiveUnderstanding, consensus, education, and finding problemsFinding the maximum number of errors
ProcessOptional preparation, then the walkthrough meetingOverview, preparation, inspection, rework, follow-up
PreparationOptionalRequired, with checklists and error-type data
RolesAuthor and participants; few fixed rolesDefined: moderator, designer, coder, tester, reader
Solutions in the meetingDiscussed, including design alternativesNot discussed; errors only
Records and metricsFew; issues notedClassified errors, reports, measurements
Follow-up of fixesLeft to the authorVerified by the moderator; reinspection if needed
RepeatabilityVaries widely (Fagan)More repeatable (Fagan)
Best forTeaching, early drafts, consensus, stakeholder reviewsCritical designs and code, where errors cost most
munotes.in168

Walkthrough, and How It Differs From an Inspection

Choosing between them

The choice follows from the ISTQB syllabus's factors from Chapter Twenty-Eight, on reviews: the criticality of the work product, the objectives, and the need for an audit trail. On ExamReg, the new help pages would go to a walkthrough with the exam cell, because the aim is to check that students will understand them; the fee and payment code would go to an inspection, because a missed error there costs students money. Many teams use both on the same work product: a walkthrough while the design is taking shape, and an inspection once it is complete.

What it does not mean

A walkthrough is not a weak inspection. It has different objectives, and for teaching and consensus it does better than an inspection would.

A walkthrough does not mean no preparation is ever done. Preparation is optional, not forbidden, and participants who prepare find more.

Tracing by hand is not the same as testing. It checks the code against a few expected results in people's heads; the cases still need to be run, as they were after this walkthrough.

Being led by the author is not a flaw. It is what makes a walkthrough good at teaching; it is also why a walkthrough should not be the only review of critical work.

Quick revision

  • Walkthrough: an author leads members of the review through a work product; participants ask questions and comment (ISO/IEC 20246). Led by the author; preparation optional (ISTQB).
  • Objectives (ISTQB v4.0.1): evaluate quality and build confidence, educate reviewers, gain consensus, generate ideas, motivate authors, detect anomalies.
  • Tracing test cases by hand is an effective way to lead a code walkthrough; the ExamReg trace found a concession that wiped out backlog fees.
  • Fagan's Table 4: the inspection has five operations with separate objectives; the walk-through has preparation and one meeting with mixed objectives.
  • Fagan's Table 5: inspections have moderator training, defined roles, checklists, error-type data, follow-up, feedback and process analysis; walk-throughs do not.
  • Walk-through results "vary widely" and are "nonrepeatable"; inspections are more repeatable; the inspected sample had 38 per cent fewer errors.
  • Use walkthroughs for teaching, drafts, consensus and stakeholders; inspections for critical work.

Test yourself

1. What is a walkthrough, and who leads it? A review in which the author leads members of the review through a work product while the participants ask questions and comment on possible issues. The author leads it, presenting the work in the order they choose.

2. What are the objectives of a walkthrough? Evaluating quality and building confidence in the work product, educating reviewers, gaining consensus, generating new ideas, motivating and enabling authors to improve, and detecting anomalies.

3. Differentiate between a walkthrough and an inspection. A walkthrough is led by the author, an inspection by a trained moderator who is never the author. A walkthrough aims at understanding and consensus as well as problems, and may discuss solutions; an inspection's meeting aims only at finding errors. A walkthrough is a meeting with optional preparation; an inspection has overview, preparation, inspection, rework and follow-up. An inspection uses defined roles, checklists, error-type data and metrics, and verifies fixes; a walkthrough usually does not, so its results are less repeatable.

munotes.in169

Walkthrough, and How It Differs From an Inspection

4. How can an author lead a code walkthrough effectively? By tracing test cases through the code by hand: participants bring test cases with expected results, and the author executes the code on paper, saying what each line does to each variable, so that a wrong value appears at the line that causes it.

5. When would you choose a walkthrough rather than an inspection? When the aim is teaching, gathering reactions to an early draft, reaching consensus, or checking a work product with non-technical stakeholders, and the work is not so critical that the cost of a full inspection is justified.

Contents This chapter on its own page

munotes.in170

Chapter Thirty-One

A Strategic Approach to Software Testing

Syllabus topic Module 1, "Software Testing Strategies: Strategic approach to software testing"

In one line

A strategic approach to testing decides, before testing starts, which levels of testing a product will get and in what order, who does each, how each is done, and how the team will know testing is complete; it tests from the small to the large, because each level catches only part of what the level before it missed.

In the wording a student can write in an examination: a test strategy is the "part of the test plan that describes the approach to testing for a specific project, test level, or test type" (ISO/IEC/IEEE 29119-2:2021). A strategic approach to software testing (1) begins with reviews and component (unit) testing and works outward through integration, system and acceptance testing, "from individual components to complete systems"; (2) gives each level its own objectives, test basis and people, with testing done at several levels of independence, from the developer to testers outside the organisation; (3) combines kinds of strategy, such as risk-based and exploratory; and (4) defines completion criteria, the "conditions under which the testing activities are considered complete" (ISO/IEC/IEEE 29119-2:2021).

What a strategy has to decide

Chapter Eleven placed the test strategy inside the test plan, under the organisation's test policy. This chapter is about what the strategy contains: the decisions that shape all the testing on a project. For ExamReg they look like this.

DecisionThe questionExamReg's answer
LevelsWhich test levels, in what order?Component, component integration, system, system integration with the payment gateway, acceptance
PeopleWho tests at each level?Developers for components and their integration; the test team for system testing; the exam cell for acceptance
TechniquesHow are tests designed at each level?Black-box techniques at every level; white-box coverage at component level; exploratory sessions before release
TypesWhich quality characteristics are tested?Functional, plus performance (the last date), security (student data) and usability (first-year students)
RegressionHow is working behaviour protected?Unit and regression suites run on every build
CompletionWhen is each level, and testing as a whole, done?Exit criteria for each level, and a release decision with the exam cell on the remaining risk

The rest of the chapter takes the decisions in order: the order of the levels, who tests, the kind of strategy, and completion.

From the small to the large

The ISTQB syllabus defines test levels as "groups of test activities that are organized and managed together", each "performed in relation to software at a given phase of development, from individual components to complete systems or, where applicable, systems of systems." A strategic approach runs through them from the smallest to the largest:

  1. Component (unit) testing checks each unit in isolation.
  2. Component integration testing checks the interfaces as units are combined.
  3. System testing checks the whole system against its specification, end to end.
  4. System integration testing checks its interfaces with other systems, such as a payment gateway.
  5. Acceptance testing validates the system against its users' needs.
munotes.in171

A Strategic Approach to Software Testing

There are two reasons for the order. The first is localisation. When a unit test fails, the fault is in that unit; when a system test fails, it could be in any unit or in any interface between them, and finding it is a longer piece of debugging. Testing the small pieces first means that by the time the large tests run, most faults that could be found in one place already have been. The second is the division of objectives. Each level looks for what it is best placed to find: a unit test for a wrong calculation, an integration test for two units that disagree about a parameter, a system test for a behaviour that only appears when everything runs together, an acceptance test for a requirement that was wrong all along. The syllabus notes that in sequential development "the exit criteria of one level are part of the entry criteria for the next level", so the levels form a chain of checkpoints.

Reviews come before all of them. A strategy that starts testing only when code exists has already given up the cheapest defects, as Chapters Twenty-Eight to Thirty on reviews, inspection and the walkthrough showed.

Worked example: why one level is never enough

ExamReg release 2.0's defect data, used throughout this book, records which activity found each of its 200 defects. Reviews found 74 before any code ran; the remaining 126 were in the code when unit testing began. The program follows those 126 through the four test levels, assuming, as a simplification, that no fix brought in a new defect.

found = {"requirements review": 16, "design review": 24, "code review": 34,
         "unit testing": 44, "integration testing": 32, "system testing": 28,
         "acceptance testing": 10}
after_release = 12
total = sum(found.values()) + after_release                     # 200, ExamReg release 2.0
remaining = total - sum(found[r] for r in ("requirements review", "design review", "code review"))
at_start = remaining
print(f"defects in the code when unit testing began: {remaining} of {total}")
for level in ("unit testing", "integration testing", "system testing", "acceptance testing"):
    caught = found[level]
    print(f"{level:<20} caught {caught:>2} of the {remaining:>3} still there ({caught / remaining:.1%})")
    remaining -= caught
print(f"{'students':<20} met the last {remaining}")
print(f"the four test levels together: {at_start - remaining} of {at_start}"
      f" ({(at_start - remaining) / at_start:.1%})")
defects in the code when unit testing began: 126 of 200
unit testing         caught 44 of the 126 still there (34.9%)
integration testing  caught 32 of the  82 still there (39.0%)
system testing       caught 28 of the  50 still there (56.0%)
acceptance testing   caught 10 of the  22 still there (45.5%)
students             met the last 12
the four test levels together: 114 of 126 (90.5%)
munotes.in172

A Strategic Approach to Software Testing

No level caught even three in five of the defects still present when it began: unit testing caught about 35 per cent, integration testing 39, system testing 56 and acceptance testing about 45. Each was far from enough on its own. Together they caught 114 of the 126, about 90.5 per cent, and with the reviews in front of them the whole strategy stopped 188 of 200 defects before students saw them. That is the argument for a strategy of several levels in one table: each level is a filter with holes, and the holes in different filters are in different places. Chapter Seventy-Nine, on defect metrics, gives these ratios their standard names.

Who tests: levels of independence

A strategy also decides who does the testing at each level. The ISTQB syllabus describes four degrees of independence: work products "can be tested by their author (no independence), by the author's peers from the same team (some independence), by testers from outside the author's team but within the organization (high independence), or by testers from outside the organization (very high independence)."

Neither extreme is right on its own. Independence helps "due to differences between the author's and the tester's cognitive biases", but "Independence is not, however, a replacement for familiarity, e.g., developers can efficiently find many defects in their own code." So the syllabus recommends mixing them: "For most projects, it is usually best to carry out testing with multiple levels of independence (e.g., developers performing component testing and component integration testing, test team performing system and system integration testing, and business representatives performing acceptance testing)." That is exactly the pattern in ExamReg's strategy above.

BenefitsDrawbacks
Independent testers (ISTQB v4.0.1)Recognise "different kinds of failures and defects compared to developers because of their different backgrounds, technical perspectives, and biases"; can "verify, challenge, or disprove assumptions" made during specification and implementationMay be "isolated from the development team", leading to poor communication or "an adversarial relationship"; developers "may lose a sense of responsibility for quality"; testers "may be seen as a bottleneck or be blamed for delays in release"

Kinds of test strategy

The 2018 Foundation syllabus, which the current version replaced, named seven common kinds of test strategy. The current syllabus no longer lists them, but the names are still in wide use and describe real choices.

KindTests are designed and chosen fromExamReg example
Analytical"an analysis of some factor (e.g., requirement or risk)"; risk-based testing is the exampleTest effort shared by product risk, as in Chapter Twenty-One, on how quality factors shape testing
Model-based"some model of some required aspect of the product", such as a state model or a business processTests derived from the login lockout's state diagram
Methodical"making systematic use of some predefined set of tests or test conditions", such as a taxonomy of likely failures or a checklistA checklist of the web-form faults the team has seen before
Process-compliant"external rules and standards" imposed on or by the organisationThe university's or a regulator's required tests, if any applied
Directed (consultative)"the advice, guidance, or instructions of stakeholders, business domain experts, or technology experts"The exam cell names the fee cases it worries about most
Regression-averse"a desire to avoid regression of existing capabilities", with reuse and automation of regression testsThe automated suite run on every build
Reactive"the component or system being tested, and the events occurring during test execution", with exploratory testing the common techniqueExploratory sessions on the new concession screens
munotes.in173

A Strategic Approach to Software Testing

The same syllabus's advice still holds: "An appropriate test strategy is often created by combining several of these types of test strategies", and its example is the very pair ExamReg uses, risk-based testing combined with exploratory testing.

When is testing complete?

Exhaustive testing is impossible, the second of the seven principles in Chapter Four, so testing never finishes by running out of tests. It finishes when agreed conditions are met. ISO/IEC/IEEE 29119-2 calls them completion criteria, "conditions under which the testing activities are considered complete", and Chapter Eleven, on writing a test plan, showed the typical exit criteria for each level: planned tests run, coverage reached, no open defect of critical severity, and so on. Three further points belong to strategy.

  1. Completion is decided per level and then for the release. Each level's exit criteria are part of the next level's entry criteria; the release decision is made on the evidence of all of them.
  2. Completion is a decision about risk. As Chapter Twenty-Three, on quality in software development, reported from the NIST study, deciding when to stop testing is the industry's hardest question, and the usual rules of thumb do not measure the risk that remains. The honest form of the decision is: the exit criteria are met, the remaining defects and risks are known, and the people who will carry them accept them.
  3. Completion can be predicted. Reliability models fit the record of failures found during testing and forecast forward; in Lyu's words, one purpose is "to predict the extra time needed to test the software to achieve a specified objective". Chapter Ninety-Two, on measuring software reliability, computes such a prediction.
munotes.in174

A Strategic Approach to Software Testing

What it does not mean

A strategy is not a list of tests. It is the set of decisions that determines which tests are written, by whom, and when they are enough.

Testing from small to large does not mean the levels never overlap. In iterative development all of them run in every iteration, and the syllabus notes that "Test levels may overlap in time."

Independent testing does not mean developers stop testing. The recommended mix has developers testing components and their integration.

Testing is not complete when the schedule runs out. It is complete when the agreed criteria are met, or when the remaining risk is knowingly accepted.

Quick revision

  • Test strategy: "part of the test plan that describes the approach to testing for a specific project, test level, or test type" (ISO/IEC/IEEE 29119-2:2021).
  • A strategy decides the levels, the people, the techniques, the types, regression and completion.
  • Test from the small to the large: component, component integration, system, system integration, acceptance, after reviews; smaller levels localise faults and each level has its own objectives.
  • ExamReg: unit 35, integration 39, system 56, acceptance 45 per cent of the defects still present; together 90.5 per cent; with reviews, 188 of 200.
  • Independence: author (none), peers (some), testers inside the organisation (high), outside it (very high); best to mix levels.
  • Kinds of strategy (ISTQB 2018): analytical, model-based, methodical, process-compliant, directed, regression-averse, reactive; usually combined.
  • Completion criteria: "conditions under which the testing activities are considered complete"; per level, then a risk-based release decision.

Test yourself

1. What is a test strategy, and what does it decide? The part of the test plan that describes the approach to testing for a project, test level or test type. It decides which test levels are used and in what order, who tests at each level, the techniques and test types, how regressions are prevented, and the criteria for completion.

2. Why does a testing strategy proceed from unit testing to system testing? Because a failure in a small test points to a small place, so faults are cheapest to locate there, and because each level has objectives it is best placed to meet: calculations in units, disagreements at interfaces in integration, whole-system behaviour in system testing, and fitness for use in acceptance testing.

3. Explain the levels of independence in testing, with their benefits and drawbacks. Testing may be done by the author, by peers, by testers outside the team, or by testers outside the organisation. Independent testers see different defects and challenge assumptions, but may become isolated, adversarial or a bottleneck, and developers may feel less responsible for quality; so a mix of levels is usually best.

4. Name four kinds of test strategy. Analytical (for example risk-based), model-based, methodical (for example checklist-based), process-compliant, directed, regression-averse and reactive (for example exploratory); any four of these.

munotes.in175

A Strategic Approach to Software Testing

5. How does a project decide that testing is complete? By completion criteria agreed in advance: exit criteria for each level such as tests run, coverage reached and no critical defect open, followed by a release decision in which the remaining risks are known and accepted by those who carry them. Reliability models can also predict how much more testing a target needs.

Contents This chapter on its own page

munotes.in176

Chapter Thirty-Two

Test Levels and Test Types

Syllabus topic Module 1, "Software Testing Strategies: Strategic approach to software testing"

In one line

A test level says where in development the testing happens and on how big a piece of the software; a test type says what the testing is about; every test sits somewhere on both, and every change calls for two more kinds of test, confirmation that the fix works and regression testing that nothing else broke.

In the wording a student can write in an examination: a test level is "one of a sequence of test stages, each of which is typically associated with the achievement of particular objectives and used to treat particular risks" (ISO/IEC/IEEE 29119-2:2021). The ISTQB syllabus names five: component (unit) testing, component integration testing, system testing, system integration testing and acceptance testing. A test type is "testing that is focused on specific quality characteristics" (the same standard), and the syllabus names four: functional, non-functional, black-box and white-box testing, all of which "can be applied to all test levels". Confirmation testing "confirms that an original defect has been successfully fixed"; regression testing "confirms that no adverse consequences have been caused by a change".

The five test levels

The ISTQB syllabus says levels are distinguished by five attributes: the test object, the test objectives, the test basis, the defects and failures looked for, and the approach and responsibilities. Taking each level through them gives the table that answers most questions about levels.

LevelTest objectMain objectiveTest basisTypical defectsUsually done by
Component (unit)One unit, in isolationEach unit works as designedDetailed design, the codeWrong calculation, a missing case, bad error handlingDevelopers
Component integrationUnits combined, and their interfacesThe units work togetherInterface and architecture designWrong parameters, mismatched data formats, wrong call orderDevelopers
SystemThe whole systemIt meets its specification, end to endRequirements specification, use casesWrong end-to-end behaviour, missing functions, performance and security failuresAn independent test team
System integrationThe system with external systemsIts interfaces with other systems workInterface agreements, protocolsMessages misread or lost between systemsTest team, with the other system's owners
AcceptanceThe system in its intended useIt is fit to be accepted and deployedBusiness needs, user requirements, contractsRequirements that were wrong or missing; a system not fit for useUsers and customers

The standards define each level in a sentence.

  • Component testing is "testing of individual hardware or software components" (IEEE 1012-2024). The syllabus adds that it "often requires specific support, such as test harnesses or unit test frameworks" and "is normally performed by developers in their development environments".
  • Integration testing is "testing in which software components, hardware components, or both are combined and tested to evaluate the interaction among them" (ISO/IEC 29110-5-1-2:2025). Component integration testing, in the syllabus's words, "is heavily dependent on the integration strategy like bottom-up, top-down or big-bang".
  • System testing is "testing conducted on a complete, integrated system to evaluate the system's compliance with its specified requirements" (ISO/IEC 29110-5-1-2:2025).
  • System integration testing "focuses on testing the interfaces of the system under test and other systems and external services" (ISTQB).
  • Acceptance testing is "formal testing conducted to enable a user, customer, or other authorized entity to determine whether to accept a system or component" (IEEE 1012-2024). Its forms, in the syllabus, are user acceptance testing, operational acceptance testing, contractual and regulatory acceptance testing, and alpha and beta testing.
munotes.in177

Test Levels and Test Types

The four test types

Where a level is a place, a type is a purpose. The ISTQB syllabus names four.

  • Functional testing "evaluates the functions that a component or system should perform". Its objective is checking "the functional completeness, functional correctness and functional appropriateness". It asks what the system does.
  • Non-functional testing "evaluates attributes other than functional characteristics of a component or system", which the syllabus calls testing "how well the system behaves": performance efficiency, compatibility, usability, reliability, security, maintainability, portability and safety, the characteristics of ISO/IEC 25010 from Chapter Twenty, on the quality model today.
  • Black-box testing "is specification-based and derives tests from documentation not related to the internal structure of the test object". ISO/IEC/IEEE 29119-1 calls the same thing specification-based testing, in which "the principal test basis is the external inputs and outputs of the test item".
  • White-box testing "is structure-based and derives tests from the system's implementation or internal structure (e.g., code, architecture, work flows, and data flows)". Its objective is "to cover the underlying structure by the tests to an acceptable level". The standard's name is structure-based testing.

Notice that the four are not one list of alternatives. Functional and non-functional divide testing by what is being evaluated; black-box and white-box divide it by where the tests come from. A single test can be functional and black-box (a fee case from the fee table) or non-functional and white-box (a timing test aimed at one slow loop). Chapter Fifty-Two, on black-box and white-box testing, starts Module 2 with the techniques of each family.

The grid: every level, every type

The syllabus says the four types "can be applied to all test levels, although the focus will be different at each level." On ExamReg the grid looks like this.

LevelFunctionalNon-functionalBlack-boxWhite-box
ComponentThe fee function returns Rs 500 for 8 days lateThe fee function answers in under a millisecondFee cases from the fee tableEvery branch of the fee function executed
Component integrationThe fee calculator passes the right total to the payment adapterThe adapter times out safely if the gateway is slowCases from the interface specificationEvery call path between the two units exercised
SystemA student registers, pays and downloads a hall ticket500 students on the last date; one student cannot see another's formScenarios from the requirementsEvery workflow step of the registration process covered
System integrationA payment made on ExamReg is confirmed by the gatewayThe link to the gateway recovers after a dropped connectionCases from the gateway's interface agreementEvery message type in the exchange covered
AcceptanceThe exam cell confirms the fee rules match their noteFirst-year students complete the form unaidedUser scenariosRarely used
munotes.in178

Test Levels and Test Types

Confirmation testing and regression testing

Two more kinds of testing are defined not by level or type but by what triggers them: a change. The syllabus says that after a change, "Testing should then also include confirmation testing and regression testing."

Confirmation testing, which ISO/IEC/IEEE 29119-2 calls retesting, is "testing performed to check that modifications made to correct a fault have successfully removed the fault". In the syllabus's version it "confirms that an original defect has been successfully fixed", by executing the tests that failed because of it and, where needed, new tests covering the change.

Regression testing is "testing performed following modifications to a test item or to its operational environment, to identify whether failures in unmodified parts of the test item occur" (ISO/IEC/IEEE 29119-1:2022). The syllabus stresses its scope: adverse consequences "could affect the same component where the change was made, other components in the same system, or even other connected systems", so "It is advisable first to perform an impact analysis to recognize the extent of the regression testing."

Worked example: a fix that passes, and a regression it causes

ExamReg's test team reports that a form submitted exactly 15 days late is refused, although the rule says it pays Rs 500. The developer fixes the first condition and, while the file is open, "tidies" another line. The program runs the confirmation test on the fix, then reruns the rest of the suite as a regression test.

def late_fee_before(days_late):          # the build tested: refuses a form 15 days late
    if days_late < 0 or days_late >= 15:
        return "refused"
    if days_late <= 0:
        return 0
    if days_late <= 7:
        return 100
    return 500

def late_fee_after(days_late):           # the developer's fix, with a tidy-up on the side
    if days_late < 0 or days_late > 15:
        return "refused"
    if days_late <= 0:
        return 0
    if days_late <= 8:                   # "tidied" while the file was open
        return 100
    return 500

suite = {"TC-01": (-1, "refused"), "TC-02": (0, 0), "TC-03": (1, 100), "TC-04": (7, 100),
         "TC-05": (8, 500), "TC-06": (15, 500), "TC-07": (16, "refused")}
failed_before = [t for t, (d, e) in suite.items() if late_fee_before(d) != e]
print("failed on the build tested:", ", ".join(failed_before))

for t in failed_before:                                   # confirmation testing
    d, e = suite[t]
    got = late_fee_after(d)
    print(f"confirmation {t} ({d} days late): expected {e}, got {got}:",
          "pass" if got == e else "FAIL")

rest = [t for t in suite if t not in failed_before]      # regression testing
broken = [t for t in rest if late_fee_after(suite[t][0]) != suite[t][1]]
print(f"regression: {len(rest)} tests rerun, {len(broken)} failed")
for t in broken:
    d, e = suite[t]
    print(f"  {t} ({d} days late): expected {e}, got {late_fee_after(d)}")
munotes.in179

Test Levels and Test Types

failed on the build tested: TC-06
confirmation TC-06 (15 days late): expected 500, got 500: pass
regression: 6 tests rerun, 1 failed
  TC-05 (8 days late): expected 500, got 100

The confirmation test passes: the reported failure is gone. A team that stopped there would ship a new defect, because the "tidy-up" moved day 8 into the Rs 100 band. Only the regression run, rerunning tests that had passed before and whose code nobody meant to change, catches it. This is why the syllabus calls regression suites "a strong candidate for automation": they are rerun after every change, they grow with every release, and the failures they catch are exactly the ones nobody is looking for. Chapter Thirty-Nine, on regression testing, smoke testing and continuous integration, runs them automatically on every build.

Confirmation testingRegression testing
QuestionIs the reported defect really fixed?Did the change break anything else?
Tests runThe tests that failed, and new tests for the changeTests that passed before, chosen by impact analysis
ScopeThe fixed defectThe changed component, other components, even other systems and the environment
Size over timeSmall, one defect at a timeGrows with every release, so it is automated

Where each kind of testing is taught in this book

The rest of Module 1 goes through the levels, and Module 2 through the techniques, in this order:

  • Static testing: the kinds of V&V and reviews, inspection and the walkthrough, Chapters Twenty-Seven to Thirty.
  • Unit testing: its purpose, techniques, frameworks and best practices, Chapters Thirty-Three to Thirty-Six.
  • Integration testing: why units fail together, the integration approaches, regression and smoke testing with continuous integration, and the challenges of integration, Chapters Thirty-Seven to Forty.
  • Validation and acceptance testing: validation testing and acceptance testing with alpha and beta, Chapters Forty-One and Forty-Two.
  • System testing: system testing end to end, the system test types, load testing and cross-browser testing, Chapters Forty-Three to Forty-Six.
  • Debugging, which follows a failure at any level, Chapter Forty-Seven.
  • Test automation: automation, browser drivers, data-driven testing and the page object model, Chapters Forty-Eight to Fifty-One.
  • Black-box, white-box and experience-based techniques in Module 2, from black-box and white-box testing in Chapter Fifty-Two to checklist-based testing in Chapter Sixty-Four.
munotes.in180

Test Levels and Test Types

What it does not mean

A test level is not a test type. "System testing" says where; "performance testing" says what; a system-level performance test is both.

White-box testing is not confined to unit testing. Structure can be covered at every level: code at component level, call paths at integration level, workflows at system level.

Confirmation testing is not regression testing. The first checks the fix; the second checks everything the fix might have disturbed.

Regression testing is not only for code changes. A change to the operating system, the database or a connected system can break unchanged code, and the definition includes changes to the "operational environment".

Quick revision

  • Test level: a test stage with its own objectives and risks. Five: component, component integration, system, system integration, acceptance.
  • Levels differ in test object, objectives, test basis, defects and failures, and approach and responsibilities.
  • Test type: testing focused on specific quality characteristics. Four: functional (what), non-functional (how well), black-box (from the specification), white-box (from the structure). All four apply at every level.
  • Confirmation testing (retesting): checks the fix removed the fault. Regression testing: checks unmodified parts still work after a change, scoped by impact analysis, and automated.
  • Worked example: the fix for day 15 passed its confirmation test; regression caught the new day-8 defect.

Test yourself

1. Name the five test levels and give the test object of each. Component testing (a single unit in isolation), component integration testing (units combined and their interfaces), system testing (the whole system), system integration testing (the system's interfaces with external systems) and acceptance testing (the system in its intended use).

2. Name the four test types and state what each evaluates. Functional testing evaluates what the system does; non-functional testing how well it behaves (performance, usability, security and so on); black-box testing derives tests from the specification, not the internal structure; white-box testing derives them from the internal structure, aiming to cover it.

3. Can a test type be applied at more than one level? Give an example. Yes; all four types apply at all levels. Performance, a non-functional type, can be tested at component level (the fee function's speed) and at system level (500 students on the last date).

4. Distinguish confirmation testing from regression testing. Confirmation testing reruns the tests that failed, and new tests for the change, to check that a defect has been fixed. Regression testing reruns tests that passed before, chosen by impact analysis, to check that the change has not broken anything else in the component, the system, connected systems or the environment.

munotes.in181

Test Levels and Test Types

5. Why is regression testing a strong candidate for automation? Because the same tests are rerun after every change, the suite grows with every release, and running it by hand each time would be slow and error-prone; automated in continuous integration, it runs on every build.

Contents This chapter on its own page

munotes.in182

Chapter Thirty-Three

Unit Testing: What It Is For

Syllabus topic Module 1, "Software Testing Strategies: Unit Testing: purpose"

In one line

Unit testing checks each smallest testable piece of a program on its own, as soon as it is written, so that a mistake is found where it was made, while it is cheap to fix and before other code is built on it.

In the wording a student can write in an examination: a unit is a "separately testable element specified in the design of a computer software component" (ISO/IEC/IEEE 24765); a unit test is the "testing of individual routines and modules by the developer or an independent tester" (ISO/IEC TR 7052:2023). The ISTQB syllabus calls the level component testing, which "focuses on testing components in isolation" and "is normally performed by developers in their development environments". Its purpose is to find defects in each unit's own logic early and in one known place; to check the unit's interface, local data, boundaries, paths and error handling; and to leave behind tests that can be rerun after every change.

What a unit is

The standards give three angles on the word. A unit is a "separately testable element specified in the design of a computer software component", a "logically separable part of a computer program", and a "software element that is not subdivided into other components or elements" (all ISO/IEC/IEEE 24765). ISO/IEC/IEEE 12207:2026 calls it a software unit: an "atomic-level software component of the software architecture that can be subjected to standalone testing".

The key phrase in every version is separately testable. In Python a unit is usually a function or a class; in Java a class or a method; in a web application a single request handler or service. ExamReg's module tree from earlier chapters gives natural units: the password check, the eligibility check, the paper selection, the fee calculator, the payment gateway adapter, the seat allocation and the PDF generator. Each can be called on its own, given inputs, and checked.

What unit testing is for

Unit testing earns its place in a strategy for four reasons.

  1. It finds defects where they are made. When a unit test fails, the fault is in that unit, a few dozen lines at most. When a system test fails, the fault could be anywhere, as Chapter Thirty-One, on the strategic approach, showed.
  2. It finds them early, when they are cheapest. Boehm and Basili's first finding was that "Finding and fixing a software problem after delivery is often 100 times more expensive than finding and fixing it during the requirements and design phase." They added a qualification that fits a project of ExamReg's size: for "small, noncritical software systems" the factor is "more like 5:1 than 100:1". Either way, the unit test is the earliest dynamic test there is.
  3. It lets the code be changed safely. A unit's tests stay with it, and rerunning them after every change is the cheapest regression testing a project has.
  4. It improves the design. A unit that is hard to test on its own is usually doing too much or depending on too much; writing its tests exposes that, which is why testability is a design property (Chapter Twenty-One, on how quality factors shape testing).
munotes.in183

Unit Testing: What It Is For

The standards stress one more thing: unit tests are normally written by the developer who wrote the unit. The ISTQB syllabus says component testing "is normally performed by developers in their development environments", and ISO/IEC TR 7052 allows "the developer or an independent tester". It is the one level where the author is expected to test their own work, because no one else knows the unit's inside as well, and the tests are needed long before a test team sees the code.

What a unit test checks

A useful checklist for testing one unit has five parts.

AspectThe questionFor ExamReg's fee calculator
InterfaceDoes the unit accept what its callers send, and return what they expect?Takes days late, backlog papers and a concession flag; returns a whole number of rupees
Local dataAre its own variables set, updated and used correctly?The running fee starts at the form fee plus backlog fees and is reduced only by the concession
BoundariesIs every edge right, at and just either side of it?0, 7, 8, 15 and 16 days late; zero backlog papers
PathsIs every branch of the unit taken at least once?With and without a concession; each late-fee band; the refusal
Error handlingIs bad input refused cleanly, with a clear message?Negative days, negative backlog papers, anything that is not a whole number

Each aspect has a chapter of its own later. Boundaries become boundary value analysis (Chapter Fifty-Five); paths become branch testing and, made exact, basis path testing (Chapter Seventy-Two); interfaces between units become integration testing (Chapter Thirty-Seven).

Worked example: unit testing the fee calculator

Here is ExamReg's fee calculator, with the fee rules fixed for this book: a form fee of Rs 800, Rs 150 for each backlog paper, the late-fee bands, forms more than 15 days late or with a negative number of days refused, and a concession that waives the form fee only. The checks below take the five aspects in turn. Each is a line of ordinary Python that calls the unit with chosen inputs and compares the answer with the rule; Chapter Thirty-Five, on writing unit tests with a framework, shows the same checks as a proper test suite.

FORM_FEE, BACKLOG_FEE = 800, 150

def total_fee(days_late, backlog_papers, concession):      # the unit under test
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    fee = FORM_FEE + BACKLOG_FEE * backlog_papers
    if concession:
        fee -= FORM_FEE                                      # a concession waives the form fee
    if days_late == 0:
        late = 0
    elif days_late <= 7:
        late = 100
    else:
        late = 500
    return fee + late

def refused(*args):
    try:
        total_fee(*args)
    except ValueError:
        return True
    return False

checks = {
    "interface: returns a whole number of rupees": isinstance(total_fee(0, 0, False), int),
    "local data: form fee plus two backlog papers": total_fee(0, 2, False) == 1100,
    "boundary: 7 days late pays Rs 100 late fee": total_fee(7, 0, False) == 900,
    "boundary: 8 days late pays Rs 500 late fee": total_fee(8, 0, False) == 1300,
    "boundary: 15 days late is accepted": total_fee(15, 0, False) == 1300,
    "boundary: 16 days late is refused": refused(16, 0, False),
    "path: no concession, on time": total_fee(0, 0, False) == 800,
    "path: concession, 3 days late, 2 backlog papers": total_fee(3, 2, True) == 400,
    "error handling: negative days refused": refused(-1, 0, False),
    "error handling: negative backlog papers refused": refused(0, -2, False),
}
for name, passed in checks.items():
    print("pass" if passed else "FAIL", " ", name)
print(f"{sum(checks.values())} of {len(checks)} checks passed;",
      "total_fee(0, -2, False) returned", total_fee(0, -2, False))
munotes.in184

Unit Testing: What It Is For

pass   interface: returns a whole number of rupees
pass   local data: form fee plus two backlog papers
pass   boundary: 7 days late pays Rs 100 late fee
pass   boundary: 8 days late pays Rs 500 late fee
pass   boundary: 15 days late is accepted
pass   boundary: 16 days late is refused
pass   path: no concession, on time
pass   path: concession, 3 days late, 2 backlog papers
pass   error handling: negative days refused
FAIL   error handling: negative backlog papers refused
9 of 10 checks passed; total_fee(0, -2, False) returned 500

Nine checks pass, and the one that fails is exactly the kind a unit test exists to catch. The calculator validates days late but never validates backlog papers, so a student record corrupted to hold -2 backlog papers produces a bill of Rs 500: the form fee of Rs 800 less Rs 300 for "negative" papers. No end-to-end system test is likely to try that input, because the form's own screen would never offer it; but a unit can be called by any other code, including a data import or a future screen, and it has to defend itself. The fix is one line at the top of the unit, refusing a negative or non-whole number of backlog papers as it already refuses bad days, and the failing check becomes the confirmation test for the fix.

Three things about the example are true of unit testing in general. The tests ran in a fraction of a second, so they can run every time the code is saved. The failure named the unit and the input exactly, so there was nothing to debug but one condition. And the checks stay with the code, so the next person to change total_fee inherits a record of what it must do.

munotes.in185

Unit Testing: What It Is For

What it does not mean

Unit testing is not debugging. The test shows that the unit fails for -2 backlog papers; finding and fixing the missing check is debugging, a separate activity (Chapter Forty-Seven).

Passing unit tests do not show that the units work together. The fee calculator can pass all its tests while the payment adapter misreads the total it returns; that is integration testing's job.

A unit is not always one function. It is the smallest separately testable element, which may be a class or a small module.

Unit tests are not a chore to add at the end. Written as the unit is written, or before it as the next chapters show, they cost least and find most.

Quick revision

  • Unit: a separately testable element of a software component; software unit: an atomic-level component that can be tested standalone.
  • Unit test: testing of individual routines and modules, usually by the developer, in the development environment.
  • Purposes: find defects where they are made and early; enable safe change; improve design and testability.
  • Late fixes are "often 100 times more expensive" than early ones, "more like 5:1" for small, noncritical systems (Boehm and Basili 2001).
  • Checklist: interface, local data, boundaries, paths, error handling.
  • Worked example: nine of ten checks passed; the unit accepted -2 backlog papers and billed Rs 500.

Test yourself

1. What is a unit, and what is unit testing? A unit is the smallest separately testable element of a program, such as a function or a class. Unit testing is the testing of individual units in isolation, normally by the developer in the development environment, to find defects in each unit's own logic.

2. State four purposes of unit testing. To find defects where they are made, in a small known place; to find them early, when fixing is cheapest; to provide tests that can be rerun after every change as regression tests; and to expose design problems, since a unit that is hard to test alone is usually doing too much.

3. What aspects of a unit should its tests check? Its interface (what it accepts and returns), its local data (its own variables), its boundaries (every edge and either side of it), its paths (every branch at least once), and its error handling (bad input refused cleanly).

4. Why are defects found by unit tests cheaper to fix than those found later? Because a failing unit test points to one small unit and one input, so diagnosis is quick, and because the defect has not yet been built upon by other code or shipped to users; Boehm and Basili found fixes after delivery often cost 100 times more than early fixes.

munotes.in186

Unit Testing: What It Is For

5. Who usually writes unit tests, and why? The developer who wrote the unit, because they know its inside best and the tests are needed as soon as the unit exists, long before an independent test team sees the code.

Contents This chapter on its own page

munotes.in187

Chapter Thirty-Four

Unit Testing Techniques: Drivers, Stubs and Test Doubles

Syllabus topic Module 1, "Software Testing Strategies: Unit Testing: purpose, techniques"

In one line

A unit rarely runs on its own: something must call it and it usually calls other things; unit testing techniques supply a driver to call the unit and check its results, and stubs or other test doubles to stand in for everything the unit depends on.

In the wording a student can write in an examination: a test driver is a "software module used to invoke a module under test and, often, provide test inputs, control and monitor execution, and report test results"; a stub is a "skeletal or special-purpose implementation of a software module, used to develop or test a module that calls or is otherwise dependent on it" (both ISO/IEC/IEEE 24765). Together they form the test harness around the unit. Test double is the general term for "any kind of pretend object used in place of a real object for testing purposes", and there are five kinds: dummy, fake, stub, spy and mock (Meszaros's vocabulary, as Fowler gives it).

Why a unit needs a harness

Take ExamReg's fee payment. The unit FeeService.pay looks up the student, works out how many days late the form is from today's date, computes the fee, charges the payment gateway and emails a receipt. Testing it for real would mean a real student database, waiting for the right date, a real payment and a real email. None of that belongs in a unit test: it is slow, it is not repeatable, and the payment would cost real money.

So the unit is tested inside a harness, which the vocabulary calls "scaffolding code written for the purpose of exercising lower-level code when the higher-level code that will ultimately exercise it is not yet available". The harness has two sides.

  • Above the unit, a driver. It plays the part of whatever will eventually call the unit: it sets up the inputs, calls the unit, and checks and reports the result. In modern practice the test itself is the driver.
  • Below the unit, stubs and other doubles. They play the parts of whatever the unit calls: the database, the clock, the gateway, the mailer. They are controlled by the test, so the unit can be exercised completely and repeatably.
A driver calls the unit under test, which calls a fake, a stub, a mock, a spy, and is given a dummy it never calls

Figure 34.1 The unit test environment: a driver above the unit, test doubles below it

The five test doubles

Fowler's essay takes its vocabulary from Gerard Meszaros, who "uses the term Test Double as the generic term for any kind of pretend object used in place of a real object for testing purposes. The name comes from the notion of a Stunt Double in movies." Meszaros then defined five kinds.

DoubleFowler's definitionIn the worked example
Dummy"passed around but never actually used. Usually they are just used to fill parameter lists."The audit log, which pay never touches but the constructor requires
Fake"actually have working implementations, but usually take some shortcut which makes them not suitable for production (an in memory database is a good example)."An in-memory student store
Stub"provide canned answers to calls made during the test, usually not responding at all to anything outside what's programmed in for the test."A clock whose today() always answers 20 November 2026
Spy"stubs that also record some information based on how they were called."A mailer that keeps every message it was asked to send
Mock"objects pre-programmed with expectations which form a specification of the calls they are expected to receive."The payment gateway, which must be charged exactly once, with the right amount
munotes.in188

Unit Testing Techniques: Drivers, Stubs and Test Doubles

The standard vocabulary has an older, looser entry for the last: a mock object is one of the "temporary dummy objects created to aid testing until the real objects become available" (ISO/IEC/IEEE 24765). In everyday speech the words are used loosely; as Fowler puts it, "all sorts of words are used: stub, mock, fake, dummy". Meszaros's names are worth learning because each says what the double does.

State verification and behaviour verification

The doubles fall into two groups by how a test uses them to decide pass or fail. Fowler draws the distinction exactly.

  • State verification "means that we determine whether the exercised method worked correctly by examining the state of the SUT and its collaborators after the method was exercised." (SUT is the system under test.) The test checks results: the amount returned, the messages the spy recorded.
  • Behaviour verification means "we instead check to see if the order made the correct calls on the warehouse", in his example; in ours, whether the fee service made the right call on the gateway. The test checks interactions.

"Of these kinds of doubles, only mocks insist upon behavior verification. The other doubles can, and usually do, use state verification."

Worked example: one unit, five doubles

Python's standard library includes unittest.mock, which, in its documentation's words, "allows you to replace parts of your system under test with mock objects and make assertions about how they have been used." The program uses it for the stub, the mock and the dummy, and writes the fake and the spy by hand, so that what each double is stays visible. The fee rules are the book's fixed ones; 10 days late with one backlog paper is Rs 800 + Rs 150 + Rs 500.

from datetime import date
from unittest.mock import Mock, sentinel

FORM_FEE, BACKLOG_FEE = 800, 150

def total_fee(days_late, backlog_papers, concession):
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    if not isinstance(backlog_papers, int) or backlog_papers < 0:
        raise ValueError("backlog papers must be a whole number, 0 or more")
    fee = FORM_FEE + BACKLOG_FEE * backlog_papers - (FORM_FEE if concession else 0)
    late = 0 if days_late == 0 else (100 if days_late <= 7 else 500)
    return fee + late

class FeeService:                                     # the unit under test
    def __init__(self, students, clock, gateway, mailer, audit_log):
        self.students, self.clock = students, clock
        self.gateway, self.mailer, self.audit_log = gateway, mailer, audit_log

    def pay(self, roll_no, last_date):
        student = self.students.find(roll_no)
        days_late = max(0, (self.clock.today() - last_date).days)
        amount = total_fee(days_late, student["backlog"], student["concession"])
        receipt = self.gateway.charge(roll_no, amount)
        self.mailer.send(student["email"], f"Paid Rs {amount}; receipt {receipt}")
        return amount

    def refund(self, roll_no, amount):                # not exercised by this test
        self.audit_log.write(f"refund {roll_no} {amount}")
        return self.gateway.refund(roll_no, amount)

class FakeStudentStore:                               # FAKE: really works, but only in memory
    def __init__(self):
        self.rows = {}
    def add(self, roll_no, **row):
        self.rows[roll_no] = row
    def find(self, roll_no):
        return self.rows[roll_no]

class MailerSpy:                                      # SPY: records how it was called
    def __init__(self):
        self.sent = []
    def send(self, to, text):
        self.sent.append((to, text))

# the DRIVER: set up the doubles, call the unit, check what happened
students = FakeStudentStore()
students.add("2026CS014", backlog=1, concession=False, email="s014@college.example")
clock = Mock()                                        # STUB: a canned answer for today()
clock.today.return_value = date(2026, 11, 20)
gateway = Mock()                                      # MOCK: the calls it receives are verified
gateway.charge.return_value = "R-1042"
mailer = MailerSpy()
service = FeeService(students, clock, gateway, mailer, audit_log=sentinel.audit_log)  # DUMMY

amount = service.pay("2026CS014", last_date=date(2026, 11, 10))
print("amount returned:", amount)                                 # state verification
gateway.charge.assert_called_once_with("2026CS014", 1450)         # behaviour verification
print("gateway charged:", gateway.charge.call_args)
print("mails sent:", len(mailer.sent), "to", mailer.sent[0][0])
print("receipt text:", mailer.sent[0][1])

class CarelessFeeService(FeeService):                 # a defect: forgets the backlog paper
    def pay(self, roll_no, last_date):
        student = self.students.find(roll_no)
        days_late = max(0, (self.clock.today() - last_date).days)
        amount = total_fee(days_late, 0, student["concession"])
        self.gateway.charge(roll_no, amount)
        return amount

gateway = Mock()
CarelessFeeService(students, clock, gateway, MailerSpy(), sentinel.audit_log).pay(
    "2026CS014", last_date=date(2026, 11, 10))
try:
    gateway.charge.assert_called_once_with("2026CS014", 1450)
except AssertionError as problem:
    print("the mock rejects the careless version:", str(problem).splitlines()[0])
    print("  actually called:", gateway.charge.call_args)
munotes.in189

Unit Testing Techniques: Drivers, Stubs and Test Doubles

amount returned: 1450
gateway charged: call('2026CS014', 1450)
mails sent: 1 to s014@college.example
receipt text: Paid Rs 1450; receipt R-1042
the mock rejects the careless version: expected call not found.
  actually called: call('2026CS014', 1300)

Each double did one job, and together they made a unit test out of something that would otherwise need a database, a calendar, a bank and a mail server.

  • The fake store answered the lookup with a real, working record, held in memory.
  • The stub clock gave the same "today" on every run, so the test is repeatable: the form is 10 days late whenever the test runs, not only on 20 November.
  • The mock gateway was never charged real money, and the test verified the call it received: exactly once, for student 2026CS014, for Rs 1,450. That is behaviour verification.
  • The spy mailer kept the receipt instead of sending it, so the test could read its address and text afterwards. That is state verification on a double.
  • The dummy filled a parameter the test never needed. sentinel supplies a unique object that fits the purpose: if the code ever did use it, the test would fail loudly.
munotes.in190

Unit Testing Techniques: Drivers, Stubs and Test Doubles

The second half of the program shows what behaviour verification is for. A careless version of pay forgets the backlog paper and charges Rs 1,300. The mock's check, assert_called_once_with, reports "expected call not found" and shows the call it actually received, pointing straight at the missing Rs 150.

Choosing doubles well

Doubles are powerful, and Fowler's essay warns about the cost of using them too freely. A test built on mocks checks how a unit talks to its collaborators, so "Mockist tests are thus more coupled to the implementation of a method. Changing the nature of calls to collaborators usually cause a mockist test to break." A few rules of thumb follow.

  • Use the real thing when it is cheap and deterministic. total_fee is a plain function; the test above calls the real one rather than doubling it.
  • Double what is slow, costly, uncontrollable or unfinished: payment, email, the clock, a network service, a unit not yet written.
  • Mock only the calls that matter to the requirement. The gateway charge matters; how many times the student store is read does not, so it is a fake, not a mock.
  • Make time and randomness injectable, as the clock is here. A unit that reads the system clock directly cannot be tested for "10 days late" without waiting ten days; this is the testability point of Chapter Twenty-One, on how quality factors shape testing.

What it does not mean

A stub is not a mock. A stub supplies answers; a mock also checks the calls it receives. The title of Fowler's essay says it: mocks aren't stubs.

Test doubles do not prove the real collaborators work. The mock gateway accepts any call the test allows; whether the real gateway does is a question for integration testing (Chapter Thirty-Seven).

A driver is not only for procedural code. In a test framework, each test method is a driver; the framework runs them all.

More doubles do not make a better test. Each one couples the test to the unit's design, so use the fewest that make the test fast and repeatable.

Quick revision

  • Test driver: invokes the module under test, provides inputs, monitors execution and reports results. Stub: a skeletal implementation standing in for a module the unit calls. Together, the test harness.
  • Test double (Meszaros, via Fowler): any pretend object used in place of a real one for testing.
  • Dummy: passed but never used. Fake: works, with a shortcut (in-memory database). Stub: canned answers. Spy: a stub that records how it was called. Mock: pre-programmed with expectations of the calls it should receive.
  • State verification: check the state after the call. Behaviour verification: check the calls made; only mocks insist on it.
  • unittest.mock: Mock, return_value, assert_called_once_with, call_args, sentinel.
  • Mock-heavy tests are coupled to the implementation: double only what is slow, costly, uncontrollable or unfinished.
munotes.in191

Unit Testing Techniques: Drivers, Stubs and Test Doubles

Test yourself

1. What are a test driver and a stub? Why are they needed in unit testing? A driver is a module that calls the unit under test with test inputs, monitors it and reports results; a stub is a simplified stand-in for a module that the unit calls. They are needed because a unit cannot run alone: its caller and its collaborators may not exist yet, or may be slow, costly or uncontrollable.

2. Name and explain the five kinds of test double. A dummy is passed but never used, only filling a parameter list. A fake has a working but simplified implementation, such as an in-memory database. A stub gives canned answers to the calls made in the test. A spy is a stub that also records how it was called. A mock is pre-programmed with expectations of the calls it should receive, and the test checks them.

3. Distinguish state verification from behaviour verification. State verification checks whether the unit worked by examining the state of the unit and its collaborators after the call, such as the amount returned. Behaviour verification checks whether the unit made the correct calls on its collaborators, such as charging the gateway once with the right amount; only mocks insist on it.

4. In the worked example, why is the clock replaced by a stub? Because the fee depends on how many days late the form is, and a unit that reads the real clock would give a different result on different days. The stub returns a fixed date, so the test is repeatable and can test 10 days late without waiting.

5. What is the risk of using mocks for everything? Mock-based tests check how the unit calls its collaborators, so they are coupled to its implementation: changing the calls, even without changing the behaviour, breaks the tests. Real objects should be used where they are cheap and deterministic.

Contents This chapter on its own page

munotes.in192

Chapter Thirty-Five

Writing Unit Tests With a Framework

Syllabus topic Module 1, "Software Testing Strategies: Unit Testing: purpose, techniques"

In one line

A unit testing framework supplies the machinery every test suite needs, so that the developer writes only the checks: test cases, fixtures for setup and cleanup, suites that group tests, a runner that runs them and reports, and assertions that decide pass or fail.

In the wording a student can write in an examination: the xUnit family of frameworks, which includes Python's unittest and Java's JUnit and TestNG, is built on four concepts, as the unittest documentation defines them. A test fixture "represents the preparation needed to perform one or more tests, and any associated cleanup actions"; a test case "is the individual unit of testing. It checks for a specific response to a particular set of inputs"; a test suite "is a collection of test cases, test suites, or both"; and a test runner "is a component which orchestrates the execution of tests and provides the outcome to the user". Assertions make each check. TestNG, the practical's tool, expresses the same ideas with annotations such as @Test and @BeforeMethod, lets tests belong to groups and carry a priority, and writes an HTML report.

Why a framework

The unit test in Chapter Thirty-Three, on what unit testing is for, was a dictionary of checks and a loop. It worked, but it would not survive a real project. One check that raised an unexpected exception would stop all the rest; every check shared the same objects, so one could disturb another; there was no way to run only some of them; and the report was whatever the loop happened to print.

A framework solves these problems once. The unittest documentation lists what it provides: it "supports test automation, sharing of setup and shutdown code for tests, aggregation of tests into collections, and independence of the tests from the reporting framework." It also tells you where the idea came from: "The unittest unit testing framework was originally inspired by JUnit". TestNG's documentation says the same of itself: it "is a testing framework inspired from JUnit and NUnit". They belong to one family, and what you learn in one carries over to the others.

The parts of an xUnit framework

PartWhat it isIn Python's unittestIn TestNG
Test case"the individual unit of testing"A method whose name starts with test in a subclass of unittest.TestCaseA method marked @Test
Test fixture"the preparation needed to perform one or more tests, and any associated cleanup actions"setUp and tearDown for every test; setUpClass and tearDownClass for a class@BeforeMethod and @AfterMethod; @BeforeClass, @AfterClass; @BeforeSuite, @BeforeTest, @BeforeGroups
Test suite"a collection of test cases, test suites, or both"unittest.TestSuite, or a loader that collects a class's testsA <suite> in testng.xml, with <test> elements and groups
Test runner"orchestrates the execution of tests and provides the outcome to the user"unittest.TextTestRunner, or python -m unittestTestNG's runner, from an IDE, the command line or a build tool
AssertionThe check that decides pass or failassertEqual, assertIn, assertRaises and the other assert methodsJava's assert, or the Assert and AssertJUnit classes
munotes.in193

Writing Unit Tests With a Framework

Two outcomes other than pass and fail are worth knowing. A skipped test is deliberately not run, with a reason recorded. An error (as distinct from a failure) means the test itself crashed with an unexpected exception, rather than an assertion finding a wrong answer.

Worked example: ExamReg's fee service under unittest

The unit is the fee service of Chapter Thirty-Four, on drivers, stubs and test doubles, now in its own module, feeservice.py, as it would be in the project.

FORM_FEE, BACKLOG_FEE = 800, 150

def total_fee(days_late, backlog_papers, concession):
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    if not isinstance(backlog_papers, int) or backlog_papers < 0:
        raise ValueError("backlog papers must be a whole number, 0 or more")
    fee = FORM_FEE + BACKLOG_FEE * backlog_papers - (FORM_FEE if concession else 0)
    late = 0 if days_late == 0 else (100 if days_late <= 7 else 500)
    return fee + late

class FeeService:
    def __init__(self, students, clock, gateway, mailer):
        self.students, self.clock, self.gateway, self.mailer = students, clock, gateway, mailer

    def pay(self, roll_no, last_date):
        student = self.students.find(roll_no)
        days_late = max(0, (self.clock.today() - last_date).days)
        amount = total_fee(days_late, student["backlog"], student["concession"])
        receipt = self.gateway.charge(roll_no, amount)
        self.mailer.send(student["email"], f"Paid Rs {amount}; receipt {receipt}")
        return amount

The test module imports it. The fixture, setUp, builds a fresh fake student store, stub clock, mock gateway and mock mailer before every test, so no test can be disturbed by another. Five test cases check the fee rules and the gateway call; one is the exam cell's new requirement, that the receipt must show the late fee separately; and one is skipped because the hall ticket PDF does not exist yet. The runner's report is printed without its timing line, which differs on every run.

import io
import unittest
from datetime import date
from unittest.mock import Mock
from feeservice import FeeService

class FakeStudentStore:
    def __init__(self, rows):
        self.rows = rows
    def find(self, roll_no):
        return self.rows[roll_no]

class FeeServiceTests(unittest.TestCase):
    def setUp(self):                                # the fixture, built afresh for every test
        self.students = FakeStudentStore({
            "2026CS014": {"backlog": 1, "concession": False, "email": "s014@college.example"},
            "2026CS027": {"backlog": 2, "concession": True, "email": "s027@college.example"}})
        self.clock = Mock()
        self.clock.today.return_value = date(2026, 11, 20)
        self.gateway = Mock()
        self.gateway.charge.return_value = "R-1042"
        self.mailer = Mock()
        self.service = FeeService(self.students, self.clock, self.gateway, self.mailer)

    def test_late_student_pays_form_backlog_and_late_fee(self):
        self.assertEqual(self.service.pay("2026CS014", date(2026, 11, 10)), 1450)

    def test_concession_waives_only_the_form_fee(self):
        self.assertEqual(self.service.pay("2026CS027", date(2026, 11, 17)), 400)

    def test_gateway_is_charged_once_with_the_total(self):
        self.service.pay("2026CS014", date(2026, 11, 10))
        self.gateway.charge.assert_called_once_with("2026CS014", 1450)

    def test_sixteen_days_late_is_refused(self):
        with self.assertRaises(ValueError):
            self.service.pay("2026CS014", date(2026, 11, 4))

    def test_receipt_shows_the_late_fee_separately(self):     # the exam cell's new requirement
        self.service.pay("2026CS014", date(2026, 11, 10))
        to, text = self.mailer.send.call_args.args
        self.assertIn("late fee Rs 500", text)

    @unittest.skip("hall ticket PDF not built yet")
    def test_hall_ticket_link_in_receipt(self):
        pass

def run(suite):
    """Run a suite and print the runner's report, less its timing line."""
    report = io.StringIO()
    unittest.TextTestRunner(stream=report, verbosity=1).run(suite)
    lines = report.getvalue().splitlines()
    print(lines[0])                                      # one mark per test: . F E s
    for line in lines:
        if line.startswith(("FAIL:", "ERROR:", "AssertionError")):
            print(line)
    print(lines[-1])                                     # OK, or FAILED with the counts

print("whole class:")
run(unittest.defaultTestLoader.loadTestsFromTestCase(FeeServiceTests))
print("smoke group, in the order given:")
run(unittest.TestSuite([FeeServiceTests("test_late_student_pays_form_backlog_and_late_fee"),
                        FeeServiceTests("test_gateway_is_charged_once_with_the_total")]))
munotes.in194

Writing Unit Tests With a Framework

whole class:
..s.F.
FAIL: test_receipt_shows_the_late_fee_separately (__main__.FeeServiceTests.test_receipt_shows_the_late_fee_separately)
AssertionError: 'late fee Rs 500' not found in 'Paid Rs 1450; receipt R-1042'
FAILED (failures=1, skipped=1)
smoke group, in the order given:
..
OK

Read the runner's first line as a row of marks, one per test in the order the loader found them, which for unittest is alphabetical: a dot for a pass, F for a failure, E for an error and s for a skip. Four tests passed, one was skipped, and one failed; and the failure did not stop the tests after it, which is the first thing a framework buys. The failure report names the test and says exactly what was wrong: the receipt reads Paid Rs 1450; receipt R-1042 and never mentions the late fee. The test is right and the code is behind the requirement; the test will stay red until the receipt is changed, which is exactly the signal it exists to give.

The second run is a suite built by hand: two chosen tests, run in the order given. That is how unittest does what TestNG calls groups and priorities, a named subset of tests run in a chosen order, and it is how a team would build a quick smoke suite to run on every build.

TestNG, the practical's framework

The practical asks students to "Create and execute automated test cases using TestNG annotations. Group test cases, define priorities, and generate HTML test reports." TestNG does with annotations what unittest does with method names and subclasses. Here is how the first test above looks in TestNG; it is an excerpt, not a whole program, and it has not been compiled here.

import org.testng.annotations.*;
import static org.testng.AssertJUnit.*;

public class FeeServiceTest {
    private FeeService service;

    @BeforeMethod
    public void setUp() { service = new FeeService(students, clock, gateway, mailer); }

    @Test(groups = {"smoke"}, priority = 1)
    public void lateStudentPaysFormBacklogAndLateFee() {
        assertEquals(1450, service.pay("2026CS014", lastDate));
    }
}

The pieces map one to one onto what the Python version did.

  • Annotations mark each method's job. The documentation defines them precisely: @BeforeMethod means "The annotated method will be run before each test method", like unittest's setUp; @BeforeClass "will be run before the first test method in the current class is invoked"; @BeforeSuite "will be run before all tests in this suite have run". As TestNG's documentation says, "it's the annotations that tell TestNG what they are", so test methods can have any names.
  • Groups. "A test method can belong to one or several groups", named in @Test(groups = {...}). A run can then include or exclude whole groups, which is how a smoke group or a slow group is run on its own.
  • Priority. The priority attribute is "The priority for this test method. Lower priorities will be scheduled first." It orders tests where the order matters; the best practice of the next chapter is to write tests that do not depend on order at all.
  • Reports. "The results of the test run are created in a file called index.html in the directory specified when launching SuiteRunner", by default test-output, and that page links to the rest of the HTML report. unittest's own runner, as its documentation says, may report through "a graphical interface, a textual interface, or return a special value"; the one in the standard library is the text runner used above.
munotes.in195

Writing Unit Tests With a Framework

Groups are chosen for a run in the suite file, testng.xml:

<suite name="ExamReg">
  <test name="Smoke">
    <groups><run><include name="smoke"/></run></groups>
    <classes><class name="examreg.FeeServiceTest"/></classes>
  </test>
</suite>

What it does not mean

A framework does not decide what to test. It runs the checks it is given; choosing them is test design, the subject of Module 2.

A skipped test is not a passed test. It records that something is not being checked, with the reason, so it can be found and finished.

A failure and an error are not the same. A failure means an assertion found a wrong answer; an error means the test crashed, which may be a fault in the code or in the test.

Groups and priorities are not a substitute for independent tests. A test that passes only when another runs first is a defect in the test suite.

Quick revision

  • xUnit concepts (unittest documentation): test fixture (preparation and cleanup), test case ("the individual unit of testing"), test suite (a collection of cases and suites), test runner (runs tests and reports); plus assertions.
  • unittest: subclass TestCase; methods named test...; setUp/tearDown; assertEqual, assertIn, assertRaises; @unittest.skip; TestSuite; TextTestRunner.
  • Runner marks: . pass, F failure, E error, s skip; a failure does not stop the other tests.
  • TestNG: annotations (@Test, @BeforeMethod, @AfterMethod, @BeforeClass, @BeforeSuite, @BeforeTest, @BeforeGroups); groups; priority (lower runs first); HTML report in test-output/index.html; groups chosen in testng.xml.
  • Worked example: 6 tests; 4 passed, 1 skipped, 1 failed (the receipt lacks the late fee); a two-test smoke suite passed.
munotes.in196

Writing Unit Tests With a Framework

Test yourself

1. Explain the terms test fixture, test case, test suite and test runner. A test fixture is the preparation needed to perform one or more tests, and the cleanup after them. A test case is the individual unit of testing, checking a specific response to a particular set of inputs. A test suite is a collection of test cases, suites or both, run together. A test runner orchestrates the running of tests and reports the outcome.

2. What does a unit testing framework provide that a hand-written loop of checks does not? Independent tests, so that one failure or crash does not stop the rest; shared setup and cleanup through fixtures; grouping into suites and selective running; standard assertions with clear failure messages; and a runner with a consistent report that a CI server can read.

3. Name four TestNG annotations and state when each method runs. @BeforeSuite runs before all tests in the suite; @BeforeClass before the first test method in the class; @BeforeMethod before each test method; @AfterMethod after each test method. @Test marks a test method.

4. What are groups and priority in TestNG? Groups are named sets that a test method can belong to, one or several, so that a run can include or exclude whole groups such as a smoke group. Priority is an attribute of @Test that orders test methods, with lower priorities scheduled first.

5. Distinguish a failure from an error in a test run. A failure means an assertion was false: the code gave a wrong answer. An error means the test raised an unexpected exception and did not finish, which may be a fault in the code or in the test itself.

Contents This chapter on its own page

munotes.in197

Chapter Thirty-Six

Unit Testing Best Practices

Syllabus topic Module 1, "Software Testing Strategies: Unit Testing: purpose, techniques, and best practices"

In one line

Good unit tests are small, fast, independent and repeatable, each checks one behaviour in the arrange, act, assert shape and is named for it, and they are written before or with the code; coverage tells you what the tests have not reached, not how good they are.

In the wording a student can write in an examination: unit testing best practices are (1) structure each test as arrange, act, assert (Wake); (2) test one behaviour per test and name the test for it; (3) keep tests independent of each other and repeatable, controlling time and randomness; (4) keep them fast, with no databases, networks or files in unit tests; (5) test the edges and invalid input; (6) keep logic out of tests; (7) write tests first, in test-driven development, where "Tests are written first, then the code is written to satisfy the tests, and then the tests and code are refactored" (ISTQB); and (8) treat coverage as a guide, since even full statement coverage "will not detect defects in all cases".

Structure: arrange, act, assert

Bill Wake named the pattern in 2001 (by his own account on the page cited) and describes its three steps:

  • "Arrange: Set up the object to be tested. We may need to surround the object with collaborators. For testing purposes, those collaborators might be test objects (mocks, fakes, etc.) or the real thing."
  • "Act: Act on the object (through some mutator). You may need to give it parameters (again, possibly test objects)."
  • "Assert: Make claims about the object, its collaborators, its parameters, and possibly (rarely!!) global state."

The shape makes a test readable at a glance: what was set up, what was done, what should be true. Wake's warning is about the alternative, a long test that arranges, acts and asserts over and over: "To understand a test like that, you have to track state over a series of activities." His advice is that "Such multi-step unit tests are usually better off being split into several tests."

He also offers a way to start: write the assert first, asking "Suppose it worked; how would I be able to tell?" The answer is the test's reason for existing.

One behaviour per test, named for it

A test should check one behaviour, so that when it fails there is one thing to look at. That is not the same as one assert line. Wake does not follow a strict one-assert rule: when one action changes several aspects of an object, he checks them together, because "the various assertions each explore a different 'dimension' of the object." The rule is one behaviour: "on-time students pay no late fee" is one behaviour, even if checking it takes two lines.

munotes.in198

Unit Testing Best Practices

Name the test for that behaviour. test_receipt_states_the_late_fee_separately reads, when it fails, as the sentence it is: the receipt does not state the late fee separately. test_3 tells the reader nothing.

Independent and repeatable

Independent. Each test should set up everything it needs and depend on no other test having run first. A framework's fixture, rebuilt before every test as in Chapter Thirty-Five, on writing unit tests with a framework, is the tool. Tests that share state pass or fail depending on their order, and a test run that is sometimes red and sometimes green teaches the team to ignore red.

Repeatable. A test should give the same result every time, on every machine. Anything that varies, today's date, a random number, the network, must be controlled. The stub clock of Chapter Thirty-Four, on drivers, stubs and test doubles, is the pattern: the unit asks a clock the test controls, so "10 days late" is 10 days late whenever the test runs.

Fast, and kept to the unit

Unit tests are run many times a day, so they must take seconds, not minutes. Wake draws the boundary: "Unit tests (for the bulk of the system) don't talk to external systems, databases, files, etc., and Arrange-Act-Assert is a pattern for unit tests." That is why the test pyramid from Chapter Seventeen, on agile, Scrum and DevOps, puts many small, fast unit tests at the base and few slow end-to-end tests at the top: a suite that is slow is a suite that is not run.

Test the edges

Most defects in a unit sit at its edges: the boundaries of a range, empty inputs, the first and last items, invalid data. Chapter Thirty-Three's checklist named them, and its worked example found the unit accepting -2 backlog papers. Each boundary and each kind of invalid input deserves its own named test; Chapter Fifty-Five, on boundary value analysis, turns this into a method.

Keep logic out of tests

A test should be straight-line code: arrange, act, assert, with no loops or conditions of its own. A test that computes its expected answer with the same kind of logic as the code it tests can share the code's mistake and pass; a test with an if in it can quietly skip its own assert. Expected values should be written out as constants, worked out by hand from the requirement: Rs 1450, not FORM_FEE + BACKLOG_FEE + 500. Where many similar cases are needed, a table of inputs and expected results, run by the framework, keeps each case visible; Chapter Fifty, on data-driven testing, shows how.

munotes.in199

Unit Testing Best Practices

Test first: test-driven development

The ISTQB syllabus describes test-driven development (TDD) as a practice that "Directs the coding through test cases (instead of extensive software design)", in which "Tests are written first, then the code is written to satisfy the tests, and then the tests and code are refactored." This book labels the three steps red (write a test for the next behaviour and watch it fail), green (write the least code that makes it pass) and refactor (improve the code's structure while every test stays green). The syllabus adds that TDD, like ATDD and BDD, "implements the principle of early testing", since "the tests are defined before the code is written".

Worked example: one TDD cycle, and a coverage trap

Chapter Thirty-Five ended with a failing test: the exam cell now wants each receipt to show the late fee separately. That failing test is the red step of a TDD cycle. The program runs the same three tests, written in the arrange, act, assert shape and named for their behaviours, against the receipt code at each step of the cycle.

Its second part shows why coverage is a guide and not a goal. The exam cell also asks for the fee per backlog paper on the receipt. The new function has two statements and no branch, so either of its two tests executes every statement in it: 100 per cent statement coverage.

import io
import unittest

PARTS = {"form fee": 800, "backlog fee": 150, "late fee": 500}      # student 2026CS014's bill

def receipt_before(parts, number):             # the code Chapter Thirty-Five's test rejected
    return f"Paid Rs {sum(parts.values())}; receipt {number}"

def receipt_green(parts, number):              # GREEN: the least code that passes
    return f"Paid Rs {sum(parts.values())}; late fee Rs {parts['late fee']}; receipt {number}"

def rupees(amount):
    return f"Rs {amount}"

def receipt_refactored(parts, number):         # REFACTOR: the same output, one place per format
    pieces = [f"Paid {rupees(sum(parts.values()))}",
              f"late fee {rupees(parts['late fee'])}",
              f"receipt {number}"]
    return "; ".join(pieces)

class ReceiptTests(unittest.TestCase):
    make_receipt = None                        # set to the version under test before each run

    def test_receipt_states_the_total_paid(self):
        parts, number = PARTS, "R-1042"                    # arrange
        text = self.make_receipt(parts, number)            # act
        self.assertIn("Paid Rs 1450", text)                # assert

    def test_receipt_states_the_late_fee_separately(self):
        parts, number = PARTS, "R-1042"
        text = self.make_receipt(parts, number)
        self.assertIn("late fee Rs 500", text)

    def test_receipt_quotes_the_receipt_number(self):
        parts, number = PARTS, "R-1042"
        text = self.make_receipt(parts, number)
        self.assertIn("receipt R-1042", text)

for stage, version in [("red, before any change", receipt_before),
                       ("green, the least code that passes", receipt_green),
                       ("refactored, same behaviour", receipt_refactored)]:
    ReceiptTests.make_receipt = staticmethod(version)
    report = io.StringIO()
    unittest.TextTestRunner(stream=report).run(
        unittest.defaultTestLoader.loadTestsFromTestCase(ReceiptTests))
    lines = report.getvalue().splitlines()
    print(f"{stage:<34} {lines[0]:<4} {lines[-1]}")

def fee_per_backlog_paper(backlog_fee, papers):    # two statements, no branch
    per_paper = backlog_fee / papers
    return round(per_paper)

print("per-paper tests pass:", fee_per_backlog_paper(300, 2) == 150,
      fee_per_backlog_paper(150, 1) == 150)
try:
    fee_per_backlog_paper(0, 0)
except ZeroDivisionError as problem:
    print("a student with no backlog papers:", type(problem).__name__)
munotes.in200

Unit Testing Best Practices

red, before any change             .F.  FAILED (failures=1)
green, the least code that passes  ...  OK
refactored, same behaviour         ...  OK
per-paper tests pass: True True
a student with no backlog papers: ZeroDivisionError

The cycle reads straight off the first three lines. Red: against the old receipt, the late-fee test fails (the F among the dots) while the other two pass. Green: one changed line makes all three pass, and no more code is written than the test demands. Refactor: the receipt is rebuilt from named pieces with one helper for amounts, a better shape for the next change, and the three tests prove the output is unchanged. The tests that drove the change stay behind as regression tests.

The last two lines are the coverage trap. Both tests pass, and together, or either alone, they execute every statement of fee_per_backlog_paper. Yet a student with no backlog papers, an entirely ordinary case, crashes the receipt with a division by zero. It is precisely the ISTQB syllabus's own example of a defect that full statement coverage can miss: exercising a statement "will not detect defects in all cases. For example, it may not detect defects that are data dependent (e.g., a division by zero that only fails when a denominator is set to zero)." Coverage told the truth, every statement ran; it just answered a different question from is the code right? An edge test, zero papers, was what was missing.

Coverage: a guide, not a goal

Coverage measures what the tests reached, and is useful exactly for that: a line or branch no test has reached is a line or branch no test has checked. What coverage cannot measure is whether the tests check the right things. A team given a coverage target can reach it with tests that execute everything and assert almost nothing. Use coverage to find the untested parts, then design tests for them with the techniques of Module 2; statement and branch coverage have their own chapters there (Chapters Fifty-Nine and Sixty).

A checklist

PracticeThe reason
Arrange, act, assertA reader sees the setup, the action and the claim at a glance
One behaviour per test, named for itA failure points at one behaviour and reads as a sentence
Independent tests, fresh fixture each timeOrder cannot change the result
Repeatable: control time, randomness, external servicesSame result on every run and machine
Fast, with no databases, networks or filesThe suite is run often enough to matter
Test edges and invalid inputThat is where defects cluster in a unit
No logic in tests; expected values as constantsA test cannot share the code's mistake
Test first (TDD): red, green, refactorEvery line of code is wanted by a test, and the tests stay as regression tests
Coverage as a guideIt finds untested code; it does not prove tested code right
munotes.in201

Unit Testing Best Practices

What it does not mean

One behaviour per test does not mean one assert line. Several asserts about one behaviour belong together.

TDD is not writing all the tests first. It is one small test at a time, each followed by the code that satisfies it and a refactoring.

Refactoring is not adding features. It changes the structure and keeps the behaviour, which the tests confirm.

High coverage is not high quality. A suite can execute everything and check little; the worked example reached every statement and missed a crash.

Quick revision

  • Arrange, act, assert (Wake, 2001): set up, act, make claims; split long multi-step tests.
  • One behaviour per test, named for it; several asserts about one behaviour are fine.
  • Independent (fresh fixture, no order dependence) and repeatable (control time, randomness, external services).
  • Fast: unit tests "don't talk to external systems, databases, files" (Wake).
  • Edges and invalid input; no logic in tests; expected values as constants.
  • TDD (ISTQB): tests first, then code to satisfy them, then refactor; implements early testing. Red, green, refactor.
  • Coverage is a guide: 100 per cent statement coverage missed a division by zero (ISTQB's own example).

Test yourself

1. What is the arrange, act, assert pattern? A structure for a unit test in three parts: arrange sets up the object under test and its collaborators, act performs the one action being tested, and assert makes claims about the result. It makes a test readable at a glance, and long tests that repeat the cycle should usually be split.

2. Why should unit tests be independent and repeatable? How is each achieved? So that a result depends only on the code under test, not on the order of tests or on the day, machine or network. Independence comes from building a fresh fixture for each test and sharing no state; repeatability from controlling time, randomness and external services with test doubles.

3. Describe one cycle of test-driven development. Write a test for the next small behaviour and run it to see it fail; write the least code that makes it pass; then refactor the code, improving its structure while all the tests stay green. The tests remain as regression tests.

4. Why is code coverage a guide and not a goal? Because it measures which code the tests executed, not whether they checked the right results. Even 100 per cent statement coverage can miss data-dependent defects, such as a division by zero that fails only when the divisor is zero; coverage should be used to find untested code, which is then tested with proper test design.

munotes.in202

Unit Testing Best Practices

5. Why should a test not contain its own logic? Because a test that computes its expected result with logic like the code's can repeat the code's mistake and pass, and conditions in a test can skip its assertions; expected values should be worked out from the requirement and written as constants.

Contents This chapter on its own page

munotes.in203

Chapter Thirty-Seven

Integration Testing: Why Units That Work Can Fail Together

Syllabus topic Module 1, "Software Testing Strategies: Integration Testing"

In one line

Integration testing checks that units which each work on their own also work together, because the commonest integration defects live not inside any unit but in the agreements between them: what is passed, in what units and format, in what order, and what happens when something goes wrong.

In the wording a student can write in an examination: integration is the "process of combining software components, hardware components, or both into an overall system" (ISO/IEC/IEEE 24765), and integration testing is testing in which components "are combined and tested to evaluate the interaction among them" (ISO/IEC 29110-5-1-2:2025). Its object is the interface, the "point at which two or more logical, physical, or both, system elements or software system elements meet and act on or communicate with each other" (ISO/IEC/IEEE 12207:2026). It finds defects that unit tests cannot: mismatched parameters, wrong units or formats, wrong call order, and mishandled failures between components. It is done incrementally, a few components at a time, rather than big bang, all at once, so that a failure points to the interface just added.

Why units that pass can fail together

A unit test checks a unit against the unit's own specification. It cannot check that two specifications agree. If the fee calculator is designed to return an amount in paise and the payment module is designed to accept rupees, both can be built exactly to their designs, both can pass every unit test, and the system can still charge a student a hundred times too much. The defect is in neither unit. It is in the space between them, which only integration testing looks at.

The 2018 ISTQB syllabus lists the typical defects of component integration testing, and every one of them is a disagreement across an interface:

  • "Incorrect data, missing data, or incorrect data encoding"
  • "Incorrect sequencing or timing of interface calls"
  • "Interface mismatch"
  • "Failures in communication between components"
  • "Unhandled or improperly handled communication failures between components"
  • "Incorrect assumptions about the meaning, units, or boundaries of the data being passed between components"

The same syllabus sets integration testing's focus accordingly: "if integrating module A with module B, tests should focus on the communication between the modules, not the functionality of the individual modules, as that should have been covered during component testing."

Worked example: two correct units, one wrong system

Chapter Two, on errors, faults and failures, listed this defect in a table: the fee page calls the payment module with the amount in paise, and the payment module expects rupees. Here it is as running code. Each unit is written to its own specification and passes the tests written from that specification; then the two are joined.

FORM_FEE, BACKLOG_FEE = 800, 150

def fee_in_paise(days_late, backlog_papers, concession):      # unit A: the fee page's calculator
    """Its design says: return the total fee in PAISE, for the page to display."""
    if days_late < 0 or days_late > 15 or backlog_papers < 0:
        raise ValueError("form not accepted")
    rupees = FORM_FEE + BACKLOG_FEE * backlog_papers - (FORM_FEE if concession else 0)
    rupees += 0 if days_late == 0 else (100 if days_late <= 7 else 500)
    return rupees * 100

class PaymentModule:                                          # unit B
    """Its interface document says: charge() takes the amount in RUPEES."""
    def __init__(self):
        self.ledger = []
    def charge(self, roll_no, amount):
        if amount <= 0:
            raise ValueError("nothing to charge")
        self.ledger.append((roll_no, amount))
        return f"R-{1041 + len(self.ledger)}"

# each unit passes the tests written from its own specification
calculator_ok = fee_in_paise(10, 1, False) == 145000          # Rs 1,450 is 145,000 paise
payments = PaymentModule()
payment_ok = payments.charge("TEST", 1450) == "R-1042" and payments.ledger == [("TEST", 1450)]
print("unit test, fee calculator:", "pass" if calculator_ok else "FAIL")
print("unit test, payment module:", "pass" if payment_ok else "FAIL")

# integration: the fee page hands its total straight to the payment module
payments = PaymentModule()
payments.charge("2026CS014", fee_in_paise(10, 1, False))
charged = payments.ledger[0][1]
print("integration test, fee page to payment:", "pass" if charged == 1450 else "FAIL",
      f"(Rs {charged} charged; the rule says Rs 1450, {charged // 1450} times less)")
munotes.in204

Integration Testing: Why Units That Work Can Fail Together

unit test, fee calculator: pass
unit test, payment module: pass
integration test, fee page to payment: FAIL (Rs 145000 charged; the rule says Rs 1450, 100 times less)

Both unit tests pass, and both are right: the calculator really does return 145,000 paise, and the payment module really does charge whatever number of rupees it is given. The integration test fails because it is the first test that checks the one thing neither unit test could, the agreement between them. A student 10 days late with one backlog paper would be charged Rs 1,45,000 instead of Rs 1,450.

Notice how the fix is chosen. Neither unit is wrong by its own design, so the fault is in the designs: the two sides never agreed on one unit of money. The lasting fix is to write the unit into the interface specification, a "document that specifies the interface characteristics of an existing or planned system or component" (ISO/IEC/IEEE 24765), and change whichever side disagrees with it; the integration test then guards the agreement from then on.

The same defect, in space: the Mars Climate Orbiter

The ExamReg defect is invented; the pattern is not. On 23 September 1999 NASA lost the Mars Climate Orbiter as it arrived at Mars. The Mishap Investigation Board's report found a single root cause: "the failure to use metric units in the coding of a ground software file, 'Small Forces,' used in trajectory models."

munotes.in205

Integration Testing: Why Units That Work Can Fail Together

The failure was an interface defect in the strict sense. The report explains that the output of one program, SM_FORCES, "as required by a MSOP Project Software Interface Specification (SIS) was to be in metric units of Newton-seconds (N-s). Instead, the data was reported in English units of pound-seconds (lbf-s)." The navigation software that read the file assumed the specified units, so it "underestimated the effect on the spacecraft trajectory by a factor of 4.45, which is the required conversion factor from force in pounds to Newtons." Over the nine-month journey the small errors added up, and at arrival "the spacecraft trajectory was approximately 170 kilometers lower than planned."

The board's findings on testing read like a list of what integration testing is for. "The Software Interface Specification (SIS) was developed but not properly used in the small forces ground software development and testing. End-to-end testing to validate the small forces ground software performance and its applicability to the specification did not appear to be accomplished." And: "The interface control process and the verification of specific ground system interfaces was not completed or was completed with insufficient rigor." Each program did its own job; the specification that joined them was written and then not tested against.

ExamReg's fee pageMars Climate Orbiter
ProducerFee calculator, in paiseSM_FORCES, in pound-seconds
ConsumerPayment module, expecting rupeesNavigation software, expecting newton-seconds
What each assumedIts own unitThe unit in the interface specification
Size of the errorA factor of 100A factor of 4.45
The test that would have caught itAn integration test across the fee-to-payment interfaceEnd-to-end testing of the small forces software against the specification

Big bang or incremental

There are two ways to bring units together.

  • Big-bang integration combines all the units at once and tests the whole. It needs no stubs or drivers, but when a test fails, the fault could be in any unit or any interface. The 2018 syllabus warns that "The greater the scope of integration, the more difficult it becomes to isolate defects to a specific component or system, which may lead to increased risk and additional time for troubleshooting."
  • Incremental integration adds a few units at a time and tests after each addition. The syllabus's advice is plain: "In order to simplify defect isolation and detect defects early, integration should normally be incremental (i.e., a small number of additional components or systems at a time) rather than 'big bang' (i.e., integrating all components or systems in one single step)."

On ExamReg the difference is concrete. The portal has seven small units under four subsystems. Integrated all at once, a failing registration could come from any of them and any of the links between them. Integrated one unit at a time, a new failure points to the unit just added and its interfaces, because everything before it was already passing. The price of incremental integration is the scaffolding it needs, stubs for units not yet added or drivers for units not yet called; the next chapter, on top-down, bottom-up and sandwich integration, counts that price for each order.

munotes.in206

Integration Testing: Why Units That Work Can Fail Together

Big bangIncremental
HowAll units combined, then testedA few units added at a time, tested after each
Stubs and driversNoneSome, depending on the order
Locating a faultHard: anywhere in the systemEasier: the newest unit and its interfaces
When defects are foundLate, all togetherEarly, one addition at a time
SuitsVery small systemsAlmost everything else

What integration testing checks

Integration tests are designed from the documents that describe the joins: the architecture, the interface specifications, sequence diagrams and protocols. ISO/IEC/IEEE 24765 defines interface testing as "testing conducted to evaluate whether systems or components pass data and control correctly to one another", and a good set of integration tests for one interface checks each of these:

  • Data: every parameter present, of the right type, in the right order, units and format; values at the boundaries of what the other side accepts.
  • Control: calls made in the right order, the right number of times, with the right results returned.
  • Failures: what the caller does when the called unit refuses, times out or is absent.
  • Shared resources: data both sides read and write, such as a student record, left consistent.

What it does not mean

Integration testing is not repeating the unit tests on a bigger build. Its tests aim at the communication between units, not at each unit's functions.

Passing unit tests do not make integration safe. Both units in the example passed; the system still failed.

An interface defect is not always in one side's code. Often both sides match their own documents, and the documents disagree.

Big bang is not always wrong. For two or three tiny units it can be the quickest route; it becomes costly as the number of interfaces grows.

Quick revision

  • Integration: combining components into an overall system. Integration testing: testing combined components to evaluate their interaction. Interface: where elements meet and act on or communicate with each other.
  • Typical integration defects (ISTQB 2018): incorrect or missing data or encoding; wrong sequencing or timing of calls; interface mismatch; communication failures and their handling; wrong assumptions about "meaning, units, or boundaries".
  • Worked example: calculator (paise) and payment module (rupees) each pass their unit tests; together they charge Rs 1,45,000 for a Rs 1,450 fee.
  • Mars Climate Orbiter (1999): pound-seconds where the interface specification required newton-seconds; trajectory effect underestimated by a factor of 4.45; about 170 km too low; end-to-end testing against the specification not accomplished.
  • Big bang (all at once, hard to locate faults) against incremental (a few at a time, faults found early and located easily; needs stubs and drivers).
munotes.in207

Integration Testing: Why Units That Work Can Fail Together

Test yourself

1. Why can units that pass their unit tests fail when integrated? Because a unit test checks a unit against its own specification, and cannot check that two specifications agree. Defects in the agreements between units, such as the units, format, order or meaning of data passed, show only when the units run together.

2. List four typical defects found by integration testing. Incorrect, missing or wrongly encoded data passed between components; incorrect sequencing or timing of interface calls; interface mismatch, such as parameters of the wrong number, type or order; and wrong assumptions about the meaning, units or boundaries of the data, as in paise passed where rupees were expected.

3. Explain the loss of the Mars Climate Orbiter as an integration failure. One ground program, SM_FORCES, produced thruster data in pound-seconds although the software interface specification required newton-seconds, and the navigation software read it as newton-seconds, underestimating its effect by a factor of 4.45. The trajectory drifted about 170 km low. The investigation found the interface specification was not properly used in testing and end-to-end testing against it was not accomplished.

4. Compare big-bang and incremental integration. Big-bang integration combines all units at once, needs no stubs or drivers, but makes faults hard to locate and finds them late. Incremental integration adds a few units at a time and tests after each, so a failure points to the newest unit and its interfaces, at the cost of writing stubs or drivers.

5. What should an integration test of one interface check? That data passes with the right parameters, types, order, units, format and boundary values; that calls happen in the right order and number with the right results; that failures such as refusals, timeouts or absence are handled; and that shared data is left consistent.

Contents This chapter on its own page

munotes.in208

Chapter Thirty-Eight

Top-Down, Bottom-Up and Sandwich Integration

Syllabus topic Module 1, "Software Testing Strategies: Integration Testing: approaches"

In one line

Incremental integration can follow a program's structure from the top down, replacing stubs with real modules, or from the bottom up, replacing drivers with real callers, or both at once meeting in the middle; the order decides what scaffolding must be written and what can be tested early.

In the wording a student can write in an examination: top-down integration starts "with the highest-level component of a hierarchy and proceeds through progressively lower-levels" (ISO/IEC/IEEE 24765), using stubs for the modules not yet integrated; it can go depth first (one branch at a time) or breadth first (one level at a time). Bottom-up integration starts "with the lowest-level components of a hierarchy and proceeds through progressively higher-levels", using drivers in place of the callers not yet integrated. Sandwich integration combines them: top-down for the upper layers and bottom-up for the lower ones, meeting at a middle layer. All three are alternatives to big-bang testing, in which elements "are combined all at once into an overall system, rather than in stages".

The structure being integrated

Every incremental order follows some structure, and for most programs the natural one is the calling hierarchy: which module calls which. ExamReg's is small enough to see whole. The portal's top module calls four subsystems, and each subsystem calls its units.

ExamReg's module tree: the portal, four subsystems (Login, Exam Form, Fee Payment, Hall Ticket) and seven units beneath them

Figure 38.1 ExamReg's calling hierarchy, the structure that top-down and bottom-up integration follow

NIST's report on structured testing lists the choices a project has: "Integration may be performed all at once, top-down, bottom-up, critical piece first, or by first integrating functional subsystems and then integrating the subsystems in separate phases using any of the basic strategies." It also gives the reason for not doing it all at once on any but the smallest system: "the system would fail in so many places at once that the debugging and retesting effort would be impractical".

Top-down integration

Top-down integration begins with the top module, the one the user meets first, and works downwards. Whatever the top module calls that is not yet integrated is replaced by a stub, a "skeletal or special-purpose implementation of a software module, used to develop or test a module that calls or is otherwise dependent on it" (ISO/IEC/IEEE 24765). As each real module arrives, it replaces its stub, and its own callees become stubs in turn.

There are two ways down the tree.

  • Depth first: follow one branch to the bottom before starting the next. On ExamReg, Login and its password check are finished before the exam form begins. The first complete feature works early.
  • Breadth first: integrate a whole level before going deeper. All four subsystems are in place, with stubbed units beneath, before any unit is added. The whole control structure works early.
munotes.in209

Top-Down, Bottom-Up and Sandwich Integration

Top-down's strength is that the program's main control flow, and something a user can see, exist from the first day. Its weakness is that the modules that do the real computation, the fee calculator and the seat allocation, arrive last, so for most of the integration their results are faked by stubs, and a stub that returns a fixed fee proves nothing about fees.

Bottom-up integration

Bottom-up integration begins with the units at the bottom, which call nothing, and works upwards. A unit whose caller is not yet integrated is exercised by a driver, a "software module used to invoke a module under test and, often, provide test inputs, control and monitor execution, and report test results" (ISO/IEC/IEEE 24765). Units are usually combined in clusters, the units one caller needs, so one driver can stand in for their caller. When the real caller arrives it replaces the driver.

Bottom-up's strength is the mirror of top-down's weakness: the computational units, the fee calculator above all, are tested with real code and real inputs early. Its weakness is that there is no program a user could recognise until the last module, the top, is integrated, and defects in the overall control flow are found last.

Sandwich integration

The two directions can be combined. Top-down integration proceeds from the top, bottom-up integration from the units, and the two meet at a chosen middle layer, which is integrated last and replaces both the stubs above it and the drivers below it. This combination is called sandwich integration, because the middle layer is tested last, between the two.

It is the practical choice on many projects, for the reason Brian Randell gave at the 1968 NATO conference about the two approaches in design: "Clearly the blind application of just one of these approaches would be quite foolish." Teams can work at both ends at once, the user-facing control flow and the computational units are both tested early, and neither needs scaffolding for the whole tree.

Worked example: orders and scaffolding for ExamReg

The program walks ExamReg's tree in each order and counts the scaffolding each needs, using one rule: when a module is integrated, each module it calls that is not yet present needs a stub (one stub per called module); and if the module that calls it is not yet present, that caller needs a driver (one driver per caller, shared by all the units it calls).

TREE = {"ExamReg": ["Login", "Exam Form", "Fee Payment", "Hall Ticket"],
        "Login": ["Password check"],
        "Exam Form": ["Eligibility check", "Paper selection"],
        "Fee Payment": ["Fee calculator", "Gateway adapter"],
        "Hall Ticket": ["Seat allocation", "PDF generator"]}
ROOT = "ExamReg"
parent = {child: p for p, children in TREE.items() for child in children}

def depth_first(m):                      # a module, then each subtree in turn
    return [m] + [x for c in TREE.get(m, []) for x in depth_first(c)]

def breadth_first():                     # level by level
    order, level = [], [ROOT]
    while level:
        order += level
        level = [c for m in level for c in TREE.get(m, [])]
    return order

def bottom_up(m):                        # every subtree before the module that calls it
    return [x for c in TREE.get(m, []) for x in bottom_up(c)] + [m]

def scaffolding(order):
    """Integrate in this order. A called module not yet present needs a STUB (one per module);
    a calling module not yet present needs a DRIVER (one per caller, shared by its callees)."""
    done, stubs, drivers = set(), set(), set()
    for m in order:
        stubs.update(c for c in TREE.get(m, []) if c not in done)
        if m in parent and parent[m] not in done:
            drivers.add(parent[m])
        done.add(m)
    return len(stubs), len(drivers)

leaves = [m for m in depth_first(ROOT) if m not in TREE]
middle = TREE[ROOT]
strategies = {"top-down, depth first": depth_first(ROOT),
              "top-down, breadth first": breadth_first(),
              "bottom-up": bottom_up(ROOT),
              "sandwich (top down to, and bottom up to, the middle layer)": [ROOT] + leaves + middle}
for name, order in strategies.items():
    stubs, drivers = scaffolding(order)
    print(f"{name}: {stubs} stubs, {drivers} drivers")
    line = "  "
    for i, m in enumerate(order):
        piece = m if i == 0 else " > " + m
        if len(line) + len(piece) > 86:              # wrap between modules, never inside a name
            print(line)
            line, piece = "    ", "> " + m
        line += piece
    print(line)
munotes.in210

Top-Down, Bottom-Up and Sandwich Integration

top-down, depth first: 11 stubs, 0 drivers
  ExamReg > Login > Password check > Exam Form > Eligibility check > Paper selection
    > Fee Payment > Fee calculator > Gateway adapter > Hall Ticket > Seat allocation
    > PDF generator
top-down, breadth first: 11 stubs, 0 drivers
  ExamReg > Login > Exam Form > Fee Payment > Hall Ticket > Password check
    > Eligibility check > Paper selection > Fee calculator > Gateway adapter
    > Seat allocation > PDF generator
bottom-up: 0 stubs, 5 drivers
  Password check > Login > Eligibility check > Paper selection > Exam Form
    > Fee calculator > Gateway adapter > Fee Payment > Seat allocation > PDF generator
    > Hall Ticket > ExamReg
sandwich (top down to, and bottom up to, the middle layer): 4 stubs, 4 drivers
  ExamReg > Password check > Eligibility check > Paper selection > Fee calculator
    > Gateway adapter > Seat allocation > PDF generator > Login > Exam Form
    > Fee Payment > Hall Ticket

The counts say what each order costs.

  • Top-down needs a stub for every module except the top: 11 of them, and no drivers, whichever way it goes down. Depth first has the login feature working after three steps; breadth first has all four subsystems wired together after five.
  • Bottom-up needs no stubs, and only 5 drivers, one for each module that calls others, because one driver can call all the units in its cluster. The fee calculator, the unit that decides money, is tested with its real code at step 6, where top-down depth first reaches it only at step 8, and breadth first at step 9, with a stub standing in for it until then.
  • Sandwich needs 4 stubs (for the four subsystems, beneath the real top module) and 4 drivers (standing in for the same four subsystems, above the real units), and the middle layer, integrated last, replaces all eight at once.
munotes.in211

Top-Down, Bottom-Up and Sandwich Integration

Stubs are usually the dearer kind. A driver only has to call a unit and check what comes back; a stub has to imitate what a missing module would have done, well enough for its caller's tests to mean something. A stub for the fee calculator that returns Rs 800 for every student lets the exam form's tests pass without saying anything about fees.

The approaches compared

Top-downBottom-upSandwichBig bang
Starts withThe top moduleThe lowest unitsBoth endsEverything
ScaffoldingStubs for called modulesDrivers for calling modulesBoth, for the middle layer onlyNone
ExamReg's count11 stubs5 drivers4 stubs and 4 driversNone
Tested earlyMain control flow; a visible skeletonComputational units with real codeBoth the control flow and the unitsNothing, until all is ready
Tested lateLow-level computation (stubbed until last)The overall control flow and the user's viewThe middle layerEverything
Locating a faultEasy: the newest moduleEasy: the newest moduleEasy, except at the final middle stepHard: anywhere
SuitsSystems where the user-facing flow is the main riskSystems whose risk is in low-level computation or hardwareLarger systems with teams working in parallelVery small systems

Choosing an order

The order should follow the risk. If ExamReg's main risk is that students cannot find their way through the portal, top-down shows the flow early; if it is that fees are wrong, bottom-up tests the calculator first with real code. The 2018 ISTQB syllabus adds a practical point: "If integration tests and the integration strategy are planned before components or systems are built, those components or systems can be built in the order required for most efficient testing." And whatever the order, it keeps the rule from the last chapter: integrate a few modules at a time, and test after each.

What it does not mean

Top-down does not mean the top module is untested until the end. It is tested first, with stubs beneath it; what arrives last is the real computation beneath.

munotes.in212

Top-Down, Bottom-Up and Sandwich Integration

Bottom-up does not need a driver for every unit. One driver standing in for a caller can exercise all the units that caller uses.

The scaffolding is not free. Stubs and drivers are code that must be written, reviewed and maintained, and a poor stub can make a test meaningless.

These orders are not only for trees. Real calling structures share modules and have cycles; the same idea, integrating along the calls with stubs or drivers at the edge, still applies.

Quick revision

  • Top-down: from the highest-level component downwards; stubs replace modules not yet integrated; depth first (one branch at a time) or breadth first (one level at a time).
  • Bottom-up: from the lowest-level components upwards; drivers replace callers not yet integrated; units combined in clusters.
  • Sandwich: top-down to a middle layer and bottom-up to it; the middle layer last.
  • Big-bang testing: all elements combined at once; faults hard to locate.
  • ExamReg: top-down 11 stubs, no drivers; bottom-up no stubs, 5 drivers; sandwich 4 stubs and 4 drivers.
  • Top-down shows the control flow early; bottom-up tests real computation early; choose the order by risk.

Test yourself

1. Explain top-down integration, with its depth-first and breadth-first forms. Integration begins with the top module and moves down the calling hierarchy, with stubs standing in for the modules not yet integrated, each replaced by the real module in turn. Depth first completes one branch before the next, so a whole feature works early; breadth first completes one level before going deeper, so the whole control structure works early.

2. Explain bottom-up integration. What does it need instead of stubs? Integration begins with the lowest-level units, which call nothing, combined in clusters and moved up the hierarchy. It needs drivers, which stand in for the callers not yet integrated, calling the units and checking their results; each is replaced when the real caller arrives.

3. Compare top-down and bottom-up integration. Top-down tests the main control flow and gives a visible skeleton early but needs stubs and tests the real computation late. Bottom-up tests the computational units with real code early and needs only drivers, but gives no working program until the end and tests the control flow last.

4. What is sandwich integration and why is it used? A combination in which the upper layers are integrated top-down and the lower layers bottom-up, meeting at a middle layer that is integrated last. It lets teams work at both ends at once, tests both the control flow and the computational units early, and needs scaffolding only around the middle layer.

munotes.in213

Top-Down, Bottom-Up and Sandwich Integration

5. How many stubs does top-down integration of ExamReg need, and why? Eleven: one for every module except the top, because each module is called by a module already integrated before it arrives, so each must first be stood in for by a stub.

Contents This chapter on its own page

munotes.in214

Chapter Thirty-Nine

Regression Testing, Smoke Testing and Continuous Integration

Syllabus topic Module 1, "Software Testing Strategies: Integration Testing: approaches"

In one line

During integration, every change is built and tested at once: a quick smoke test decides whether the build is fit to test at all, a regression suite checks that nothing that used to work has broken, and continuous integration runs both automatically on every change, so a defect is found within minutes of the change that caused it.

In the wording a student can write in an examination: regression testing is "testing performed following modifications to a test item or to its operational environment, to identify whether failures in unmodified parts of the test item occur" (ISO/IEC/IEEE 29119-1:2022). A smoke test is, in Steve McConnell's words, "a relatively simple check to see whether the product 'smokes' when it runs": a short test that "should exercise the entire system from end to end" and decide whether the build is stable enough to test further. Continuous integration is a "technique that continually merges artifacts, including source code updates from all developers on a team, into a shared mainline to build and test the developed system" (IEEE 2675-2021). Jenkins, the practical's tool, runs it: a Pipeline, written in a Jenkinsfile, builds and tests every change in stages and keeps the history of every build.

Integration never stops

Chapters Thirty-Seven and Thirty-Eight treated integration as a sequence: add a unit, test, add the next. On a real project the sequence never ends. Developers change units every day, each change is integrated with everyone else's, and each integration can break something that worked the day before. Two kinds of testing keep that process under control. Smoke testing asks, quickly, whether today's build works at all. Regression testing asks, thoroughly, whether it still does everything it did yesterday.

The daily build and smoke test

In 1996 Steve McConnell described a practice he reported as common at Microsoft and some other shrink-wrap software companies: "Every file is compiled, linked, and combined into an executable program every day, and the program is then put through a 'smoke test,' a relatively simple check to see whether the product 'smokes' when it runs."

He gave four benefits. The first: "It minimizes integration risk." The second, that it reduces the risk of low quality: "You bring the system to a known, good state, and then you keep it there." The third is the one testers value most: "It supports easier defect diagnosis. When the product is built and tested every day, it's easy to pinpoint why the product is broken on any given day. If the product worked on Day 17 and is broken on Day 18, something that happened between the two builds broke the product." The fourth: it improves morale.

His rules for the smoke test itself are worth learning as written.

munotes.in215

Regression Testing, Smoke Testing and Continuous Integration

  • It "should exercise the entire system from end to end. It does not have to be exhaustive, but it should be capable of exposing major problems."
  • It "should be thorough enough that if the build passes, you can assume that it is stable enough to be tested more thoroughly."
  • "The smoke test must evolve as the system evolves."
  • And without it the practice is worthless: "The daily build has little value without the smoke test."

A build that fails the smoke test is broken, and McConnell's rule is that "fixing it becomes top priority."

Smoke testing and regression testing compared

Smoke testRegression test
QuestionIs this build working at all, and worth testing further?Does everything that worked before still work?
SizeA few end-to-end casesMany cases, growing with every release
WhenFirst, on every buildAfter the smoke test passes, on every build or change
TimeMinutesAs long as the suite needs; kept fast by automation
A failure meansThe build is broken: stop and fix or revertA specific behaviour has regressed: find the change that did it

Continuous integration

McConnell's daily build is the ancestor of continuous integration (CI), which shortens the cycle from a day to every change. Martin Fowler's article defines the practice: "each member of a team merges their changes into a codebase together with their colleagues changes at least daily. Each of these integrations is verified by an automated build (including test) to detect integration errors as quickly as possible."

His article sets out the practices that make it work. Among them:

  • Put everything in a version controlled mainline, and automate the build.
  • Make the build self-testing: the automated build runs the tests.
  • Everyone pushes commits to the mainline every day, and every push to mainline triggers a build.
  • Fix broken builds immediately. He quotes Kent Beck: "nobody has a higher priority task than fixing the build", and gives the usual remedy: "Usually the best way to fix the build is to revert the latest commit from the mainline, taking the system back to the last-known good build."
  • Keep the build fast, because "The whole point of Continuous Integration is to provide rapid feedback."
  • Everyone can see what's happening.

The 2018 ISTQB syllabus connects this back to integration testing: CI "has become common practice", and "Such continuous integration often includes automated regression testing, ideally at multiple test levels."

Worked example: ExamReg's build history

The program plays the part of a CI server for six commits to ExamReg's late-fee code. Every commit is built and smoke tested first; only if the smoke test passes is the regression suite, every value of days late from 0 to 16, run against the rule. The last good build is remembered, which is what a build history is for.

munotes.in216

Regression Testing, Smoke Testing and Continuous Integration

def rule(days):                                       # the late-fee rule every build must keep
    if days < 0 or days > 15:
        return "refused"
    return 0 if days == 0 else (100 if days <= 7 else 500)

def late_fee_41(days):                                # commit 41: the fee bands
    if days < 0 or days > 15:
        return "refused"
    return 0 if days == 0 else (100 if days <= 7 else 500)

def late_fee_43(days):                                # commit 43: a "tidy-up" of the bands
    if days < 0 or days > 15:
        return "refused"
    return 0 if days == 0 else (100 if days <= 8 else 500)

def late_fee_45(days):                                # commit 45: a constant renamed, one use missed
    if days < 0 or days > 15:
        return "refused"
    return 0 if days == 0 else (LATE_BAND_ONE if days <= 7 else 500)

history = [(41, "fee bands", late_fee_41), (42, "concession screen", late_fee_41),
           (43, "tidy fee bands", late_fee_43), (44, "revert 43", late_fee_41),
           (45, "rename fee constant", late_fee_45), (46, "finish the rename", late_fee_41)]

def smoke(fee):                                       # a few end-to-end cases, run first
    try:
        return fee(0) == 0 and fee(3) == 100
    except Exception:
        return False

def regression(fee):                                  # every value of days late, 0 to 16
    return [d for d in range(0, 17) if fee(d) != rule(d)]

last_good = None
for build, message, fee in history:
    if not smoke(fee):
        print(f"#{build} {message:<20} smoke FAIL   BROKEN: fix or revert before anything else")
        continue
    failed = regression(fee)
    if failed:
        days = ", ".join(map(str, failed))
        print(f"#{build} {message:<20} smoke pass   regression FAIL on day {days}:"
              f" look at what changed after #{last_good}")
    else:
        print(f"#{build} {message:<20} smoke pass   regression 17 of 17   good")
        last_good = build
#41 fee bands            smoke pass   regression 17 of 17   good
#42 concession screen    smoke pass   regression 17 of 17   good
#43 tidy fee bands       smoke pass   regression FAIL on day 8: look at what changed after #42
#44 revert 43            smoke pass   regression 17 of 17   good
#45 rename fee constant  smoke FAIL   BROKEN: fix or revert before anything else
#46 finish the rename    smoke pass   regression 17 of 17   good

Each line is what a team would see in its build history.

  • Builds 41 and 42 pass both stages. Commit 42 changed a different screen, and the regression suite proves it did not disturb the fees.
  • Build 43 passes the smoke test, because the portal works and the common cases are right, but the regression suite catches the "tidied" band: 8 days late now costs Rs 100. The history says where to look: build 42 was good, so the defect is in commit 43. McConnell's Day 17 and Day 18, shrunk to one commit.
  • Build 44 reverts commit 43, Fowler's standard remedy, and the mainline is green again while the tidy-up is redone properly.
  • Build 45 fails the smoke test: a renamed constant was missed in one place, and the fee page crashes for any late form. The build is broken, and nothing else should be tested or merged until it is fixed.
  • Build 46 finishes the rename and is green.
munotes.in217

Regression Testing, Smoke Testing and Continuous Integration

Notice what the two kinds of test caught. The smoke test caught the crash in seconds but would never have noticed the day-8 band; the regression suite caught the band but would have been wasted on a build that could not run. Each is there for what the other misses.

Jenkins: continuous integration in the practical

The practical asks students to "Configure Jenkins to execute Selenium automation scripts automatically. Generate build reports and analyze test execution history." Jenkins's documentation describes the parts.

  • Pipeline: "a suite of plugins which supports implementing and integrating continuous delivery pipelines into Jenkins"; a Pipeline's code "defines your entire build process, which typically includes stages for building an application, testing it and then delivering it."
  • Jenkinsfile: the Pipeline "is written into a text file (called a Jenkinsfile) which in turn can be committed to a project's source control repository", so the build process is "versioned and reviewed like any other code."
  • Stage: "a conceptually distinct subset of tasks performed through the entire Pipeline (e.g. 'Build', 'Test' and 'Deploy' stages)".
  • Step: "A single task", such as sh, which runs a shell command, or junit, a step "for aggregating test reports".

A Jenkinsfile for the build history above, in the declarative syntax the documentation's example uses, looks like this. The make targets are placeholders for whatever commands run the project's build, smoke tests and regression tests; for the practical they would start the Selenium suite.

pipeline {
    agent any
    stages {
        stage('Build') {
            steps { sh 'make' }
        }
        stage('Smoke test') {
            steps { sh 'make smoke' }
        }
        stage('Regression tests') {
            steps {
                sh 'make check'
                junit 'reports/**/*.xml'
            }
        }
    }
}

Each push runs the stages in order; a failing stage stops the run and marks the build, and the junit step turns the test results into reports Jenkins can show. The list of builds, with each one's result, is the build history: the same record the worked example printed, which is how a team reads "look at what changed after #42" off a screen.

What it does not mean

A smoke test is not a quick version of the whole suite. It is a deliberately small end-to-end check of whether the build is fit to test, and it grows only as the system grows.

munotes.in218

Regression Testing, Smoke Testing and Continuous Integration

Regression testing is not retesting the fix. That is confirmation testing; regression testing checks everything else the change might have disturbed.

Continuous integration is not a tool. It is a practice of integrating and testing every change; Jenkins is one tool that automates it.

A red build is not a disaster. It is the system working: a defect found minutes after it was made, with the change that caused it already known.

Quick revision

  • Regression testing: after a change, check that unmodified parts still work; grows with every release, so it is automated.
  • Smoke test (McConnell 1996): a simple end-to-end check of whether the build "smokes"; passing means it is stable enough to test more thoroughly; failing means the build is broken and fixing it is top priority.
  • Daily build benefits: less integration risk, less risk of low quality, easier defect diagnosis (Day 17 good, Day 18 broken), better morale.
  • Continuous integration (Fowler; IEEE 2675-2021): everyone merges at least daily, every push triggers an automated self-testing build, broken builds fixed immediately (often by revert), builds kept fast, results visible.
  • Jenkins: Pipeline, Jenkinsfile in source control, stages, steps (sh, junit), build history.
  • Worked example: #43 passed smoke but failed regression on day 8; #45 failed smoke; the last good build points at the culprit commit.

Test yourself

1. Distinguish smoke testing from regression testing. A smoke test is a small, quick, end-to-end check run first on every build to decide whether the build works well enough to test further; failing it means the build is broken. Regression testing runs a large and growing suite, after the smoke test passes, to check that changes have not broken anything that used to work.

2. What is the daily build and smoke test, and what are its benefits? Every day the whole program is compiled, linked and combined, then given a smoke test. It minimises integration risk, reduces the risk of low quality by keeping the system in a known good state, makes defects easy to diagnose because a failure must come from the day's changes, and improves morale.

3. Define continuous integration and list four of its practices. A practice in which every team member merges changes into a shared mainline at least daily, and each integration is verified by an automated build with tests. Practices: keep everything in version-controlled mainline; automate the build and make it self-testing; trigger a build on every push; fix broken builds immediately; keep the build fast; let everyone see the results.

4. How does a build history help locate a defect? It records the result of every build. If build 42 passed and build 43 failed, the defect came in with commit 43, so the search is narrowed to one change, which can be reverted while it is fixed.

munotes.in219

Regression Testing, Smoke Testing and Continuous Integration

5. What are a Jenkins Pipeline, a Jenkinsfile, a stage and a step? A Pipeline is Jenkins's model of the whole build, test and delivery process; a Jenkinsfile is the text file, kept in source control, in which the Pipeline is written; a stage is a distinct group of tasks such as Build or Test; and a step is a single task within a stage, such as running a shell command or collecting test reports.

Contents This chapter on its own page

munotes.in220

Chapter Forty

The Challenges of Integration Testing

Syllabus topic Module 1, "Software Testing Strategies: Integration Testing: approaches and challenges"

In one line

Integration testing is harder than unit testing because its object, the interface, is shared by people, teams and even companies: specifications are incomplete or not followed, scaffolding costs effort and can lie, faults are hard to locate, environments differ from production, outside services change without notice, and data and timing problems appear only when parts run together.

In the wording a student can write in an examination: the main challenges of integration testing are (1) interface specifications that are incomplete, ambiguous or not followed; (2) the cost and fidelity of stubs and drivers; (3) locating the fault when a combination fails; (4) test environments and configuration that differ from production; (5) third-party and external services that the team does not control; (6) data and timing problems, including shared data and concurrency; and (7) ownership of interfaces that fall between teams. Each has a known answer: reviewed interface specifications with tests written from them, doubles checked against the real thing, incremental and continuous integration, production-like environments, sandboxes and recorded replies from providers, realistic data and time control, and clear end-to-end ownership.

1. Interfaces specified badly, or not followed

An integration test can only check an interface against what was agreed, and on many projects the agreement is thin: a meeting, an email, a comment in the code. Worse, a good specification may exist and not be used. The Mars Climate Orbiter board found exactly that: "The Software Interface Specification (SIS) was developed but not properly used in the small forces ground software development and testing", and "The interface control process and the verification of specific ground system interfaces was not completed or was completed with insufficient rigor."

The answer. Write the interface down, with its units, formats, ranges, error codes and call order; review it with both sides (Chapter Twenty-Eight's reviews apply to interface documents as much as to requirements); and write integration tests from it, so that the document and the tests cannot drift apart. The board's own recommendation was a system verification matrix covering "all project requirements including all Interface Control Documents (ICDs)".

2. Stubs and drivers cost effort, and can lie

Incremental integration needs scaffolding, and Chapter Thirty-Eight counted it: eleven stubs for top-down integration of ExamReg. Every stub is code to write and maintain. A subtler cost is fidelity. A stub is written from someone's belief about how the real module behaves, and if that belief is wrong, or the real module changes, every test that uses the stub passes while the real integration fails.

Worked example: a stub that went out of date. ExamReg's team wrote a stub of the payment gateway from the provider's first version, where a reply carried a status and a receipt. The provider has since upgraded the service, and its replies now carry a state and a receipt_no. The gateway and its field names are invented for this example; the pattern is common wherever a team depends on a service it does not own.

munotes.in221

The Challenges of Integration Testing

class GatewayStub:                            # written by the team against the gateway's version 1
    def charge(self, roll_no, rupees):
        return {"status": "ok", "receipt": "R-TEST"}

class GatewayVersion2:                        # the real service after its provider's upgrade
    def charge(self, roll_no, rupees):
        return {"state": "SUCCESS", "receipt_no": "R-2208"}

def pay(gateway, roll_no, rupees):            # ExamReg's side of the interface
    reply = gateway.charge(roll_no, rupees)
    if reply["status"] != "ok":
        raise RuntimeError("payment refused")
    return reply["receipt"]

print("integration test with the stub: receipt", pay(GatewayStub(), "2026CS014", 1450), "(passes)")

# a check that the stub still looks like the real thing: compare the fields of its reply
# with a reply recorded from the provider's current sandbox
recorded_reply = {"state": "SUCCESS", "receipt_no": "R-2208"}
stub_reply = GatewayStub().charge("2026CS014", 1450)
print("stub replies with:     ", ", ".join(sorted(stub_reply)))
print("provider replies with: ", ", ".join(sorted(recorded_reply)))
print("stub still faithful:", set(stub_reply) == set(recorded_reply))

try:
    pay(GatewayVersion2(), "2026CS014", 1450)
except KeyError as missing:
    print("against the real gateway in staging: KeyError", missing)
integration test with the stub: receipt R-TEST (passes)
stub replies with:      receipt, status
provider replies with:  receipt_no, state
stub still faithful: False
against the real gateway in staging: KeyError 'status'

The integration test with the stub passes, and would go on passing forever, because the stub still answers the way the old gateway did. The second check is the answer to the problem: compare the stub against a reply recorded from the real service. Its fields no longer match, which tells the team to update the stub, and the code, before the real gateway is ever called. The last line shows what would happen otherwise: against the real, upgraded gateway, ExamReg's payment code crashes looking for a status that no longer exists.

The answer. Keep doubles as simple as the tests allow; check them regularly against the real service, as above; and run a smaller set of integration tests against the real thing in a test environment, so that a lying stub is caught before production does it.

3. Locating the fault

When two units fail together, the fault may be in either, in the interface between them, or in a third unit that fed one of them bad data. The 2018 ISTQB syllabus names the difficulty: "The greater the scope of integration, the more difficult it becomes to isolate defects to a specific component or system, which may lead to increased risk and additional time for troubleshooting."

The answer. Integrate incrementally, as Chapter Thirty-Seven on integration testing recommends, so that a new failure points to the newest addition; integrate continuously, with the continuous integration of Chapter Thirty-Nine, so that the build history names the change that broke things; and log what crosses each interface, so that the data at the moment of failure can be read afterwards.

munotes.in222

The Challenges of Integration Testing

4. Environments and configuration

Integration tests run in a test environment, and every difference between it and production is a place for a defect to hide. Fowler puts it directly: "If we test in a different environment, every difference results in a risk that what happens under test won't happen in production." The ISTQB syllabus says the same of system integration testing, which "requires suitable test environments preferably similar to the operational environment."

Configuration is part of the environment. Knight Capital's loss in 2012 (Chapter Three, on why software must be tested) began when new code reached seven of eight servers: a deployment and configuration difference that no test of the code alone could catch.

The answer. Build test environments to mimic production as closely as possible. Fowler's list is specific: "Use the same database software, with the same versions, use the same version of the operating system", and he judges that "the price is usually small compared to hunting down a single bug that crawled out of the hole created by environment mismatches." Keep configuration in version control, deploy to test and production by the same automated steps, and check after deployment that every server received the same build.

5. Services you do not control

ExamReg depends on a payment gateway it does not own, and many systems depend on dozens of such services. Their owners change them, rate-limit test calls, and rarely provide a test copy that behaves exactly like the live one. NIST's 2002 study named interoperability as one of the main inadequacies of testing: "The integration of applications is a difficult and uncertain process", and it posed the underlying question: "if application A and application B interoperate and if application B and application C interoperate, what are the prospects of applications A and C interoperating?"

The answer. Use the provider's sandbox for system integration testing; record real replies and check the team's doubles against them, as in the worked example; pin the version of the interface the system is built against, and watch for the provider's change notices; and test the failure paths, a slow, refusing or absent service, because an outside service will certainly fail sometimes.

6. Data and timing

Some integration defects need the right data or the right timing to appear. Units share data: a student record written by the exam form and read by the fee payment must mean the same thing to both. And units run concurrently: two requests for the last seat in a hall, or two payments for one form, can interleave in ways no single-threaded test produces. Chapter Three, on why software must be tested, described both kinds. Therac-25's overdoses came from a race between an operator's edits and the software's setup; the Patriot's clock drifted only after the system had run far longer than any test.

munotes.in223

The Challenges of Integration Testing

The answer. Integration test data should be realistic in volume and variety, and consistent across units; time should be controllable in tests (the stub clock of Chapter Thirty-Four); concurrent paths should be tested deliberately, with many simultaneous requests; and long-running behaviour needs long runs, the soak tests of Chapter Forty-Five, on load testing.

7. Interfaces between teams

An interface is often the boundary between two teams, and a boundary no one owns is where problems wait. The Orbiter board found this too: communication across project elements had broken down, with "a failure to elevate concerns with full end-to-end problem ownership." The 2018 ISTQB syllabus's advice is about people as much as tests: "Ideally, testers performing system integration testing should understand the system architecture, and should have influenced integration planning."

The answer. Give every interface an owner and every end-to-end flow someone responsible for it; involve testers in integration planning; and make integration problems visible to everyone, as a shared build history does.

The challenges and their answers

ChallengeWhat goes wrongThe usual answer
Interface specificationsIncomplete, ambiguous, or written and not followedReviewed specifications with units, formats and errors; tests written from them
Stubs and driversEffort to build; doubles that no longer match the real thingSimple doubles; checks against recorded real replies; some tests against the real thing
Locating the faultA combined failure could be anywhereIncremental and continuous integration; logs at interfaces
Environments and configurationTest and production differProduction-like environments; configuration in version control; the same automated deployment everywhere
External servicesChanges, limits, no faithful test copySandboxes; recorded replies; pinned versions; tests of failure paths
Data and timingShared data misread; races; long-running driftRealistic, consistent data; controllable time; concurrency and soak tests
OwnershipInterfaces between teams belong to no oneAn owner for every interface; testers in integration planning

What it does not mean

These challenges are not reasons to skip integration testing. Each is a reason integration defects survive when integration testing is thin, as the Orbiter's did.

A passing test with doubles is not a passing integration. It shows that the unit works with the doubles; only a test with the real partner shows that the two work together.

A production-like environment is not a luxury. Fowler's judgement is that it usually costs less than the bugs that environment differences hide.

munotes.in224

The Challenges of Integration Testing

Timing problems are not rare just because they are hard to reproduce. They are hard to reproduce precisely because they depend on timing, which is why they must be tested for deliberately.

Quick revision

  • Seven challenges: interface specifications, stubs and drivers, locating the fault, environments and configuration, external services, data and timing, ownership.
  • Orbiter (NASA 1999): interface specification "not properly used"; interface verification incomplete; no "full end-to-end problem ownership".
  • Worked example: a stub from the gateway's version 1 kept passing; a check against a recorded reply showed its fields no longer matched; the real version 2 made ExamReg crash with a KeyError.
  • "The greater the scope of integration, the more difficult it becomes to isolate defects" (ISTQB 2018).
  • Test in a clone of production: "every difference results in a risk" (Fowler).
  • Interoperability: A with B and B with C says little about A with C (NIST 2002).

Test yourself

1. List and explain five challenges of integration testing. Interface specifications that are incomplete or not followed, so tests have nothing reliable to check against; the cost and fidelity of stubs and drivers; locating the fault when combined units fail; test environments and configuration that differ from production; and external services that change without the team's control. Data and timing problems and unclear ownership of interfaces are two more.

2. How can a stub make an integration test pass wrongly, and how is that prevented? A stub answers the way its author believed the real module behaves; if that belief is wrong or the real module changes, tests with the stub pass while the real integration fails. It is prevented by keeping stubs simple, checking them regularly against replies recorded from the real service, and running some integration tests against the real thing.

3. Why should the integration test environment resemble production, and how is that achieved? Because every difference between them is a risk that behaviour seen in testing will not happen in production. It is achieved by using the same software versions, operating system, libraries and configuration, keeping configuration in version control, and deploying to both by the same automated steps.

4. What made the Mars Climate Orbiter failure an integration testing failure? The interface specification required metric units, one program produced English units, and the program reading its output assumed the specified units. The investigation found the specification was not properly used in development and testing, interface verification was incomplete, and end-to-end testing against the specification was not accomplished.

5. How do teams deal with third-party services in integration testing? By testing against the provider's sandbox, recording real replies and checking their doubles against them, pinning the version of the interface they depend on and watching for change notices, and testing the failure paths, such as a slow, refusing or absent service.

Contents This chapter on its own page

munotes.in225

Chapter Forty-One

Validation Testing

Syllabus topic Module 1, "Software Testing Strategies: Validation Testing: ensuring adherence to user requirements"

In one line

Validation testing checks the finished, integrated software against what its users need: each user requirement has validation criteria, each is traced to tests that are run on the complete system, and every departure from a requirement is recorded as a deviation and resolved before release.

In the wording a student can write in an examination: validation is "confirmation, through the provision of objective evidence, that the requirements for a specific intended use or application have been fulfilled" (ISO/IEC/IEEE 12207:2026). Validation testing tests the complete system, usually black-box and from the users' point of view, against validation criteria: the acceptance criteria a system "is required to satisfy to be accepted by a user, customer, or other authorized entity" (ISO/IEC 33202:2024). A requirements traceability matrix links each need to its requirements and each requirement to its tests, so gaps show in both directions. Each failed requirement is a deviation, a "departure from a specified requirement", fixed, accepted under a waiver, or taken back to the users. A configuration review (audit) confirms that what was tested is exactly what will be released, with its documentation.

Where validation testing fits

By this point in the strategic approach of Chapter Thirty-One, the units have been tested and integrated. What exists is a complete system, and the question changes from does each part work and fit the others? to does the whole do what its users need? That is validation in the sense of Chapter Twenty-Six, on verification and validation: checking against the intended use, not only against a design document.

Validation testing is done on the whole system, from outside, in the terms the users used: registering for an exam, paying a fee, downloading a hall ticket. It is the testers' last systematic check before the users' own acceptance testing, the subject of the next chapter.

Validation criteria

A requirement can only be validated if everyone agrees what fulfilling it means. The criteria come in several kinds, and a validation plan covers all of them:

  • Functional: every function the users asked for works as they described it (the fee rules, the paper choices, the hall ticket's contents).
  • Behavioural and performance: the response and capacity they need (the hall ticket opening on a phone within 3 seconds; 500 students on the last date).
  • Content and data: what the system shows is accurate and complete (subjects, dates, seat numbers).
  • Usability and accessibility: the intended users can do their tasks (a first-year student unaided; a student using a screen reader).
  • Documentation: the user guide and the administrator's guide describe the system as built.

Each criterion should be stated with a number or a checkable condition, the rule from Chapter Twenty-One, on how quality factors shape testing: a criterion nobody can check is a criterion nobody can meet.

munotes.in226

Validation Testing

The requirements traceability matrix

The tool that keeps validation complete is the traceability matrix, a "matrix that records the relationship between two or more products of the development process" (ISO/IEC/IEEE 24765). For requirements, ISO/IEC/IEEE 29148 defines the requirements traceability matrix (RTM) as a "structured information artifact that links requirements to their higher-level requirements or needs or to lower-level implementation", and the link must work both ways: bidirectional traceability is an "association among two or more logical entities that is discernible in either direction".

The two directions answer different questions.

  • Forward, from each need to its requirements and from each requirement to its tests: is every need met, and is every requirement tested? The ISTQB syllabus's example: "Traceability of test cases to requirements can verify that the requirements are covered by test cases."
  • Backward, from each test to its requirement and from each requirement to its need: is every test testing something someone asked for, and does every requirement serve a real need?

The syllabus adds the wider benefits: "good traceability makes it possible to determine the impact of changes, facilitates audits, and helps meet IT governance criteria."

Worked example: ExamReg's traceability matrix

Chapter Five, on the basic test process, traced ExamReg's fee rules to their tests in one direction. Here is the full matrix for release 2.0's validation: five user needs gathered from students and the exam cell, eight requirements, nine validation tests with their results, and the one waiver the exam cell has signed. The program checks it in both directions.

needs = {"N1": "register for exams online", "N2": "the right fee from every student",
         "N3": "know the seat before the exam", "N4": "records stay private",
         "N5": "usable with a screen reader"}
requirements = {                                  # requirement: (the need it serves, text)
    "R1": ("N1", "log in with roll number and password"),
    "R2": ("N1", "choose regular and backlog papers"),
    "R3": ("N2", "fee computed by the fee rules"),
    "R4": ("N2", "pay through the gateway; receipt emailed"),
    "R5": ("N3", "hall ticket shows seat, subjects and dates"),
    "R6": ("N3", "hall ticket opens on a phone within 3 seconds"),
    "R7": ("N4", "no student can open another student's form"),
    "R8": (None, "export registrations to a spreadsheet"),
}
tests = {                                         # validation test: (requirement, result)
    "VT-01": ("R1", "pass"), "VT-02": ("R2", "pass"), "VT-03": ("R3", "pass"),
    "VT-04": ("R3", "pass"), "VT-05": ("R4", "pass"), "VT-06": ("R5", "pass"),
    "VT-07": ("R6", "fail"), "VT-08": ("R7", "pass"), "VT-09": (None, "pass"),
}
waivers = {"R6": "4.2 s on a basic phone; accepted by the exam cell for this session, fix in 2.1"}

def status(r):
    results = [result for requirement, result in tests.values() if requirement == r]
    if not results:
        return "NOT TESTED"
    return "validated" if all(x == "pass" for x in results) else "FAILS"

print("forward, need by need:")
for n, text in needs.items():
    served_by = [r for r, (need, _) in requirements.items() if need == n]
    found = ", ".join(f"{r} {status(r)}" for r in served_by) or "NO REQUIREMENT"
    print(f"  {n} {text}: {found}")
print("backward checks:")
print("  requirements serving no need:", ", ".join(r for r, (n, _) in requirements.items() if n is None))
print("  requirements with no test:   ", ", ".join(r for r in requirements if status(r) == "NOT TESTED"))
print("  tests tracing to nothing:    ", ", ".join(t for t, (r, _) in tests.items() if r is None))
print("deviations:")
for r, (_, text) in requirements.items():
    if status(r) == "FAILS":
        print(f"  {r} {text}: " + (f"waiver, {waivers[r]}" if r in waivers else "NO DECISION"))
munotes.in227

Validation Testing

forward, need by need:
  N1 register for exams online: R1 validated, R2 validated
  N2 the right fee from every student: R3 validated, R4 validated
  N3 know the seat before the exam: R5 validated, R6 FAILS
  N4 records stay private: R7 validated
  N5 usable with a screen reader: NO REQUIREMENT
backward checks:
  requirements serving no need: R8
  requirements with no test:    R8
  tests tracing to nothing:     VT-09
deviations:
  R6 hall ticket opens on a phone within 3 seconds: waiver, 4.2 s on a basic phone; accepted by the exam cell for this session, fix in 2.1

Five findings come out of one small matrix, and each needs a different response.

  1. N5 has no requirement. Students who use screen readers need the form to work for them, and nothing in the requirements says so; so nothing was built or tested for it. Only the forward direction, from needs, could reveal this. The response is a new requirement, a design change and tests: validation has found a gap in the requirements themselves.
  2. R8 serves no recorded need. Either someone wanted the spreadsheet export and the need was never written down, or it is work nobody asked for. The response is to ask, and either record the need or drop the requirement.
  3. R8 has no test. Whatever its origin, it is being released unchecked.
  4. VT-09 traces to nothing. A test that checks no requirement is either testing something nobody asked for, or testing a real requirement that was never written, which is the more useful discovery.
  5. R6 fails, under a waiver. The hall ticket takes 4.2 seconds on a basic phone against a requirement of 3. That is a deviation, and it has been resolved by a decision, not by a fix.

Deviations and how they are resolved

Every failed validation test records a deviation, a "departure from a specified requirement" (ISO/IEC/IEEE 24765). A deviation cannot simply be left open at release; it has one of three resolutions.

munotes.in228

Validation Testing

  • Fix it before release, and rerun the validation test, with the regression tests described among the test levels and test types of Chapter Thirty-Two.
  • Accept it under a waiver: a "written authorization to accept a configuration item or other designated item which ... is found to depart from specified requirements, but is nevertheless considered suitable for use as is or after rework by an approved method" (ISO/IEC/IEEE 24765). The waiver is written, signed by the people entitled to accept the risk, and says when the deviation will be fixed. R6's is the example.
  • Change the requirement, with the users, when validation shows the requirement itself was wrong or unrealistic.

What is not allowed is the fourth option, the most common in practice: a deviation that nobody decides about. The program prints NO DECISION for exactly that case.

The configuration review

Validation tests prove something about the build that was tested. The configuration review, or configuration audit, makes sure that build is the one being released, complete and correctly documented. ISO/IEC/IEEE 24765 defines a configuration audit as "in configuration management, an independent examination of the configuration status to compare with the actual configuration", and there are two classic kinds.

  • A functional configuration audit (FCA) is an "evaluation to help ensure that the product meets baseline functional and performance capabilities" (INCOSE Systems Engineering Handbook, 2023): in effect, that the validation results cover the baseline requirements.
  • A physical configuration audit (PCA) is an "evaluation to ensure that the operational system or product conforms to the operational and configuration documentation" (INCOSE, 2023): that the release as built matches its documents.

For ExamReg release 2.0 the review asks: is the build to be installed the same build, 2.0.4, that passed validation? Are all the configuration items, code, database scripts, fee configuration, user guide and administrator's guide, at the versions listed in the release note? Does the user guide describe the fee rules the software applies? Knight Capital's loss in 2012 (Chapter Three, on why software must be tested) is the lesson of skipping this check: the code tested was not the code running on every server.

Validation testing and acceptance testing

The two overlap, and the difference is who does them and why. Validation testing is the testers' systematic check that the complete system meets the requirements, planned from the traceability matrix. Acceptance testing, next, is the users' and customer's own decision whether to accept it, in their own environment: user acceptance, operational acceptance, alpha and beta testing. A good validation test run is what gives acceptance testing a system worth the users' time.

munotes.in229

Validation Testing

What it does not mean

Validation testing is not re-running the unit and integration tests. It tests the whole system against user requirements, in the users' terms.

A traceability matrix is not paperwork for its own sake. Its forward direction finds untested requirements and unmet needs; its backward direction finds tests and features nobody asked for.

A waiver is not a way of hiding a failure. It records a known deviation, who accepted the risk, and when it will be fixed.

Passing validation does not prove the requirements are complete. The screen-reader need was missing from them; only tracing from needs found it.

Quick revision

  • Validation: objective evidence that the requirements for the intended use are fulfilled; validation testing tests the complete system against user requirements.
  • Validation criteria: functional, performance, content, usability and accessibility, documentation; each measurable. Acceptance criteria: what a system must satisfy to be accepted.
  • RTM (ISO/IEC/IEEE 29148): links requirements to needs and to implementation and tests; bidirectional traceability.
  • Forward finds unmet needs and untested requirements; backward finds requirements serving no need and tests tracing to nothing.
  • Deviation: departure from a specified requirement; resolved by fixing, a waiver, or changing the requirement; never left undecided.
  • Configuration review: FCA (meets baseline functional and performance capabilities), PCA (conforms to its documentation); the build tested is the build released.

Test yourself

1. What is validation testing, and when is it done? Testing of the complete, integrated system against the users' requirements, usually black-box and in the users' terms, to confirm with objective evidence that the requirements for its intended use are fulfilled. It is done after integration testing and before the users' acceptance testing.

2. What are validation criteria? Give three kinds. The measurable conditions a system must meet for each requirement to count as fulfilled. Kinds include functional (each function works as described), performance (response and capacity), content accuracy, usability and accessibility, and documentation.

3. What is a requirements traceability matrix, and what does it reveal in each direction? A structured record linking needs to requirements and requirements to tests and results. Forward, it reveals needs with no requirement and requirements with no test; backward, it reveals requirements that serve no recorded need and tests that trace to no requirement.

4. What is a deviation, and how may it be resolved? A departure from a specified requirement found in testing. It may be fixed and retested, accepted under a written waiver that records who accepted the risk and when it will be fixed, or resolved by changing the requirement with the users; it must not be left undecided.

5. What is the purpose of a configuration review before release? To confirm that the build being released is exactly the one that passed validation, that all configuration items are at the listed versions, and that the documentation matches the product: a functional audit against the baseline requirements and a physical audit against the documentation.

Contents This chapter on its own page

munotes.in230

Chapter Forty-Two

Acceptance Testing: Alpha, Beta and User Acceptance

Syllabus topic Module 1, "Software Testing Strategies: Validation Testing: ensuring adherence to user requirements"

In one line

Acceptance testing is the users' and customer's own test of the finished system, to decide whether to accept it: user acceptance testing asks whether users can do their work with it, operational acceptance testing whether it can be run, contractual and regulatory acceptance testing whether it meets the contract and the law, and alpha and beta testing try it with real users at the developer's site and then at their own.

In the wording a student can write in an examination: acceptance testing is "formal testing conducted to enable a user, customer, or other authorized entity to determine whether to accept a system or component" (IEEE 1012-2024). Its common forms are user acceptance testing (UAT), operational acceptance testing (OAT), contractual and regulatory acceptance testing, and alpha and beta testing. Alpha testing is the "first stage of testing before a product is considered ready for commercial or operational use" and is done at the developer's site, not by the development team; a beta test is the "second stage of testing when a product is in limited production use" (ISO/IEC/IEEE 24765), done by users at their own locations. It ends in sign-off, the acquirer's formal acceptance of the product.

What acceptance testing is for

By acceptance testing, the system has been verified level by level and validated against its requirements by the testers. Acceptance testing moves the decision to the people who will own and use it. The 2018 ISTQB syllabus gives its objectives: "Establishing confidence in the quality of the system as a whole", "Validating that the system is complete and will work as expected", and "Verifying that functional and non-functional behaviors of the system are as specified".

It adds a point that surprises students: "Defects may be found during acceptance testing, but finding defects is often not an objective, and finding a significant number of defects during acceptance testing may in some cases be considered a major project risk." Acceptance testing is a demonstration and a decision; a system that arrives at it full of defects was not ready to be there.

The forms of acceptance testing

User acceptance testing (UAT). In the 2018 syllabus's words, UAT "is typically focused on validating the fitness for use of the system by intended users in a real or simulated operational environment", and its main objective is "building confidence that the users can use the system to meet their needs, fulfill requirements, and perform business processes with minimum difficulty, cost, and risk." On ExamReg, students register, choose papers and pay, and exam cell clerks process the forms, using the scenarios of their own working day.

Operational acceptance testing (OAT). The acceptance testing "by operations or systems administration staff", usually "in a (simulated) production environment". Its tests "focus on operational aspects", including "Testing of backup and restore", "Installing, uninstalling and upgrading", "Disaster recovery", "User management", "Maintenance tasks", data load and migration, security checks and performance. Its objective is confidence "that the operators or system administrators can keep the system working properly for the users in the operational environment, even under exceptional or difficult conditions." On ExamReg, the college's IT cell restores last night's backup, upgrades from release 1.9, and adds and removes clerk accounts.

munotes.in231

Acceptance Testing: Alpha, Beta and User Acceptance

Contractual and regulatory acceptance testing. "Contractual acceptance testing is performed against a contract's acceptance criteria for producing custom-developed software. Acceptance criteria should be defined when the parties agree to the contract." ExamReg was built for the college by a software house, so the contract's criteria decide whether the college accepts it. "Regulatory acceptance testing is performed against any regulations that must be adhered to, such as government, legal, or safety regulations", sometimes with results witnessed or audited by the regulator.

Alpha and beta testing. These, the syllabus explains, "are typically used by developers of commercial off-the-shelf (COTS) software who want to get feedback from potential or existing users, customers, and/or operators before the software product is put on the market."

  • "Alpha testing is performed at the developing organization's site, not by the development team, but by potential or existing customers, and/or operators or an independent test team."
  • "Beta testing is performed by potential or existing customers, and/or operators at their own locations. Beta testing may come after alpha testing, or may occur without any preceding alpha testing having occurred."

Beta testing's special value is its environment. One of its objectives, in the syllabus, is "the detection of defects related to the conditions and environment(s) in which the system will be used, especially when those conditions and environment(s) are difficult to replicate by the development team." ExamReg's equivalent is a pilot: one department's students register for real, on their own phones and home connections, a week before the portal opens to the whole college. Those phones and networks are exactly what the development team cannot reproduce, and the 4.2-second hall ticket found in the validation testing of Chapter Forty-One is the kind of finding a beta produces.

The forms side by side

FormWho testsWhereQuestion it answersExamReg example
User acceptance testingIntended usersReal or simulated operational environmentCan users do their work with it?Students register and pay; clerks process forms
Operational acceptance testingOperations and system administrators(Simulated) production environmentCan it be run and kept running?Backup and restore, upgrade, account management
Contractual acceptance testingUsers or independent testersAs the contract saysDoes it meet the contract's acceptance criteria?The college checks the software house's delivery
Regulatory acceptance testingUsers or independent testers, sometimes witnessed by a regulatorAs the regulation requiresDoes it comply with the regulations?Any rule that applies to handling students' personal data
Alpha testingCustomers, operators or an independent team, not the developersThe developer's siteDoes it work for real users, under watch?Clerks and a few students at the software house
Beta testingCustomers or operatorsTheir own locationsDoes it work in real conditions?One department's pilot on its own phones
munotes.in232

Acceptance Testing: Alpha, Beta and User Acceptance

Acceptance criteria, and writing them first

Acceptance is only a decision if its criteria are known in advance. Acceptance criteria are the "criteria that a system or component is required to satisfy to be accepted by a user, customer, or other authorized entity" (ISO/IEC 33202:2024), and for contractual acceptance the syllabus insists they "should be defined when the parties agree to the contract."

Acceptance test-driven development (ATDD) takes the idea to its end. In the current ISTQB syllabus, ATDD "Derives tests from acceptance criteria as part of the system design process", and "Tests are written before the part of the application is developed to satisfy the tests." The acceptance tests then exist before the code, and the final acceptance run is the last of many.

Worked example: the exam cell's sign-off

The contract between the college and the software house lists ExamReg's acceptance criteria. After UAT with students and clerks and OAT with the IT cell, the exam cell checks each criterion against the results, and signs off only if all are met. The program applies the criteria to release 2.0's results, including the two defects still open.

uat = {"critical": (8, 8), "other": (11, 12)}           # scenarios passed, run: students and clerks
oat = {"backup and restore": "pass", "upgrade from release 1.9": "pass",
       "add and remove clerk accounts": "pass", "restore minutes": 40}
open_defects = [("DR-311", "major", "hall ticket takes 4.2 s on a basic phone", "waiver W-07"),
                ("DR-318", "minor", "hall ticket font small when printed", None)]

checks = [
    ("UAT: every critical scenario passes", uat["critical"][0] == uat["critical"][1]),
    ("UAT: at least 90% of other scenarios pass", uat["other"][0] / uat["other"][1] >= 0.90),
    ("OAT: backup, restore, upgrade, accounts pass",
     all(v == "pass" for v in oat.values() if isinstance(v, str))),
    ("OAT: a full restore within 60 minutes", oat["restore minutes"] <= 60),
    ("no critical defect open", not any(sev == "critical" for _, sev, _, _ in open_defects)),
    ("every open major defect has a waiver", all(w for _, sev, _, w in open_defects if sev == "major")),
]
for text, met in checks:
    print(f"{'met    ' if met else 'NOT MET'}  {text}")
if all(met for _, met in checks):
    carried = [f"{d} ({sev}, {w or 'to fix in 2.1'})" for d, sev, _, w in open_defects]
    print("sign-off: ACCEPTED, carrying", "; ".join(carried))
else:
    print("sign-off: NOT ACCEPTED")
munotes.in233

Acceptance Testing: Alpha, Beta and User Acceptance

met      UAT: every critical scenario passes
met      UAT: at least 90% of other scenarios pass
met      OAT: backup, restore, upgrade, accounts pass
met      OAT: a full restore within 60 minutes
met      no critical defect open
met      every open major defect has a waiver
sign-off: ACCEPTED, carrying DR-311 (major, waiver W-07); DR-318 (minor, to fix in 2.1)

Every criterion is met, so the exam cell signs off, and the sign-off is honest about what it carries: the slow hall ticket, under the waiver the exam cell granted in validation, and a minor printing defect, both to be fixed in release 2.1. Two things about the example matter more than its numbers. The criteria were written into the contract before the testing, so nobody could argue them into shape afterwards. And acceptance did not require zero defects; it required that no critical defect be open and that every major one be knowingly accepted by the people who carry its risk.

Sign-off

The standards give the final step its formal meaning. Acceptance is the "action by an authorized representative of the acquirer by which the acquirer assumes ownership of products as partial or complete performance of an agreement" (ISO/IEC/IEEE 24748-5:2017). An older definition in the general vocabulary describes the classic contractual case: an acceptance test is a "test of a system or functional unit usually performed by the purchaser on his premises after installation with the participation of the vendor to help ensure that the contractual requirements are met" (ISO/IEC 2382). A sign-off records who accepted, what was accepted (the exact release), against which criteria, and with which known deviations and waivers.

What it does not mean

Acceptance testing is not a hunt for defects. Its purpose is confidence and a decision; many defects found at this stage mean the system arrived too early.

Alpha testing is not the developers testing their own system. It happens at the developer's site but is done by customers, operators or an independent team.

Beta testing is not unplanned. Its participants, duration, feedback channel and exit criteria are planned like any other test.

Acceptance does not mean the system is perfect. It means the owners have accepted it, with its known deviations recorded.

Quick revision

  • Acceptance testing: formal testing to let a user, customer or other authorised entity decide whether to accept a system (IEEE 1012-2024).
  • Objectives: confidence in the whole system, validation that it is complete and works as expected, verification of functional and non-functional behaviour; finding defects is often not an objective.
  • UAT: fitness for use by intended users. OAT: operability (backup and restore, install and upgrade, disaster recovery, user management, maintenance). Contractual: against the contract's acceptance criteria. Regulatory: against regulations.
  • Alpha: at the developer's site, not by the developers. Beta: by users at their own locations, in real conditions.
  • Acceptance criteria agreed in advance; ATDD writes the acceptance tests first.
  • Acceptance: the acquirer assumes ownership; sign-off records the release, criteria and carried deviations.
munotes.in234

Acceptance Testing: Alpha, Beta and User Acceptance

Test yourself

1. What is acceptance testing, and how does it differ from system testing? Formal testing that lets the users, customer or other authorised body decide whether to accept the system. System testing is done by the testers to find defects against the specification; acceptance testing is done by or for the owners to build confidence and make the acceptance decision, and finding many defects there is itself a warning.

2. Describe four forms of acceptance testing. User acceptance testing checks that the intended users can do their work with the system. Operational acceptance testing checks that operators can run it: backup and restore, installation and upgrade, disaster recovery, user management. Contractual acceptance testing checks the contract's acceptance criteria. Regulatory acceptance testing checks compliance with applicable regulations, sometimes witnessed by the regulator.

3. Distinguish alpha testing from beta testing. Alpha testing is done at the developing organisation's site by potential or existing customers, operators or an independent test team, not by the developers. Beta testing is done by customers or operators at their own locations, in real conditions, and can find defects that depend on environments the developers cannot reproduce.

4. Why should acceptance criteria be agreed before testing? So that acceptance is a decision against known conditions rather than an argument afterwards; for contractual acceptance they should be agreed when the contract is made, and in ATDD they become tests before the code is written.

5. What does a sign-off record? That an authorised representative of the acquirer has accepted the product, which exact release was accepted, against which acceptance criteria, and with which known deviations and waivers carried into use.

Contents This chapter on its own page

munotes.in235

Chapter Forty-Three

System Testing: The Whole System, End to End

Syllabus topic Module 1, "Software Testing Strategies: System Testing: comprehensive end to end testing"

In one line

System testing tests the complete, integrated system as a whole, against its specification, by running real end-to-end tasks through every part at once and checking its non-functional qualities in an environment as close to production as possible; it finds the defects that live in the flow between parts, which no smaller test can see.

In the wording a student can write in an examination: system testing is "testing conducted on a complete, integrated system to evaluate the system's compliance with its specified requirements" (ISO/IEC 29110-5-1-2:2025). The ISTQB syllabus describes it as focusing "on the overall behavior and capabilities of an entire system or product, often including functional testing of end-to-end tasks and the non-functional testing of quality characteristics." It is based on the requirements, use cases and system manuals; it is usually done by an independent test team; it runs in a test environment that "should ideally correspond to the final target or production environment"; and it includes system integration testing of the interfaces with external systems.

What system testing is for

Every earlier level tested less than the whole: a unit, a pair of units, a subsystem. System testing is the first time the entire system runs as its users will use it, from the first screen to the last result. The 2018 ISTQB syllabus lists its objectives:

  • "Reducing risk"
  • "Verifying whether the functional and non-functional behaviors of the system are as designed and specified"
  • "Validating that the system is complete and will work as expected"
  • "Building confidence in the quality of the system as a whole"
  • "Finding defects"
  • "Preventing defects from escaping to higher test levels or production"

It adds that "System testing often produces information that is used by stakeholders to make release decisions", which is why its results, not the unit tests', are what a release meeting reads.

What it is tested against, and what it finds

Test basis. System tests are designed from documents that describe the whole system: the syllabus lists "System and software requirement specifications (functional and non-functional)", "Risk analysis reports", "Use cases", "Epics and user stories", "Models of system behavior", "State diagrams", and "System and user manuals". A use case, in one standard's words, is a "description of behavioral requirements of a system and its interaction with a user" (ISO/IEC/IEEE 26515:2018), and each use case is a natural source of end-to-end scenarios.

Typical defects. The syllabus's list is the reason system testing exists:

  • "Incorrect calculations"
  • "Incorrect or unexpected system functional or non-functional behavior"
  • "Incorrect control and/or data flows within the system"
  • "Failure to properly and completely carry out end-to-end functional tasks"
  • "Failure of the system to work properly in the system environment(s)"
  • "Failure of the system to work as described in system and user manuals"
munotes.in236

System Testing: The Whole System, End to End

The third and fourth items are the defects no other level can reach: each part may do its job and the flow between them may still be wrong.

End-to-end testing

System testing's characteristic test is the end-to-end scenario: one complete task, from the user's first action to the system's last effect, run through every part it touches. A scenario is a "step-by-step description of a series of events that occur concurrently or sequentially" (ISO/IEC/IEEE 24765). For ExamReg the central scenario is a registration: log in, pass the eligibility check, choose papers, see the fee, pay, receive the hall ticket. The scenario succeeds only if every step works and every hand-over between steps carries the right thing.

Good end-to-end scenarios are chosen, not just the happy path. For each use case, a system test set covers:

  • the main flow that most users follow;
  • the alternative flows (a concession, backlog papers, a late form);
  • the exception flows (an ineligible student, a form refused as too late, a payment declined);
  • and, for the next chapter, the non-functional conditions under which each must still work.

Worked example: five students, end to end

The program is ExamReg reduced to its five steps and joined together as one system, with the payment provider's sandbox standing in for the real gateway; the sandbox declines one student's card, as test sandboxes let a tester arrange. Five scenarios cover the main flow, one alternative flow (a concession student, late, with backlog papers) and three exception flows, each with the end state the requirements expect: the form's result, the payment, and whether a hall ticket was issued.

FORM_FEE, BACKLOG_FEE = 800, 150

def total_fee(days_late, backlog_papers, concession):
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    fee = FORM_FEE + BACKLOG_FEE * backlog_papers - (FORM_FEE if concession else 0)
    return fee + (0 if days_late == 0 else (100 if days_late <= 7 else 500))

class SandboxGateway:                         # the provider's test sandbox: declines one card
    def charge(self, roll_no, rupees):
        return roll_no != "2026CS040"

class ExamReg:                                # the whole system, every module integrated
    def __init__(self, students, gateway):
        self.students, self.gateway = students, gateway
        self.forms, self.hall_tickets, self.next_seat = {}, {}, 101

    def submit_form(self, roll_no, backlog_papers, days_late):
        student = self.students[roll_no]
        if not student["term_granted"]:
            return "not eligible"
        try:
            fee = total_fee(days_late, backlog_papers, student["concession"])
        except ValueError:
            return "refused"
        self.forms[roll_no] = {"fee": fee, "status": "submitted"}
        return fee

    def pay(self, roll_no):
        form = self.forms[roll_no]
        form["status"] = "paid" if self.gateway.charge(roll_no, form["fee"]) else "declined"
        return form["status"]

    def issue_hall_ticket(self, roll_no):
        if roll_no in self.forms:             # a seat for every student with a form
            self.hall_tickets[roll_no] = self.next_seat
            self.next_seat += 1
        return self.hall_tickets.get(roll_no)

students = {"2026CS014": {"term_granted": True, "concession": False},
            "2026CS027": {"term_granted": True, "concession": True},
            "2026CS031": {"term_granted": False, "concession": False},
            "2026CS040": {"term_granted": True, "concession": False},
            "2026CS052": {"term_granted": True, "concession": False}}
scenarios = [  # (student, backlog papers, days late, expected: form result, payment, hall ticket?)
    ("2026CS014", 0, 0, (800, "paid", True)),
    ("2026CS027", 2, 3, (400, "paid", True)),
    ("2026CS031", 0, 0, ("not eligible", None, False)),
    ("2026CS040", 1, 10, (1450, "declined", False)),
    ("2026CS052", 0, 16, ("refused", None, False)),
]

def describe(form, payment, ticket):
    return (f"form {form}, payment {payment or 'not attempted'},"
            f" hall ticket {'issued' if ticket else 'none'}")

system = ExamReg(students, SandboxGateway())
for roll_no, backlog, days, expected in scenarios:
    result = system.submit_form(roll_no, backlog, days)
    payment = system.pay(roll_no) if isinstance(result, int) else None
    ticket = system.issue_hall_ticket(roll_no) is not None
    actual = (result, payment, ticket)
    print(f"{roll_no}: {describe(*actual)}: {'pass' if actual == expected else 'FAIL'}")
    if actual != expected:
        print(f"           expected {describe(*expected)}")
munotes.in237

System Testing: The Whole System, End to End

2026CS014: form 800, payment paid, hall ticket issued: pass
2026CS027: form 400, payment paid, hall ticket issued: pass
2026CS031: form not eligible, payment not attempted, hall ticket none: pass
2026CS040: form 1450, payment declined, hall ticket issued: FAIL
           expected form 1450, payment declined, hall ticket none
2026CS052: form refused, payment not attempted, hall ticket none: pass

Four scenarios pass, and the one that fails is the one a unit test would never run. Student 2026CS040's card is declined, correctly, and the form's status says so; but the hall-ticket module issues a seat anyway, because it checks that a form exists, not that it has been paid for. Every module did its own job: the fee was right, the gateway's answer was recorded, the seat was allocated. The defect is in the flow between payment and hall ticket, "Incorrect control and/or data flows within the system", in the syllabus's words, and a student who never paid would sit the examination. It needed a scenario that went all the way through with an exception in the middle, which is exactly what end-to-end system testing designs for.

The test environment

A test environment is the "environment containing facilities, hardware, software, firmware, and procedures needed to conduct a test" (ISO/IEC/IEEE 29119-2:2021). For system testing it matters more than at any earlier level, because the system is tested with all its parts and dependencies, and "The test environment should ideally correspond to the final target or production environment" (ISTQB 2018). For ExamReg that means the same web server, database version and configuration as the college's live portal; the payment provider's sandbox in place of the live gateway; realistic student records, with personal data replaced; and a clock that can be set, so that the last date and the late bands can be tested on any day. Chapter Forty, on the challenges of integration testing, gave the reasons for each.

munotes.in238

System Testing: The Whole System, End to End

System integration testing

The syllabus names a companion level: system integration testing "focuses on testing the interfaces of the system under test and other systems and external services", and "requires suitable test environments preferably similar to the operational environment." ExamReg's external partners are the payment gateway and the email service that sends receipts. System integration tests check the conversations with them: a payment approved, declined, timed out; a receipt sent, bounced, delayed. The worked example's declined card is a system integration condition that led straight to a system defect, which is why the two levels are usually planned together.

Who does system testing

The syllabus says that "System testing is typically carried out by independent testers who rely heavily on specifications", and warns of the cost of poor specifications at this level: "Defects in specifications (e.g., missing user stories, incorrectly stated business requirements, etc.) can lead to a lack of understanding of, or disagreements about, expected system behavior." Its remedy is the one this book has repeated since Chapter Eighteen, on the role of testing in each phase: "Early involvement of testers in user story refinement or static testing activities, such as reviews, helps to reduce the incidence of such situations."

What it does not mean

System testing is not the sum of the unit tests. It tests flows across the whole system, which no unit test sees; the hall-ticket defect passed every unit test.

End to end does not mean only the happy path. The most revealing scenarios contain an exception partway through.

System testing is not acceptance testing. Testers do it against the specification to find defects; users do acceptance testing to decide whether to accept.

A test environment is not the production environment. It should resemble production closely, and every difference is a risk to be named.

Quick revision

  • System testing: testing a complete, integrated system against its specified requirements (ISO/IEC 29110-5-1-2:2025); end-to-end functional tasks plus non-functional quality characteristics (ISTQB).
  • Objectives: reduce risk, verify functional and non-functional behaviour, validate completeness, build confidence, find defects, stop them escaping; informs release decisions.
  • Test basis: requirements, risk reports, use cases, user stories, behaviour and state models, manuals.
  • Typical defects: incorrect calculations, wrong behaviour, incorrect control and data flows, failed end-to-end tasks, environment failures, behaviour unlike the manuals.
  • End-to-end scenarios: main, alternative and exception flows through every part.
  • Worked example: a declined payment still received a hall ticket; only an end-to-end exception scenario showed it.
  • Test environment as close to production as possible; system integration testing of external interfaces; usually done by independent testers.

Test yourself

1. What is system testing, and what is it based on? Testing of a complete, integrated system to evaluate its compliance with its specified requirements, both functional and non-functional. It is based on requirement specifications, use cases and user stories, risk analyses, models of system behaviour, and system and user manuals.

munotes.in239

System Testing: The Whole System, End to End

2. List four typical defects that system testing finds. Incorrect calculations; incorrect control or data flows within the system, such as a hall ticket issued for an unpaid form; failure to carry out end-to-end tasks completely; and failure of the system to work in its environment or as its manuals describe.

3. What is an end-to-end scenario, and how should scenarios be chosen? A complete task from the user's first action to the system's final effect, run through every part it touches. Scenarios should cover the main flow, the alternative flows and the exception flows of each use case, since the flows with an exception partway through are the most revealing.

4. Why must the system test environment resemble production? Because the whole system is tested with all its parts and dependencies, and any difference in software versions, configuration, data or external services is a risk that behaviour in testing will differ from behaviour in use.

5. Distinguish system testing from system integration testing. System testing tests the behaviour and qualities of the whole system against its specification. System integration testing tests the interfaces between the system and other systems and external services, such as a payment gateway or an email service.

Contents This chapter on its own page

munotes.in240

Chapter Forty-Four

Recovery, Security, Stress, Performance and Deployment Testing

Syllabus topic Module 1, "Software Testing Strategies: System Testing: comprehensive end to end testing"

In one line

A system that does the right thing can still fail its users by not recovering from a crash, not keeping out an attacker, collapsing under a crowd, being too slow, or breaking when it is installed; the system test types test each of these, each with its own goal and method.

In the wording a student can write in an examination: recovery testing forces the system to fail and checks that it restores itself and its data (backup and recovery testing measures "the degree to which system state can be restored from backup within specified parameters of time, cost, completeness, and accuracy"). Security testing evaluates whether a test item and its data "are protected so that unauthorized persons or systems cannot use, read, or modify them, and authorized persons or systems are not denied access to them". Stress testing evaluates behaviour "under conditions of loading above anticipated or specified capacity requirements, or of resource availability below minimum specified requirements". Performance testing evaluates "the degree to which a test item accomplishes its designated functions within given constraints of time and other resources". Deployment testing checks that the system installs, upgrades, configures and rolls back correctly in its target environment. (Definitions from ISO/IEC/IEEE 24765 and 29119.)

Recovery testing

Goal. Most systems will fail at some point: a power cut, a crashed process, a lost disk. Recovery testing checks the system's recoverability, its "capability ... in the event of an interruption or a failure to recover the data directly affected and re-establish the desired state of the system" (ISO/IEC 25010:2023). The standard vocabulary defines recovery itself as the "restoration of a system, program, database, or other system resource to a state in which it can perform required functions".

Method. Make the system fail on purpose, at chosen points, and check what state it comes back to. The questions are always the same: is any data lost, is any operation done twice or half-done, and how long does recovery take? A related check, backup and recovery testing, restores the system from its backups and measures the result against "specified parameters of time, cost, completeness, and accuracy". The 2018 ISTQB syllabus places backup and restore and disaster recovery among the operational acceptance tests of Chapter Forty-Two.

Worked example: a crash in the middle of a payment. ExamReg charges the student through the gateway, then records the form as paid. What if the server stops between those two steps? The program injects a crash at two points, restarts the system, lets the student press Pay again, and checks the one thing that matters: the student was charged exactly once and the form is marked paid. It tests two designs. The plain one does the two steps and nothing else. The journaled one writes a note, payment started, before the money moves and payment done after, and at every restart it asks the gateway about any payment started but not done. (A production system would use database transactions and the provider's own safeguards; the journal here is the idea in its smallest form.)

munotes.in241

Recovery, Security, Stress, Performance and Deployment Testing

class Crash(Exception):
    """The server stops at this point: power cut, killed process."""

class Gateway:                                   # the provider keeps its own record of charges
    def __init__(self):
        self.charges = []
    def charge(self, roll_no, rupees):
        self.charges.append((roll_no, rupees))
    def charged(self, roll_no):
        return sum(1 for r, _ in self.charges if r == roll_no)

def pay_plain(db, journal, gateway, roll_no, fee, crash_at=None):
    if db.get(roll_no) == "paid":
        return
    if crash_at == "before charge":
        raise Crash
    gateway.charge(roll_no, fee)
    if crash_at == "after charge":
        raise Crash
    db[roll_no] = "paid"

def pay_journaled(db, journal, gateway, roll_no, fee, crash_at=None):
    if db.get(roll_no) == "paid":
        return
    journal.append(("start", roll_no))           # written before the money moves
    if crash_at == "before charge":
        raise Crash
    gateway.charge(roll_no, fee)
    if crash_at == "after charge":
        raise Crash
    db[roll_no] = "paid"
    journal.append(("done", roll_no))

def recover(db, journal, gateway):               # run at every restart
    started = {r for kind, r in journal if kind == "start"}
    finished = {r for kind, r in journal if kind == "done"}
    for roll_no in started - finished:           # ask the provider what really happened
        db[roll_no] = "paid" if gateway.charged(roll_no) else "unpaid"

for design, pay in [("plain", pay_plain), ("journaled", pay_journaled)]:
    for crash_at in ["before charge", "after charge"]:
        db, journal, gateway = {}, [], Gateway()
        try:
            pay(db, journal, gateway, "2026CS014", 1450, crash_at)
        except Crash:
            recover(db, journal, gateway)        # restart; the plain design has no journal to read
        pay(db, journal, gateway, "2026CS014", 1450)     # the student presses Pay again
        n = gateway.charged("2026CS014")
        verdict = "pass" if n == 1 and db.get("2026CS014") == "paid" else f"FAIL: charged {n} times"
        print(f"{design:<10} crash {crash_at:<14} {verdict}")
plain      crash before charge  pass
plain      crash after charge   FAIL: charged 2 times
journaled  crash before charge  pass
journaled  crash after charge   pass

A crash before the charge is harmless in both designs: no money moved, and the student simply pays again. A crash after the charge is where the plain design fails: the gateway took Rs 1,450, the form still says unpaid, and when the student presses Pay again they are charged a second time. The journaled design survives the same crash, because at restart it finds a payment started but not done, asks the gateway, learns the money was taken, and records the form as paid before the student returns. No functional test would have found this; every function works. Only a test that stops the system at the wrong moment does.

Security testing

Goal. ISO/IEC/IEEE 29119-2 defines security testing as a "test type conducted to evaluate the degree to which a test item, and associated data and information, are protected so that unauthorized persons or systems cannot use, read, or modify them, and authorized persons or systems are not denied access to them." Both halves count: keeping the wrong people out, and letting the right people in.

munotes.in242

Recovery, Security, Stress, Performance and Deployment Testing

Method. Security testing looks for vulnerabilities, each a "potential flaw or weakness in software design or implementation that can be exercised (accidentally triggered or intentionally exploited) and result in harm to the system" (ISO/IEC 23643:2020). Typical activities are checks of authentication and access control, attempts to read or change other users' data, malformed and malicious input, scans for known vulnerabilities in the components used, and penetration testing, in which a tester attacks the system as an intruder would.

ExamReg. A student logged in as 2026CS014 changes the form number in the page's address to another student's and presses Enter: the portal must refuse. A password is tried wrongly three times: the account must lock for 30 minutes, as the lockout rule requires. A roll number containing database commands is typed into the login box: it must be treated as text. And, the other half, a genuine student on the last date must not be locked out by any of these defences.

Stress testing

Goal. Stress testing is the "type of performance efficiency testing conducted to evaluate a test item's behavior under conditions of loading above anticipated or specified capacity requirements, or of resource availability below minimum specified requirements" (ISO/IEC/IEEE 29119-1:2022). The question is not whether the system copes with its specified load, but what happens beyond it: does it slow down gracefully, refuse new users politely, or crash and lose data?

Method. Push the load past the specified capacity in steps, or take resources away (memory, database connections, network bandwidth), and watch for the breaking point and the manner of breaking. Then remove the stress and check that the system recovers without being restarted by hand.

ExamReg. The requirement is 500 simultaneous users on the last date. A stress test drives 800, then 1,200, and checks that beyond capacity the portal shows a please try again in a few minutes page, that no half-submitted form is saved, that no payment is taken twice, and that the portal returns to normal once the load falls. Its companion, robustness, is the "degree to which a system or component can function correctly in the presence of invalid inputs or stressful environmental conditions" (ISO/IEC/IEEE 24765).

Performance testing

Goal. Performance testing evaluates "the degree to which a test item accomplishes its designated functions within given constraints of time and other resources" (ISO/IEC/IEEE 29119-2:2021): response time, throughput and resource use, against measurable requirements.

munotes.in243

Recovery, Security, Stress, Performance and Deployment Testing

Method. Performance testing is a family, and the standards name its members.

KindWhat it evaluates (ISO/IEC/IEEE 29119-1 and 24765)ExamReg example
Load testingBehaviour "under anticipated conditions of varying load, usually between anticipated conditions of low, typical, and peak usage"50, 200 and 500 users; the fee page within 2 seconds
Stress testingBehaviour above the specified capacity, or with too few resources800 and 1,200 users
Capacity testing"the level at which increasing load ... compromises a test item's ability to sustain required performance"Find the number of users at which the fee page passes 2 seconds
Volume testingThe ability "to process specified volumes of data"A term's 20,000 registrations in the database, and the exam cell's full report
Endurance testingWhether it "can sustain a required load continuously for a specified period of time"Typical load for the whole two-week window

Endurance testing, often called soak testing, is the one the Patriot story of Chapter Three, on why software must be tested, argues for: some failures appear only after long, continuous running. Load testing, with the practical's tool, is the next chapter.

Deployment testing

Goal. Deployment is the "stage of a project in which a system is put into operation and transition issues are resolved" (IEEE 2675-2021). Deployment testing checks that stage: that the system can be installed, configured, upgraded and, if necessary, rolled back in its real environment. It tests installability, the "capability of a product to be effectively and efficiently installed successfully or uninstalled in a specified environment" (ISO/IEC 25010:2023), and, where a system must run in several environments, portability: portability testing evaluates "the ease with which a test item can be transferred from one hardware or software environment to another" (ISO/IEC/IEEE 29119-2:2021).

Method. Install from the release package, exactly as operations will, on an environment like production; upgrade from the previous release with real data; check the configuration on every server; and rehearse the rollback. Knight Capital's loss in 2012, told in Chapter Three on why software must be tested, is the case for all of it: the code was right, and one of eight servers never received it.

ExamReg. Install release 2.0 on a copy of the college's server; upgrade a copy of last term's database from release 1.9 and check that every registration survived; confirm that both web servers report build 2.0.4; roll back to 1.9 and forward again.

The five types compared

TypeGoalHow it is doneExamReg example
RecoveryThe system and its data come back correctly after a failureForce failures at chosen points; restore from backups; time itCrash between charging and recording a payment
SecurityUnauthorised use is prevented; authorised use is not deniedAccess-control checks, malicious input, vulnerability scans, penetration testsChanging the form number in the address
StressGraceful behaviour beyond capacity, and recovery afterwardsLoad above capacity; remove resources1,200 users against a 500-user requirement
PerformanceFunctions done within time and resource limitsLoad, capacity, volume and endurance tests against measurable targetsThe fee page within 2 seconds for 500 users
DeploymentInstalls, upgrades, configures and rolls back correctlyInstall and upgrade from the release package in a production-like environmentUpgrade from 1.9 with last term's data
munotes.in244

Recovery, Security, Stress, Performance and Deployment Testing

What it does not mean

These are not optional extras after functional testing. A portal that computes every fee correctly and double-charges after a crash, or falls over on the last date, has failed its users.

Stress testing is not load testing. Load testing stays within the anticipated range; stress testing goes beyond it on purpose.

Security testing is not only about keeping attackers out. The definition also requires that authorised users are not denied access.

Recovery is not the same as a backup. A backup is data; recovery is the tested ability to return to a correct state with it, within a stated time.

Quick revision

  • Recovery testing: force failures; check data and state are restored; backup and recovery testing measures restoration within time, cost, completeness and accuracy; recoverability (ISO/IEC 25010:2023).
  • Security testing: unauthorised persons cannot use, read or modify; authorised persons are not denied (ISO/IEC/IEEE 29119-2); look for vulnerabilities.
  • Stress testing: load above specified capacity or resources below minimum (ISO/IEC/IEEE 29119-1); watch how it breaks and recovers.
  • Performance testing: functions within constraints of time and resources; family: load, stress, capacity, volume, endurance.
  • Deployment testing: install, configure, upgrade, roll back; installability and portability.
  • Worked example: a crash after the charge made the plain design charge twice; the journaled design recovered.

Test yourself

1. What is recovery testing? Describe how it is done. Testing that the system and its data return to a correct state after a failure. It is done by deliberately making the system fail at chosen points, such as between charging a payment and recording it, then restarting and checking that nothing is lost, duplicated or half-done, and by restoring from backups and measuring the time, completeness and accuracy of the restoration.

2. Define security testing and give three ExamReg security tests. Testing that evaluates whether the system and its data are protected so that unauthorised persons cannot use, read or modify them, while authorised persons are not denied access. ExamReg tests: a student cannot open another student's form by changing the form number in the address; three wrong passwords lock the account for 30 minutes; database commands typed into the login box are treated as text.

munotes.in245

Recovery, Security, Stress, Performance and Deployment Testing

3. Distinguish stress testing from load testing. Load testing evaluates behaviour under anticipated loads, from low to peak. Stress testing evaluates behaviour beyond the specified capacity, or with fewer resources than specified, to see how the system breaks and whether it recovers.

4. Name four kinds of performance testing, with an example of each. Load testing (50, 200 and 500 users against the fee page's 2-second target); capacity testing (the number of users at which the target is first missed); volume testing (a term's 20,000 registrations in the database); endurance testing (typical load sustained for the whole two-week window).

5. What does deployment testing check, and why does it matter? That the system installs, is configured, upgrades with real data and can be rolled back correctly in its target environment. It matters because a correct program can still fail in use if it is deployed wrongly, as when one of Knight Capital's eight servers never received the new code.

Contents This chapter on its own page

munotes.in246

Chapter Forty-Five

Load Testing: Users, Ramp-Up, Throughput and Error Rate

Syllabus topic Practical, "Load Testing Using Apache JMeter"

In one line

A load test drives a system with a planned number of simulated users and measures how it responds: in JMeter, a thread group sets the users, how quickly they arrive (the ramp-up) and how often each repeats (the loop count), and the results are read as response times, percentiles, throughput and the percentage of errors.

In the wording a student can write in an examination: load testing evaluates "the behavior of a test item under anticipated conditions of varying load, usually between anticipated conditions of low, typical, and peak usage" (ISO/IEC/IEEE 29119-1:2022). In Apache JMeter, the practical's tool, a thread group "controls the number of threads JMeter will use to execute your test": each thread is one simulated user; the ramp-up period "tells JMeter how long to take to 'ramp-up' to the full number of threads chosen"; and the loop count sets how many times each thread repeats the test. The results are read from reports such as the Aggregate Report: the average, median and 90% line of response time, the error % ("Percent of requests with errors") and the throughput, where "Throughput = (number of requests) / (total time)".

The practical's task

The practical asks students to "Design and execute load testing scenarios for a web application. Configure thread groups, ramp-up time, and loop count. Analyze response time, throughput, and error percentage using graphical reports." This chapter explains each of those terms from JMeter's own manual and computes each measure from a results file, so that the numbers JMeter shows in the lab can be read with understanding.

Load, stress, endurance and spikes

Chapter Forty-Four placed load testing in the performance family. For ExamReg the differences look like this:

  • Load test: 50, 200 and 500 simultaneous students, the range the requirements anticipate, against the target of the fee page in 2 seconds.
  • Stress test: 800 and 1,200 students, beyond capacity, to see how the portal fails and whether it recovers.
  • Endurance (soak) test: a typical load kept up for many hours, to find slow leaks and the long-running faults of Chapter Three, on why software must be tested.
  • A spike: a sudden jump in load. JMeter's manual uses the word for a load problem a badly configured test can create by accident: with 100 threads and a ramp-up of 0, "all the threads would start at the same time, and it would produce an unwanted spike of the load." For ExamReg a spike is also real: the minute hall tickets are released, many students ask for them at once, and a test should reproduce that on purpose.

The thread group: threads, ramp-up and loop count

JMeter's manual describes the thread group as the start of every test: "Thread group elements are the beginning points of any test plan", and its controls let a tester "Set the number of threads", "Set the ramp-up period" and "Set the number of times to execute the test".

munotes.in247

Load Testing: Users, Ramp-Up, Throughput and Error Rate

  • Threads. "Each thread will execute the test plan in its entirety and completely independently of other test threads. Multiple threads are used to simulate concurrent connections to your server application." A thread is one simulated user.
  • Ramp-up period. The manual's own example: "If 10 threads are used, and the ramp-up period is 100 seconds, then JMeter will take 100 seconds to get all 10 threads up and running. Each thread will start 10 (100/10) seconds after the previous thread was begun." Its advice cuts both ways: "Ramp-up needs to be long enough to avoid too large a work-load at the start of a test, and short enough that the last threads start running before the first ones finish (unless one wants that to happen)." Its rule of thumb: "Start with Ramp-up = number of threads and adjust up or down as needed."
  • Loop count. "By default, the thread group is configured to loop once through its elements." A loop count of 5 makes each thread run the test five times, so the plan sends threads × loops requests for each sampler.

Reading the results

JMeter's glossary defines the measurements precisely.

  • Elapsed time: "JMeter measures the elapsed time from just before sending the request to just after the last response has been received." This is the response time the reports use.
  • Latency: measured "from just before sending the request to just after the first response has been received".
  • Median: "a number which divides the samples into two equal halves".
  • 90% Line: "the value below which 90% of the samples fall". The Aggregate Report puts it in words a student can use: "90 % of the samples took no more than this time. The remaining samples took at least as long as this." It also reports the 95% and 99% lines.
  • Error %: "Percent of requests with errors".
  • Throughput: "Throughput is calculated as requests/unit of time. The time is calculated from the start of the first sample to the end of the last sample." So "Throughput = (number of requests) / (total time)".

The percentiles matter more than the average. An average can look healthy while a tenth of the students wait far too long; the 90% line says how long nine students in ten waited at most, which is closer to what a requirement like the fee page within 2 seconds means.

Worked example: reading one test's results

A tester runs a first load test on ExamReg's fee page with a thread group of 20 threads, a ramp-up of 20 seconds and a loop count of 2, and saves the results. JMeter can write its results to a CSV file; the one below is this book's illustration of such a file, reduced to the five columns the report needs: when each request started (in milliseconds from the start of the test), its elapsed time, its label, the server's response code and whether it succeeded.

munotes.in248

Load Testing: Users, Ramp-Up, Throughput and Error Rate

start_ms,elapsed_ms,label,response_code,success
0,380,fee page,200,true
380,410,fee page,200,true
1000,417,fee page,200,true
1417,463,fee page,200,true
2000,454,fee page,200,true
2454,516,fee page,200,true
3000,491,fee page,200,true
3491,569,fee page,200,true
4000,528,fee page,200,true
4528,622,fee page,200,true
5000,565,fee page,200,true
5565,675,fee page,200,true
6000,602,fee page,200,true
6602,728,fee page,200,true
7000,639,fee page,200,true
7639,431,fee page,200,true
8000,676,fee page,200,true
8676,484,fee page,200,true
9000,413,fee page,200,true
9413,537,fee page,200,true
10000,450,fee page,200,true
10450,590,fee page,200,true
11000,487,fee page,200,true
11487,643,fee page,200,true
12000,524,fee page,200,true
12524,696,fee page,200,true
13000,561,fee page,200,true
13561,749,fee page,200,true
14000,598,fee page,200,true
14598,452,fee page,200,true
15000,635,fee page,200,true
15635,505,fee page,200,true
16000,672,fee page,200,true
16672,558,fee page,200,true
17000,409,fee page,200,true
17409,120,fee page,503,false
18000,446,fee page,200,true
18446,664,fee page,200,true
19000,95,fee page,503,false
19095,717,fee page,200,true

The program works out what the thread group means, then computes an Aggregate Report row from the file by the definitions above, and finally counts how many requests were ever in flight at the same moment.

import csv
import math

threads, ramp_up_s, loops = 20, 20, 2                # the thread group, as configured
print(f"thread group: {threads} threads, ramp-up {ramp_up_s} s, loop count {loops}:"
      f" a thread starts every {ramp_up_s / threads:.1f} s, {threads * loops} requests planned")

with open("fee-page-results.csv") as f:
    samples = [(int(r["start_ms"]), int(r["elapsed_ms"]), r["success"] == "true")
               for r in csv.DictReader(f)]

times = sorted(e for _, e, _ in samples)
n = len(times)
def line(pct):                                       # "pct % of the samples took no more than this"
    return times[math.ceil(pct / 100 * n) - 1]

errors = sum(1 for _, _, ok in samples if not ok)
first_start = min(s for s, _, _ in samples)
last_end = max(s + e for s, e, _ in samples)
throughput = n / ((last_end - first_start) / 1000)   # requests / total time
print(f"samples {n}, average {sum(times) / n:.0f} ms, median {line(50)} ms,"
      f" 90% line {line(90)} ms, 95% line {line(95)} ms, 99% line {line(99)} ms")
print(f"min {times[0]} ms, max {times[-1]} ms, error {errors / n:.1%},"
      f" throughput {throughput:.2f} requests/s")

events = sorted([(s, 1) for s, _, _ in samples] + [(s + e, -1) for s, e, _ in samples])
in_flight = peak = 0
for _, change in events:
    in_flight += change
    peak = max(peak, in_flight)
print(f"most requests in flight at once: {peak} (the plan had {threads} threads)")
thread group: 20 threads, ramp-up 20 s, loop count 2: a thread starts every 1.0 s, 40 requests planned
samples 40, average 529 ms, median 528 ms, 90% line 676 ms, 95% line 717 ms, 99% line 749 ms
min 95 ms, max 749 ms, error 5.0%, throughput 2.02 requests/s
most requests in flight at once: 2 (the plan had 20 threads)
munotes.in249

Load Testing: Users, Ramp-Up, Throughput and Error Rate

Most of the report is what a student expects to read. Forty samples; an average and a median near 530 milliseconds; a 90% line of 676 ms, so nine requests in ten took no more than that; two failures with response code 503, an error rate of 5 per cent; and a throughput of about 2 requests a second.

The last line is the real finding. The plan had 20 threads, but at no moment were more than 2 requests in flight, because each thread finished its two requests in about a second while the next thread started a second later. The manual's warning describes exactly this: ramp-up must be "short enough that the last threads start running before the first ones finish". This test never came near 20 simultaneous students, so it says nothing about the portal's behaviour with 20, let alone the 500 of the requirement. The next run should shorten the ramp-up, raise the loop count, or set a duration, and the concurrency should be checked again before anyone reads the response times as an answer.

Two other cautions follow from the manual. The 503 responses may be the server refusing work under load, or a fault that has nothing to do with load; the report counts them, and a tester must look at them. And throughput depends on how the test was built: the Aggregate Report notes that "If other samplers and timers are in the same thread, these will increase the total time, and therefore reduce the throughput value."

A load test plan for ExamReg

QuestionExamReg's answer
What load is anticipated?Up to 500 simultaneous students on the last date
Which pages?Login, exam form, fee page, payment hand-off, hall ticket
Threads and ramp-up50, 200 and 500 threads, with ramp-ups short enough that all are running together
DurationLong enough at each level for response times to settle, checked by the in-flight count
Targets90% line of the fee page within 2 seconds; error % below 1 per cent
EnvironmentA copy of production, with the payment provider's sandbox
ReadingAggregate Report per page; response time against active threads, graphed

What it does not mean

The number of threads is not the number of simultaneous users. It is the most that can be active; the ramp-up, loop count and response times decide how many actually overlap, as the worked example showed.

The average response time is not the students' experience. Percentiles, the 90% and 95% lines, say how long most students waited at worst.

munotes.in250

Load Testing: Users, Ramp-Up, Throughput and Error Rate

A load test is not a stress test. Load testing stays within anticipated load; stress testing goes beyond it on purpose.

A high error percentage is not always the server's fault. A test that sends malformed requests, or runs out of test data, produces errors too; each kind must be looked at.

Quick revision

  • Load testing: behaviour under anticipated load, from low to peak (ISO/IEC/IEEE 29119-1:2022).
  • Thread group: threads (simulated users), ramp-up period (time to start them all; delay = ramp-up ÷ threads), loop count (repetitions per thread).
  • Ramp-up "long enough to avoid too large a work-load at the start", "short enough that the last threads start running before the first ones finish"; start with ramp-up = threads.
  • Elapsed time, latency, median, 90% line (90 per cent of samples took no more), error %, throughput = requests ÷ total time.
  • Worked example: 40 samples, median 528 ms, 90% line 676 ms, 5 per cent errors, about 2 requests/s, and never more than 2 requests in flight, so the test did not load the portal as planned.

Test yourself

1. Explain the thread group's settings: number of threads, ramp-up period and loop count. The number of threads is the number of simulated users, each running the test plan independently. The ramp-up period is how long JMeter takes to start all the threads, so each starts ramp-up divided by threads seconds after the previous one. The loop count is how many times each thread repeats the test.

2. If 30 threads are configured with a ramp-up period of 120 seconds, when does each thread start? Each successive thread starts 4 seconds after the previous one, the 120 seconds shared among 30 threads, so the last starts 116 seconds into the test.

3. What are the 90% line, error % and throughput in a JMeter report? The 90% line is the response time that 90 per cent of samples did not exceed. Error % is the percentage of requests that failed. Throughput is the number of requests divided by the total time from the start of the first sample to the end of the last.

4. Why is the 90% line often more useful than the average? Because an average can look acceptable while a tenth of users wait far too long; the 90% line states the worst wait for nine users in ten, which matches requirements such as a page appearing within 2 seconds.

5. A test with 20 threads shows good response times, but no more than 2 requests were ever in flight. What went wrong, and what should be changed? The ramp-up was too long for requests so short: each thread finished before the next few started, so the server never carried 20 users at once. The ramp-up should be shortened, the loop count raised or a duration set, and the concurrency checked before the response times are trusted.

Contents This chapter on its own page

munotes.in251

Chapter Forty-Six

Cross-Browser and Compatibility Testing

Syllabus topic Practical, "Cross-Browser Testing Using Selenium Grid"

In one line

A web application must work in every browser, platform, screen and network its users bring, and there are too many combinations to test them all; compatibility testing chooses a sample that covers every pair of choices, and Selenium Grid runs the same automated tests on many browsers and machines at once.

In the wording a student can write in an examination: compatibility is the "capability of a product to exchange information with other products, or to perform its required functions while sharing the same common environment and resources" (ISO/IEC 25010:2023), and compatibility testing measures "the degree to which a test item can function satisfactorily alongside other independent products in a shared environment (co-existence), and where necessary, exchanges information with other systems or components (interoperability)" (ISO/IEC/IEEE 29119-1:2022). Cross-browser testing is its most common form for web applications: the same tests run on each supported browser and platform. Because the combinations multiply, testers use pairwise testing, in which "test cases are designed to execute all possible discrete combinations of each pair of input parameters" (ISO/IEC TR 29119-11:2020). Selenium Grid runs WebDriver tests on many browser and machine combinations from one entry point, in Standalone, Hub and Node, or Distributed mode.

Why browsers and platforms differ

ExamReg's students register from college computers running Windows, lab machines running Ubuntu, and their own Android phones, in Chrome, Firefox or Edge, on fast campus networks and slow mobile data. Each browser draws pages, runs scripts and handles forms in its own way, and each platform and screen changes the layout. A fee page that looks right in Chrome on a laptop can hide its Pay button off the edge of a phone screen, or refuse to load its date picker in another browser. None of that is a defect in ExamReg's logic; all of it stops a student registering.

ISO/IEC 25010 names the two halves of the quality at stake:

  • Co-existence: the "capability of a product to perform its required functions efficiently while sharing a common environment and resources with other products, without detrimental impact on any other product".
  • Interoperability: the "capability of a product to exchange information with other products and mutually use the information that has been exchanged".

For a web portal, the browser and platform are the shared environment, and the payment gateway and email service are the products it exchanges information with.

The compatibility matrix

A compatibility test starts by listing the parameters that vary and the values each can take. For ExamReg, the practical's three browsers and the platforms students actually use:

ParameterValues
BrowserChrome, Firefox, Edge
PlatformWindows 11, Ubuntu, Android
ScreenPhone-sized, desktop-sized
NetworkFast, slow

Every combination is 3 × 3 × 2 × 2 = 36 configurations, and each would need the whole regression suite. Add a fifth parameter, and the count multiplies again. Testing all of them is the exhaustive testing that Chapter Four's second principle rules out.

munotes.in252

Cross-Browser and Compatibility Testing

Pairwise testing

Pairwise testing rests on an assumption: that a compatibility defect is usually triggered by one value (a browser that mishandles a control) or by two values together (a browser that misbehaves only on small screens), and seldom needs three specific values at once. Combinatorial testing is the "class of specification-based test design techniques based on exercising combinations of parameter-value (P-V) pairs" (ISO/IEC/IEEE 29119-1:2022), and its commonest form, pairwise testing, chooses configurations so that every pair of values appears together in at least one of them.

Worked example: from 36 configurations to 9

The program lists every configuration and every pair of values the matrix contains, then chooses configurations greedily: at each step, the one that covers the most pairs not yet covered, until none is left. It then compares the run time of the regression suite on one machine against a Grid, for the practical's comparison of execution time.

from itertools import combinations, product

params = {"browser": ["Chrome", "Firefox", "Edge"],
          "platform": ["Windows 11", "Ubuntu", "Android"],
          "screen": ["phone-sized", "desktop-sized"],
          "network": ["fast", "slow"]}
values = list(params.values())
every_config = list(product(*values))

def pairs_in(config):                        # every pair of parameter values one config tests
    return {((i, config[i]), (j, config[j])) for i, j in combinations(range(len(config)), 2)}

all_pairs = set().union(*(pairs_in(c) for c in every_config))
chosen, uncovered = [], set(all_pairs)
while uncovered:                             # greedy: take the config that covers most new pairs
    best = max(every_config, key=lambda c: len(pairs_in(c) & uncovered))
    chosen.append(best)
    uncovered -= pairs_in(best)

print(f"every combination: {len(every_config)} configurations; pairs of values: {len(all_pairs)}")
print(f"pairwise selection: {len(chosen)} configurations cover every pair")
for n, config in enumerate(chosen, 1):
    print(f"  {n}. " + ", ".join(config))

suite_seconds = {"Chrome": 142, "Firefox": 168, "Edge": 151}   # one run of the regression suite
print(f"suite on one machine, one browser after another: {sum(suite_seconds.values())} s;"
      f" on a Grid, one node per browser: {max(suite_seconds.values())} s")
every combination: 36 configurations; pairs of values: 37
pairwise selection: 9 configurations cover every pair
  1. Chrome, Windows 11, phone-sized, fast
  2. Chrome, Ubuntu, desktop-sized, slow
  3. Firefox, Android, phone-sized, slow
  4. Edge, Android, desktop-sized, fast
  5. Firefox, Windows 11, desktop-sized, fast
  6. Edge, Windows 11, phone-sized, slow
  7. Firefox, Ubuntu, phone-sized, fast
  8. Chrome, Android, phone-sized, fast
  9. Edge, Ubuntu, phone-sized, fast
suite on one machine, one browser after another: 461 s; on a Grid, one node per browser: 168 s

Nine configurations cover all 37 pairs of values, a quarter of the 36. Nine is also the least any selection could manage: browser and platform alone make 3 × 3 = 9 pairs, and each configuration contains only one of them. Every browser meets every platform, every screen size and every network speed; every platform meets both screens and both networks. If Firefox breaks only on phone-sized screens, configuration 3 or 7 will show it.

munotes.in253

Cross-Browser and Compatibility Testing

The honest limit is the other side of the same fact: a defect that needs three particular values together, say Edge on Ubuntu on a slow network, may not be in the nine. Pairwise testing is a bet on the assumption above, and it is made knowingly. Where a particular triple is known to be risky, it is added by hand.

The last line answers the practical's second question. Run one browser after another on one machine, the suite takes 142 + 168 + 151 = 461 seconds; run on a Grid with a node for each browser, the three runs go in parallel and the whole takes as long as the slowest, 168 seconds.

Selenium Grid

Selenium Grid runs WebDriver tests on remote machines, so that one test suite can drive many browsers on many platforms from a single entry point. Its documentation describes three ways to deploy it.

  • Standalone: "Standalone combines all Grid components seamlessly into one", in a single process on a single machine, and "is also the easiest mode to spin up a Selenium Grid." Its listed uses include running quick suites before pushing code and a simple Grid inside a CI tool such as Jenkins.
  • Hub and Node: described as "the most used role because it allows to" combine different machines in one Grid, including "Machines with different operating systems and/or browser versions", to have "a single entry point to run WebDriver tests in different environments", and to scale capacity without tearing the Grid down. The Hub receives the tests; each Node is a machine with browsers, which on startup "will detect the available drivers that it can use".
  • Distributed: "each component is started separately, and ideally on different machines."

Inside, the documentation says, "Grid is composed by six different components", and a Hub is made of five of them, "Router, Distributor, Session Map, New Session Queue, and Event Bus":

ComponentWhat it does (Selenium documentation)
Router"redirects new session requests to the queue, and redirects running sessions requests to the Node running that session"
New Session Queue"adds new session requests to a queue, which will be queried by the Distributor"
Distributor"queries the New Session Queue for new session requests, and assigns them to a Node when the capabilities match"
Session Map"maps session IDs to the Node where the session is running"
Event Bus"enables internal communication between different Grid components"
NodeRuns the browser sessions; it detects the drivers available on its machine
munotes.in254

Cross-Browser and Compatibility Testing

A test asks the Grid for a browser by its capabilities, for example Firefox on a particular platform; the Distributor finds a Node that has them and starts the session there. That is how the practical's instruction to "Execute automation scripts across multiple browsers (Chrome, Firefox, Edge) using Selenium Grid" is carried out: one suite, three requests for different capabilities, three Nodes working at once.

Recording compatibility issues

A compatibility failure is only useful if it says where it happened. Each defect report from cross-browser testing should record the browser and its version, the platform and its version, the screen size and the network, alongside the steps and the expected and actual results; without them, a developer on a different machine will often be unable to reproduce the failure. Chapter Seventy-Six, on writing a defect report, sets out the full report.

What it does not mean

Cross-browser testing is not running the suite once in the developer's favourite browser. It is the same tests on every supported combination, chosen deliberately.

Pairwise selection does not guarantee every defect is found. It guarantees that every pair of values is tried; defects needing three specific values may slip through.

Selenium Grid does not write tests. It runs existing WebDriver tests on remote browsers; the tests themselves are the subject of Chapter Forty-Nine, on driving a browser.

Compatibility is not only about browsers. It includes co-existence with other products in the same environment and interoperability with the systems a product exchanges data with.

Quick revision

  • Compatibility (ISO/IEC 25010:2023): exchanging information with other products, or working while sharing an environment; two sub-characteristics, co-existence and interoperability.
  • Compatibility testing (ISO/IEC/IEEE 29119-1:2022); cross-browser testing runs the same tests on each supported browser and platform.
  • A matrix of browser × platform × screen × network gave 36 configurations.
  • Pairwise testing covers every pair of values; greedy selection reached 9, the minimum (3 browsers × 3 platforms); triples may be missed.
  • Selenium Grid: Standalone, Hub and Node, Distributed; components Router, Distributor, Session Map, New Session Queue, Event Bus, Node; capabilities choose the Node.
  • Execution time: 461 s one browser after another, 168 s in parallel on a Grid.

Test yourself

1. What is compatibility testing? Distinguish co-existence from interoperability. Testing that measures how well a product functions alongside other products in a shared environment and exchanges information with other systems. Co-existence is working efficiently while sharing an environment and its resources without harming other products; interoperability is exchanging information with other products and using the information exchanged.

2. Why cannot every browser and platform combination be tested, and what technique reduces them? Because the combinations multiply: three browsers, three platforms, two screen sizes and two network speeds already give 36 configurations, each needing the whole suite. Pairwise testing reduces them by choosing configurations so that every pair of values appears together at least once, here nine configurations.

munotes.in255

Cross-Browser and Compatibility Testing

3. What is the weakness of pairwise selection? It guarantees coverage of every pair of values, not of every combination of three or more, so a defect that appears only with three particular values together may not be in the selection; known risky combinations are added by hand.

4. Describe the Hub and Node arrangement of Selenium Grid. A Hub is the single entry point that receives WebDriver test requests; it contains the Router, Distributor, Session Map, New Session Queue and Event Bus. Nodes are machines with browsers and drivers that register with the Hub; the Distributor assigns each new session to a Node whose capabilities match the browser and platform the test asked for.

5. Why does Selenium Grid shorten the time for cross-browser testing? Because the suite runs on several Nodes in parallel instead of on one machine in turn, so the total time is that of the slowest browser's run rather than the sum of all of them: 168 seconds instead of 461 in the worked example.

Contents This chapter on its own page

munotes.in256

Chapter Forty-Seven

Debugging: From a Failure Back to Its Fault

Syllabus topic Module 1, "Software Testing Strategies: Strategic approach to software testing"

In one line

Testing shows that a failure happens; debugging finds the fault that causes it, removes it, and makes sure the removal broke nothing else; it proceeds by reproducing the failure, narrowing down its cause by evidence rather than guesswork, fixing it, and retesting.

In the wording a student can write in an examination: to debug is "to detect, locate, and correct faults in a computer program" (ISO/IEC/IEEE 24765). In the ISTQB syllabus's words, "Testing and debugging are separate activities": when testing triggers a failure, "debugging is concerned with finding causes of this failure (defects), analyzing these causes, and eliminating them", through "Reproduction of a failure", "Diagnosis (finding the defect)" and "Fixing the defect". Diagnosis follows one of three broad approaches: brute force (collect all the evidence and search it), backtracking (trace backwards from the symptom to the cause), and cause elimination (form hypotheses and test them away, down to bisection). The fix is followed by confirmation testing and regression testing, because a fix is itself a change that can introduce a new fault.

Testing and debugging are different jobs

Chapter One, on what software testing is, separated the two, and the difference matters in practice.

TestingDebugging
Starts fromRequirements and expected resultsAn observed failure or a reported defect
AimShow that failures occur, or find defectsFind the fault behind a failure, and remove it
ResultA pass or fail verdict, and a defect reportA corrected program
Can be planned in advanceYes: test cases are designed before they are runOnly partly: where the fault lies is unknown until it is found
Usually done byTesters, and developers for their own unitsThe developer who owns the code
Needs knowledge of the codeNot for black-box testingAlways

The word itself has a story. Chapter Two, on errors, faults and failures, told of the moth found in the Harvard Mark II in 1947 and taped into its log book; "bug" and "debug", the museum's record says, "soon became a standard part of the language of computer programmers."

Step 1: reproduce the failure

A failure that cannot be made to happen again cannot be diagnosed with confidence, and a fix for it cannot be confirmed. The first job is a reliable way to produce it: the same inputs, the same data, the same environment and, if timing matters, the same sequence of events. The defect report from testing (Chapter Seventy-Six, on writing a defect report) is where this starts, which is why a good report gives exact steps, data and environment.

Then make the failing case as small as it can be while still failing. If a registration with nine papers, two backlog papers and a concession charges the wrong fee, does one paper with a concession still fail? Every detail removed is a place the fault is not.

munotes.in257

Debugging: From a Failure Back to Its Fault

Step 2: diagnose, three ways

Diagnosis is the hard part, and there are three broad ways to go about it.

Brute force. Collect as much evidence as possible and search it: add print statements or log lines everywhere, take a dump, a "display of some aspect of a computer program's execution state, usually the contents of internal storage or registers", or record a trace, a "record of the execution of a computer program, showing the sequence of instructions executed, the names and values of variables, or both" (both ISO/IEC/IEEE 24765). It needs the least thought and the most effort, and it drowns the one relevant fact in thousands of irrelevant ones. It is the usual first attempt and the right last resort.

Backtracking. Start where the failure shows and work backwards through the code, asking at each step where the wrong value came from. The fee shown is Rs 100 too low: which line set the fee last? What were its inputs? Where did those come from? In a small unit backtracking is quick and precise; in a large program the number of paths leading to a statement grows until tracing them all by hand is impractical.

Cause elimination. Form hypotheses about the cause and design an experiment that rules each in or out. Is it the concession? Rerun without the concession. Is it the backlog fee? Rerun with no backlog papers. Each experiment halves, or better, the space where the fault can be. Its sharpest form is bisection: when a test passed on an old version and fails on today's, test the version halfway between, and keep halving.

Most real debugging mixes the three: a little brute-force logging to see the shape of the problem, a hypothesis, an experiment, some backtracking once the fault is cornered in a few lines.

Worked example: bisection over a project's history

ExamReg's regression suite, run on today's build, finds that a concession student three days late with two backlog papers is charged Rs 100 instead of Rs 400. The same test passed at the start of the term, forty commits ago. Nobody knows which change broke it.

Git provides this as a command. Its documentation describes it: "This command uses a binary search algorithm to find which commit in your project's history introduced a bug. You use it by first telling it a 'bad' commit that is known to contain the bug, and a 'good' commit that is known to be before the bug was introduced." And if a script can decide good or bad, git bisect run does the searching automatically. The program below does the same thing to a simulated history of forty versions of the fee code, running the failing test on each version it chooses.

munotes.in258

Debugging: From a Failure Back to Its Fault

FORM_FEE, BACKLOG_FEE = 800, 150

def make_version(commit):                    # the fee code as it stood after each commit
    def total_fee(days_late, backlog_papers, concession):
        fee = FORM_FEE + BACKLOG_FEE * backlog_papers
        if concession:
            fee = 0 if commit >= 27 else fee - FORM_FEE     # commit 27 "simplified" this line
        return fee + (0 if days_late == 0 else (100 if days_late <= 7 else 500))
    return total_fee

history = [make_version(c) for c in range(40)]     # commits 0 to 39; 39 is today's build

def good(commit):                                  # the failing test, run on one version
    return history[commit](3, 2, True) == 400

lo, hi, runs = 0, 39, 0                            # commit 0 known good, commit 39 known bad
while hi - lo > 1:
    mid = (lo + hi) // 2
    runs += 1
    verdict = "good" if good(mid) else "bad"
    print(f"test commit {mid:>2}: {verdict}")
    if verdict == "good":
        lo = mid
    else:
        hi = mid
print(f"first bad commit: {hi}, found with {runs} test runs among {len(history)} commits")
test commit 19: good
test commit 29: bad
test commit 24: good
test commit 26: good
test commit 27: bad
first bad commit: 27, found with 5 test runs among 40 commits

Five test runs name commit 27 as the first bad one, where testing the 38 commits between the two known ends one by one could have taken 38. Each run halves the commits that could be responsible, so even a history of a thousand commits needs only about ten. Bisection does not say what is wrong with commit 27, but it turns somewhere in forty changes into somewhere in this one change, and one change is small enough to read. Reading it shows the "simplification": the concession now sets the fee to zero instead of subtracting the form fee, wiping out the backlog fees, the very defect Chapter Thirty's walkthrough found on paper.

Tools that help

  • A breakpoint is a "point in a computer program at which execution can be suspended to permit manual or automated monitoring of program performance or results" (ISO/IEC/IEEE 24765). A debugger lets a developer stop at a breakpoint, inspect every variable and step through the code line by line.
  • Instrumentation, "devices or instructions installed or inserted into hardware or software to monitor the operation of a system or component", includes logging: the log written during testing (Chapter Nine, on test execution) is often the best evidence a debugger has.
  • Traces and dumps, above, for brute-force searching.
  • Version history, for bisection and for reading what changed.
munotes.in259

Debugging: From a Failure Back to Its Fault

Step 3: fix it, and make sure the fix is the only change

A fix is a change to the program, and changes are how faults get in. Lyu's handbook names this plainly as one of the two ways software reliability falls: "incorrect modifications to the software" (Chapter Twenty-Three, on quality in software development). Chapter Thirty-Two, on test levels and test types, ran exactly such a case: a fix for day 15 that passed its confirmation test and broke day 8.

Before a fix is written, three questions are worth asking:

  1. Is the same mistake made elsewhere? If one form forgot to validate its input, others may have too.
  2. What could this fix break? The answer decides which regression tests to run, the impact analysis of Chapter Eighteen, on the role of testing in each phase.
  3. What would have prevented it? A review checklist item, a unit test, a clearer requirement: the answer goes back to quality assurance, and to the root cause, the "source of a defect such that if it is removed, the defect is decreased or removed" (ISO/IEC/IEEE 24765).

After the fix, the ISTQB syllabus sets out what follows: "Subsequent confirmation testing checks whether the fixes resolved the problem", preferably "done by the same person who performed the initial test", and "Subsequent regression testing can also be performed, to check whether the fixes are causing failures in other parts of the test object."

What it does not mean

Debugging is not testing. Testing finds that a failure happens; debugging finds and removes its cause. They are done by different people with different knowledge.

Debugging is not guessing. Changing code until the symptom goes away can hide a fault instead of removing it; each change should test a stated hypothesis.

A fix is not finished when the failing test passes. It is finished when the regression tests also pass, and the question of prevention has been asked.

Bisection does not explain a fault. It finds the change that introduced it; someone still has to read that change and understand it.

Quick revision

  • Debug: "to detect, locate, and correct faults in a computer program" (ISO/IEC/IEEE 24765). Testing and debugging are separate activities (ISTQB).
  • Process: reproduce (reliably, then minimise), diagnose, fix; then confirmation and regression testing.
  • Brute force: collect dumps, traces and logs and search them; most effort, least thought.
  • Backtracking: from the symptom backwards to the cause; good in small units.
  • Cause elimination: hypotheses and experiments; bisection halves the search each time (git bisect: "a binary search algorithm to find which commit ... introduced a bug").
  • Worked example: 5 test runs found commit 27 among 40.
  • Tools: breakpoints, debuggers, instrumentation and logs, traces, dumps, version history.
  • Fixes can introduce faults ("incorrect modifications", Lyu): ask where else, what it could break, and what would have prevented it; find the root cause.
munotes.in260

Debugging: From a Failure Back to Its Fault

Test yourself

1. Distinguish testing from debugging. Testing starts from requirements and expected results and shows that failures occur, producing a verdict and a defect report; it can be planned in advance and, for black-box testing, needs no knowledge of the code. Debugging starts from a failure, finds the fault that causes it and removes it, producing a corrected program; it is done by the developer who owns the code.

2. What are the steps of the debugging process? Reproduce the failure reliably and reduce it to the smallest failing case; diagnose it, finding the defect; fix the defect; then confirm the fix with the failing test and run regression tests to check that nothing else broke.

3. Explain brute force, backtracking and cause elimination. Brute force collects all available evidence, such as logs, traces and dumps, and searches it for the fault. Backtracking starts at the point where the failure appears and traces backwards through the code to where the wrong value came from. Cause elimination forms hypotheses about the cause and designs experiments that rule them in or out, narrowing the search until the fault is found.

4. How does bisection find the change that introduced a defect? Given one version known to be good and a later one known to be bad, it tests the version halfway between, and keeps whichever half contains the change from good to bad, halving the range each time. Forty commits need about five test runs; Git automates this with git bisect.

5. Why can fixing one fault introduce another, and how is the risk controlled? Because a fix is itself a change to the program, and it can disturb behaviour elsewhere or be wrong in itself. The risk is controlled by checking where else the same mistake occurs, analysing what the fix could affect, running confirmation and regression tests, and finding the root cause so the fault is prevented rather than only patched.

Contents This chapter on its own page

munotes.in261

Chapter Forty-Eight

Test Automation: What to Automate and What Not To

Syllabus topic Practical, "Creation of Test Suite Using Selenium IDE"

In one line

Test automation lets a machine run tests again and again at almost no cost per run, but writing and maintaining the automated tests costs a great deal, so automation pays only for tests that are run often, are stable, and can be checked mechanically; everything that needs judgement stays with people.

In the wording a student can write in an examination: test automation is the use of tools to execute tests, compare results and report them. The ISTQB syllabus lists its potential benefits, among them "Time saved by reducing repetitive manual work", "Prevention of simple human errors through greater consistency and repeatability", "Reduced test execution times to provide earlier defect detection, faster feedback and faster time to market" and "More time for testers to design new, deeper and more effective tests"; and its risks, among them "Unrealistic expectations about the benefits of a tool" and "Using a test tool when manual testing is more appropriate." Record and playback, as in Selenium IDE, captures a tester's actions in a browser as a script that can be replayed; it is quick to start but brittle, and serious suites are exported and restructured. Automation follows the test pyramid: many automated unit tests, fewer service tests, few end-to-end UI tests.

What automation buys

The ISTQB syllabus's full list of potential benefits is short and worth knowing:

  • "Time saved by reducing repetitive manual work (e.g., execute regression tests, re-enter the same test data, compare expected results vs actual results, and check against coding standards)"
  • "Prevention of simple human errors through greater consistency and repeatability"
  • "More objective assessment (e.g., coverage) and providing measures that are too complicated for humans to determine"
  • "Easier access to information about testing to support test management and test reporting"
  • "Reduced test execution times to provide earlier defect detection, faster feedback and faster time to market"
  • "More time for testers to design new, deeper and more effective tests"

Every item is about repetition. A tool does the same thing the same way every time, which is exactly what regression testing needs: the syllabus calls regression suites "a strong candidate for automation" because they are run after every change and grow with every release (Chapter Thirty-Two, on test levels and test types).

What automation costs

The same section opens with a warning: "Simply acquiring a tool does not guarantee success. Each new tool will require effort to achieve real and lasting benefits (e.g., for tool introduction, maintenance and training)." Among the risks it lists:

  • "Unrealistic expectations about the benefits of a tool (including functionality and ease of use)."
  • "Inaccurate estimations of time, costs, effort required to introduce a tool, maintain test scripts and change the existing manual test process."
  • "Using a test tool when manual testing is more appropriate."
  • "Relying on a tool too much, e.g., ignoring the need of human critical thinking."
  • Dependence on a vendor, or on open-source software that may be abandoned, and a tool "not compatible with the development platform".
munotes.in262

Test Automation: What to Automate and What Not To

The cost that surprises teams most is maintenance. An automated test is code. When the application's pages change, the tests that drive them break even though the application works, and someone must repair them before the suite means anything again.

Worked example: when does automating ExamReg's regression suite pay?

ExamReg's regression suite has 120 test cases. Run by hand, at about three minutes each, one run takes about 6 hours. Automating it takes, say, 60 hours to write and debug once, and each automated run then costs a quarter of an hour to start and read, plus whatever it takes to repair the scripts the application's changes have broken since the last run. These hours are this book's illustration; the program shows how the answer depends on the one number teams most often forget, the upkeep.

import math

manual_per_run = 6.0          # hours: 120 regression cases by hand, about 3 minutes each
build = 60.0                  # hours: writing and debugging the automated suite, once
run_per_run = 0.25            # hours: starting the automated run and reading its report

def break_even(upkeep_per_run):
    """Runs after which automation has cost no more than doing the same runs by hand."""
    saving = manual_per_run - run_per_run - upkeep_per_run
    return math.ceil(build / saving) if saving > 0 else None

for label, upkeep in [("well-structured scripts", 0.75), ("recorded, brittle scripts", 4.0),
                      ("scripts repaired almost every run", 5.9)]:
    n = break_even(upkeep)
    print(f"{label:<34} upkeep {upkeep:>4} h/run: "
          + (f"pays for itself after {n} runs" if n else "never pays for itself"))

print("runs   by hand   automated (well-structured)")
for n in (1, 5, 12, 20, 40):
    print(f"{n:>4} {n * manual_per_run:>9.1f} {build + n * (run_per_run + 0.75):>11.1f}")
well-structured scripts            upkeep 0.75 h/run: pays for itself after 12 runs
recorded, brittle scripts          upkeep  4.0 h/run: pays for itself after 35 runs
scripts repaired almost every run  upkeep  5.9 h/run: never pays for itself
runs   by hand   automated (well-structured)
   1       6.0        61.0
   5      30.0        65.0
  12      72.0        72.0
  20     120.0        80.0
  40     240.0       100.0

With well-structured scripts, which need about three quarters of an hour of repair per run, automation costs more for the first eleven runs and breaks even at the twelfth; by the fortieth run it has cost 100 hours against 240 by hand. If ExamReg's suite runs on every build through continuous integration (Chapter Thirty-Nine), twelve runs come in a few days, and the investment is repaid almost at once.

With brittle scripts the picture changes completely. Four hours of repair every run pushes break-even out to the thirty-fifth run, and if the scripts need repairing nearly as long as the manual run itself, automation never pays at all. That is the ISTQB syllabus's risk of "Inaccurate estimations of time, costs, effort required to ... maintain test scripts" in numbers, and it is why the next three chapters, on waits, data-driven testing and the page object model, are about making scripts that survive change.

munotes.in263

Test Automation: What to Automate and What Not To

Record and playback: Selenium IDE

The practical begins with Selenium IDE, which its project describes as "Open source record and playback test automation for the web". The tester works through the application in the browser; the IDE records each click and keystroke as a command; replaying the recording repeats the actions and checks what the tester asked it to check. Its page lists the features that make recorded tests less fragile: "Selenium IDE records multiple locators for each element it interacts with. If one locator fails during playback, the others will be tried until one is successful"; a recorded test can "re-use one test case inside of another" with the run command, "allowing you to re-use your login logic in multiple places throughout a suite"; and its tests can run "on any browser/OS combination in parallel" through a command-line runner.

Record and playback is the fastest way to start automating, and a good way to learn what a browser test does. Its limits are the reasons the practical then asks students to "export the automation script in WebDriver format":

  • Brittleness. A recording is tied to the page as it was on the day it was recorded; a renamed button or a moved field can break it.
  • Timing. A recording replays at its own pace; a page that loads more slowly than it did during recording can make a step fail. Chapter Forty-Nine's waits deal with this.
  • Hard-wired data. A recording repeats the same inputs; testing many inputs means many recordings, or the data-driven testing of Chapter Fifty.
  • Duplication. Every recording that logs in repeats the login steps; when the login page changes, every one breaks. The page object model of Chapter Fifty-One puts each page's details in one place.

What to automate, and what not to

AutomateKeep manual
Regression suites run after every changeExploratory testing, which designs the next test from the last result (Chapter Sixty-Three)
Smoke tests run on every buildUsability: whether a first-year student understands the form
Unit and service tests, the base of the pyramidTests of features that are still changing every week
Data-heavy checks: every fee for every day and concessionOne-off checks that will never be repeated
Performance and load tests, which need many simultaneous usersJudgements about look and feel, wording and helpfulness
Checks across many browsers and platformsAnything whose expected result a machine cannot decide
munotes.in264

Test Automation: What to Automate and What Not To

The rule under the table is the ISTQB syllabus's warning against "Relying on a tool too much, e.g., ignoring the need of human critical thinking." A tool checks what it was told to check. Deciding what to check, and noticing what nobody thought to check, remains the tester's work.

The test automation pyramid

Chapter Seventeen, on agile, Scrum and DevOps, introduced the test pyramid: many small, fast unit tests at the base, fewer service tests in the middle, few broad UI tests at the top. It is also the shape of a good automation portfolio. Unit tests are cheap to write, fast to run and rarely break for reasons unrelated to a real defect; end-to-end UI tests through a browser are slow, and they break whenever a page changes. A suite built mostly from recorded UI tests is the inverted pyramid, and the break-even arithmetic above explains why it disappoints.

The test pyramid: many unit tests, fewer service tests, few UI tests

Figure 48.1 The test pyramid: automate mostly at the base, where tests are small, fast and stable

What it does not mean

Automation does not replace testers. It replaces repetition; the design of tests, exploration and judgement remain human work.

Automated tests are not free after they are written. They must be maintained as the application changes, and that upkeep decides whether automation pays.

Record and playback is not a finished automation suite. It is a starting point, to be exported and restructured for anything that must last.

More automated UI tests are not always better. A suite should be mostly unit and service tests, with a few end-to-end tests at the top.

Quick revision

  • Benefits (ISTQB v4.0.1, 6.2): less repetitive work, fewer simple human errors, objective measures, easier reporting, faster feedback, more time for deeper tests.
  • Risks: unrealistic expectations, underestimated costs of introduction and maintenance, automating what should be manual, over-reliance, vendor or open-source dependence, platform incompatibility.
  • Worked example: 60 hours to build, 6 hours by hand per run; break-even after 12 runs with well-structured scripts, 35 with brittle ones, never if every run needs repair.
  • Selenium IDE: record and playback; multiple locators per element; reusable test cases; export to WebDriver.
  • Record and playback's limits: brittleness, timing, hard-wired data, duplication.
  • Automate the repetitive and mechanical; keep exploration, usability and judgement manual; follow the test pyramid.

Test yourself

1. State four benefits and four risks of test automation. Benefits: time saved on repetitive work such as regression tests; fewer simple human errors through consistency; faster execution, giving earlier defect detection and feedback; and more time for testers to design deeper tests. Risks: unrealistic expectations of the tool; underestimated effort to introduce it and maintain scripts; automating tests that should be manual; and over-reliance on the tool at the expense of critical thinking.

munotes.in265

Test Automation: What to Automate and What Not To

2. Why does the maintenance of automated tests decide whether automation pays? Because automation costs a large amount once and a small amount per run, and saves the cost of a manual run each time. Maintenance is part of the cost per run; if scripts break often and take long to repair, the saving per run shrinks, and the number of runs needed to recover the building cost grows, possibly without limit.

3. What is record and playback, and what are its limits? A way of automating browser tests by recording a tester's actions as a script and replaying them, as Selenium IDE does. Its limits are brittleness when pages change, failures when timing differs, inputs hard-wired into each recording, and steps such as login duplicated across many recordings.

4. Which tests should be automated, and which should not? Automate tests that are run often and checked mechanically: regression and smoke suites, unit and service tests, data-heavy checks, load tests and cross-browser runs. Keep manual the tests that need judgement or change constantly: exploratory testing, usability, look and feel, one-off checks, and features still being redesigned.

5. How does the test pyramid guide automation? It says to automate mostly at the base, with many small, fast, stable unit tests; fewer tests at the service level; and only a few end-to-end tests through the user interface, which are slow and break whenever pages change.

Contents This chapter on its own page

munotes.in266

Chapter Forty-Nine

Driving a Browser: WebDriver, Locators and Waits

Syllabus topic Practical, "Web Element Interaction and Synchronization Handling"

In one line

WebDriver is a standard protocol through which a test program drives a real browser; the test finds elements on the page with locators, acts on them, and asserts on what happens, and because pages change after they load, it must wait for the elements it needs, preferably with explicit waits for exact conditions.

In the wording a student can write in an examination: WebDriver "is a remote control interface that enables introspection and control of user agents. It provides a platform- and language-neutral wire protocol as a way for out-of-process programs to remotely instruct the behavior of web browsers" (W3C WebDriver specification). A local end, a language library such as Selenium's, sends commands to a remote end, the browser's driver, within a session. Elements are found with locators: the specification defines five location strategies (css selector, link text, partial link text, tag name, xpath), and Selenium supports eight (adding class name, id and name). Because a test can run faster than a page changes, it must synchronise: an implicit wait is "a global setting that applies to every element location call for the entire session"; an explicit wait is a loop that polls "for a specific condition to evaluate as true" up to a timeout. The two must not be mixed.

The practical's tasks

This chapter covers two of the practical's exercises. The first: "Write an automation script to perform login on a specified web page. Verify successful login using assertions such as URL validation, welcome message validation, and logout visibility check." The second: "Develop a script to automate interaction with various web elements (textbox, radio button, checkbox, dropdown, alert, and frames). Implement implicit and explicit waits for synchronization handling."

WebDriver: a standard remote control for browsers

WebDriver is now a W3C specification, which says of itself that it "is derived from the popular Selenium WebDriver browser automation framework". Its opening describes its purpose: "It is primarily intended to allow web authors to write tests that automate a user agent from a separate controlling process". The specification describes the protocol as communication between two ends.

  • The local end "represents the client side of the protocol, which is usually in the form of language-specific libraries providing an API on top of the WebDriver protocol". Selenium's Java, Python and other bindings are local ends.
  • The remote end "hosts the server side of the protocol". It is either an intermediary node, which acts as a proxy (Selenium Grid's Hub, from the cross-browser testing of Chapter Forty-Six, is one), or an endpoint node, "implemented by a user agent or a similar program": the browser's own driver.

A test opens a session with the remote end, sends commands in it (go to this address, find this element, click it, read its text) and closes it. Because the protocol is a standard, one test program can drive Chrome, Firefox and Edge, which is what makes the cross-browser testing of Chapter Forty-Six possible.

munotes.in267

Driving a Browser: WebDriver, Locators and Waits

Locators: finding elements on the page

Before a test can type into a field or read a message, it must find the element. The specification defines an element location strategy as "an enumerated attribute deciding what technique should be used to search for elements", and lists five:

Specification's strategyKeywordFinds elements by
CSS selectorcss selectorA CSS selector, the language used to style pages
Link text selectorlink textThe exact visible text of a link
Partial link text selectorpartial link textPart of a link's visible text
Tag nametag nameThe element's tag, such as input
XPath selectorxpathAn XPath expression over the page's structure

Selenium's documentation lists "8 traditional location strategies", adding three that are the commonest in practice: class name ("Locates elements whose class name contains the search value"), id ("Locates elements whose ID attribute matches the search value") and name ("Locates elements whose NAME attribute matches the search value").

Choosing a locator is choosing how the test will break when the page changes. An id that the developers keep stable is the sturdiest; a long xpath that walks from the top of the page through every nested element breaks the moment anything above the target moves. A good test uses the most stable locator available, and a team that wants stable tests asks its developers for stable ids on the elements tests need.

Why tests must wait

Selenium's documentation names the problem: "Perhaps the most common challenge for browser automation is ensuring that the web application is in a state to execute a particular Selenium command as desired. The processes often end up in a race condition where sometimes the browser gets into the right state first (things work as intended) and sometimes the Selenium code executes first (things do not work as intended). This is one of the primary causes of flaky tests." A page's script may add an element a moment after the page loads, or reveal it only after a click, and "An element must be both present and displayed on the page in order for Selenium to interact with it."

The obvious remedy, a pause of a fixed length, is poor: "Because the code can't know exactly how long it needs to wait, this can fail when it doesn't sleep long enough. Alternately, if the value is set too high and a sleep statement is added in every place it is needed, the duration of the session can become prohibitive." Selenium provides two better mechanisms.

munotes.in268

Driving a Browser: WebDriver, Locators and Waits

  • Implicit wait: "a global setting that applies to every element location call for the entire session". Every search for an element keeps trying for up to the set time before reporting that the element is missing, and "as soon as the element is located, the driver will return the element reference".
  • Explicit wait: "loops added to the code that poll the application for a specific condition to evaluate as true before it exits the loop and continues to the next command in the code. If the condition is not met before a designated timeout value, the code will give a timeout error." Its advantage is precision: "explicit waits are a great choice to specify the exact condition to wait for in each place it is needed."

The documentation's warning is emphatic: "Do not mix implicit and explicit waits. Doing so can cause unpredictable wait times. For example, setting an implicit wait of 10 seconds and an explicit wait of 15 seconds could cause a timeout to occur after 20 seconds."

Worked example: four ways to wait for a welcome message

The program simulates what the practical does with a real browser. After ExamReg's login, the home page's script shows the welcome message 1.8 seconds later and the logout link 2.4 seconds later, as a real page might. A simulated clock replaces real time, so the numbers are the same on every run. The test tries four ways of reading the welcome message, then performs the practical's three login assertions with explicit waits.

class NoSuchElement(Exception):
    pass

class Clock:                                  # simulated time, so every run gives the same numbers
    def __init__(self):
        self.now = 0.0
    def sleep(self, seconds):
        self.now += seconds

class FakeBrowser:                            # stands in for a real browser behind WebDriver
    def __init__(self, clock):
        self.clock, self.url, self.elements = clock, "https://examreg.college.example/login", {}
    def log_in(self):                         # the page answers after a short, variable delay
        self.url = "https://examreg.college.example/home"
        t = self.clock.now
        self.elements = {("id", "welcome"): ("Welcome, 2026CS014", t + 1.8),
                         ("id", "logout"): ("Log out", t + 2.4)}
    def find_element(self, by, value):
        text, appears_at = self.elements.get((by, value), (None, float("inf")))
        if self.clock.now < appears_at:
            raise NoSuchElement(f"{by}={value}")
        return text

def explicit_wait(browser, clock, by, value, timeout=10.0, poll=0.5):
    """Poll for the element until it is there or the timeout passes."""
    waited = 0.0
    while True:
        try:
            return browser.find_element(by, value)
        except NoSuchElement:
            if waited >= timeout:
                raise
            clock.sleep(poll)
            waited += poll

def attempt(strategy):
    clock = Clock()
    browser = FakeBrowser(clock)
    browser.log_in()
    try:
        if strategy == "explicit wait":
            text = explicit_wait(browser, clock, "id", "welcome")
        else:
            clock.sleep({"no wait": 0, "sleep 1 s": 1, "sleep 5 s": 5}[strategy])
            text = browser.find_element("id", "welcome")
        return f"found '{text}' after {clock.now:.1f} s"
    except NoSuchElement as missing:
        return f"FAIL, {missing} not there yet at {clock.now:.1f} s"

for strategy in ["no wait", "sleep 1 s", "sleep 5 s", "explicit wait"]:
    print(f"{strategy:<14} {attempt(strategy)}")

clock = Clock()                               # the practical's login check, with explicit waits
browser = FakeBrowser(clock)
browser.log_in()
checks = {"URL is the home page": browser.url.endswith("/home"),
          "welcome message names the student": "2026CS014" in explicit_wait(browser, clock, "id", "welcome"),
          "logout link is visible": explicit_wait(browser, clock, "id", "logout") == "Log out"}
for name, ok in checks.items():
    print(f"assert {name}: {'pass' if ok else 'FAIL'}")
munotes.in269

Driving a Browser: WebDriver, Locators and Waits

no wait        FAIL, id=welcome not there yet at 0.0 s
sleep 1 s      FAIL, id=welcome not there yet at 1.0 s
sleep 5 s      found 'Welcome, 2026CS014' after 5.0 s
explicit wait  found 'Welcome, 2026CS014' after 2.0 s
assert URL is the home page: pass
assert welcome message names the student: pass
assert logout link is visible: pass

The four strategies behave exactly as the documentation predicts. With no wait the test looks at once and fails; with a one-second sleep it looks too early and fails. On a real machine these two would fail on some runs and pass on others, depending on how fast the page loaded that time: a flaky test. A five-second sleep passes, but spends 5 seconds where 2 were enough, on every run, at every place it is used. The explicit wait polls every half second and returns at 2.0 seconds, as soon as the message is there, and would give up with a clear error after 10 seconds if it never appeared.

The three assertions are the practical's login check: the address is the home page, the welcome message names the student, and the logout link is visible. Each element is read through an explicit wait, so the check is both fast and stable.

The same waits in Selenium

In Selenium's Python binding, an implicit wait is one line, set once for the session:

driver.implicitly_wait(2)

An explicit wait names its condition where it is needed. The documentation's form waits with a lambda; here the condition is the welcome message being displayed:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.wait import WebDriverWait

wait = WebDriverWait(driver, timeout=10)
wait.until(lambda _: driver.find_element(By.ID, "welcome").is_displayed())

A suite should choose one approach and use it consistently; the documentation's warning against mixing them is about exactly the unpredictable timings a mixed suite produces.

Interacting with other elements

The second exercise's elements each have a characteristic trap:

  • Text boxes: clear any existing value before typing, and assert on the value after.
  • Radio buttons and check boxes: check the selected state before and after clicking, since clicking a selected check box clears it.
  • Drop-down lists: select by visible text or value, not by position, which changes when options are added.
  • Alerts: a browser alert blocks the page until it is accepted or dismissed; the test must switch to it.
  • Frames: an element inside a frame cannot be found until the test switches into that frame, and must switch back afterwards.
munotes.in270

Driving a Browser: WebDriver, Locators and Waits

Every one of them can appear late, which is why the waits above apply to all of them.

What it does not mean

WebDriver is not Selenium. Selenium is a project whose libraries are local ends; WebDriver is the W3C standard protocol they, and the browsers' drivers, speak.

A fixed sleep is not synchronisation. It is a guess about timing, too short on a slow day and wasteful on a fast one.

An implicit wait does not fix every timing problem. It waits only for elements to be found, not for a condition such as a value changing or a spinner disappearing; explicit waits handle those.

A passing browser test is not proof of a stable test. A test that passes today and fails tomorrow with no change to the application is flaky, and usually a waiting problem.

Quick revision

  • WebDriver (W3C): a remote control interface and wire protocol for browsers; local end (language library), remote end (intermediary node, such as a Grid hub, or endpoint node, the browser's driver); commands in a session.
  • Location strategies (specification): css selector, link text, partial link text, tag name, xpath. Selenium adds class name, id and name: 8 locators. Prefer stable ids.
  • Flaky tests come from a race condition between the test and the page.
  • Implicit wait: global, every element search; explicit wait: polls for a specific condition until a timeout. Do not mix them.
  • Worked example: no wait and a 1 s sleep failed; a 5 s sleep passed slowly; the explicit wait passed at 2.0 s.
  • Login check: assert the URL, the welcome message and the logout link's visibility.

Test yourself

1. What is WebDriver? Explain the local end and the remote end. A W3C-standard remote control interface and wire protocol that lets a separate program control a web browser. The local end is the client side, usually a language library such as Selenium's, which sends commands; the remote end is the server side that carries them out, either an intermediary node such as a Grid hub or an endpoint node, the browser's driver.

2. Name the locator strategies Selenium supports. Which is usually the most robust, and why? Class name, css selector, id, name, link text, partial link text, tag name and xpath. A stable id is usually the most robust, because it does not change when the page's layout or text changes, unlike a long XPath or link text.

3. Why do browser tests need synchronisation, and why is a fixed sleep a poor solution? Because pages change after they load, and a test may try to use an element before it exists or is displayed, a race that makes tests flaky. A fixed sleep guesses the delay: too short and the test still fails on slow runs, too long and every run wastes time wherever it is used.

munotes.in271

Driving a Browser: WebDriver, Locators and Waits

4. Distinguish an implicit wait from an explicit wait. An implicit wait is a global session setting that makes every element search keep trying for up to a set time. An explicit wait is written where it is needed and polls for one specific condition, such as an element being displayed, until it is true or a timeout passes. They should not be mixed, because the combination causes unpredictable wait times.

5. Write the assertions for a login check, as the practical asks. After logging in, assert that the current URL is the expected home page, that the welcome message is present and names the logged-in user, and that the logout link is visible, reading each element through an explicit wait.

Contents This chapter on its own page

munotes.in272

Chapter Fifty

Data-Driven Testing

Syllabus topic Practical, "Data-Driven Testing Using Excel Integration"

In one line

Data-driven testing separates what a test does from the values it does it with: one test script reads its inputs and expected results from a table, such as a spreadsheet, and runs once for every row, so adding a case means adding a row, not writing code; keyword-driven testing goes further and puts the steps themselves in a table of action words.

In the wording a student can write in an examination: the data-driven test approach "separates out the test inputs and expected results, usually into a spreadsheet, and uses a more generic test script that can read the input data and execute the same test script with different data" (ISTQB 2018). In the keyword-driven test approach, "a generic script processes keywords describing the actions to be taken (also called action words), which then calls keyword scripts to process the associated test data"; ISO/IEC/IEEE 29119-5:2024 defines keyword-driven testing as "testing using test cases composed from keywords". Both let people who do not program contribute tests. In Python, unittest's subTest runs one test over many rows and reports each failing row separately; in Java, TestNG's @DataProvider supplies rows and Apache POI reads and writes Excel files.

The practical's task

The practical asks students to "Write a program to update 10 student records in an Excel file and validate changes on the web application. Perform data reading, writing, and verification using Apache POI library." Apache POI describes itself as "the Java API for Microsoft Documents": it lets a Java program open an Excel workbook, read its cells and write new values. The test logic of the exercise does not depend on POI; the same pattern, read the rows, run the test for each, write the results back, is shown below in Python with a CSV file, which is what a spreadsheet is saved as when only the data matters.

Separating the test from its data

A test has two parts: what it does (the steps and checks) and what it does them with (inputs and expected results). Writing both into one script means a new case needs a new script, even when it differs only in its numbers. Separating them gives one script and a table.

The benefits follow directly.

  • New cases cost a row. The exam cell can add a fee case for next session's concession rule without touching the test code.
  • Non-programmers can contribute. The 2018 syllabus notes that with these approaches "testers who are not familiar with the scripting language can also contribute by creating test data and/or keywords for these predefined scripts."
  • The data can be reviewed as data. A table of days late, backlog papers, concessions and expected fees can be checked against the fee rules by someone who knows the rules, which is a review of the test basis.
  • The same data can drive tests at different levels: a unit test of the fee calculator and a browser test of the fee page can read the same sheet.
munotes.in273

Data-Driven Testing

Worked example: ten rows, one test

The fee cases for ExamReg sit in a sheet, one row per case, with its identifier, its inputs and the fee the rules give. The unit under test is build 43 from Chapter Thirty-Nine, on continuous integration, whose "tidied" band charges Rs 100 for a form 8 days late.

case,days_late,backlog_papers,concession,expected
F01,0,0,no,800
F02,7,0,no,900
F03,8,0,no,1300
F04,15,1,no,1450
F05,16,0,no,refused
F06,3,2,yes,400
F07,0,1,yes,150
F08,10,0,yes,500
F09,-1,0,no,refused
F10,1,3,no,1350

The test module reads the sheet and runs one test method over every row, using unittest's subTest. The documentation explains why: "When there are very small differences among your tests, for instance some parameters, unittest allows you to distinguish them inside the body of a test method using the subTest() context manager." After the run, the results are written back as a second sheet, the "writing" half of the practical's task.

import csv
import io
import unittest

FORM_FEE, BACKLOG_FEE = 800, 150

def total_fee(days_late, backlog_papers, concession):      # build 43, with its "tidied" band
    if days_late < 0 or days_late > 15:
        return "refused"
    late = 0 if days_late == 0 else (100 if days_late <= 8 else 500)
    return FORM_FEE + BACKLOG_FEE * backlog_papers - (FORM_FEE if concession else 0) + late

def run_row(row):                                          # one row of the sheet, through the unit
    return str(total_fee(int(row["days_late"]), int(row["backlog_papers"]),
                         row["concession"] == "yes"))

with open("fee-cases.csv") as sheet:                       # the test data, kept out of the code
    CASES = list(csv.DictReader(sheet))

class FeeRules(unittest.TestCase):
    def test_every_row_of_the_sheet(self):                 # ONE test, driven by every row
        for row in CASES:
            with self.subTest(case=row["case"]):
                self.assertEqual(run_row(row), row["expected"])

report = io.StringIO()
result = unittest.TextTestRunner(stream=report).run(
    unittest.defaultTestLoader.loadTestsFromTestCase(FeeRules))
for line in report.getvalue().splitlines():
    if line.startswith(("FAIL:", "AssertionError")):
        print(line)
print(f"{len(CASES)} rows run by one test; {len(result.failures)} failing")

with open("fee-results.csv", "w", newline="") as out:     # results written back, as a sheet
    writer = csv.writer(out)
    writer.writerow(["case", "expected", "actual", "result"])
    for row in CASES:
        actual = run_row(row)
        writer.writerow([row["case"], row["expected"], actual,
                         "pass" if actual == row["expected"] else "FAIL"])
print(open("fee-results.csv").read(), end="")
FAIL: test_every_row_of_the_sheet (__main__.FeeRules.test_every_row_of_the_sheet) (case='F03')
AssertionError: '900' != '1300'
10 rows run by one test; 1 failing
case,expected,actual,result
F01,800,800,pass
F02,900,900,pass
F03,1300,900,FAIL
F04,1450,1450,pass
F05,refused,refused,pass
F06,400,400,pass
F07,150,150,pass
F08,500,500,pass
F09,refused,refused,pass
F10,1350,1350,pass

One test method ran all ten rows. Row F03, 8 days late, failed, and the failure report names the row by its identifier: build 43 charged Rs 900 where the rules give Rs 1,300. The other nine rows ran and passed even though F03 failed first, which is what subTest is for; the documentation notes that without it, "execution would stop after the first failure". The results sheet puts the verdict beside each row, where the exam cell can read it without reading any code.

munotes.in274

Data-Driven Testing

Keyword-driven testing

Data-driven testing puts the data in a table. Keyword-driven testing puts the steps there too. ISO/IEC/IEEE 29119-5 defines a keyword as "one or more words used as a reference to a specific set of actions intended to be performed during the execution of one or more test cases". An automation engineer codes each keyword once in a keyword library; a test case is then a table of keywords with their data, which anyone who knows the application can write.

The program below is ExamReg's login test written that way, against a simulated login page so that it runs anywhere.

class FakeLoginPage:                          # a login page, simulated so the test runs anywhere
    def __init__(self):
        self.fields, self.page, self.message = {}, "login", ""
    def type_into(self, field, text):
        self.fields[field] = text
    def press(self, button):
        if button == "Log in":
            if self.fields == {"roll": "2026CS014", "password": "test-pass-1"}:
                self.page, self.message = "home", "Welcome, 2026CS014"
            else:
                self.message = "Wrong roll number or password"

page = FakeLoginPage()
library = {                                   # the keyword library: each action word, coded once
    "enter": lambda field, text: page.type_into(field, text),
    "press": lambda button: page.press(button),
    "check page": lambda name: page.page == name,
    "check message": lambda text: page.message == text,
}
test_case = [                                 # a keyword-driven test case: a table, no code
    ("enter", "roll", "2026CS014"),
    ("enter", "password", "test-pass-1"),
    ("press", "Log in"),
    ("check page", "home"),
    ("check message", "Welcome, 2026CS014"),
]
for step, (keyword, *data) in enumerate(test_case, 1):
    outcome = library[keyword](*data)
    verdict = "" if outcome is None else ("  pass" if outcome else "  FAIL")
    print(f"{step}. {keyword}: {', '.join(data)}{verdict}")
1. enter: roll, 2026CS014
2. enter: password, test-pass-1
3. press: Log in
4. check page: home  pass
5. check message: Welcome, 2026CS014  pass

The test case is five rows of plain words. If the login page changes, only the code behind enter and press changes, once, and every test written in keywords keeps working. That is the same idea as the page object model of the next chapter, reached from the other side: hide how an action is done behind a name that says what it does.

Data-driven testing in TestNG

In the practical's Java setting, TestNG supports the same pattern directly. Its documentation lists "Support for data-driven testing (with @DataProvider)" among its main features: a method marked @DataProvider returns the rows, and a @Test method that names it is run once for each row. The rows can come from anywhere, including an Excel sheet read with Apache POI.

The pitfalls

  • Wrong data looks like a wrong program. A mistake in an expected value makes a correct program fail, or, worse, a wrong program pass. The sheet is part of the test and needs the same review as code.
  • Rows without names are hard to trace. A failure reported as row 7 is slow to act on; an identifier such as F03, traced to a requirement, is not.
  • More rows are not better tests. A thousand rows chosen without a technique test the same things again and again; the rows should come from equivalence partitioning and boundary value analysis (Chapters Fifty-Four and Fifty-Five).
  • Rows that depend on each other. If row 5 relies on a record created by row 4, the order becomes part of the test and one failure spreads to the next.
  • Spreadsheet formatting. Numbers stored as text, dates reformatted by the spreadsheet program, trailing spaces: the script must read the data the way it was meant, which is one reason simple formats such as CSV are popular for test data.
munotes.in275

Data-Driven Testing

What it does not mean

Data-driven testing is not a test design technique. It is a way of running many cases through one script; which cases to run is decided by the techniques of Module 2.

A spreadsheet of test data is not a test plan. It holds inputs and expected results; objectives, scope and criteria are elsewhere.

Keyword-driven testing is not testing without programming. Someone must code and maintain the keyword library; the table is what non-programmers write.

One failing row is not one failing test to ignore. Each row is a test case, and a failure is reported and handled like any other.

Quick revision

  • Data-driven: test inputs and expected results in a table, "usually into a spreadsheet"; one generic script runs every row (ISTQB 2018).
  • Keyword-driven: tests composed from keywords (action words) coded once in a keyword library (ISO/IEC/IEEE 29119-5:2024).
  • Benefits: new cases cost a row; non-programmers contribute; data reviewable; data reusable across levels.
  • Python: subTest runs one test over many rows and reports each failure separately. Java: TestNG @DataProvider; Apache POI reads and writes Excel.
  • Worked example: 10 rows, 1 test method; F03 (8 days late) failed, Rs 900 against Rs 1,300; results written back as a sheet.
  • Pitfalls: wrong data, untraceable rows, rows without technique, dependent rows, spreadsheet formatting.

Test yourself

1. What is data-driven testing, and what are its benefits? A test approach that separates test inputs and expected results into a table, such as a spreadsheet, and uses one generic script to run the same test with every row. New cases cost only a row, testers who do not program can add cases, the data can be reviewed against the rules, and the same data can drive tests at several levels.

2. Distinguish data-driven from keyword-driven testing. Data-driven testing puts the test data in a table while the steps stay in code. Keyword-driven testing also puts the steps in a table, as keywords or action words, each coded once in a keyword library, so that whole test cases can be written without programming.

munotes.in276

Data-Driven Testing

3. What does unittest's subTest provide in a data-driven test? It lets one test method run over many rows while reporting each failing row separately, with the row's parameters, and without stopping at the first failure.

4. What does Apache POI do in the practical's exercise? It is the Java library that reads and writes Microsoft Office files, so the test program can read student records and test data from an Excel workbook, update them, and write results back.

5. List three pitfalls of data-driven testing. Errors in the data, which make correct programs fail or wrong programs pass; rows chosen without a test design technique, which repeat the same test many times; and rows that depend on each other or on spreadsheet formatting, which make results unreliable.

Contents This chapter on its own page

munotes.in277

Chapter Fifty-One

The Page Object Model

Syllabus topic Practical, "Page Object Model (POM) Framework Implementation"

In one line

A page object is a class that represents a page, or a part of one, and offers the actions a user can take there as methods named for what they do; tests call those methods and never touch the page's HTML, so when the page changes, only the page object changes.

In the wording a student can write in an examination: the Page Object Model (POM) is "a Design Pattern that has become popular in test automation for enhancing test maintenance and reducing code duplication. A page object is an object-oriented class that serves as an interface to a page of your AUT" (application under test) (Selenium documentation). In Martin Fowler's words, "A page object wraps an HTML page, or fragment, with an application-specific API, allowing you to manipulate page elements without digging around in the HTML." Its advantages are "a clean separation between the test code and page-specific code, such as locators" and "a single repository for the services or operations the page offers". Page objects expose services, not internals; they generally make no assertions, which belong in the tests; and their methods return other page objects, so a test reads as a user's journey.

The problem it solves

Chapter Forty-Eight showed what brittle scripts cost: automation that needs repair after every change may never pay for itself. The commonest cause of that brittleness is the one Fowler names: "if you write tests that manipulate the HTML elements directly your tests will be brittle to changes in the UI." When every test that logs in contains the login page's three locators, a redesign of the login page breaks every one of those tests, and each must be found and repaired.

The Selenium documentation states the remedy's benefit in a sentence: "if the UI changes for the page, the tests themselves don't need to change, only the code within the page object needs to change. Subsequently, all changes to support that new UI are located in one place."

What a page object looks like

Fowler's rule of thumb is that a page object "should allow a software client to do anything and see anything that a human can", through "an interface that's easy to program to and hides the underlying widgetry in the window." His guidance on the interface is concrete: "to access a text field you should have accessor methods that take and return a string, check boxes should use booleans, and buttons should be represented by action oriented method names." And the test of a good one: "A good rule of thumb is to imagine changing the concrete control - in which case the page object interface shouldn't change."

The Selenium documentation summarises the rules:

munotes.in278

The Page Object Model

  • "The public methods represent the services that the page or component offers"
  • "Try not to expose the internals of the page or component"
  • "Generally don't make assertions"
  • "Methods return other Page Objects, Page Component Objects, or optionally themselves (for fluent syntax)"
  • "Need not represent an entire page all the time"
  • "Different results for the same action are modelled as different methods"

Two of these deserve a word. On assertions, the documentation is firm: "Page objects themselves should never make verifications or assertions. This is part of your test and should always be within the test's code". Fowler agrees ("I favor having no assertions in page objects"), allowing only checks of a page's invariants, such as that the browser is on the right page when the page object is created. On size, "a Page Object need not represent an entire page": a navigation bar that appears everywhere can be one component object, and "The essential principle is that there is only one place in your test suite with knowledge of the structure of the HTML of a particular (part of a) page."

Worked example: a redesign of the login page

ExamReg's login page is redesigned in release 2: its three fields and button get new ids. The program runs the same three login tests written two ways, first with the locators inside every test, then through a LoginPage and a HomePage page object, against both releases. A simulated browser stands in for the real one, keeping WebDriver's method names, so the program runs anywhere.

class NoSuchElement(Exception):
    pass

class FakeDriver:                             # a browser, simulated so the tests run anywhere
    def __init__(self, login_ids):
        self.login_ids, self.page, self.typed = login_ids, "login", {}
    def find_element(self, by, value):
        on_page = self.login_ids.values() if self.page == "login" else ["welcome"]
        if value not in on_page:
            raise NoSuchElement(value)
        return FakeElement(self, value)

class FakeElement:
    def __init__(self, driver, element_id):
        self.driver, self.id = driver, element_id
    def send_keys(self, text):
        self.driver.typed[self.id] = text
    def click(self):
        ids, typed = self.driver.login_ids, self.driver.typed
        if typed.get(ids["roll"]) == "2026CS014" and typed.get(ids["password"]) == "test-pass-1":
            self.driver.page = "home"
    @property
    def text(self):
        return "Welcome, 2026CS014" if self.id == "welcome" else ""

RELEASE_1 = {"roll": "roll", "password": "pwd", "button": "login-btn"}
RELEASE_2 = {"roll": "roll-number", "password": "password", "button": "sign-in"}   # a redesign

# --- tests WITHOUT page objects: every test knows the page's HTML ---------------------------
def raw_login_succeeds(driver):
    driver.find_element("id", "roll").send_keys("2026CS014")
    driver.find_element("id", "pwd").send_keys("test-pass-1")
    driver.find_element("id", "login-btn").click()
    return "2026CS014" in driver.find_element("id", "welcome").text

def raw_wrong_password_refused(driver):
    driver.find_element("id", "roll").send_keys("2026CS014")
    driver.find_element("id", "pwd").send_keys("wrong")
    driver.find_element("id", "login-btn").click()
    return driver.page == "login"          # stands for checking the page address

def raw_empty_roll_refused(driver):
    driver.find_element("id", "roll").send_keys("")
    driver.find_element("id", "pwd").send_keys("test-pass-1")
    driver.find_element("id", "login-btn").click()
    return driver.page == "login"          # stands for checking the page address

# --- the same tests WITH page objects: only the page classes know the HTML --------------------
class LoginPage:
    ROLL, PASSWORD, BUTTON = ("id", "roll"), ("id", "pwd"), ("id", "login-btn")
    def __init__(self, driver):
        self.driver = driver
    def log_in_as(self, roll_no, password):   # a service the page offers, named by intention
        self.driver.find_element(*self.ROLL).send_keys(roll_no)
        self.driver.find_element(*self.PASSWORD).send_keys(password)
        self.driver.find_element(*self.BUTTON).click()
        return HomePage(self.driver) if self.driver.page == "home" else self

class HomePage:
    def __init__(self, driver):
        self.driver = driver
    def welcome_text(self):
        return self.driver.find_element("id", "welcome").text

def pom_login_succeeds(driver):
    home = LoginPage(driver).log_in_as("2026CS014", "test-pass-1")
    return isinstance(home, HomePage) and "2026CS014" in home.welcome_text()

def pom_wrong_password_refused(driver):
    return isinstance(LoginPage(driver).log_in_as("2026CS014", "wrong"), LoginPage)

def pom_empty_roll_refused(driver):
    return isinstance(LoginPage(driver).log_in_as("", "test-pass-1"), LoginPage)

def run(tests, release):
    passed = 0
    for test in tests:
        try:
            passed += test(FakeDriver(release))
        except NoSuchElement:
            pass                               # a locator no longer matches the page
    return f"{passed} of {len(tests)} pass"

RAW = [raw_login_succeeds, raw_wrong_password_refused, raw_empty_roll_refused]
POM = [pom_login_succeeds, pom_wrong_password_refused, pom_empty_roll_refused]
print("release 1, without page objects:", run(RAW, RELEASE_1))
print("release 1, with page objects:   ", run(POM, RELEASE_1))
print("release 2, without page objects:", run(RAW, RELEASE_2))
print("release 2, with page objects:   ", run(POM, RELEASE_2))
LoginPage.ROLL, LoginPage.PASSWORD, LoginPage.BUTTON = (
    ("id", "roll-number"), ("id", "password"), ("id", "sign-in"))   # the one repair
print("release 2, after repairing LoginPage only:", run(POM, RELEASE_2))
munotes.in279

The Page Object Model

release 1, without page objects: 3 of 3 pass
release 1, with page objects:    3 of 3 pass
release 2, without page objects: 0 of 3 pass
release 2, with page objects:    0 of 3 pass
release 2, after repairing LoginPage only: 3 of 3 pass

On release 1 both styles pass, and nothing yet shows the difference. On release 2 both fail, as they must: the page really changed. The difference is in the repair. The tests without page objects contain the login page's three locators nine times, three in each test, and every occurrence must be found and changed; in a real suite, with dozens of tests that log in, the count runs into the hundreds. The page-object tests contain them three times, in LoginPage alone, and the one repair the program makes, three lines in one class, brings all three tests back. The tests themselves were not touched; they still read as intentions: log in as this student, and check the welcome.

Notice also the page objects' shape, which follows the rules above. log_in_as is a service named for what a user does, not for which boxes are filled; it returns a HomePage when the login succeeds and the LoginPage itself when it does not, which the documentation's rule "Different results for the same action are modelled as different methods" would split into two methods in a larger suite. And neither page object makes an assertion: HomePage.welcome_text returns the text, and the test decides whether it is right.

POM with TestNG, in the practical

The practical asks students to "Design and implement automation framework using Page Object Model (POM). Create reusable page classes and execute modular test cases." The Java shape is the same as the Python one: a class for each page or component, its locators as private fields, its services as public methods returning page objects; and TestNG test classes (Chapter Thirty-Five, on writing unit tests with a framework) that create the page objects in a @BeforeMethod fixture and make their assertions in @Test methods. Data-driven tests (Chapter Fifty) combine naturally with it: the data provider supplies the rows, and the page objects carry them through the pages.

munotes.in280

The Page Object Model

What it does not mean

A page object is not one class per web page. Fowler notes the name misleads: page objects should be built "for the significant elements on a page", so a page may have several.

A page object is not a place for assertions. It provides access to the page; the test decides what is correct.

POM does not make tests immune to change. When the page changes, the page object must still be repaired; the gain is that the repair happens once.

POM is not only for Selenium. Any test that drives a user interface, a mobile app's screens included, can hide the interface's details behind objects named for what users do.

Quick revision

  • POM (Selenium documentation): a design pattern for test maintenance and less duplication; a page object is a class serving as the interface to a page of the application under test.
  • Fowler: a page object "wraps an HTML page, or fragment, with an application-specific API"; it should let a client "do anything and see anything that a human can"; test it by imagining the concrete control changed.
  • Advantages: separation of test code from page-specific code such as locators; a single repository of the page's services; UI changes repaired in one place.
  • Rules: public methods are services; hide internals; generally no assertions; return other page objects; need not be a whole page; different results, different methods.
  • Worked example: after a redesign, nine inline locators across three tests against three in one page object; one repair restored all three page-object tests.

Test yourself

1. What is the Page Object Model, and what problem does it solve? A design pattern in which each page, or significant part of one, is represented by a class that offers the user's actions as methods and hides its locators and layout. It solves the brittleness of tests that manipulate HTML directly: when a page changes, only its page object needs repairing, not every test that uses the page.

2. State two advantages of POM as the Selenium documentation gives them. A clean separation between test code and page-specific code such as locators and layout; and a single repository for the services or operations the page offers, instead of having them scattered through the tests. Both mean that changes for a new UI are made in one place.

munotes.in281

The Page Object Model

3. Should a page object contain assertions? Explain. Generally no. Assertions are part of the test and belong in the test's code; the page object provides access to the page's data and services. The accepted exception is a check of the page's invariants, such as confirming the browser is on the expected page when the object is created.

4. What should a page object's methods return, and why? Other page objects, component objects, or the page object itself, so that a test can follow the user's journey from page to page; results such as text or states are returned as fundamental types for the test to assert on.

5. In the worked example, why did the page-object tests need only one repair after the redesign? Because the login page's locators appeared only in the LoginPage class; the tests called its log_in_as method and never named an element, so changing the three locators in that one class repaired every test that logs in.

Contents This chapter on its own page

munotes.in282

Module II

Black-box, white-box and experience-based test design, software metrics and complexity, defect management, the quality movement, SQA and reliability, ISO 9000, formal technical reviews, quality costs and the seven basic quality tools

munotes.in

Chapter Fifty-Two

Black-Box and White-Box Testing

Syllabus topic Module 2, "White Box Testing and Black Box Testing"

In one line

Black-box testing chooses its tests from what the software is supposed to do, without looking inside it; white-box testing chooses them from the code itself; the first cannot see code that nobody asked for, the second cannot see a requirement that nobody coded, and so a sound test effort uses both.

In the wording a student can write in an examination: black-box testing designs tests from the specified behaviour of the software. In the ISTQB syllabus's words, "Black-box test techniques (also known as specification-based techniques) are based on an analysis of the specified behavior of the test object without reference to its internal structure." White-box testing designs tests from its code: "White-box test techniques (also known as structure-based techniques) are based on an analysis of the test object's internal structure and processing." The names come from two kinds of box that ISO/IEC/IEEE 24765 defines: a black box is a "system or component whose inputs, outputs, and general function are known but whose contents or implementation are unknown or irrelevant", and a glass box is one "whose internal contents or implementation are known". Black-box testing finds behaviour that is missing or wrong against the requirements, but cannot say how much of the code it ran; white-box testing measures coverage and finds code that no requirement explains, but "if the software does not implement one or more requirements, white-box testing may not detect the resulting defects of omission" (ISTQB). Grey-box testing, a working term rather than a standard one, tests from outside with partial knowledge of the inside.

Two boxes, and their other names

MU's Module 2 opens on a pair of words borrowed from engineering. ISO/IEC/IEEE 24765 defines them as properties of the thing being examined: a black box is one whose inputs, outputs and general function are known while its contents are "unknown or irrelevant", and a glass box is one whose contents are known. Testing calls the glass box a white box; both words mean that the tester can see in.

The same two approaches go under several names, and a question paper may use any of them:

Black-box testingWhite-box testing
Specification-based testing (ISTQB v4.0.1, ISO/IEC/IEEE 29119-1)Structure-based testing (ISTQB v4.0.1, ISO/IEC/IEEE 29119-1)
Behavioural or behaviour-based testing (the 2018 ISTQB syllabus)Structural testing (the 2018 ISTQB syllabus, and MU's own heading)
Functional testing, in the first of its two senses belowGlass-box testing, after ISO/IEC/IEEE 24765's glass box

ISO/IEC/IEEE 29119-1:2022 gives the formal definitions. Specification-based testing is "testing in which the principal test basis is the external inputs and outputs of the test item, commonly based on a specification, rather than its implementation in source code or executable software". Structure-based testing is "dynamic testing in which the tests are derived from an examination of the structure of the test item".

munotes.in283

Black-Box and White-Box Testing

A trap in the word functional. ISO/IEC/IEEE 24765 records two meanings of functional testing. The first is "testing that ignores the internal mechanism of a system or component and focuses solely on the outputs generated in response to selected inputs and execution conditions", which is black-box testing by another name, and the sense MU uses when it writes of functional or specification-based testing as black box. The second is "testing conducted to evaluate the compliance of a system or component with specified functional requirements", which is about what is tested (functions, as against performance or usability, the test types of Chapter Thirty-Two), not about how the tests are chosen. The two senses cross. A load test designed from a response-time requirement is black-box but not functional in the second sense; a unit test of the fee calculation is functional in the second sense and can perfectly well be designed white-box.

What each tester sees

Take one function from ExamReg, the fee calculation, and give it to two testers.

The black-box tester is given the requirement and the way to call the function, and nothing else. The requirement is ExamReg's fee rule:

RuleValue
Form feeRs 800, waived for a student with a fee concession
Backlog papersRs 150 each; a negative number of backlog papers is refused
Late feeRs 0 on or before the last date; Rs 100 for 1 to 7 days late; Rs 500 for 8 to 15 days late
More than 15 days lateThe form is not accepted
Days late0 for any form on or before the last date, so a negative number is invalid and refused

and the call is total_fee(days_late, backlog_papers, concession), which returns the fee in rupees or refuses the form by raising ValueError.

The white-box tester is given all of that and the code as well:

FORM_FEE, BACKLOG_FEE = 800, 150

def total_fee(days_late, backlog_papers, concession):
    if days_late < 0:
        raise ValueError("form not accepted")
    if not isinstance(backlog_papers, int) or backlog_papers < 0:
        raise ValueError("backlog papers must be a whole number, 0 or more")
    fee = FORM_FEE + BACKLOG_FEE * backlog_papers
    if backlog_papers >= 8:
        fee -= BACKLOG_FEE
    if concession:
        fee -= FORM_FEE
    if days_late == 0:
        late = 0
    elif days_late <= 7:
        late = 100
    else:
        late = 500
    return fee + late

The function has two defects. Neither tester is told what they are, and this chapter does not name them until the tests have found them.

Worked example, part 1: the black-box tester

Working from the table alone, the black-box tester writes one test for each rule: on time with no backlog, a few days late with backlog papers, a concession student in the Rs 500 band, and the three kinds of refusal. The expected result of every test is read off the table.

munotes.in284

Black-Box and White-Box Testing

from examreg_fee import total_fee

def run(args):
    try:
        return total_fee(*args)
    except ValueError:
        return "refused"

# (days late, backlog papers, concession) -> expected result, each from one rule of the table
black_box = [((0, 0, False), 800),        # on time, no backlog: the form fee only
             ((3, 2, False), 1200),       # 1 to 7 days late: Rs 100; two backlog papers
             ((10, 1, True), 650),        # 8 to 15 days late: Rs 500; concession waives the form fee
             ((16, 0, False), "refused"), # more than 15 days late: form not accepted
             ((-1, 0, False), "refused"), # a negative number of days is invalid
             ((0, -1, False), "refused")] # a negative number of backlog papers is invalid

failed = 0
for args, expected in black_box:
    actual = run(args)
    verdict = "pass" if actual == expected else "FAIL"
    failed += verdict == "FAIL"
    print(f"{str(args):<18} expected {str(expected):<8} got {str(actual):<8} {verdict}")
print(f"black-box: {len(black_box)} tests, {failed} failed")
(0, 0, False)      expected 800      got 800      pass
(3, 2, False)      expected 1200     got 1200     pass
(10, 1, True)      expected 650      got 650      pass
(16, 0, False)     expected refused  got 1300     FAIL
(-1, 0, False)     expected refused  got refused  pass
(0, -1, False)     expected refused  got refused  pass
black-box: 6 tests, 1 failed

Five tests pass. The fourth fails: a form 16 days late should be refused, and the function charges the form fee and the Rs 500 late fee instead, Rs 1300 in all. The requirement says what should happen after 15 days, the tester read the requirement, and the failure followed. Nothing about the code was needed to find it.

What this tester cannot say is how much of the code those six tests exercised. The ISTQB syllabus is blunt about it: "Performing only black-box testing does not provide a measure of actual code coverage."

Worked example, part 2: coverage, and the white-box tester

Measuring coverage needs the code, so it is the white-box tester's first job. The program below runs the same six tests again, this time with a tracer, built on Python's sys.settrace, recording which statements of total_fee execute; Chapter Fifty-Eight, on structural testing, explains how such a tracer works. Then it runs the white-box tester's own tests, designed by reading the code: one test to make each if go each way it can, and one to reach each raise.

The expected result of every white-box test still comes from the requirement table, never from the code. A test that took its expected answer from the code could only confirm that the code does what it does.

munotes.in285

Black-Box and White-Box Testing

import ast, inspect, sys
from examreg_fee import total_fee

def statements(func):
    """Line number -> text of every statement in the body of func."""
    first = func.__code__.co_firstlineno
    source = inspect.getsource(func).splitlines()
    return {node.lineno + first - 1: source[node.lineno - 1].strip()
            for node in ast.walk(ast.parse("\n".join(source)).body[0])
            if isinstance(node, ast.stmt) and not isinstance(node, ast.FunctionDef)}

def run_traced(args, ran):
    """Call total_fee(*args), adding the line numbers it executes to the set ran."""
    def on_line(frame, event, arg):
        if event == "line":
            ran.add(frame.f_lineno)
        return on_line
    sys.settrace(lambda frame, event, arg:
                 on_line if frame.f_code is total_fee.__code__ else None)
    try:
        return total_fee(*args)
    except ValueError:
        return "refused"
    finally:
        sys.settrace(None)

def report(name, tests):
    ran, every, lines = set(), statements(total_fee), []
    for args, expected in tests.items():
        actual = run_traced(args, ran)
        if actual != expected:
            lines.append(f"FAIL {args}: expected {expected}, got {actual}")
    failed = len(lines)
    lines += [f"never executed: {every[n]}" for n in sorted(every.keys() - ran)]
    print(f"{name}: {len(tests)} tests, {failed} failed, "
          f"statements executed {len(every.keys() & ran)} of {len(every)}")
    for line in lines:
        print("   ", line)

black_box = {(0, 0, False): 800, (3, 2, False): 1200, (10, 1, True): 650,
             (16, 0, False): "refused", (-1, 0, False): "refused", (0, -1, False): "refused"}
white_box = {(0, 0, False): 800,           # every if false; days_late == 0
             (5, 8, True): 1300,           # backlog_papers >= 8, concession and the elif all true
             (9, 0, False): 1300,          # the final else
             (-1, 0, False): "refused",    # the first raise
             (0, -1, False): "refused"}    # the second raise
report("black-box suite", black_box)
report("white-box suite", white_box)
report("both together", black_box | white_box)
black-box suite: 6 tests, 1 failed, statements executed 14 of 15
    FAIL (16, 0, False): expected refused, got 1300
    never executed: fee -= BACKLOG_FEE
white-box suite: 5 tests, 1 failed, statements executed 15 of 15
    FAIL (5, 8, True): expected 1300, got 1150
both together: 8 tests, 2 failed, statements executed 15 of 15
    FAIL (16, 0, False): expected refused, got 1300
    FAIL (5, 8, True): expected 1300, got 1150

The first line is the ISTQB syllabus's point made concrete. The black-box tests, every one of them derived properly from the requirement, executed 14 of the function's 15 statements. The one they never reached is fee -= BACKLOG_FEE, which runs only when backlog_papers >= 8. No rule in the table mentions eight backlog papers, so no black-box technique would ever pick that value out: to a tester who sees only the requirement, eight backlog papers is simply more of the same.

The white-box tester, reading the code, sees the condition and writes a test that makes it true: five days late, eight backlog papers, a concession. The table says the concession removes the form fee, so the fee should be 8 × 150 + 100 = 1300 rupees. The function charges 1150, because it quietly takes one backlog paper's fee off for anyone with eight or more. No requirement asks for that line; it is code nobody specified, and only a tester who could see it could have aimed a test at it.

munotes.in286

Black-Box and White-Box Testing

Now look at what the white-box suite did not do. It executed all 15 statements, every one, and it contains no test of a form 16 days late, because nothing in the code suggests one. There is no statement that handles a form more than 15 days late, so there is nothing for coverage to count. Complete statement coverage, and the defect the black-box tester found with one test is invisible to it. The ISTQB syllabus names exactly this: "A corresponding weakness is that if the software does not implement one or more requirements, white-box testing may not detect the resulting defects of omission" (it cites Watson 1996).

The last line is the practical conclusion: the eight different tests of the two suites together find both defects and execute every statement.

What each finds, and what each misses

The example is small, but it has the shape of the general case.

Black-box testing finds behaviour that is missing or wrong against the requirement: a rule never implemented, as here, or a fee miscalculated, or a boundary in the wrong place. Its tests can be written as soon as the requirement exists, before any code, and in the ISTQB syllabus's words (Chapter Eight quoted them in full, on the three families of test design techniques) they remain useful "if the implementation changes, but the required behavior stays the same". What it misses is code no requirement mentions, and it cannot measure how much of the code it has run. Its tests are only as good as the requirement they come from: a vague requirement gives vague tests.

White-box testing reaches every part of the code, including the parts nobody specified: leftover rules, debugging shortcuts, dead code, unexpected paths. The ISTQB syllabus puts its strength this way: "A fundamental strength that all white-box test techniques share is that the entire software implementation is taken into account during testing, which facilitates defect detection even when the software specification is vague, outdated or incomplete." It also gives a number, which black-box testing cannot: "White-box coverage measures provide an objective measurement of coverage and the necessary information to allow additional tests to be generated to increase this coverage, and subsequently increase confidence in the code." What it misses is what the code leaves out, and its tests can only be designed once the design or the code exists, and must be redesigned when the code is rewritten.

munotes.in287

Black-Box and White-Box Testing

White-box testing without running anything. Because it works from structure, it can be applied before the code runs at all. The syllabus notes that "White-box test techniques can be used in static testing (e.g., during dry runs of code)", and that they are well suited to reviewing code not yet ready for execution, pseudocode, and other logic "which can be modeled with a control flow graph", the graph Chapter Fifty-Eight, on structural testing, draws. The walkthrough of Chapter Thirty, which traced the fee code by hand, was white-box testing of exactly this kind.

What they share. Both need an oracle, the expected result, and in both it comes from the requirement. The difference is only in how the inputs are chosen: from the requirement's rules, or from the code's statements and branches.

Black-box and white-box testing compared

Black-box testingWhite-box testing
Also calledSpecification-based, behavioural, functional (first sense)Structure-based, structural, glass-box
The tester seesInputs, outputs and the requirementThe code as well as the requirement
Tests are derived fromRequirements, specifications, user stories, interfacesThe code or detailed design: statements, branches, paths
Knowledge neededWhat the software should doAlso how it is built: the language, and coverage tools
Expected results come fromThe requirementThe requirement, never the code
Tests can be designedAs soon as the requirement existsOnly after the design or the code exists
After the code is rewrittenStill valid if the behaviour is unchangedMust be redesigned
Coverage is measured overItems of the requirement: partitions, boundaries, rules, transitionsItems of the code: statements, branches
FindsMissing and wrong behaviourCode no test reached, and code no requirement explains
MissesCode that no requirement mentionsRequirements the code never implemented
Test levelsAll levelsAll levels, not only unit testing
MU's techniquesEquivalence partitioning, boundary value analysis, decision tables, state transitionsStatement testing, branch testing

Two rows need a word. The knowledge row follows the 2018 ISTQB syllabus, which says white-box test design "may involve special skills or knowledge, such as the way the code is built, how data is stored (e.g., to evaluate possible database queries), and how to use coverage tools and to correctly interpret their results." And the levels row follows ISO/IEC/IEEE 29119-1, whose note on structure-based testing says it "is not restricted to use at component level and can be used at all levels, e.g. menu item coverage as part of a system test": a system test that makes sure every menu item of ExamReg has been opened at least once is measuring the coverage of a structure too.

Grey-box testing

Between the two sits a term that students meet but that has no standard definition. SEVOCAB's whole vocabulary of 5,314 terms, which carries ISO/IEC/IEEE 24765, and both editions of the ISTQB syllabus were searched for grey and gray, and neither word appears in any of them.

munotes.in288

Black-Box and White-Box Testing

In its usual sense, as the encyclopedia article on it describes, grey-box testing is testing "where the tester has partial knowledge of the internal structure of the software". The tests are run from outside, through the interfaces a user or another system would use, but they are chosen with some knowledge of the inside; the article lists knowledge such as "software architecture, Unified Modeling Language (UML) models, Web Services Description Language (WSDL) information and state models".

For ExamReg, a grey-box tester might know from the design that the hall ticket is issued only when the payment record changes to paid, and that the fee page's calculator works in paise while the payment module takes rupees, the mismatch Chapter Thirty-Seven's integration test caught. That tester still tests through the web pages, as a student would, and has no need to read the code; but chooses the tests with that knowledge (a payment the gateway accepts but that never reaches the record, an amount large enough for a unit mistake to show) and checks the payment record afterwards, which a tester who knew nothing of the inside would have no reason to do. That is the sense in which it is half of each.

Using the two together

The worked example also shows a sensible order, the one the syllabus's sentence on coverage measures implies. Design the black-box tests first, from the requirement, since they can be written before the code exists and survive changes to it. Run them with a coverage measurement. Read what they did not execute, and design white-box tests for that.

Each statement the black-box tests never reached is a question. Either it serves a requirement the black-box tests missed, in which case a test is added (and perhaps the requirement is clarified), or it serves no requirement at all, in which case it is a defect or it should be deleted. In the worked example it was the second: fee -= BACKLOG_FEE for eight or more backlog papers was a rule nobody had asked for.

The third family, experience-based testing (Chapter Sixty-One), comes after both, and the syllabus calls it "complementary" to them: it looks for what neither a requirement nor the code suggests.

What it does not mean

Black-box testing is not testing without thought. It is as systematic as white-box testing; the next five chapters are its techniques.

White-box testing does not test the code against itself. Its inputs are chosen from the code, but its expected results come from the requirement, or it could never fail.

munotes.in289

Black-Box and White-Box Testing

100 per cent statement coverage does not mean the requirement is met. The white-box suite above executed every statement and missed a whole rule.

Grey-box testing is not a third family of techniques. ISTQB's third family is experience-based testing; grey-box describes how much the tester knows of the inside, not a way of choosing tests.

Functional testing is not always black-box testing. In ISO/IEC/IEEE 24765's second sense it means testing of the functional requirements, which can be designed either way.

Quick revision

  • Black-box (specification-based): tests from the specified behaviour "without reference to its internal structure" (ISTQB v4.0.1); a black box is known by its inputs, outputs and function only (ISO/IEC/IEEE 24765).
  • White-box (structure-based, structural, glass-box): tests from "the test object's internal structure and processing" (ISTQB v4.0.1); a glass box's contents are known.
  • Functional testing has two senses (ISO/IEC/IEEE 24765): ignoring the internal mechanism, which is black-box; and testing against functional requirements, which is a test type.
  • Worked example: the black-box suite (6 tests) found the missing refusal after 15 days but executed only 14 of 15 statements; the white-box suite (5 tests) executed all 15 and found a rule nobody asked for (Rs 1150 charged for Rs 1300), but missed the omission; the two together found both.
  • White-box strength: the whole implementation is considered, even when the specification is "vague, outdated or incomplete"; weakness: "defects of omission" (ISTQB, citing Watson 1996).
  • Expected results come from the requirement in both.
  • Grey-box: testing from outside with "partial knowledge of the internal structure"; a working term, in no standard vocabulary.
  • Order: black-box tests first, a coverage measurement, then white-box tests for what was not reached.

Test yourself

1. Distinguish black-box testing from white-box testing. Black-box testing derives tests from the specified behaviour of the software, its requirements and interfaces, without reference to its internal structure; white-box testing derives them from the internal structure, the code's statements and branches. Black-box tests can be designed as soon as the requirements exist and stay valid when the code is rewritten; white-box tests need the design or the code and change with it. Black-box testing finds missing and wrong behaviour but cannot measure code coverage or see code no requirement mentions; white-box testing measures coverage and finds unreached or unspecified code but cannot find requirements that were never implemented.

2. What is a defect of omission, and why can white-box testing miss it? A defect of omission is a requirement the software does not implement at all, such as a missing check. White-box tests are derived from the code, and coverage counts only code that exists; if there is no code for a requirement, nothing in the structure points to it, and a white-box suite can reach 100 per cent coverage without ever testing it.

munotes.in290

Black-Box and White-Box Testing

3. Why measure the code coverage of a black-box test suite, and what does the measurement show? Because black-box testing alone gives no measure of how much code was exercised. The measurement lists the statements or branches the tests never reached; each is either a requirement the tests missed or code no requirement explains, and white-box tests are then designed for it.

4. In white-box testing, where does the expected result of a test come from, and why? From the requirement, as in black-box testing. Only the choice of inputs comes from the code. If the expected result were taken from the code, the test would only confirm that the code does what it does and could never fail.

5. What is grey-box testing? Give an example. Testing from outside, through the software's interfaces, by a tester with partial knowledge of its inside, such as its architecture, data design or state models. It is a working term with no standard definition. For example, a tester who knows that ExamReg issues a hall ticket only when the payment record becomes paid tests through the web pages but chooses a payment that the gateway accepts and the record never receives, and checks the record afterwards.

6. Give the two meanings of functional testing. Testing that ignores the internal mechanism and looks only at the outputs produced for chosen inputs, which is black-box testing; and testing that evaluates compliance with the specified functional requirements, as against non-functional ones such as performance, which can be designed black-box or white-box.

Contents This chapter on its own page

munotes.in291

Chapter Fifty-Three

Specification-Based Testing

Syllabus topic Module 2, "Functional/Specification, based Testing as Black Box"

In one line

Specification-based testing takes everything a test needs from the specification: what to test (the test conditions), how much to test (the coverage items), and what the right answer is (the expected result); the code is never consulted, so the tests can be written before it exists, and they can never be better than the specification they come from.

In the wording a student can write in an examination: specification-based testing, which MU calls functional or specification-based testing and places under black box, is "testing in which the principal test basis is the external inputs and outputs of the test item, commonly based on a specification, rather than its implementation in source code or executable software" (ISO/IEC/IEEE 29119-1:2022). The test basis is the "information used as the basis for designing and implementing test cases" (ISO/IEC/IEEE 29119-2), and a specification is an "information item that identifies, in a complete, precise, and verifiable manner, the requirements, design, behavior, or other expected characteristics of a system, service, or process" (ISO/IEC/IEEE 15289). From it the tester derives test conditions, a test model, coverage items and test cases whose expected results are the "observable predicted behavior of the test item under specified conditions based on its specification or another source" (ISO/IEC/IEEE 29119-4). MU names four specification-based techniques: equivalence partitioning, boundary value analysis, decision table testing and state transition testing.

What counts as a specification

Anything that says what the software should do, precisely enough to decide whether it did it, can be a test basis. The common ones:

Test basisWhat it gives the tester
A software requirements specification (SRS): "documentation of the essential requirements (functions, performance, design constraints, and attributes) of the software and its external interfaces" (IEEE 1012-2024)Rules, limits and required responses, statement by statement
A use case: "in UML, a complete task of a system that provides a measurable result of value for an actor" (ISO/IEC/IEEE 24765)Sequences of interactions, with their alternative and exception flows
A user story with its acceptance criteria (ISTQB v4.0.1, section 4.5)Conditions the story must meet, as scenarios or as rules
Business rule tablesCombinations of conditions and the outcome each selects
State modelsWhat the system does in each state, for each event
Interface specificationsFormats, ranges and error responses at a boundary between systems

The ISTQB syllabus gives the user story's usual form, "As a [role], I want [goal to be accomplished], so that I can [resulting business value for the role]", followed by acceptance criteria, which it says "may be viewed as the test conditions that should be exercised by the tests". It names the two common ways of writing them: "Scenario-oriented (e.g., Given/When/Then format used in BDD, see section 2.1.3)" and "Rule-oriented (e.g., bullet point verification list, or tabulated form of input-output mapping)". Chapter Forty-Two, on acceptance testing, showed ExamReg's acceptance criteria at work; to a tester, each form is simply a specification in a particular shape.

munotes.in292

Specification-Based Testing

Complete, precise, verifiable

The three adjectives in the definition of a specification are exactly the three ways a specification can fail its tester.

  • Complete. A rule the specification leaves out is a test condition nobody derives. Chapter Fifty-Two's black-box tester found the missing refusal after 15 days only because the requirement stated the limit.
  • Precise. A rule that can be read two ways gives two expected results, and a test cannot pass against both. Chapter Twenty-Eight's review of the first draft of ExamReg's fee rule found this in up to a week late.
  • Verifiable, which ISO/IEC/IEEE 15289 defines as "can be checked for correctness by a person or tool". The fee page must be fast cannot be checked; the fee page responds within two seconds can.

Designing tests is therefore also a review of the specification. The ISTQB syllabus makes it part of test analysis: "The test basis and the test object are also evaluated to identify defects they may contain and to assess their testability." And of user stories in particular: "If a stakeholder does not know how to test a user story, this may indicate that the user story is not clear enough, or that it does not reflect something valuable to them, or that the stakeholder just needs help in testing (Wake 2003)."

From a specification to a test case

The ISTQB syllabus splits the work into two activities. Test analysis "answers the question 'what to test?' in terms of measurable coverage criteria"; test design "answers the question 'how to test?'". In ISO/IEC/IEEE 29119's terms, four things are produced along the way.

  1. Test conditions. A test condition is a "testable aspect of a component or system, such as a function, transaction, feature, quality attribute, or structural element identified as a basis for testing" (ISO/IEC/IEEE 29119-2). From the fee rule: the late fee for a form 1 to 7 days late, the refusal of a form more than 15 days late.
  2. A test model. A test model is a "representation of a test item that is used during the test case design process" (ISO/IEC/IEEE 29119-4): the partitions of the days late, a decision table of the fee conditions, a state diagram of the login. Choosing the model is choosing the technique.
  3. Coverage items. The technique says which parts of the model must each be exercised: every partition, every boundary, every rule, every transition. Coverage is then measured over the specification, not over the code.
  4. Test cases. For each coverage item, inputs that exercise it and the expected result the specification gives for those inputs.
munotes.in293

Specification-Based Testing

The last step is where the specification acts as the oracle, the source that decides pass or fail (Chapter Seven, on what a test case is, defined it). The worked example below uses it in two ways: read by a person, one test at a time, and written as a program that can judge any number of tests.

Worked example, part 1: one test for each atomic requirement

The simplest specification-based technique tests each requirement once. ISO/IEC/IEEE 29119-4 calls it requirements-based testing, a "specification-based test case design technique based on exercising atomic requirements". An atomic requirement says one thing; a compound one hides several. ExamReg's fee rule, split into atomic requirements, becomes ten:

IdAtomic requirement
R1The form fee is Rs 800
R2A fee concession waives the form fee only
R3Each backlog paper costs Rs 150
R4A negative number of backlog papers is refused
R5A number of backlog papers that is not whole is refused
R6There is no late fee on or before the last date (0 days late)
R7The late fee is Rs 100 for 1 to 7 days late
R8The late fee is Rs 500 for 8 to 15 days late
R9A form more than 15 days late is refused
R10A negative number of days late is refused

The developers' code is printed here only because the programs below must import it. The tester never reads it.

def total_fee(days_late, backlog_papers, concession):
    """ExamReg's fee calculation, as delivered for testing."""
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    if not isinstance(backlog_papers, int) or backlog_papers < 0:
        raise ValueError("backlog papers must be a whole number, 0 or more")
    fee = (0 if concession else 800) + 150 * backlog_papers
    if days_late == 0:
        late = 0
    elif days_late < 7:
        late = 100
    else:
        late = 500
    return fee + late

One test per atomic requirement, each expected result read off its rule:

from fee_impl import total_fee

atomic = ["R1", "R2", "R3", "R4", "R5", "R6", "R7", "R8", "R9", "R10"]

# (requirement, inputs as (days late, backlog papers, concession), expected result from the rule)
tests = [("R1", (0, 0, False), 800),         ("R2", (0, 2, True), 300),
         ("R3", (0, 3, False), 1250),        ("R4", (0, -1, False), "refused"),
         ("R5", (0, 1.5, False), "refused"), ("R6", (0, 1, False), 950),
         ("R7", (4, 0, False), 900),         ("R8", (10, 0, False), 1300),
         ("R9", (16, 0, False), "refused"),  ("R10", (-2, 0, False), "refused")]

def run(args):
    try:
        return total_fee(*args)
    except ValueError:
        return "refused"

covered = set()
for req, args, expected in tests:
    actual = run(args)
    covered.add(req)
    print(f"{req:<4} {str(args):<17} expected {str(expected):<8} got {str(actual):<8}"
          f" {'pass' if actual == expected else 'FAIL'}")
print(f"requirement coverage: {len(covered)} of {len(atomic)} atomic requirements"
      f" = {100 * len(covered) / len(atomic):.0f}%")
munotes.in294

Specification-Based Testing

R1   (0, 0, False)     expected 800      got 800      pass
R2   (0, 2, True)      expected 300      got 300      pass
R3   (0, 3, False)     expected 1250     got 1250     pass
R4   (0, -1, False)    expected refused  got refused  pass
R5   (0, 1.5, False)   expected refused  got refused  pass
R6   (0, 1, False)     expected 950      got 950      pass
R7   (4, 0, False)     expected 900      got 900      pass
R8   (10, 0, False)    expected 1300     got 1300     pass
R9   (16, 0, False)    expected refused  got refused  pass
R10  (-2, 0, False)    expected refused  got refused  pass
requirement coverage: 10 of 10 atomic requirements = 100%

Every atomic requirement has been exercised and every test has passed. By the measure this technique defines, testing is complete: 100 per cent requirement coverage.

It is not complete, because the code is wrong. The trouble is the grain of the coverage items. R7 is one coverage item, Rs 100 for 1 to 7 days late, and one test at 4 days covers it; but seven different values satisfy it, and nothing in requirements-based testing says which of them to try. A coverage measure is only as fine as the items it counts. The four techniques MU names are, among other things, finer ways of cutting the same specification into coverage items.

Worked example, part 2: the specification as a program

A person reading the specification can judge one test at a time. Written as a program, the specification can judge any number. Below, the tester writes ExamReg's fee rule as a function, from the specification alone, without looking at the developers' code: an executable oracle. Then the program tries 200 inputs chosen at random, which ISO/IEC/IEEE 29119-4 recognises as a technique in its own right, random testing, a "specification-based test design technique based on generating test cases to exercise randomly selected test item inputs". Each result is judged by the oracle.

import random
from fee_impl import total_fee

def specified_fee(days_late, backlog_papers, concession):
    """The oracle: ExamReg's fee rule written as code by the tester, from the specification
    alone, independently of the developers' code."""
    if not 0 <= days_late <= 15:
        return "refused"
    if not isinstance(backlog_papers, int) or backlog_papers < 0:
        return "refused"
    late = 0 if days_late == 0 else 100 if days_late <= 7 else 500
    return (0 if concession else 800) + 150 * backlog_papers + late

def run(args):
    try:
        return total_fee(*args)
    except ValueError:
        return "refused"

rng = random.Random(53)                     # a fixed seed, so the run can be repeated
inputs = [(rng.randint(-3, 18), rng.randint(-1, 9), rng.choice([True, False]))
          for _ in range(200)]
failures = [args for args in inputs if run(args) != specified_fee(*args)]

print(f"random tests: {len(inputs)}, judged by the specification: {len(failures)} failed")
for args in failures[:3]:
    print(f"   {args}: specified {specified_fee(*args)}, got {run(args)}")
print("days late in every failing test:", sorted({args[0] for args in failures}))
sevens = [args for args in inputs if args[0] == 7]
print(f"tests that happened to use 7 days late: {len(sevens)}, of which "
      f"{sum(args[1] < 0 for args in sevens)} also had a negative number of backlog papers")

p_seven = 1 / 22                             # days late drawn from -3 to 18: 22 values
for n in (10, 50, 200):
    print(f"chance that {n:>3} random tests never try 7 days late: {(1 - p_seven) ** n:.4f}")
munotes.in295

Specification-Based Testing

random tests: 200, judged by the specification: 4 failed
   (7, 0, True): specified 100, got 500
   (7, 4, True): specified 700, got 1100
   (7, 2, False): specified 1200, got 1600
days late in every failing test: [7]
tests that happened to use 7 days late: 6, of which 2 also had a negative number of backlog papers
chance that  10 random tests never try 7 days late: 0.6280
chance that  50 random tests never try 7 days late: 0.0977
chance that 200 random tests never try 7 days late: 0.0001

The oracle found what the ten requirement tests missed. Four of the 200 tests failed, and every failing test was 7 days late: the specification says 1 to 7 days late costs Rs 100, and the code charges Rs 500 at exactly 7 days. (Six tests happened to use 7 days late; the other two also had a negative number of backlog papers, so the specification and the code both refused them, correctly.) The defect is a boundary written as < 7 where the rule means up to and including 7. It is the very question Chapter Twenty-Eight's review raised about the first draft of the rule, whether a form exactly 7 days late pays Rs 100 or Rs 500: the specification was corrected to say 1 to 7 days, and this code follows the other reading.

Two lessons sit in this output, one for each half of the technique.

An executable specification is a powerful oracle. Once the rule is a program, a test costs nothing to judge, and hundreds or thousands of inputs can be tried. But the oracle must be written from the specification by someone other than the developer, or from a separate reading of it. An oracle copied from the developers' code would share its mistake, agree with it everywhere, and pass it.

Random inputs find a narrow defect only by volume. The defect lives at one value out of the 22 the program drew from. The last three lines compute the chance that a batch of random tests never tries that value: about 63 per cent for 10 tests, about 10 per cent for 50, and 0.0001 for 200. Random testing with a good oracle works when tests are cheap enough to run in their hundreds. When each test costs a person's time, the tests must be aimed. Chapter Fifty-Five's boundary value analysis puts a test on 7 and on 8 days by rule, the first time, every time.

munotes.in296

Specification-Based Testing

The four techniques as four test models

Each of MU's four techniques reads the specification for a different kind of structure, builds a different test model from it, and counts a different kind of coverage item.

TechniqueThe test modelCoverage itemsWhat in a specification calls for it
Equivalence partitioning (Chapter Fifty-Four)Partitions, each a "class of inputs or outputs that are expected to be treated similarly by the test item"Each partitionValues that fall into ranges or groups treated alike
Boundary value analysis (Chapter Fifty-Five)The boundaries of equivalence partitionsEach boundary valueRanges with limits: 1 to 7 days, 8 to 15 days
Decision table testing (Chapter Fifty-Six)A decision table: "tabular representation of decision rules between causes (inputs described as Boolean conditions) and effects (outputs described as Boolean expressions)"Each decision ruleSeveral conditions combining to choose an outcome
State transition testing (Chapter Fifty-Seven)A state model: states, events and transitionsEach transitionBehaviour that depends on what happened before

The quoted definitions are ISO/IEC/IEEE 29119-4's. The same standard lists more specification-based techniques than MU names: requirements-based testing and random testing, both used above, and scenario testing, "based on exercising sequences of interactions between the test item and other systems", which is how Chapter Forty-Three, on system testing, drove ExamReg end to end.

Strengths and limits

Strengths.

  • The tests can be designed as soon as the specification exists, before any code, and they stay valid when the code is rewritten.
  • The tester needs to understand the domain (fees, forms, the exam cell's rules), not the programming language.
  • The specification supplies the expected results, so every test has an oracle from the start.
  • Designing tests reviews the specification, and a defect found in a requirement costs far less to fix than one found after delivery, as Chapter Ninety-Six, on why reviews pay, shows with Boehm and Basili's figures.

Limits.

  • The tests are only as good as the specification: what it leaves out, no test derives; what it leaves vague, no test can judge.
  • Code that no requirement mentions is never aimed at (Chapter Fifty-Two, on black-box and white-box testing).
  • Coverage is measured over the specification, and at a coarse grain it can reach 100 per cent while a defect remains, as part 1 showed.
munotes.in297

Specification-Based Testing

What it does not mean

Specification-based testing does not need a long formal document. Any source that says precisely what the right result is can serve: a user story's acceptance criteria, a table of rules, a state model.

It is not only for functional requirements. A response-time requirement or a usability requirement is a specification too, and tests derived from it are specification-based.

100 per cent requirement coverage is not complete testing. Each requirement was exercised once, and the defect at 7 days survived.

Random testing is not careless testing. It is a named technique in ISO/IEC/IEEE 29119-4, and it is only as good as the oracle that judges its results.

Quick revision

  • Specification-based testing (ISO/IEC/IEEE 29119-1:2022): the principal test basis is the external inputs and outputs, from a specification, not the implementation.
  • A specification is complete, precise and verifiable (ISO/IEC/IEEE 15289); each word is a way it can fail the tester.
  • Test bases: SRS, use cases, user stories with acceptance criteria (scenario-oriented or rule-oriented), rule tables, state models, interface specifications.
  • Test analysis answers "what to test?", test design "how to test?" (ISTQB); the products are test conditions, a test model, coverage items and test cases with expected results.
  • Worked example: 10 atomic requirements, 10 tests, 100% requirement coverage, all passed, defect missed; 200 random tests judged by an executable specification found it at 7 days.
  • Random tests find a narrow defect by volume: 10 tests miss 7 days with a chance of about 0.63, 200 tests about 0.0001.
  • MU's four techniques are four test models: partitions, boundaries, decision rules, transitions.

Test yourself

1. What is specification-based testing? Why is it called black-box testing? Testing in which tests are derived from the external inputs and outputs of the test item as a specification describes them, not from its implementation. It is black-box testing because the tester does not look inside: the code is neither read nor needed, and the tests are designed from what the software is supposed to do.

2. List four kinds of test basis a specification-based tester can use. A software requirements specification; use cases, with their main and alternative flows; user stories with acceptance criteria, written as Given/When/Then scenarios or as rule lists; business rule tables; state models; and interface specifications. Any four.

3. What is a test oracle in specification-based testing, and how can a specification be made into one that judges many tests? The oracle is the source of the expected result, here the specification. Written as a program by the tester from the specification alone, independently of the developers' code, it can compute the expected result for any input, so any number of tests, including randomly generated ones, can be judged automatically.

munotes.in298

Specification-Based Testing

4. Explain requirements-based testing. What is its weakness? It derives one or more tests for each atomic requirement, a requirement that states one thing, and measures coverage as the proportion of atomic requirements exercised. Its weakness is the grain of its coverage items: a requirement such as a fee for 1 to 7 days late counts as covered after one test, although many values satisfy it and a defect at one of them can remain, as in the worked example.

5. Why must a specification be complete, precise and verifiable for testing? A missing rule produces no test condition, so its absence is never tested; an imprecise rule gives more than one expected result, so a test against it cannot be judged; and a rule that cannot be checked by a person or tool gives no test at all.

6. Name MU's four specification-based techniques and the test model each builds. Equivalence partitioning builds partitions of inputs or outputs treated alike; boundary value analysis builds the boundaries of those partitions; decision table testing builds a table of conditions and the actions each rule selects; state transition testing builds a model of states, events and transitions.

Contents This chapter on its own page

munotes.in299

Chapter Fifty-Four

Equivalence Partitioning

Syllabus topic Module 2, "Black box: Equivalence Partitioning"

In one line

Equivalence partitioning divides every input (and output) into groups of values that the specification says must be treated alike, and tests one value from each group, on the reasoning that if one value of a group shows a defect, any other would have shown it too; it turns thousands of possible values into a handful of tests.

In the wording a student can write in an examination: an equivalence partition is a "class of inputs or outputs that are expected to be treated similarly by the test item" (ISO/IEC/IEEE 29119-4), and equivalence partitioning is a "test design technique in which test cases are designed to exercise equivalence partitions by using one or more representative members of each partition" (ISO/IEC/IEEE 29119-1). In the ISTQB syllabus's words, "if a test case, that tests one value from an equivalence partition, detects a defect, this defect should also be detected by test cases that test any other value from the same partition. Therefore, one test for each partition is sufficient." Partitions are valid or invalid; they "must not overlap and must be non-empty sets". Coverage is the number of partitions exercised divided by the number identified. With several inputs, Each Choice coverage exercises every partition of every input at least once, and invalid partitions are tested one at a time, so that one defect cannot hide another.

The idea

Take the simplest question a result portal answers: given a student's percentage of marks, what grade is it? A percentage recorded to one decimal place can be any of the 1001 values from 0.0 to 100.0, and testing every one is testing the same few rules a thousand times. The specification does not treat those values one by one. It treats them in groups: every value from 80.0 up to, but not including, 90.0 gets the same grade. If the program handles 85 correctly, the specification gives no reason to think 84.3 is handled differently; and if 85 is handled wrongly, 84.3 almost certainly is too.

That is the whole of the technique. Each group is an equivalence partition, the name saying that within it every value is equivalent for testing, and one representative stands for the rest.

The ISTQB syllabus lists where partitions can be found: "inputs, outputs, configuration items, internal values, time-related values, and interface parameters". The partitions "may be continuous or discrete, ordered or unordered, finite or infinite".

Valid and invalid partitions

A partition of values the program should process is valid; a partition of values it should reject is invalid. Both must be tested, because a program that computes every correct grade and also accepts a percentage of 120 is defective.

The ISTQB syllabus admits that the line between them is drawn differently in different places: "valid values may be interpreted as those that should be processed by the test object or as those for which the specification defines their processing. Invalid values may be interpreted as those that should be ignored or rejected by the test object or as those for which no processing is defined in the test object specification." Whatever the convention, the tester writes down which partitions are which, so the expected result of every test is settled before it runs.

munotes.in300

Equivalence Partitioning

Worked problem 1: MU's own grade table

MU prints the rule in the circular that carries this syllabus, as its table of letter grades and grade points:

% of marksLetter gradeGrade point
90.0 to 100O (Outstanding)10
80.0 to < 90.0A+ (Excellent)9
70.0 to < 80.0A (Very Good)8
60.0 to < 70.0B+ (Good)7
55.0 to < 60.0B (Above Average)6
50.0 to < 55.0C (Average)5
40.0 to < 50.0P (Pass)4
Below 40.0F (Fail)0
Ab (Absent)Ab (Absent)0

Notice how the table is written: 80.0 to < 90.0, not 80 to 90. Every value belongs to exactly one row, which is the syllabus's rule that partitions "must not overlap". The table is a partition of the input already, and it gives the tester nine valid partitions, one of them not a number at all: a student who was absent has no percentage, and the input is the mark Ab.

The table does not say what to do with values no student can have, so the tester adds the invalid partitions the specification leaves implicit and records the decision: a percentage below 0, a percentage above 100, and an input that is not a number. Below 40.0 is read as 0 to below 40.0, because nothing below 0 is a percentage. That gives twelve partitions and twelve tests, one value from the middle of each.

The function under test is printed so the program can import it; the tester works from MU's table.

def grade(percent):
    """Letter grade and grade point for a percentage of marks, or for "Ab" (absent)."""
    if percent == "Ab":
        return ("Ab", 0)
    if not isinstance(percent, (int, float)) or percent < 0:
        raise ValueError("not a percentage of marks")
    if percent >= 90.0:
        return ("O", 10)
    if percent > 80.0:
        return ("A+", 9)
    if percent >= 70.0:
        return ("A", 8)
    if percent >= 60.0:
        return ("B+", 7)
    if percent >= 50.0:
        return ("C", 5)
    if percent >= 40.0:
        return ("P", 4)
    return ("F", 0)
from mu_grade import grade

# (partition, valid?, one value from inside it, expected result from MU's table)
partitions = [("90.0 to 100",     "valid",   95,        ("O", 10)),
              ("80.0 to < 90.0",  "valid",   85,        ("A+", 9)),
              ("70.0 to < 80.0",  "valid",   75,        ("A", 8)),
              ("60.0 to < 70.0",  "valid",   65,        ("B+", 7)),
              ("55.0 to < 60.0",  "valid",   57.5,      ("B", 6)),
              ("50.0 to < 55.0",  "valid",   52.5,      ("C", 5)),
              ("40.0 to < 50.0",  "valid",   45,        ("P", 4)),
              ("0 to < 40.0",     "valid",   20,        ("F", 0)),
              ("absent",          "valid",   "Ab",      ("Ab", 0)),
              ("below 0",         "invalid", -5,        "refused"),
              ("above 100",       "invalid", 120,       "refused"),
              ("not a number",    "invalid", "seventy", "refused")]

def run(value):
    try:
        letter, point = grade(value)
        return f"{letter} ({point})"
    except ValueError:
        return "refused"

failed = 0
for name, kind, value, expected in partitions:
    want = expected if expected == "refused" else f"{expected[0]} ({expected[1]})"
    got = run(value)
    failed += got != want
    print(f"{name:<16} {kind:<8} {value!r:<10} expected {want:<8} got {got:<8}"
          f" {'pass' if got == want else 'FAIL'}")
print(f"partition coverage: {len(partitions)} of {len(partitions)} partitions, {failed} failed")
munotes.in301

Equivalence Partitioning

90.0 to 100      valid    95         expected O (10)   got O (10)   pass
80.0 to < 90.0   valid    85         expected A+ (9)   got A+ (9)   pass
70.0 to < 80.0   valid    75         expected A (8)    got A (8)    pass
60.0 to < 70.0   valid    65         expected B+ (7)   got B+ (7)   pass
55.0 to < 60.0   valid    57.5       expected B (6)    got C (5)    FAIL
50.0 to < 55.0   valid    52.5       expected C (5)    got C (5)    pass
40.0 to < 50.0   valid    45         expected P (4)    got P (4)    pass
0 to < 40.0      valid    20         expected F (0)    got F (0)    pass
absent           valid    'Ab'       expected Ab (0)   got Ab (0)   pass
below 0          invalid  -5         expected refused  got refused  pass
above 100        invalid  120        expected refused  got O (10)   FAIL
not a number     invalid  'seventy'  expected refused  got refused  pass
partition coverage: 12 of 12 partitions, 2 failed

Twelve tests, 100 per cent partition coverage, and two failures, each of a kind equivalence partitioning is built to catch.

A missing partition. A student with 57.5 per cent should get B (Above Average), grade point 6, and gets C. The code has no branch for the B band at all: its test for C begins at 50.0, so every value from 55.0 to below 60.0 falls into C. One representative value from the B partition was enough to show it, and any other value from the partition would have shown the same.

A missing rejection. A percentage of 120 should be refused and is given O. The code checks for values below 0 and never for values above 100. The invalid partition exists only because the tester wrote it down; a test set built from the valid rows of MU's table alone would never have tried it.

munotes.in302

Equivalence Partitioning

The code has a third defect, which none of these twelve tests touches. Every value tested sat in the middle of its partition, where, if the partitions are right, one value is as good as another. The places where that assumption fails are the edges, and Chapter Fifty-Five, on boundary value analysis, tests them.

Coverage

The ISTQB syllabus defines the measure: "Coverage is measured as the number of partitions exercised by at least one test case, divided by the total number of identified partitions, and is expressed as a percentage." To reach 100 per cent, "test cases must exercise all identified partitions (including invalid partitions)". The twelve tests above reached 12 of 12.

Two words in that definition deserve attention. Identified: coverage is measured against the partitions the tester found, so a partition nobody identified (had the tester not thought of percentages above 100) lowers nothing and is never tested. And including invalid partitions: a suite that tests only valid values cannot reach 100 per cent.

Several inputs: Each Choice coverage

Most functions take more than one input, and each input has its own partitions. ExamReg's fee rule has three inputs:

InputValid partitionsInvalid partitions
Days late0; 1 to 7; 8 to 15Negative; more than 15
Backlog papersWhole numbers, 0 or moreNegative whole numbers; numbers that are not whole
ConcessionYes; noNone

That is ten partitions in all. Every combination of them would be 5 × 3 × 2 = 30 tests; working through combinations in a disciplined way is what decision tables do, in Chapter Fifty-Six. Equivalence partitioning asks for less. The ISTQB syllabus calls the simplest criterion Each Choice coverage (after Ammann and Offutt), and it "requires test cases to exercise each partition from each set of partitions at least once". Since one test takes one value for each input, it covers one partition of each at once, and the smallest set that achieves Each Choice coverage has as many tests as the input with the most partitions: here five.

One invalid value per test

The smallest set has a trap in it, and the program below springs it. The version under test checks for negative backlog papers but has forgotten to refuse a number that is not whole.

def total_fee(days_late, backlog_papers, concession):        # the version under test
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    if backlog_papers < 0:
        raise ValueError("backlog papers cannot be negative")
    fee = (0 if concession else 800) + 150 * backlog_papers
    late = 0 if days_late == 0 else (100 if days_late <= 7 else 500)
    return fee + late

# the partitions of each input: (name, valid?, the test that a value belongs to it)
partitions = {
    "days late": [("negative", False, lambda d: d < 0),     ("0", True, lambda d: d == 0),
                  ("1 to 7", True, lambda d: 1 <= d <= 7),  ("8 to 15", True, lambda d: 8 <= d <= 15),
                  ("over 15", False, lambda d: d > 15)],
    "backlog papers": [("negative whole", False, lambda b: b == int(b) and b < 0),
                       ("whole, 0 or more", True, lambda b: b == int(b) and b >= 0),
                       ("not whole", False, lambda b: b != int(b))],
    "concession": [("yes", True, lambda c: c is True), ("no", True, lambda c: c is False)],
}

def check(name, tests):
    covered, combined, lines = set(), 0, []
    for args, expected in tests:
        invalid = 0
        for (inp, parts), value in zip(partitions.items(), args):
            for part, valid, contains in parts:
                if contains(value):
                    covered.add((inp, part))
                    invalid += not valid
        combined += invalid > 1
        try:
            actual = total_fee(*args)
        except ValueError:
            actual = "refused"
        lines.append(f"   {str(args):<18} expected {str(expected):<8} got {str(actual):<8}"
                     f" {'pass' if actual == expected else 'FAIL'}")
    total = sum(len(parts) for parts in partitions.values())
    print(f"{name}: {len(tests)} tests, Each Choice coverage {len(covered)} of {total},"
          f" {combined} test(s) with more than one invalid value")
    print("\n".join(lines))

check("smallest set", [((-2, -1, True), "refused"), ((20, 1.5, False), "refused"),
                       ((0, 2, True), 300), ((4, 0, False), 900), ((10, 1, True), 650)])
check("one invalid value per test",
      [((-2, 0, False), "refused"), ((20, 0, False), "refused"), ((0, -1, False), "refused"),
       ((0, 1.5, False), "refused"), ((0, 2, True), 300), ((4, 0, False), 900),
       ((10, 1, True), 650)])
munotes.in303

Equivalence Partitioning

smallest set: 5 tests, Each Choice coverage 10 of 10, 2 test(s) with more than one invalid value
   (-2, -1, True)     expected refused  got refused  pass
   (20, 1.5, False)   expected refused  got refused  pass
   (0, 2, True)       expected 300      got 300      pass
   (4, 0, False)      expected 900      got 900      pass
   (10, 1, True)      expected 650      got 650      pass
one invalid value per test: 7 tests, Each Choice coverage 10 of 10, 0 test(s) with more than one invalid value
   (-2, 0, False)     expected refused  got refused  pass
   (20, 0, False)     expected refused  got refused  pass
   (0, -1, False)     expected refused  got refused  pass
   (0, 1.5, False)    expected refused  got 1025.0   FAIL
   (0, 2, True)       expected 300      got 300      pass
   (4, 0, False)      expected 900      got 900      pass
   (10, 1, True)      expected 650      got 650      pass

Both sets reach Each Choice coverage of 10 of 10. The smallest set passes completely. The second finds the defect: one and a half backlog papers is charged Rs 1025.0 instead of being refused.

The smallest set had a test for the not whole partition, (20, 1.5, False), and it passed, because the form was refused for being 20 days late before the backlog papers were ever examined. The second invalid value never had a chance to fail. This is fault masking, a "condition in which one fault prevents the detection of another" (ISO/IEC/IEEE 24765). The 2018 edition of the ISTQB syllabus gives the rule for partitions: invalid partitions "should be tested individually, i.e., not combined with other invalid equivalence partitions, to ensure that failures are not masked." The current syllabus says the same of state transitions: "Testing only one invalid transition in a single test case helps to avoid defect masking". The price is two extra tests, and it buys a test that can actually fail.

munotes.in304

Equivalence Partitioning

The procedure, in five steps

  1. List the inputs and outputs the specification names, including the ones that are not numbers (the Ab mark, the concession).
  2. Partition each into valid and invalid partitions, from the specification's own groupings, and add the invalid partitions it leaves implicit; check that no two partitions overlap and none is empty.
  3. Choose one value from each partition, usually from its middle, away from the edges that the next chapter tests.
  4. Combine the values into test cases: with several inputs, reach Each Choice coverage, and put only one invalid value in any test.
  5. Write each expected result from the specification, run, and measure coverage as partitions exercised over partitions identified.

Strengths and limits

Strengths. Few tests replace many: twelve tests for a grade function that accepts a thousand valid percentages. Every partition of the specification is tried, so a missing rule, like the B band, shows itself. Invalid partitions make rejection part of the test design instead of an afterthought. And coverage is a clear number.

Limits. The technique rests on the partitions being right. If the code splits a partition the specification treats as one (a special case at some value inside it), one representative will usually miss it. It tests the middles and says nothing about the edges, which is where the next chapter looks. And Each Choice coverage ignores combinations of partitions, which is where decision tables look.

What it does not mean

Equivalent does not mean equal. The values of a partition are equivalent for testing, because the specification treats them alike; they still produce different outputs (every percentage from 80.0 to below 90.0 is A+, and every backlog count gives a different fee).

One value per partition is not a law. It is the minimum. A tester may choose more, and should, where a partition is large or the risk is high.

Invalid partitions are not optional. Coverage counts them, and the defects they find (the grade given for 120 per cent) are defects a user will meet.

Each Choice coverage is not combination coverage. It exercises every partition, not every combination of partitions.

munotes.in305

Equivalence Partitioning

Quick revision

  • Equivalence partition: a "class of inputs or outputs that are expected to be treated similarly by the test item" (ISO/IEC/IEEE 29119-4); partitions must not overlap and must be non-empty (ISTQB).
  • One value per partition, because a defect shown by one value should be shown by any other in the same partition.
  • Valid partitions are processed, invalid ones rejected; conventions vary, so write them down.
  • Coverage = partitions exercised ÷ partitions identified, including the invalid ones.
  • MU's grade table (page 107 of its circular): 9 valid partitions (8 bands and Ab) plus 3 invalid; 12 tests found the missing B band (57.5 gave C) and the missing check above 100 (120 gave O).
  • Each Choice coverage: every partition of every input at least once; the smallest set has as many tests as the input with most partitions.
  • One invalid value per test, or one fault masks another: (20, 1.5, False) passed, (0, 1.5, False) failed.

Test yourself

1. What is equivalence partitioning? On what assumption does it rest? A black-box test design technique that divides the inputs and outputs of a test item into partitions of values the specification treats alike, and tests at least one representative value from each. It assumes that if one value of a partition reveals a defect, any other value of the same partition would too, so one test per partition is sufficient.

2. Distinguish valid from invalid partitions, with an example of each from MU's grade table. A valid partition holds values the program should process, such as percentages from 80.0 to below 90.0, which give A+; an invalid partition holds values it should reject, such as a percentage above 100 or an input that is not a number.

3. How is equivalence partition coverage measured? As the number of partitions exercised by at least one test divided by the total number of identified partitions, including invalid ones, expressed as a percentage.

4. Derive equivalence partitions and one test value for each for MU's grade table. Valid: 90.0 to 100 (95, O), 80.0 to below 90.0 (85, A+), 70.0 to below 80.0 (75, A), 60.0 to below 70.0 (65, B+), 55.0 to below 60.0 (57.5, B), 50.0 to below 55.0 (52.5, C), 40.0 to below 50.0 (45, P), 0 to below 40.0 (20, F) and absent (Ab). Invalid: below 0 (-5), above 100 (120), not a number (seventy). Twelve tests.

5. What is Each Choice coverage, and how many tests does it need for ExamReg's fee rule? Coverage in which every partition of every input is exercised by at least one test, without regard to combinations. The fee rule's inputs have 5, 3 and 2 partitions, so at least 5 tests are needed; with invalid partitions tested one at a time, 7.

munotes.in306

Equivalence Partitioning

6. Why should invalid partitions be tested one at a time? Because when a test holds two invalid values, the program may reject it for the first and never examine the second, so a defect in handling the second is masked. In the worked example, a test that was both 20 days late and had 1.5 backlog papers passed, while a test with only the 1.5 found the defect.

Contents This chapter on its own page

munotes.in307

Chapter Fifty-Five

Boundary Value Analysis

Syllabus topic Module 2, "Black box: ... Boundary Value Analysis"

In one line

Boundary value analysis tests the edges of ordered partitions, the smallest and largest value of each and the values just across, because the edge is where a programmer writes < for <= or puts a limit one step out of place; it finds the defects that equivalence partitioning, testing the middles, walks past.

In the wording a student can write in an examination: boundary value analysis (BVA) is a "specification-based test design technique based on exercising the boundaries of equivalence partitions" (ISO/IEC/IEEE 29119-1:2022). A boundary value is a "data value that corresponds to a minimum or maximum input, internal, or output value specified for a system or component" (ISO/IEC/IEEE 24765). In the ISTQB syllabus's words, "BVA can only be used for ordered partitions. The minimum and maximum values of a partition are its boundary values." In 2-value BVA, "for each boundary value there are two coverage items: this boundary value and its closest neighbor belonging to the adjacent partition"; in 3-value BVA, "this boundary value and both its neighbors". Coverage is the number of those coverage items exercised divided by the number identified. BVA finds boundaries that are "misplaced to positions above or below their intended positions or are omitted altogether".

Why the edges

Chapter Fifty-Four's twelve equivalence partitioning tests of MU's grade table each took a value from the middle of its partition, and the chapter warned that the code had a third defect none of them touched. The reason is in the ISTQB syllabus: "BVA focuses on the boundary values of the partitions because developers are more likely to make errors with these boundary values."

The errors have a familiar shape. A specification says 80.0 to < 90.0 is A+, and the program must turn it into comparisons: >= 80.0 and < 90.0. Each comparison can be written with the wrong operator, > for >=, or the wrong constant, 7 where 8 was meant, and either mistake moves the boundary by exactly one step. Every value in the middle of the partition still gets the right answer. Only the value at the edge, or the one just across it, gets the wrong one. A test at 85 cannot tell >= 80.0 from > 80.0; a test at 80.0 can.

Ordered partitions only

A boundary needs an order: a smallest and a largest value, and neighbours on either side. Percentages, days and numbers of papers are ordered. The Ab mark for an absent student is not a point on the percentage scale, and a concession is yes or no, with nothing in between. For those, equivalence partitioning is the whole technique; BVA has nothing to add.

munotes.in308

Boundary Value Analysis

Finding the boundary values, and the question of precision

For whole numbers, a boundary value's neighbour is one step away: the neighbour of 7 days is 8 days. For a continuous quantity, one step is the precision the values are recorded in, and the specification has to say what it is.

MU's table prints percentages to one decimal place (80.0 to < 90.0) and grade point averages to two (8.00 to < 9.00). If the percentages really are recorded to one decimal place, the value just below 80.0 is 79.9. If the portal computed them to two, it would be 79.99, and a test at 79.9 would leave 79.91 to 79.99 untested. A tester who cannot find the precision in the specification asks, and writes the answer down. This chapter takes the table at its word: one decimal place, a step of 0.1.

With the step settled, each partition of MU's table has a lowest and a highest value: F is 0.0 to 39.9, P is 40.0 to 49.9, and so on up to O at 90.0 to 100.0. Between every pair of adjacent partitions there are two boundary values, the highest of one and the lowest of the next, and those pairs are exactly the 2-value coverage items. At the two ends of the valid range the invalid partitions supply the neighbours: -0.1 below 0.0, and 100.1 above 100.0.

Worked example 1: the grade table's third defect

The same grade function that Chapter Fifty-Four's equivalence partitioning tests ran against, unchanged:

def grade(percent):
    """Letter grade and grade point for a percentage of marks, or for "Ab" (absent)."""
    if percent == "Ab":
        return ("Ab", 0)
    if not isinstance(percent, (int, float)) or percent < 0:
        raise ValueError("not a percentage of marks")
    if percent >= 90.0:
        return ("O", 10)
    if percent > 80.0:
        return ("A+", 9)
    if percent >= 70.0:
        return ("A", 8)
    if percent >= 60.0:
        return ("B+", 7)
    if percent >= 50.0:
        return ("C", 5)
    if percent >= 40.0:
        return ("P", 4)
    return ("F", 0)

The program writes MU's table as ordered partitions, derives the boundary values from them instead of typing them, and judges each result against the table.

from mu_grade import grade

# MU's ordered partitions at the table's precision of 0.1: (lowest, highest, expected result)
partitions = [(None, -0.1, "refused"),                     # invalid: below 0
              (0.0, 39.9, "F (0)"), (40.0, 49.9, "P (4)"), (50.0, 54.9, "C (5)"),
              (55.0, 59.9, "B (6)"), (60.0, 69.9, "B+ (7)"), (70.0, 79.9, "A (8)"),
              (80.0, 89.9, "A+ (9)"), (90.0, 100.0, "O (10)"),
              (100.1, None, "refused")]                    # invalid: above 100

def expected(value):
    for low, high, result in partitions:
        if (low is None or value >= low) and (high is None or value <= high):
            return result

def run(value):
    try:
        letter, point = grade(value)
        return f"{letter} ({point})"
    except ValueError:
        return "refused"

# 2-value BVA: every partition's minimum and maximum (open ends have none)
values = sorted({v for low, high, _ in partitions for v in (low, high) if v is not None})
print(f"2-value boundary values ({len(values)}):", " ".join(f"{v:.1f}" for v in values))
failed = [v for v in values if run(v) != expected(v)]
for v in failed:
    print(f"   {v:.1f}: expected {expected(v)}, got {run(v)}")
print(f"{len(failed)} of {len(values)} failed")

# 3-value BVA: each boundary value and both its neighbours, one step of 0.1 away
three = sorted({round(v + step, 1) for v in values for step in (-0.1, 0, 0.1)})
failed3 = [v for v in three if run(v) != expected(v)]
print(f"3-value: {len(three)} values, {len(failed3)} failed:", " ".join(f"{v:.1f}" for v in failed3))
munotes.in309

Boundary Value Analysis

2-value boundary values (18): -0.1 0.0 39.9 40.0 49.9 50.0 54.9 55.0 59.9 60.0 69.9 70.0 79.9 80.0 89.9 90.0 100.0 100.1
   55.0: expected B (6), got C (5)
   59.9: expected B (6), got C (5)
   80.0: expected A+ (9), got A (8)
   100.1: expected refused, got O (10)
4 of 18 failed
3-value: 36 values, 7 failed: 55.0 55.1 59.8 59.9 80.0 100.1 100.2

Eighteen boundary values, and four failures that between them show all three of the function's defects.

  • 80.0 gets A instead of A+. This is the third defect. The code says percent > 80.0 where MU's table says 80.0 and above, so the boundary sits one step too high and a student with exactly 80.0 per cent is given a grade point of 8 instead of 9. Chapter Fifty-Four's equivalence partitioning test at 85 could not see it; the test at 80.0 is the only kind that can. It is the ISTQB syllabus's first typical defect, a boundary "misplaced to positions above or below their intended positions".
  • 55.0 and 59.9 get C instead of B. The missing B band shows itself again, now at both of its edges.
  • 100.1 gets O instead of being refused. The upper limit of the valid range is not in the code at all, the syllabus's other typical defect, a boundary "omitted altogether".

The 3-value run tests 36 values and fails 7, but the three extra failures (55.1, 59.8 and 100.2) sit inside the same faulty partitions and reveal no new defect. Here 3-value BVA doubled the tests and found nothing that 2-value had missed. The next example shows when the third value earns its place.

Worked example 2: two values or three

ExamReg's late fee has integer days, so the step is 1. Its ordered partitions are: negative (invalid), exactly 0, 1 to 7, 8 to 15, and more than 15 (invalid). The version under test is the one Chapter Fifty-Three's specification-based tests ran against, which charges Rs 500 at exactly 7 days. The program derives the 2-value and 3-value test sets from the partitions and runs both, then runs one value from each partition, as equivalence partitioning would, for comparison.

munotes.in310

Boundary Value Analysis

It then reproduces the ISTQB syllabus's own example of the difference between the two versions. A decision meant to accept every x up to and including 10 has been written as a test for x equal to 10, and the syllabus observes that "no test data derived from the 2-value BVA (x = 10, x = 11) can detect the defect. However, x = 9, derived from the 3-value BVA, is likely to detect it."

def late_fee(days_late):                     # the version under test, from Chapter 53
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    if days_late == 0:
        return 0
    elif days_late < 7:
        return 100
    return 500

def specified(days_late):                    # the oracle, from ExamReg's fee rule
    if not 0 <= days_late <= 15:
        return "refused"
    return 0 if days_late == 0 else 100 if days_late <= 7 else 500

def run(days_late):
    try:
        return late_fee(days_late)
    except ValueError:
        return "refused"

# ordered partitions of days late (whole days): (lowest, highest); None is an open end
partitions = [(None, -1), (0, 0), (1, 7), (8, 15), (16, None)]
two = sorted({v for p in partitions for v in p if v is not None})
three = sorted({v + d for v in two for d in (-1, 0, 1)})

def report(name, values, got, want):
    bad = [f"{v} (expected {want(v)}, got {got(v)})" for v in values if got(v) != want(v)]
    print(f"{name} {values}: {len(bad)} failed" + (": " + "; ".join(bad) if bad else ""))

report("late fee, 2-value BVA", two, run, specified)
report("late fee, 3-value BVA", three, run, specified)
report("late fee, one value per partition", [-5, 0, 4, 10, 20], run, specified)

def accepted(x):                             # specified: accept x <= 10
    return x == 10                           # ISTQB's example of a defect: = where <= was meant

report("ISTQB's example, 2-value", [10, 11], accepted, lambda x: x <= 10)
report("ISTQB's example, 3-value", [9, 10, 11], accepted, lambda x: x <= 10)
late fee, 2-value BVA [-1, 0, 1, 7, 8, 15, 16]: 1 failed: 7 (expected 100, got 500)
late fee, 3-value BVA [-2, -1, 0, 1, 2, 6, 7, 8, 9, 14, 15, 16, 17]: 1 failed: 7 (expected 100, got 500)
late fee, one value per partition [-5, 0, 4, 10, 20]: 0 failed
ISTQB's example, 2-value [10, 11]: 0 failed
ISTQB's example, 3-value [9, 10, 11]: 1 failed: 9 (expected True, got False)

Read the first three lines together. The 2-value set, seven tests, puts a test on 7 and on 8 by rule, and the test at 7 fails: the defect that Chapter Fifty-Three's two hundred random tests, judged by the specification, found only because six of them happened to land on 7, found here the first time, with a test designed to find it. One value from the middle of each partition (4 for the Rs 100 band, 10 for the Rs 500 band) passes everything. The 3-value set, thirteen tests, finds the same single defect and nothing more.

munotes.in311

Boundary Value Analysis

The last two lines are the syllabus's example, run. A comparison written as x == 10 agrees with x <= 10 at 10 (both true) and at 11 (both false), so the two 2-value tests pass. Only a value on the valid side away from the boundary, 9, exposes it, and only 3-value BVA includes 9. The difference is not in where the boundary is but in what the comparison does next to it. "3-value BVA is more rigorous than 2-value BVA as it may detect defects overlooked by 2-value BVA", at the price of more tests.

Why did the late fee not need the third value? Its partition 1 to 7 has two boundaries, and 2-value BVA tests both ends, 1 and 7. A defect of the == 7 kind there would fail at 1. The partition of values up to 10 in the syllabus's example has no lower end to test, which is exactly where the third value matters.

Coverage

The ISTQB syllabus measures each version against its own coverage items. For 2-value BVA, "Coverage is measured as the number of boundary values that were exercised, divided by the total number of identified boundary values"; for 3-value BVA, "the number of boundary values and their neighbors exercised, divided by the total number of identified boundary values and their neighbors". In the worked examples, full 2-value coverage took 18 tests for the grade table and 7 for the late fee; full 3-value coverage took 36 and 13.

2-value BVA3-value BVA
Coverage items per boundary valueThe value and its closest neighbour in the adjacent partitionThe value and both its neighbours
References the syllabus givesCraig 2002, Myers 2011Koomen 2006, O'Regan 2019
Tests: MU's grade table1836
Tests: ExamReg's late fee713
FindsBoundaries misplaced by a step, or omittedThe same, and some wrong comparisons next to a boundary
CostsFewer testsUp to twice as many

The procedure

  1. Start from the ordered partitions equivalence partitioning found (Chapter Fifty-Four, on equivalence partitioning).
  2. Settle the step: 1 for whole numbers; for anything continuous, the precision the specification states, or the answer to the question the tester asks.
  3. List the boundary values: the lowest and highest value of every partition, including the edges of the invalid partitions that bound the valid range.
  4. For 3-value BVA, add each boundary value's neighbours one step either side.
  5. Write each expected result from the specification, run, and measure coverage.
munotes.in312

Boundary Value Analysis

Boundaries that are not written as numbers

The 24765 definition speaks of input, internal and output values, and boundaries hide in places a table does not show.

  • Output boundaries. The largest fee ExamReg can ever charge is set by the largest number of backlog papers a student may have. The fee rule does not give one, so BVA cannot even list the boundary. Asking for it is itself a finding: a limit the specification never stated.
  • Time boundaries. On or before the last date has an edge in time: a form submitted at the last minute of the last date must be charged nothing, and one a minute later must be charged the Rs 100 fee. The ISTQB syllabus includes "time-related values" among the things partitions, and so boundaries, can be found for.
  • Size boundaries. A name field that holds 60 characters has boundaries at 60 and 61 characters, and at 0.

Strengths and limits

Strengths. BVA aims at the place where defects in comparisons actually show, and it finds them with a test designed to find them rather than by luck. It adds only a few tests to equivalence partitioning, and its coverage is measurable. And listing boundaries forces questions a specification often leaves open: the precision, the largest value, the last minute.

Limits. It needs ordered partitions. It inherits every mistake in the partitions: a partition nobody identified has no boundaries to test. It tests one input at a time, and a defect that needs a boundary value of one input and a particular value of another is a combination, which is the subject of the next chapter.

What it does not mean

BVA does not replace equivalence partitioning. It builds on the same partitions and tests their edges; the middles are still equivalence partitioning's.

A boundary is not always printed as a number. On or before the last date is a boundary in time, and a field's length is a boundary in size.

3-value BVA is not always better value. It may find more; in both worked examples on ExamReg and MU's table it found nothing that 2-value had missed, at up to twice the cost.

Full BVA coverage does not prove the partitions were right. It proves only that their edges, as identified, were tried.

Quick revision

  • BVA (ISO/IEC/IEEE 29119-1): exercising the boundaries of equivalence partitions; only ordered partitions; boundary values are each partition's minimum and maximum (ISTQB).
  • Why: "developers are more likely to make errors with these boundary values"; defects where boundaries are "misplaced" or "omitted altogether".
  • 2-value: the boundary value and its closest neighbour in the adjacent partition. 3-value: the boundary value and both neighbours. Coverage = items exercised ÷ items identified.
  • Precision: the step between a boundary and its neighbour must be known; MU's table prints 0.1 for percentages.
  • MU's grade table: 18 boundary values found all three defects, including 80.0 graded A instead of A+, which the middles had missed.
  • Late fee: 2-value tests 7 and 8 and finds the defect at 7 that one value per partition misses; 3-value (13 tests) finds nothing more there.
  • ISTQB's example: x <= 10 written as x == 10 passes 10 and 11 and fails only at 9, a 3-value test.
munotes.in313

Boundary Value Analysis

Test yourself

1. What is boundary value analysis, and why is it used? A black-box technique that tests the boundary values of ordered equivalence partitions, each partition's minimum and maximum and their neighbours. It is used because programmers most often go wrong at the edges, writing the wrong comparison or putting a limit one step out, and such defects give right answers everywhere except at or next to the boundary.

2. Distinguish 2-value and 3-value BVA. In 2-value BVA each boundary value gives two coverage items, the value and its closest neighbour in the adjacent partition; in 3-value BVA it gives three, the value and both its neighbours. 3-value BVA needs more tests and can find defects 2-value misses, such as a comparison that is right at the boundary but wrong next to it.

3. Derive the 2-value boundary values for MU's grade table, taking percentages to one decimal place. -0.1 and 0.0; 39.9 and 40.0; 49.9 and 50.0; 54.9 and 55.0; 59.9 and 60.0; 69.9 and 70.0; 79.9 and 80.0; 89.9 and 90.0; 100.0 and 100.1. Eighteen values.

4. Derive the 2-value and 3-value test values for ExamReg's late fee (whole days; refused if negative or over 15; Rs 0 at 0 days, Rs 100 for 1 to 7, Rs 500 for 8 to 15). 2-value: -1, 0, 1, 7, 8, 15, 16. 3-value: -2, -1, 0, 1, 2, 6, 7, 8, 9, 14, 15, 16, 17.

5. Why can BVA not be applied to every input? Because it needs an order: a boundary is the smallest or largest value of a partition, with neighbours on either side. Inputs whose partitions are unordered, such as a yes or no concession or the Ab mark for an absent student, have no boundaries and are tested by equivalence partitioning alone.

6. A comparison x <= 10 has been written as x == 10. Which test values find the defect, and why does 2-value BVA miss it? 2-value BVA tests 10 and 11; the faulty comparison is true at 10 and false at 11, just as the correct one is, so both tests pass. 3-value BVA adds 9, where the correct comparison is true and the faulty one false, so the test at 9 fails.

Contents This chapter on its own page

munotes.in314

Chapter Fifty-Six

Decision Table Testing

Syllabus topic Module 2, "Black box: ... Decision Table Testing"

In one line

A decision table lists the conditions a rule depends on and the actions it takes, one column for every combination of conditions that matters, so that a tester can see every case the rule covers, test each one, and notice the combinations a specification forgot or contradicted.

In the wording a student can write in an examination: a decision table is a "table of all contingencies that are to be considered in the description of a problem together with the action to be taken" (ISO 5806), or, in the software testing standard's words, a "tabular representation of decision rules between causes (inputs described as Boolean conditions) and effects (outputs described as Boolean expressions)" (ISO/IEC/IEEE 29119-4). Decision table testing is a "specification-based test case design technique based on exercising decision rules in a decision table" (ISO/IEC/IEEE 29119-4). In the ISTQB syllabus's words, decision tables are used "for testing the implementation of requirements that specify how different combinations of conditions result in different outcomes". The conditions and actions are the rows; "Each column corresponds to a decision rule that defines a unique combination of conditions, along with the associated actions". The coverage items are the feasible columns, and one test per column gives 100 per cent coverage.

Why combinations need their own technique

Equivalence partitioning and boundary value analysis look at one input at a time. Chapter Fifty-Four's seven equivalence partitioning tests of the fee rule tried every partition of every input, but only five combinations of them, and a rule that depends on how the conditions combine can be wrong in a combination nobody tried. The fee rule is exactly such a rule: whether the form fee is charged depends on the concession, whether a late fee is charged and which one depends on the days late, and whether anything is charged depends on the form being accepted at all.

The ISTQB syllabus names the strength: the technique "provides a systematic approach to identify all the combinations of conditions, some of which might otherwise be overlooked. It also helps to find any gaps or contradictions in the requirements."

The parts of a decision table

  • Conditions are the questions the rule asks, written so that each has a definite answer: Is the student 8 or more days late?
  • Actions are what the system does: charge the form fee, refuse the form.
  • Rules are the columns: each is one combination of answers to the conditions, with the actions that combination requires.

The syllabus gives the usual notation. For conditions, "'T' (true) means that the condition is satisfied. 'F' (false) means that the condition is not satisfied", a dash means the condition is irrelevant for that rule, and "'N/A' means that the condition is infeasible for a given rule". For actions, "'X' means that the action should occur. Blank means that the action should not occur."

munotes.in315

Decision Table Testing

A table whose conditions and actions are all true or false is a limited-entry table. In an extended-entry table, "some or all the conditions and actions may also take on multiple values (e.g., ranges of numbers, equivalence partitions, discrete values)". The fee rule could have one extended-entry condition, days late: 0, 1 to 7, or 8 to 15, in place of the two Boolean ones used below. Both describe the same rule; the limited-entry form is the one to learn first.

Building the table for ExamReg's fee rule

Step 1: the conditions. Five questions decide every fee:

ConditionQuestion
C1Is the number of days late from 0 to 15?
C2Is the number of backlog papers a whole number, 0 or more?
C3Does the student have a concession?
C4Is the student at least 1 day late?
C5Is the student at least 8 days late?

Step 2: the actions. Refuse the form; charge the form fee of Rs 800; charge Rs 150 per backlog paper; charge the late fee of Rs 100; charge the late fee of Rs 500.

Step 3: the full table. Five conditions, each true or false, give 2 × 2 × 2 × 2 × 2 = 32 columns. "A full decision table has enough columns to cover every combination of conditions."

Step 4: remove the infeasible columns. Some combinations cannot happen. No number of days is below 1 and at least 8, so every column with C4 false and C5 true is infeasible; nor can a form 1 to 7 days late fail C1. "The table can be simplified by deleting columns containing infeasible combinations of conditions."

Step 5: merge the columns where a condition does not matter. A form outside 0 to 15 days is refused whatever the other conditions say, so every feasible column with C1 false merges into one rule, with a dash in every other condition. The same happens for an invalid number of backlog papers. "The table can also be minimized by merging columns, in which some conditions do not affect the outcome, into a single column."

The result is eight rules. The table is written in a file the programs share, because it will serve both as the thing to be checked and as the oracle.

# ExamReg's fee rule as a limited-entry decision table.
# A condition is T (true), F (false) or "-" (it does not matter for that rule).
CONDITIONS = ["C1 days late is 0 to 15",
              "C2 backlog papers is a whole number, 0 or more",
              "C3 the student has a concession",
              "C4 at least 1 day late",
              "C5 at least 8 days late"]
ACTIONS = ["refuse the form", "form fee Rs 800", "Rs 150 per backlog paper",
           "late fee Rs 100", "late fee Rs 500"]

RULES = {   #  C1   C2   C3   C4   C5      actions (numbers into ACTIONS)
    "R1": (("F", "-", "-", "-", "-"), {0}),
    "R2": (("T", "F", "-", "-", "-"), {0}),
    "R3": (("T", "T", "F", "F", "F"), {1, 2}),
    "R4": (("T", "T", "F", "T", "F"), {1, 2, 3}),
    "R5": (("T", "T", "F", "T", "T"), {1, 2, 4}),
    "R6": (("T", "T", "T", "F", "F"), {2}),
    "R7": (("T", "T", "T", "T", "F"), {2, 3}),
    "R8": (("T", "T", "T", "T", "T"), {2, 4}),
}

def conditions(days_late, backlog_papers, concession):
    """The five conditions for one set of inputs, as T or F."""
    facts = [0 <= days_late <= 15,
             isinstance(backlog_papers, int) and backlog_papers >= 0,
             concession, days_late >= 1, days_late >= 8]
    return tuple("T" if fact else "F" for fact in facts)

def rule_for(conds):
    """Every rule whose conditions match; a sound table gives exactly one."""
    return [name for name, (rule, _) in RULES.items()
            if all(r in ("-", c) for r, c in zip(rule, conds))]

def fee(actions, backlog_papers):
    """The fee the table's actions add up to, or "refused"."""
    if 0 in actions:
        return "refused"
    amounts = {1: 800, 2: 150 * backlog_papers, 3: 100, 4: 500}
    return sum(amounts[a] for a in actions)
munotes.in316

Decision Table Testing

Worked example, part 1: checking the table before trusting it

A decision table is a model, and a model can be wrong: a rule can be missing, two rules can claim the same case, or a rule's actions can disagree with the specification. The program below prints the table in the usual layout and then checks it against the fee rule written directly as code, over every combination of 24 values of days late, four numbers of backlog papers and both answers to the concession: 192 inputs.

from itertools import product
from fee_table import CONDITIONS, ACTIONS, RULES, conditions, rule_for, fee

names = list(RULES)
print(" " * 48 + " ".join(f"{n:>3}" for n in names))
for i, text in enumerate(CONDITIONS):
    print(f"{text:<48}" + " ".join(f"{RULES[n][0][i]:>3}" for n in names))
for a, text in enumerate(ACTIONS):
    row = f"{text:<48}" + " ".join(f"{'X' if a in RULES[n][1] else '':>3}" for n in names)
    print(row.rstrip())

def specified(days_late, backlog_papers, concession):        # the fee rule, FINDINGS 5.1.1
    if not 0 <= days_late <= 15:
        return "refused"
    if not isinstance(backlog_papers, int) or backlog_papers < 0:
        return "refused"
    late = 0 if days_late == 0 else 100 if days_late <= 7 else 500
    return (0 if concession else 800) + 150 * backlog_papers + late

# every combination of inputs over a range wide enough to reach every partition
inputs = list(product(range(-3, 21), [-1, 0, 2, 1.5], [False, True]))
seen = {conditions(*args) for args in inputs}
unmatched = [args for args in inputs if len(rule_for(conditions(*args))) != 1]
wrong = [args for args in inputs
         if fee(RULES[rule_for(conditions(*args))[0]][1], args[1]) != specified(*args)]
print(f"full table: {2 ** len(CONDITIONS)} columns, {len(seen)} feasible,"
      f" {2 ** len(CONDITIONS) - len(seen)} infeasible")
print(f"{len(inputs)} inputs checked: {len(unmatched)} matched by no rule or by two,"
      f" {len(wrong)} where the table disagrees with the specification")
munotes.in317

Decision Table Testing

                                                 R1  R2  R3  R4  R5  R6  R7  R8
C1 days late is 0 to 15                           F   T   T   T   T   T   T   T
C2 backlog papers is a whole number, 0 or more    -   F   T   T   T   T   T   T
C3 the student has a concession                   -   -   F   F   F   T   T   T
C4 at least 1 day late                            -   -   F   T   T   F   T   T
C5 at least 8 days late                           -   -   F   F   T   F   F   T
refuse the form                                   X   X
form fee Rs 800                                           X   X   X
Rs 150 per backlog paper                                  X   X   X   X   X   X
late fee Rs 100                                               X           X
late fee Rs 500                                                   X           X
full table: 32 columns, 20 feasible, 12 infeasible
192 inputs checked: 0 matched by no rule or by two, 0 where the table disagrees with the specification

Read the table column by column. R1 and R2 are the two refusals, merged from every column where their deciding condition fails. R3 to R5 are the students without a concession: on time, 1 to 7 days late, 8 to 15 days late. R6 to R8 are the same three for students with a concession, and the only difference is the missing X against the form fee, which is the whole of the concession rule. Every rule charges Rs 150 per backlog paper, even when the number of papers is 0, because that action is always the same formula.

The last two lines are the table's own test. Of the 32 columns of the full table, 20 are feasible, the combinations some real input produces, and 12 are infeasible. Every one of the 192 inputs is matched by exactly one rule, so the table has no gaps (no case without a rule) and no overlaps (no case with two rules that might contradict each other); and in every case the table's actions add up to the fee the specification gives. Only now is it fit to be an oracle.

Worked example, part 2: one test per rule

The coverage items are the rules: one test each, with the expected result read off the table. The version under test has a defect that only one combination reveals. The program runs the eight rule tests, and for comparison the seven equivalence partitioning tests of Chapter Fifty-Four, which reached full Each Choice coverage of the same fee rule.

munotes.in318

Decision Table Testing

from fee_table import RULES, conditions, rule_for, fee

def total_fee(days_late, backlog_papers, concession):        # the version under test
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    if not isinstance(backlog_papers, int) or backlog_papers < 0:
        raise ValueError("backlog papers must be a whole number, 0 or more")
    fee = 150 * backlog_papers
    if not concession:
        fee += 800
    if days_late >= 8:
        fee += 500
    elif days_late >= 1 and not concession:
        fee += 100
    return fee

def run(args):
    try:
        return total_fee(*args)
    except ValueError:
        return "refused"

# the feasible columns of the full table: every combination of conditions some input produces
feasible = {conditions(d, b, c) for d in range(-3, 21) for b in (-1, 0, 2, 1.5) for c in (False, True)}

def test_set(name, tests):
    failed = []
    for args in tests:
        (rule,) = rule_for(conditions(*args))
        expected = fee(RULES[rule][1], args[1])            # the table is the oracle
        if run(args) != expected:
            failed.append(f"{rule} {args}: expected {expected}, got {run(args)}")
    rules = {rule_for(conditions(*args))[0] for args in tests}
    columns = {conditions(*args) for args in tests}
    print(f"{name}: {len(tests)} tests, rules exercised {len(rules)} of {len(RULES)},"
          f" full-table columns {len(columns)} of {len(feasible)}, {len(failed)} failed")
    for line in failed:
        print("   FAIL", line)

test_set("one test per rule", [(20, 0, False), (3, -1, False), (0, 2, False), (4, 1, False),
                               (10, 0, False), (0, 1, True), (4, 2, True), (12, 0, True)])
test_set("Chapter 54's seven tests", [(-2, 0, False), (20, 0, False), (0, -1, False),
                                      (0, 1.5, False), (0, 2, True), (4, 0, False), (10, 1, True)])
one test per rule: 8 tests, rules exercised 8 of 8, full-table columns 8 of 20, 1 failed
   FAIL R7 (4, 2, True): expected 400, got 300
Chapter 54's seven tests: 7 tests, rules exercised 5 of 8, full-table columns 6 of 20, 0 failed

The test for R7 fails: a student with a concession, 4 days late, with 2 backlog papers, should pay 2 × 150 + 100 = 400 rupees and is charged 300. The code has quietly extended the concession to the Rs 100 late fee (elif days_late >= 1 and not concession), something no rule says. Only one combination shows it, concession and 1 to 7 days late, and the decision table has a column for exactly that combination.

The equivalence partitioning tests pass completely. Each of them tried a concession and each tried the band of 1 to 7 days late, but never the two together: they exercised only 5 of the 8 rules, and R7 was not among them. Each Choice coverage, by its definition, "does not take into account combinations of partitions"; a decision table is nothing but combinations.

munotes.in319

Decision Table Testing

Coverage, full and minimized

The ISTQB syllabus defines the measure: "the coverage items are the columns containing feasible combinations of conditions", and "Coverage is measured as the number of exercised columns, divided by the total number of feasible columns". Measured against the minimized table, the eight rule tests achieve 8 of 8, 100 per cent. Measured against the full table they exercise 8 of its 20 feasible columns. The difference is the merged columns: a test for R1 at 20 days late stands for every refused column, including a form at minus 2 days, on the argument that the conditions merged away do not affect the outcome.

That argument is the table's, not a fact about the code. If the code refused a form at minus 2 days by a different path, perhaps a different error message, the minimized table would not see it. Where the risk justifies it, a tester can test the full table's feasible columns instead, or combine techniques: the invalid partitions of Chapter Fifty-Four's equivalence partitioning test both refusals separately.

When there are too many conditions

The syllabus warns that "the number of rules grows exponentially with the number of conditions". Each condition doubles the full table: 5 conditions give 32 columns, 10 give 1024, and 20 give over a million. Its remedies are the two used here and one more: "a minimized decision table or a risk-based approach may be used", testing the rules where a defect would cost most. A table that grows past what can be tested is also a message about the specification: a rule depending on ten independent conditions at once is a rule users will find hard to understand, and it may be worth splitting.

The procedure

  1. List the conditions the specification's rule depends on, each with a definite true or false answer (or a small set of values, for an extended-entry table).
  2. List the actions.
  3. Build the full table: one column per combination of conditions.
  4. Delete the infeasible columns, and merge those in which a condition does not affect the actions, marking it with a dash.
  5. Fill in the actions of each rule from the specification, and check the table: every case matched by exactly one rule, and every rule's actions as the specification says. A gap or a contradiction found here is a defect in the specification.
  6. Write one test per rule, with the expected result read off the table, and measure coverage over the feasible columns.

What it does not mean

A decision table is not only a test technique. It is a way of writing a rule down, and it is useful to analysts and developers before any test exists; building one is often how a specification's gaps are first noticed.

munotes.in320

Decision Table Testing

One test per rule is not every combination of input values. A rule stands for every input that satisfies its conditions; the values within it are chosen with equivalence partitioning and boundary value analysis.

A minimized table is not the same test as the full table. Merged columns are tested by one representative, on the assumption that the merged conditions really do not matter.

Infeasible is not the same as irrelevant. An infeasible combination cannot happen (below 1 day late and at least 8); an irrelevant condition can take either value without changing the actions.

Quick revision

  • Decision table: conditions and actions as rows, rules as columns (ISO 5806, ISO/IEC/IEEE 29119-4); for requirements where "different combinations of conditions result in different outcomes" (ISTQB).
  • Notation: T, F, a dash for irrelevant, N/A for infeasible; X for an action that occurs, blank for one that does not. Limited-entry (all true or false) against extended-entry (several values).
  • ExamReg's fee rule: 5 conditions, a full table of 32 columns, 20 feasible, 12 infeasible, minimized to 8 rules; checked over 192 inputs with no gap, no overlap and no disagreement.
  • Coverage items: the feasible columns; one test per rule, 8 tests.
  • The concession defect (Rs 300 charged where 400 is due) showed only in rule R7; Chapter Fifty-Four's seven equivalence partitioning tests, with full Each Choice coverage, exercised 5 of the 8 rules and missed it.
  • Strength: every combination is considered, and gaps and contradictions in the requirement come to light. Weakness: exponential growth; remedies are minimization and risk.

Test yourself

1. What is a decision table? Describe its parts. A table that records a rule as the conditions it depends on and the actions it takes, with one column, a decision rule, for each combination of conditions. The condition rows are marked T, F, or a dash where the condition does not matter; the action rows are marked X where the action occurs. It is used for requirements in which different combinations of conditions lead to different outcomes.

2. Distinguish a limited-entry from an extended-entry decision table. In a limited-entry table every condition and action is true or false. In an extended-entry table some conditions or actions take several values, such as ranges or equivalence partitions, for example days late as 0, 1 to 7 or 8 to 15.

3. How is a full decision table reduced, and what must be checked afterwards? Columns with infeasible combinations of conditions are deleted, and columns whose actions do not depend on some condition are merged into one, with a dash for that condition. Afterwards the table is checked for completeness (every possible case matched by a rule), for no overlaps (no case matched by two rules), and against the specification (each rule's actions correct).

munotes.in321

Decision Table Testing

4. Construct a decision table for ExamReg's fee rule and derive the tests. Conditions: days late 0 to 15; backlog papers a whole number, 0 or more; concession; at least 1 day late; at least 8 days late. Actions: refuse; form fee Rs 800; Rs 150 per backlog paper; late fee Rs 100; late fee Rs 500. Rules: R1 refuse if days late is out of range; R2 refuse if backlog papers are invalid; R3 to R5 without a concession, on time, 1 to 7 days and 8 to 15 days late, each with the form and backlog fees and the matching late fee; R6 to R8 the same with a concession, without the form fee. One test per rule, for example (20, 0, no), (3, -1, no), (0, 2, no), (4, 1, no), (10, 0, no), (0, 1, yes), (4, 2, yes) and (12, 0, yes).

5. How is decision table coverage measured? As the number of feasible columns exercised by the tests divided by the total number of feasible columns, as a percentage; a minimized table's rules are its columns.

6. Why can a set of tests with full Each Choice coverage miss a defect that decision table testing finds? Because Each Choice coverage requires every partition of every input to be exercised, but not in combination. A defect that appears only when two conditions hold together, such as a concession student 1 to 7 days late, is reached only by a test with that combination, and a decision table has a rule for each combination that matters.

Contents This chapter on its own page

munotes.in322

Chapter Fifty-Seven

State Transition Testing

Syllabus topic Module 2, "Black box: ... State Transition, Testing"

In one line

When what a system does depends on what has already happened, the specification is a set of states and the events that move between them, and state transition testing drives the system through sequences of events, checking the state after each one: every state, every valid transition, and every transition that must be refused.

In the wording a student can write in an examination: a state is a "condition that characterizes the behavior of a function, subfunction or element at a point in time" (ISO/IEC/IEEE 29148), and a transition is a "change from one state to another state or the same state" (ISO/IEC 11411). A state diagram "depicts the states that a system or component can assume, and shows the events or circumstances that cause or result from a change from one state to another" (ISO/IEC/IEEE 24765). State transition testing is a "specification-based test case design technique based on exercising transitions in a state model" (ISO/IEC/IEEE 29119-4). In the ISTQB syllabus, a transition "is initiated by an event, which may be additionally qualified by a guard condition", and is labelled "event [guard condition] / action"; a state table is "a model equivalent to a state diagram" that also shows the invalid transitions. Its three coverage criteria are all states, valid transitions (0-switch) and all transitions, in increasing strength.

When the answer depends on the past

Every technique so far has treated a function as a machine that turns inputs into an output: the same days late, backlog papers and concession always give the same fee. Many behaviours are not like that. Type the correct password into ExamReg and the answer depends on what happened before: if the account was locked a minute ago, the correct password is refused. The input is the same; the history is different.

A finite state machine, a "computational model consisting of a finite number of states and transitions between those states, possibly with accompanying actions" (ISO/IEC/IEEE 24765), captures that history in a single word, the current state. Testing it means testing not single inputs but sequences of them. The ISTQB syllabus puts it this way: "A test case based on a state diagram or state table is usually represented as a sequence of events, which results in a sequence of state changes (and actions, if needed). One test case may, and usually will, cover several transitions between states."

ExamReg's login rule

The specification (the book's own, fixed in its record of decisions) says:

  • Three consecutive wrong passwords lock the account for 30 minutes.
  • While the account is locked, every attempt, with the right password or a wrong one, is refused, and the lock is not extended.
  • After 30 minutes the account unlocks, and the count of failures starts again from 0.
  • A correct password before the third failure logs the student in and resets the count to 0.
  • Logging out returns to the start.
munotes.in323

State Transition Testing

Five states are enough to describe it: ready, 1 failure, 2 failures, locked and logged in. There are four events: a correct password, a wrong one, 30 minutes passing, and log out.

ExamReg's login rule as a state diagram: ready, 1 failure, 2 failures, locked and logged in

Figure 57.1 ExamReg's login rule: the ten valid transitions, each labelled event / action

The diagram shows the valid transitions only, labelled in the syllabus's syntax, event then action: wrong / lock from 2 failures to locked. Two of them return to the state they left (correct and wrong while locked); a transition may do that, and the definition allows for it ("or the same state").

Guard conditions are the other half of the syntax. The same rule could be drawn with a single state, logged out, and a count of failures, where a wrong password is two transitions with guards: wrong [failures < 2] / add a failure, which stays logged out, and wrong [failures = 2] / lock, which goes to locked. Both models describe the same behaviour. The five-state version makes every step visible, which is why it is used here; the guarded version stays small when the limit is ten attempts instead of three.

The state table, and the transitions that must be refused

The diagram cannot show what happens when an event arrives in a state that has no arrow for it. The state table can. Its "rows represent states, and its columns represent events", its cells hold the target state, and, in the syllabus's words, "In contrast to the state diagram, the state table explicitly shows invalid transitions, which are represented by empty cells."

The table is written in a file the programs share:

# ExamReg's login rule as a state table: (state, event) -> (next state, action).
# A (state, event) pair with no entry is an INVALID transition: the event is refused
# and the state does not change.
STATES = ["ready", "1 failure", "2 failures", "locked", "logged in"]
EVENTS = ["correct", "wrong", "30 minutes", "log out"]
TABLE = {
    ("ready", "correct"):      ("logged in", "welcome"),
    ("ready", "wrong"):        ("1 failure", "try again"),
    ("1 failure", "correct"):  ("logged in", "welcome"),
    ("1 failure", "wrong"):    ("2 failures", "try again"),
    ("2 failures", "correct"): ("logged in", "welcome"),
    ("2 failures", "wrong"):   ("locked", "locked for 30 minutes"),
    ("locked", "correct"):     ("locked", "account locked"),
    ("locked", "wrong"):       ("locked", "account locked"),
    ("locked", "30 minutes"):  ("ready", "unlocked"),
    ("logged in", "log out"):  ("ready", "goodbye"),
}

def expected_states(events, start="ready"):
    """Walk the table: the state after each event, and the (state, event) pairs used."""
    state, states, pairs = start, [], []
    for e in events:
        pairs.append((state, e))
        state = TABLE.get((state, e), (state, None))[0]
        states.append(state)
    return states, pairs
munotes.in324

State Transition Testing

from login_rule import STATES, EVENTS, TABLE

print((f"{'state':<12}" + "".join(f"{e:<13}" for e in EVENTS)).rstrip())
for s in STATES:
    row = f"{s:<12}" + "".join(f"{TABLE.get((s, e), ('',))[0]:<13}" for e in EVENTS)
    print(row.rstrip())
cells = len(STATES) * len(EVENTS)
print(f"{cells} cells: {len(TABLE)} valid transitions, {cells - len(TABLE)} invalid (blank)")
state       correct      wrong        30 minutes   log out
ready       logged in    1 failure
1 failure   logged in    2 failures
2 failures  logged in    locked
locked      locked       locked       ready
logged in                                          ready
20 cells: 10 valid transitions, 10 invalid (blank)

Five states and four events make twenty cells: ten valid transitions and ten blanks. Each blank is a question the specification answers with refuse, and change nothing: logging out when nobody is logged in, entering a password when already logged in, the 30 minutes passing when the account is not locked. A blank cell is not a cell nobody thought about. Filling in the table is how a tester makes sure somebody did, and a blank the specification does not account for is a finding in itself.

Three levels of coverage

The ISTQB syllabus discusses three criteria.

  • All states coverage: "the coverage items are the states", measured as "the number of exercised states divided by the total number of states".
  • Valid transitions coverage, "also called 0-switch coverage": "the coverage items are single valid transitions", and every valid transition must be exercised.
  • All transitions coverage: "the coverage items are all the transitions shown in a state table", so the tests "must exercise all the valid transitions and attempt to execute invalid transitions". And, as with invalid partitions: "Testing only one invalid transition in a single test case helps to avoid defect masking".

The syllabus ranks them: "All states coverage is weaker than valid transitions coverage, because it can typically be achieved without exercising all the transitions. Valid transitions coverage is the most widely used coverage criterion. Achieving full valid transitions coverage guarantees full all states coverage. Achieving full all transitions coverage guarantees both full all states coverage and full valid transitions coverage and should be a minimum requirement for mission and safety-critical software."

Worked example: the login rule, tested three ways

The version under test keeps a count of failures and two flags. It has two defects. The program runs three test sets against it, one designed for each criterion, walks the state table to work out the expected state after every event, and stops a test at the first state that differs.

  • All states: one test, wrong, wrong, wrong, 30 minutes, correct, which visits all five states.
  • Valid transitions: four tests that between them take every one of the ten arrows.
  • All transitions: those four, plus one test for each of the ten blank cells: a short sequence that reaches the state, then the one invalid event.
munotes.in325

State Transition Testing

from login_rule import STATES, EVENTS, TABLE, expected_states

class Login:                                            # the version under test
    def __init__(self):
        self.failures, self.locked, self.logged_in = 0, False, False

    def state(self):
        if self.logged_in:
            return "logged in"
        if self.locked:
            return "locked"
        return ["ready", "1 failure", "2 failures"][self.failures]

    def event(self, e):
        if e == "correct" and not self.logged_in:
            self.logged_in, self.failures, self.locked = True, 0, False
        elif e == "wrong" and not self.logged_in and not self.locked:
            self.failures += 1
            if self.failures == 3:
                self.locked, self.failures = True, 0
        elif e == "30 minutes" and self.locked:
            self.locked = False
        elif e == "log out":
            self.logged_in, self.locked = False, False
        # anything else is refused, and nothing changes

def run(name, tests):
    states, valid, invalid, failures = set(["ready"]), set(), set(), []
    for events in tests:
        want, pairs = expected_states(events)
        states.update(want)
        valid.update(p for p in pairs if p in TABLE)
        invalid.update(p for p in pairs if p not in TABLE)
        login = Login()
        for i, (e, w) in enumerate(zip(events, want)):
            login.event(e)
            if login.state() != w:
                failures.append(f"{events}: after '{e}' in state '{pairs[i][0]}',"
                                f" expected '{w}', got '{login.state()}'")
                break
    print(f"{name}, {len(tests)} test(s): states {len(states)} of {len(STATES)},"
          f" valid transitions {len(valid)} of {len(TABLE)},"
          f" invalid {len(invalid)} of {len(STATES) * len(EVENTS) - len(TABLE)}; {len(failures)} failed")
    for f in failures:
        print("   FAIL", f)

all_states = [["wrong", "wrong", "wrong", "30 minutes", "correct"]]
valid_transitions = [["correct", "log out"], ["wrong", "correct"], ["wrong", "wrong", "correct"],
                     ["wrong", "wrong", "wrong", "wrong", "correct", "30 minutes"]]
reach = {"ready": [], "1 failure": ["wrong"], "2 failures": ["wrong", "wrong"],
         "locked": ["wrong", "wrong", "wrong"], "logged in": ["correct"]}
invalid_tests = [reach[s] + [e] for s in STATES for e in EVENTS if (s, e) not in TABLE]

run("all states", all_states)
run("valid transitions", valid_transitions)
run("all transitions", valid_transitions + invalid_tests)
all states, 1 test(s): states 5 of 5, valid transitions 5 of 10, invalid 0 of 10; 0 failed
valid transitions, 4 test(s): states 5 of 5, valid transitions 10 of 10, invalid 0 of 10; 1 failed
   FAIL ['wrong', 'wrong', 'wrong', 'wrong', 'correct', '30 minutes']: after 'correct' in state 'locked', expected 'locked', got 'logged in'
all transitions, 14 test(s): states 5 of 5, valid transitions 10 of 10, invalid 10 of 10; 2 failed
   FAIL ['wrong', 'wrong', 'wrong', 'wrong', 'correct', '30 minutes']: after 'correct' in state 'locked', expected 'locked', got 'logged in'
   FAIL ['wrong', 'wrong', 'wrong', 'log out']: after 'log out' in state 'locked', expected 'locked', got 'ready'

All states coverage: 5 of 5 states, and nothing found. One test visited every state and passed. It used only 5 of the 10 valid transitions, and the ones it skipped are where both defects are. This is the syllabus's point that all states coverage "can typically be achieved without exercising all the transitions".

munotes.in326

State Transition Testing

Valid transitions coverage: 10 of 10, and the first defect. The fourth test locks the account, tries a wrong password (refused, correctly), and then the right one, and the student is logged in. The code checks the lock when a password is wrong and forgets it when the password is right. A lock exists to stop someone guessing; with this defect the guesser simply carries on through the 30 minutes, every wrong guess refused without delay and the first right one let in, so the lock slows nobody down. Only a test that takes the arrow correct out of locked could find it.

All transitions coverage: 10 of 10 valid and 10 of 10 invalid, and the second defect. Logging out is not possible while the account is locked, since nobody is logged in, and the specification says the event must be refused. The code's log-out branch clears the lock as well as the login, so a request to log out, which any browser can send, unlocks the account at once. No valid transition goes near it; only the attempt at an invalid one found it. The same test set that found it tried each invalid transition in its own test, so neither defect could hide the other.

The pattern is the syllabus's ranking made concrete: each stronger criterion found everything the weaker one did, and more. For a login lock, a security control, the syllabus's advice that all transitions coverage "should be a minimum requirement for mission and safety-critical software" is the right standard.

The procedure

  1. Identify the states from the specification: the distinct situations in which the system behaves differently. Name each for what it means (locked, not state 4).
  2. Identify the events, and the actions and guard conditions that go with them.
  3. Draw the state diagram of the valid transitions, then build the state table of every state and event, and decide each blank: refused, or a missing requirement.
  4. Choose a coverage criterion by risk: all states at the least, valid transitions as the usual choice, all transitions for anything safety- or security-critical.
  5. Write test cases as event sequences from the start state, each with the expected state (and action) after every event, putting at most one invalid event in each test.
  6. Run, check the state after every event, and measure coverage over states, valid transitions and invalid transitions.

What it does not mean

A state is not a screen. 1 failure and 2 failures show the same login page; they are different states because the next wrong password does different things.

munotes.in327

State Transition Testing

A test case is not one event. It is a sequence from a known start, and the state must be checked after each event, not only at the end: a defect in the middle of a sequence can be undone by the events after it.

The invalid transitions are not optional. They are half of the table here, and the more dangerous of the two defects was behind one of them.

Full valid transitions coverage does not test every path. It takes each arrow at least once; longer combinations of arrows are a stronger criterion still.

Quick revision

  • State (ISO/IEC/IEEE 29148), transition "to another state or the same state" (ISO/IEC 11411), state diagram (ISO/IEC/IEEE 24765); state transition testing exercises "transitions in a state model" (ISO/IEC/IEEE 29119-4).
  • Label syntax: "event [guard condition] / action" (ISTQB). A state table shows the invalid transitions as empty cells.
  • A test case is a sequence of events from a start state, checked after every event.
  • Coverage: all states is weaker than valid transitions (0-switch, the most used), which is weaker than all transitions (valid ones exercised, invalid ones attempted, one per test).
  • ExamReg's login: 5 states, 4 events, 20 cells, 10 valid, 10 invalid.
  • Worked example: all states (1 test) found nothing; valid transitions (4 tests) found the correct password accepted while locked; all transitions (14 tests) also found that logging out while locked unlocks the account.

Test yourself

1. What is state transition testing, and when is it used? A black-box technique that models the system as states, events and transitions, and derives tests as sequences of events that exercise the model's states and transitions. It is used when the system's response to an input depends on its history, such as logins, orders, workflows and devices with modes.

2. Draw the state diagram and state table for ExamReg's login lockout. States: ready, 1 failure, 2 failures, locked, logged in. Transitions: a wrong password moves ready to 1 failure, 1 failure to 2 failures, and 2 failures to locked; a correct password moves ready, 1 failure or 2 failures to logged in; while locked, correct and wrong both stay locked and are refused; 30 minutes moves locked to ready; log out moves logged in to ready. In the table the other ten cells are blank: invalid transitions, refused with no change.

3. Distinguish the three coverage criteria of state transition testing. All states coverage requires every state to be visited; valid transitions (0-switch) coverage requires every valid transition to be exercised; all transitions coverage requires every valid transition to be exercised and every invalid transition to be attempted. Each is stronger than the one before, and full valid transitions coverage guarantees full all states coverage.

munotes.in328

State Transition Testing

4. Why should each invalid transition be tested in a separate test case? Because one defect can mask another: once a test has gone wrong at one invalid event, the state it is in is no longer the one expected, and the next invalid event is being tried in the wrong state, so its result means nothing. One invalid transition per test keeps each result interpretable.

5. What is a guard condition? Give an example. A condition that must be true for an event to cause a particular transition, written in square brackets after the event. For example, with a count of failures, wrong [failures = 2] / lock moves the account to locked, while wrong [failures < 2] / add a failure leaves it logged out.

6. In the worked example, why did all states coverage miss both defects? Because it can be reached with a single test that visits every state without taking every transition: the test never tried a correct password while locked, and never tried to log out while locked, which are the two places where the code was wrong.

Contents This chapter on its own page

munotes.in329

Chapter Fifty-Eight

Structural Testing: Seeing Inside the Code

Syllabus topic Module 2, "Structural Testing as White Box"

In one line

Structural testing designs tests from the structure of the code, drawn as a control flow graph of statements and the transfers of control between them, and measures how much of that structure the tests have exercised: its statements, its branches, and, in stronger criteria, its conditions and paths.

In the wording a student can write in an examination: structure-based testing, which MU calls structural testing and places under white box, is "dynamic testing in which the tests are derived from an examination of the structure of the test item" (ISO/IEC/IEEE 29119-1:2022). The structure most often used is the control flow graph: in the words of NIST Special Publication 500-235, "Each flow graph consists of nodes and edges. The nodes represent computational statements or expressions, and the edges represent transfer of control between nodes." Each structural technique names its coverage items (statements, branches, conditions, paths) and measures test coverage, the "degree, expressed as a percentage, to which specified test coverage items have been exercised by a test case or test cases" (ISO/IEC/IEEE 29119-2). The criteria form a hierarchy: branch coverage subsumes statement coverage, and the condition and path criteria go further still. MU names two, statement testing and branch testing.

What structure means

Chapter Fifty-Two, on black-box and white-box testing, set the two approaches side by side. This chapter opens the white box. Its structure can be the architecture of a system (which components call which), the menu tree of an application, or the flow of data from where a value is set to where it is used; ISO/IEC/IEEE 29119-1 notes that structure-based testing "is not restricted to use at component level". But the structure the ISTQB syllabus concentrates on, and MU with it, is the code's control flow: the "sequence in which operations are performed during the execution of a computer program" (ISO/IEC/IEEE 24765). The syllabus explains why it covers statement and branch testing only: "Because of their popularity and simplicity".

The control flow graph

NIST 500-235, McCabe and Watson's report on structured testing, gives the idea its standard form: "Control flow graphs describe the logic structure of software modules." A module has one entry point, the "point in a software module at which execution of the module can begin", and an exit, the "point in a software module at which execution of the module can terminate" (both ISO/IEC/IEEE 24765). Between them the graph has a node for each statement and an edge for each way control can pass from one statement to the next.

Here is a function from ExamReg that checks an exam form before it is accepted, kept in a file of its own so that the programs below can import it:

munotes.in330

Structural Testing: Seeing Inside the Code

def check_form(papers, days_late):
    """The problems with an exam form: an empty list means it can be accepted."""
    problems = []
    if not papers:
        problems.append("no papers selected")
    for code in papers:
        if len(code) != 7:
            problems.append(f"bad paper code {code}")
    if days_late > 15:
        problems.append("more than 15 days late")
    return problems

and here is its control flow graph, one numbered node per statement:

The control flow graph of check_form: nine nodes, four of them decisions, and twelve edges

Figure 58.1 The control flow graph of check_form: T and F mark the two edges out of each decision

Four of the nodes are decisions, "statements in which a choice between two or more possible outcomes controls which set of actions will result" (ISO/IEC/IEEE 29119-1): the two if statements, the for loop (another paper to check, or none left) and the if inside it. Each has two edges out, marked T and F. The loop shows as back edges, from the end of the loop body to node 4, which asks again whether there is another paper. Node 1 is the entry and node 9 the exit. Nine nodes and twelve edges; Chapter Seventy, on cyclomatic complexity, will count them again.

Paths, and why nobody tests them all

A path is "a sequence of instructions that are performed in the execution of a computer program" (ISO/IEC/IEEE 24765), and NIST 500-235 notes that "Each possible execution path of a software module has a corresponding path from the entry to the exit node of the module's control flow graph." Testing every path would test every way the module can run. But the loop in check_form goes round once for each paper, so a form with one paper, two papers or twenty takes a different path each time, and in the report's words, "most programs with a loop have a potentially infinite number of paths, which are not subject to exhaustive testing even in theory."

So structural testing, like every other technique, chooses a finite set of coverage items out of an infinite space. The criteria differ in what they choose.

The hierarchy of coverage criteria

CriterionCoverage items (definition)Where in this book
Statement testing"exercising executable statements in the source code of the test item" (29119-4)Chapter Fifty-Nine
Branch testing"exercising branches in the control flow of the test item" (29119-4); in 24765, "each outcome of each decision point"Chapter Sixty
Decision testing"exercising decision outcomes in the control flow of the test item" (29119-1)Chapter Sixty
Branch condition testing"exercising Boolean values of the conditions within decisions and the decision outcomes" (29119-4)Chapter Sixty, in outline
Branch condition combination testing"exercising combinations of Boolean values of conditions within a decision" (29119-4)Chapter Sixty, in outline
MC/DC testing"demonstrating that a single Boolean condition within a decision can independently affect the outcome of the decision" (29119-4)Chapter Sixty, on branch coverage, in outline
Data flow testing"exercising definition-use pairs" (29119-4)Named only
Structured (basis path) testinga basis set of independent paths, as many as the cyclomatic complexity (NIST 500-235)Chapter Seventy-Two
Path testing"all or selected paths through a computer program" (24765)All paths: impossible with loops
munotes.in331

Structural Testing: Seeing Inside the Code

The criteria are ordered by subsumption: one criterion subsumes another when every test set that satisfies the first also satisfies the second. The ISTQB syllabus states the first step: "Branch coverage subsumes statement coverage. This means that any set of test cases achieving 100% branch coverage also achieves 100% statement coverage (but not vice versa)." NIST 500-235 states a later one: "structured testing subsumes branch and statement coverage testing". Moving down the table buys thoroughness with more tests.

A small coverage tracer

To measure coverage, something must watch the code run. Python offers the hook: sys.settrace calls a function of our choosing every time the interpreter starts a new line. The file below uses it to measure statement and branch coverage of one function; this chapter and the next two use it. It reads the function's source with the ast module to find its statements and its decisions (every if, elif, while and for), records the line numbers each test executes, and decides which way each decision went by where control went next: into the decision's body is its True branch, anywhere else is its False branch.

"""A small coverage tracer: which statements and which branches of one function ran.

It watches one function with sys.settrace, which calls back on every new line executed.
A decision is an if, elif, while or for; its True branch is taken when the next line
executed is inside its body, its False branch when it is anywhere else (or the function
returns). Limits: one statement per line, and no recursion; that is all it was built for.
"""
import ast
import inspect
import sys

def structure(func):
    """Statements {line: text} and decisions {line: (first, last line of the body)}."""
    first = func.__code__.co_firstlineno
    source = inspect.getsource(func).splitlines()
    tree = ast.parse("\n".join(source)).body[0]
    statements, decisions = {}, {}
    for node in ast.walk(tree):
        docstring = isinstance(node, ast.Expr) and isinstance(node.value, ast.Constant)
        if isinstance(node, ast.stmt) and node is not tree and not docstring:
            statements[node.lineno + first - 1] = source[node.lineno - 1].strip()
        if isinstance(node, (ast.If, ast.While, ast.For)):
            decisions[node.lineno + first - 1] = (node.body[0].lineno + first - 1,
                                                  node.body[-1].end_lineno + first - 1)
    return statements, decisions

def trace(func, args):
    """Call func(*args); return the line numbers it executed, in order."""
    lines = []
    def on_line(frame, event, arg):
        if event == "line":
            lines.append(frame.f_lineno)
        return on_line
    sys.settrace(lambda frame, event, arg: on_line if frame.f_code is func.__code__ else None)
    try:
        func(*args)
    except Exception:
        pass                                   # a test's verdict is not the tracer's business
    finally:
        sys.settrace(None)
    return lines

def measure(func, tests):
    """Statement and branch coverage of func over tests (a list of argument tuples)."""
    statements, decisions = structure(func)
    ran, taken = set(), set()
    for args in tests:
        lines = trace(func, args)
        ran.update(lines)
        for i, line in enumerate(lines):
            if line in decisions:
                low, high = decisions[line]
                after = lines[i + 1] if i + 1 < len(lines) else None
                taken.add((line, after is not None and low <= after <= high))
    branches = {(d, outcome) for d in decisions for outcome in (True, False)}
    return {"statements": (len(ran & statements.keys()), len(statements)),
            "branches": (len(taken), len(branches)),
            "statements missed": [statements[n] for n in sorted(statements.keys() - ran)],
            "branches missed": [f"{statements[d]} -> {outcome}"
                                for d, outcome in sorted(branches - taken)]}
munotes.in332

Structural Testing: Seeing Inside the Code

Worked example: coverage, one test at a time

The program adds three tests to check_form one at a time and measures after each, printing what is still missing. Each new test was chosen by reading what the previous measurement said had not run.

from coverage_tracer import measure
from check_form import check_form

tests = []
for new_test in [(["USCS501", "USCS502"], 3), ([], 20), (["CS501"], 0)]:
    tests.append(new_test)
    c = measure(check_form, tests)
    s, b = c["statements"], c["branches"]
    print(f"after {len(tests)} test(s): statements {s[0]} of {s[1]}, branches {b[0]} of {b[1]}")
    for text in c["statements missed"]:
        print("    statement never run:", text)
    for text in c["branches missed"]:
        print("    branch never taken: ", text)
after 1 test(s): statements 6 of 9, branches 5 of 8
    statement never run: problems.append("no papers selected")
    statement never run: problems.append(f"bad paper code {code}")
    statement never run: problems.append("more than 15 days late")
    branch never taken:  if not papers: -> True
    branch never taken:  if len(code) != 7: -> True
    branch never taken:  if days_late > 15: -> True
after 2 test(s): statements 8 of 9, branches 7 of 8
    statement never run: problems.append(f"bad paper code {code}")
    branch never taken:  if len(code) != 7: -> True
after 3 test(s): statements 9 of 9, branches 8 of 8

The first test, a valid form with two paper codes, three days late, is the kind of test a black-box tester writes first, and it executes 6 of the 9 statements and 5 of the 8 branches. It took every decision's normal way: the form had papers, the codes were valid, the lateness was allowed. The three statements it never ran are the three ways of finding a problem, which are exactly the lines where a mistake would go unnoticed if only valid forms were ever tried.

The measurement says what to test next. The second test, an empty form 20 days late, makes two of the missing branches go the other way and runs two of the missing statements. What remains is the check on a paper code's length, and the third test, a form with the five-character code CS501, reaches it. After three tests, 9 of 9 statements and 8 of 8 branches: 100 per cent on both criteria.

munotes.in333

Structural Testing: Seeing Inside the Code

Two things about this example matter beyond it. First, the tests came from the structure: each was designed to make a particular decision go a particular way. Their expected results, though, still come from the specification (Chapter Fifty-Two, on black-box and white-box testing, explained why), and this program checks only coverage, not correctness. Second, 100 per cent did not take many tests: three tests covered every statement and every branch of a function with a loop. Coverage is a measure of what the tests reached, and the next two chapters show, each with a defect, what reaching is and is not worth.

What the tracer cannot see

The tracer counts lines, and a line can hide decisions of its own.

  • A conditional expression, late = 100 if days_late <= 7 else 500, is one line and one statement, but it holds a decision. The tracer sees the line run and cannot tell which way it went. Written as an if statement, the same logic has two branches it can see.
  • A compound condition, if days_late < 0 or days_late > 15:, is one decision, but it holds two conditions, "Boolean expression containing no Boolean operators" (ISO/IEC/IEEE 29119-4). Branch coverage asks only that the whole decision go both ways; the criteria of condition coverage ask about each condition inside it. Chapter Sixty returns to this.

Real coverage tools work from the compiled code and see more; but every coverage figure is a count of some items, and knowing which items a tool counts is part of reading its figure.

What it does not mean

Structural testing does not replace specification-based testing. It finds code no test reached; it cannot find a requirement no code implements, the defect of omission that Chapter Fifty-Two's black-box tester found and its white-box tester missed.

100 per cent coverage does not mean the code is correct. It means every coverage item ran at least once, with whatever inputs happened to reach it.

A control flow graph is not a flowchart of the requirements. It is drawn from the code as written, mistakes included.

Path coverage is not the next step after branch coverage in practice. With a loop there are infinitely many paths; the practical criteria choose a finite set, such as a basis set of independent paths.

Quick revision

  • Structure-based (structural) testing (ISO/IEC/IEEE 29119-1): tests "derived from an examination of the structure of the test item".
  • Control flow graph (NIST 500-235): nodes for statements or expressions, edges for transfers of control; one entry, one exit; decisions have T and F edges; loops show as back edges.
  • Paths: "most programs with a loop have a potentially infinite number of paths" (NIST 500-235), so criteria choose finite coverage items.
  • Hierarchy: statement, branch (decision), branch condition, condition combination, MC/DC, and path criteria; branch coverage subsumes statement coverage (ISTQB); structured testing subsumes both (NIST 500-235).
  • The tracer: sys.settrace records the lines run; ast finds statements and decisions; the next line decides which branch was taken.
  • Worked example: 1 valid test gave 6 of 9 statements and 5 of 8 branches; 3 tests gave 100 per cent of both.
  • One line can hide a decision (a conditional expression) or several conditions (and, or).
munotes.in334

Structural Testing: Seeing Inside the Code

Test yourself

1. What is structural testing? Why is it called white-box testing? Testing whose test cases are derived from the internal structure of the test item, most often the control flow of its code, and whose thoroughness is measured as coverage of that structure. It is white-box testing because the tester must see inside the item, its code or design, to derive the tests.

2. What is a control flow graph? Draw one for check_form. A graph of a module's logic in which nodes represent statements or expressions and edges represent the transfers of control between them, from one entry node to one exit node. For check_form: nine nodes (the initial assignment, the empty-form decision and its append, the loop decision, the code-length decision and its append, the lateness decision and its append, the return), with T and F edges out of the four decisions and back edges from the loop body to the loop decision; twelve edges in all.

3. Why can all paths through a program not be tested? Because a loop can be repeated any number of times, and each number of repetitions is a different path, so most programs with a loop have a potentially infinite number of paths.

4. Arrange the structural coverage criteria in order of strength. Statement coverage; branch or decision coverage; condition criteria (branch condition, branch condition combination, MC/DC); basis path (structured) testing; all paths. Branch coverage subsumes statement coverage, and structured testing subsumes both.

5. How can coverage be measured? Explain the tracer's method. By instrumenting or observing the program as the tests run and recording which coverage items execute. The tracer uses Python's trace hook to record every line of the function that runs; it finds the statements and decisions from the source, and for each decision records the branch taken by checking whether the next line executed lies inside the decision's body.

munotes.in335

Structural Testing: Seeing Inside the Code

6. What can a line-based measurement of coverage miss? Decisions and conditions inside a single line: a conditional expression is one statement containing a decision, and a compound condition joined by and or or is one decision containing several conditions, whose individual values are not measured.

Contents This chapter on its own page

munotes.in336

Chapter Fifty-Nine

Statement Testing and Statement Coverage

Syllabus topic Module 2, "White Box: Statement testing"

In one line

Statement testing designs tests to execute every executable statement of the code at least once, and statement coverage is the percentage of statements the tests have executed; 100 per cent guarantees that no statement has gone completely untested, and it is the weakest of the structural criteria, because it says nothing about the choices the code makes or the values it meets.

In the wording a student can write in an examination: statement testing is a "structure-based test case design technique based on exercising executable statements in the source code of the test item" (ISO/IEC/IEEE 29119-4). In the ISTQB syllabus's words, "the coverage items are executable statements", and "Coverage is measured as the number of statements exercised by the test cases divided by the total number of executable statements in the code, and is expressed as a percentage." At 100 per cent, "each statement with a defect will be executed, which may cause a failure demonstrating the presence of the defect". But statement coverage "may not detect defects that are data dependent (e.g., a division by zero that only fails when a denominator is set to zero)", and "100% statement coverage does not ensure that all the decision logic has been tested".

What counts as a statement

The measure divides by "the total number of executable statements", so the first question is what to count. A statement, in ISO/IEC/IEEE 24765's general definition, is "a meaningful expression that defines data, specifies program actions, or directs the assembler or compiler". Statement coverage counts only the executable ones, the ones that do something when the program runs.

The tracer this book uses (Chapter Fifty-Eight, on structural testing, built it) counts every statement in a function's body except its docstring. That includes the headers of compound statements (an if, a for), because evaluating the condition is itself executed, and it excludes the lines that are only punctuation of the language: the def line, a bare else:, comments and blank lines.

Different tools count differently: some count lines, some statements, some the compiled instructions underneath. The same tests can therefore show different percentages in different tools. A statement coverage figure means something only together with the rule that counted it, and figures from two tools should never be compared directly.

Designing tests for statement coverage

Statement testing works from the control flow graph: a node is a statement, so the task is to find a few paths from the entry that between them pass through every node, and inputs that make the program take them. The syllabus sets the aim as "to design test cases that exercise statements in the code until an acceptable level of coverage is achieved". Usually that level is 100 per cent. When it is not reachable, the unreached statement is worth a question of its own: code that no input can execute is dead code, and dead code is a finding.

munotes.in337

Statement Testing and Statement Coverage

Here are two functions from ExamReg. The first computes the late fee, and the second summarises a paper's result for the exam cell: a pass percentage, and a flag that sends a paper with fewer than 40 per cent passing to review. The specification adds a rule for a paper nobody appeared for: it has no percentage, and it goes to review.

def late_fee(days_late):
    """The late fee, in rupees, for a form days_late days after the last date (0 to 15)."""
    if days_late > 0:
        fee = 100
    if days_late > 7:
        fee = 500
    return fee

def result_summary(passed, appeared):
    """A paper's pass percentage, and whether it goes to the exam cell for review."""
    percent = passed / appeared * 100
    review = False
    if percent < 40:
        review = True
    return round(percent, 1), review

Read late_fee as a graph. Its five statements lie on one path if both conditions are true, so one test with more than 7 days late executes all of them. result_summary is the same: one test with a pass rate below 40 per cent runs every statement. So statement testing asks for one test of each, designed like this: 10 days late (expected Rs 500), and 10 passes out of 40 (expected 25.0 per cent, sent to review).

The tracer, as Chapter Fifty-Eight wrote it for structural testing; each chapter's programs run on their own, so the file is repeated here unchanged:

"""A small coverage tracer: which statements and which branches of one function ran.

It watches one function with sys.settrace, which calls back on every new line executed.
A decision is an if, elif, while or for; its True branch is taken when the next line
executed is inside its body, its False branch when it is anywhere else (or the function
returns). Limits: one statement per line, and no recursion; that is all it was built for.
"""
import ast
import inspect
import sys

def structure(func):
    """Statements {line: text} and decisions {line: (first, last line of the body)}."""
    first = func.__code__.co_firstlineno
    source = inspect.getsource(func).splitlines()
    tree = ast.parse("\n".join(source)).body[0]
    statements, decisions = {}, {}
    for node in ast.walk(tree):
        docstring = isinstance(node, ast.Expr) and isinstance(node.value, ast.Constant)
        if isinstance(node, ast.stmt) and node is not tree and not docstring:
            statements[node.lineno + first - 1] = source[node.lineno - 1].strip()
        if isinstance(node, (ast.If, ast.While, ast.For)):
            decisions[node.lineno + first - 1] = (node.body[0].lineno + first - 1,
                                                  node.body[-1].end_lineno + first - 1)
    return statements, decisions

def trace(func, args):
    """Call func(*args); return the line numbers it executed, in order."""
    lines = []
    def on_line(frame, event, arg):
        if event == "line":
            lines.append(frame.f_lineno)
        return on_line
    sys.settrace(lambda frame, event, arg: on_line if frame.f_code is func.__code__ else None)
    try:
        func(*args)
    except Exception:
        pass                                   # a test's verdict is not the tracer's business
    finally:
        sys.settrace(None)
    return lines

def measure(func, tests):
    """Statement and branch coverage of func over tests (a list of argument tuples)."""
    statements, decisions = structure(func)
    ran, taken = set(), set()
    for args in tests:
        lines = trace(func, args)
        ran.update(lines)
        for i, line in enumerate(lines):
            if line in decisions:
                low, high = decisions[line]
                after = lines[i + 1] if i + 1 < len(lines) else None
                taken.add((line, after is not None and low <= after <= high))
    branches = {(d, outcome) for d in decisions for outcome in (True, False)}
    return {"statements": (len(ran & statements.keys()), len(statements)),
            "branches": (len(taken), len(branches)),
            "statements missed": [statements[n] for n in sorted(statements.keys() - ran)],
            "branches missed": [f"{statements[d]} -> {outcome}"
                                for d, outcome in sorted(branches - taken)]}
munotes.in338

Statement Testing and Statement Coverage

Worked example, part 1: 100 per cent, and every test passes

from coverage_tracer import measure
from examreg_checks import late_fee, result_summary

def call(func, args):
    return f"{func.__name__}({', '.join(map(repr, args))})"

# (function, tests designed to execute every statement, with results expected from the specification)
plan = [(late_fee, [((10,), 500)]),
        (result_summary, [((10, 40), (25.0, True))])]
for func, tests in plan:
    for args, expected in tests:
        got = func(*args)
        print(f"{call(func, args)}: expected {expected}, got {got},"
              f" {'pass' if got == expected else 'FAIL'}")
    s = measure(func, [args for args, _ in tests])["statements"]
    print(f"   statement coverage {s[0]} of {s[1]} = {100 * s[0] / s[1]:.0f}%")
late_fee(10): expected 500, got 500, pass
   statement coverage 5 of 5 = 100%
result_summary(10, 40): expected (25.0, True), got (25.0, True), pass
   statement coverage 5 of 5 = 100%

Both functions: every statement executed, 100 per cent statement coverage, every test passed. By the criterion of this chapter, both are fully tested. Both are defective.

Worked example, part 2: what 100 per cent missed

Two more inputs, one for each function, both entirely ordinary: a form submitted on the last date itself, and a paper for which every registered student was absent.

from coverage_tracer import measure
from examreg_checks import late_fee, result_summary

def call(func, args):
    return f"{func.__name__}({', '.join(map(repr, args))})"

for func, args, expected in [(late_fee, (0,), 0), (result_summary, (0, 0), (None, True))]:
    try:
        got = func(*args)
    except Exception as problem:
        got = f"crash: {type(problem).__name__}"
    print(f"{call(func, args)}: expected {expected}, got {got}")

c = measure(late_fee, [(10,)])
print(f"late_fee's one test: statements {c['statements'][0]} of {c['statements'][1]},"
      f" branches {c['branches'][0]} of {c['branches'][1]}")
for text in c["branches missed"]:
    print("   branch never taken:", text)
late_fee(0): expected 0, got crash: UnboundLocalError
result_summary(0, 0): expected (None, True), got crash: ZeroDivisionError
late_fee's one test: statements 5 of 5, branches 2 of 4
   branch never taken: if days_late > 0: -> False
   branch never taken: if days_late > 7: -> False
munotes.in339

Statement Testing and Statement Coverage

Both crash, one for each of the two weaknesses the ISTQB syllabus names.

The decision logic that was never tested. A form on time should pay nothing. But late_fee gives fee a value only inside the two if statements, and when both conditions are false no statement assigns it, so the return finds no value and Python stops with UnboundLocalError. Every statement of the function had been executed; what had never happened was the absence of the first assignment. The path on which the defect lies, where days_late > 0 is false, contains no statement of its own, so statement coverage never asks for it. The last lines of the output say so in the tracer's words: the one test covered 5 of 5 statements but only 2 of the 4 branches, and the two it never took are both False. This is the syllabus's second weakness, that full statement coverage "may not exercise all the branches". The next chapter makes those branches its target.

The value that was never tried. result_summary divides by the number of students who appeared, and when none did the division fails with ZeroDivisionError. The statement had run, with 40 as the divisor; the defect lives not in a path but in a value. This is the syllabus's own example of a data-dependent defect, "a division by zero that only fails when a denominator is set to zero". No structural criterion, branch coverage included, would have required a test with zero: there is no branch here to cover. A boundary value analysis of students appeared would, because 0 is the smallest value the input can take (Chapter Fifty-Five, on boundary value analysis), and so would the tester's habit of trying zero, which Chapter Sixty-Two, on error guessing, turns into a technique.

So what is statement coverage worth?

A great deal, at the bottom end. The syllabus states the guarantee: at 100 per cent, "each statement with a defect will be executed, which may cause a failure demonstrating the presence of the defect". The converse is the practical use. A statement that no test has executed has not been tested at all, and whatever defect it holds will be met first by a user. In Chapter Fifty-Eight's structural testing of check_form, the first test left all three of its problem-reporting statements unexecuted; a measurement of statement coverage points at them by name.

What it cannot do is certify anything. Reaching every statement once proves that no statement was left out of the testing, not that the testing of each was adequate, and the two crashes above came from functions at 100 per cent.

munotes.in340

Statement Testing and Statement Coverage

Statement coverage in numbers

Exam questions on statement coverage ask one of three things.

  • Compute the coverage. Statement coverage = statements executed ÷ executable statements × 100. If a program has 20 executable statements and the tests execute 17, statement coverage is 85 per cent.
  • Find the fewest tests for 100 per cent. Trace paths through the control flow graph so that together they pass through every statement node. late_fee needs one test; a function with an if and an else, each containing a statement, needs at least two, because no single run executes both.
  • Show that 100 per cent is not enough. Give a test set with full statement coverage and an input it misses that fails: exactly the worked example above.

What it does not mean

100 per cent statement coverage does not mean the code is tested. It means each statement ran at least once, with whatever values reached it.

A path with no statements is still a path. The False side of an if with no else has nothing to execute, and statement coverage never visits it.

Coverage is not the same percentage in every tool. It depends on what the tool counts as a statement.

Statement testing does not choose the expected results. The test at 10 days late was designed to reach statements; its expected Rs 500 came from the fee rule.

Quick revision

  • Statement testing (ISO/IEC/IEEE 29119-4): "exercising executable statements in the source code of the test item".
  • Statement coverage = statements exercised ÷ executable statements, as a percentage (ISTQB); what counts as executable depends on the tool.
  • Design from the control flow graph: paths that together pass through every statement node.
  • 100 per cent guarantees that every defective statement was executed at least once, and so may have failed.
  • Weaknesses (ISTQB): data-dependent defects, such as division by zero; branches not exercised.
  • Worked example: one test gave 100 per cent statement coverage of each function; late_fee(0) crashed (no value assigned on the untested False path: 2 of 4 branches), and result_summary(0, 0) crashed (division by zero).

Test yourself

1. Define statement testing and statement coverage. Statement testing is a white-box technique that designs test cases to execute the executable statements of the code. Statement coverage is the number of executable statements exercised by the tests divided by the total number of executable statements, expressed as a percentage.

2. A module has 40 executable statements, and the tests execute 34. What is the statement coverage? 85 per cent: 34 of the 40 statements.

3. What does 100 per cent statement coverage guarantee, and what does it not? It guarantees that every executable statement has been executed at least once, so a defect in any statement has had a chance to cause a failure. It does not guarantee that every decision outcome has been exercised, so a defect on a path with no statements (the false side of an if without an else) can be missed, nor that the statements met the values that make them fail, such as a zero divisor.

munotes.in341

Statement Testing and Statement Coverage

4. Give a test set with 100 per cent statement coverage of late_fee, and show a defect it misses. One test, 10 days late, executes all five statements and passes, returning Rs 500. The input 0 days late, which should give Rs 0, makes the function crash, because neither assignment to fee is executed and the return has no value.

5. Why can structural coverage not find a division by zero in result_summary? Because the defect depends on a value, not on a path: the division statement is executed by any test, and no branch or statement requires the divisor to be zero. Only a test with zero students appeared finds it, which boundary value analysis or error guessing would choose.

6. Why may different tools report different statement coverage for the same tests? Because they count different things as executable statements: lines, language statements or compiled instructions, with or without compound statement headers. A coverage figure is meaningful only with its counting rule.

Contents This chapter on its own page

munotes.in342

Chapter Sixty

Branch Testing and Branch Coverage

Syllabus topic Module 2, "White Box: ... Branch testing"

In one line

Branch testing designs tests so that every decision in the code goes every way it can, true and false, and branch coverage measures the fraction of those outcomes the tests have taken; it catches what statement coverage misses on the empty side of an if, and it in turn misses what hides inside a compound condition, which the condition criteria go after.

In the wording a student can write in an examination: a branch is a "computer program construct in which one of two or more alternative sets of program statements is selected for execution" (ISO/IEC/IEEE 24765), and branch testing is "testing designed to execute each outcome of each decision point in a computer program" (ISO/IEC/IEEE 24765), or a "structure-based test case design technique based on exercising branches in the control flow of the test item" (ISO/IEC/IEEE 29119-4). In the ISTQB syllabus's words, "A branch is a transfer of control between two nodes in the control flow graph", either "unconditional (i.e., straight-line code) or conditional (i.e., a decision outcome)", and "Coverage is measured as the number of branches exercised by the test cases divided by the total number of branches". Branch coverage subsumes statement coverage: "any set of test cases achieving 100% branch coverage also achieves 100% statement coverage (but not vice versa)". It "may not detect defects requiring the execution of a specific path in a code", and it does not look inside compound conditions, which condition coverage and MC/DC do.

Branches, decisions and outcomes

In the control flow graph of Chapter Fifty-Eight, on structural testing, every edge is a transfer of control. The ISTQB syllabus calls every such edge a branch, whether it is unconditional (one statement simply follows the next) or conditional (one of the two edges out of a decision). "Conditional branches typically correspond to a true or false outcome from an 'if...then' decision, an outcome from a switch/case statement, or a decision to exit or continue in a loop."

The standards also name the conditional kind separately. A decision outcome is the "result of a decision that determines the branch to be executed" (ISO/IEC/IEEE 29119-4), and decision testing is a "structure-based test case design technique based on exercising decision outcomes in the control flow of the test item" (ISO/IEC/IEEE 29119-1). Counting all edges and counting decision outcomes give different percentages for the same tests below 100 per cent, and the same answer at 100: in code with one entry, if every outcome of every decision has been taken, every statement has run, and an unconditional edge is taken whenever the statement before it runs. The tracer of Chapter Fifty-Eight, on structural testing, counts decision outcomes, two for each if, elif, while and for, and so does this chapter.

munotes.in343

Branch Testing and Branch Coverage

Why branch coverage is stronger: the empty side of an if

The syllabus states the relation: "When 100% branch coverage is achieved, all branches in the code, unconditional and conditional, are exercised by test cases", and "Branch coverage subsumes statement coverage." The reason is simple. Every statement lies on some branch, so taking every branch runs every statement. The converse fails because a branch can contain no statement at all. An if without an else has a False branch that goes straight to the next statement; statement coverage never needs it, and branch coverage always does.

Chapter Fifty-Nine ended on exactly that branch. Its one test of late_fee executed all 5 statements and took only 2 of the 4 branches, and the two it never took were both False.

The tracer and the functions, as the earlier chapters wrote them:

"""A small coverage tracer: which statements and which branches of one function ran.

It watches one function with sys.settrace, which calls back on every new line executed.
A decision is an if, elif, while or for; its True branch is taken when the next line
executed is inside its body, its False branch when it is anywhere else (or the function
returns). Limits: one statement per line, and no recursion; that is all it was built for.
"""
import ast
import inspect
import sys

def structure(func):
    """Statements {line: text} and decisions {line: (first, last line of the body)}."""
    first = func.__code__.co_firstlineno
    source = inspect.getsource(func).splitlines()
    tree = ast.parse("\n".join(source)).body[0]
    statements, decisions = {}, {}
    for node in ast.walk(tree):
        docstring = isinstance(node, ast.Expr) and isinstance(node.value, ast.Constant)
        if isinstance(node, ast.stmt) and node is not tree and not docstring:
            statements[node.lineno + first - 1] = source[node.lineno - 1].strip()
        if isinstance(node, (ast.If, ast.While, ast.For)):
            decisions[node.lineno + first - 1] = (node.body[0].lineno + first - 1,
                                                  node.body[-1].end_lineno + first - 1)
    return statements, decisions

def trace(func, args):
    """Call func(*args); return the line numbers it executed, in order."""
    lines = []
    def on_line(frame, event, arg):
        if event == "line":
            lines.append(frame.f_lineno)
        return on_line
    sys.settrace(lambda frame, event, arg: on_line if frame.f_code is func.__code__ else None)
    try:
        func(*args)
    except Exception:
        pass                                   # a test's verdict is not the tracer's business
    finally:
        sys.settrace(None)
    return lines

def measure(func, tests):
    """Statement and branch coverage of func over tests (a list of argument tuples)."""
    statements, decisions = structure(func)
    ran, taken = set(), set()
    for args in tests:
        lines = trace(func, args)
        ran.update(lines)
        for i, line in enumerate(lines):
            if line in decisions:
                low, high = decisions[line]
                after = lines[i + 1] if i + 1 < len(lines) else None
                taken.add((line, after is not None and low <= after <= high))
    branches = {(d, outcome) for d in decisions for outcome in (True, False)}
    return {"statements": (len(ran & statements.keys()), len(statements)),
            "branches": (len(taken), len(branches)),
            "statements missed": [statements[n] for n in sorted(statements.keys() - ran)],
            "branches missed": [f"{statements[d]} -> {outcome}"
                                for d, outcome in sorted(branches - taken)]}
munotes.in344

Branch Testing and Branch Coverage

def late_fee(days_late):
    """The late fee, in rupees, for a form days_late days after the last date (0 to 15)."""
    if days_late > 0:
        fee = 100
    if days_late > 7:
        fee = 500
    return fee

def result_summary(passed, appeared):
    """A paper's pass percentage, and whether it goes to the exam cell for review."""
    percent = passed / appeared * 100
    review = False
    if percent < 40:
        review = True
    return round(percent, 1), review

Worked example 1: from 2 of 4 branches to 4 of 4

Branch testing reads the missing outcomes and designs a test for them. Both False outcomes happen together when the student is not late at all, so the second test is a form on the last date, expected Rs 0. A third test, 3 days late, is the kind a tester adds for the Rs 100 band; the program shows what it adds to coverage.

from coverage_tracer import measure
from examreg_checks import late_fee

def run(days):
    try:
        return late_fee(days)
    except Exception as problem:
        return f"crash: {type(problem).__name__}"

tests = []
for days, expected in [(10, 500), (0, 0), (3, 100)]:
    tests.append((days,))
    c = measure(late_fee, tests)
    got = run(days)
    print(f"add late_fee({days}): expected {expected}, got {got}"
          f" {'pass' if got == expected else 'FAIL'}; now statements"
          f" {c['statements'][0]} of {c['statements'][1]}, branches {c['branches'][0]} of {c['branches'][1]}")
add late_fee(10): expected 500, got 500 pass; now statements 5 of 5, branches 2 of 4
add late_fee(0): expected 0, got crash: UnboundLocalError FAIL; now statements 5 of 5, branches 4 of 4
add late_fee(3): expected 100, got 100 pass; now statements 5 of 5, branches 4 of 4

The test that branch coverage asked for is the test that fails. At 0 days neither assignment runs, and the function crashes instead of returning Rs 0. Two tests reach 100 per cent branch coverage (the combination not late but more than 7 days is impossible, so two tests take all four outcomes), and the third adds nothing to the count, though it is still a good test of the Rs 100 band. The fix the failure points to is an else, or an initial fee = 0, on the path that had no statement.

What branch coverage misses

The ISTQB syllabus names the first limit: "exercising a branch with a test case will not detect defects in all cases. For example, it may not detect defects requiring the execution of a specific path in a code." Taking every branch at least once is not taking every combination of branches, and a defect that shows only when a particular True here follows a particular False there can survive 100 per cent.

munotes.in345

Branch Testing and Branch Coverage

The second limit is inside the decisions. A decision may combine several conditions, each a "Boolean expression containing no Boolean operators" (ISO/IEC/IEEE 29119-4), with and and or. Branch coverage asks only that the whole decision be true once and false once. It does not ask why it was false.

Worked example 2: a defect inside a condition

Here is ExamReg's fee rule for one student, without the backlog papers: the form fee of Rs 800, waived by a concession, plus the late fee, which applies to everyone. The code has the defect Chapter Fifty-Six's decision table caught, written this time as part of a condition.

def total_fee(days_late, concession):
    """Form fee Rs 800, waived with a concession, plus the late fee (days_late 0 to 15)."""
    fee = 800
    if concession:
        fee = 0
    if days_late >= 8:
        fee = fee + 500
    elif days_late >= 1 and not concession:
        fee = fee + 100
    return fee

The elif is one decision with two conditions: c1, days_late >= 1, and c2, not concession. The program measures the branch coverage of a test set designed to take all six branch outcomes, and prints, for every test that reaches the elif, the value of each condition.

from coverage_tracer import measure
from fee_rule import total_fee

def specified(days_late, concession):                    # the fee rule: the late fee applies to all
    late = 0 if days_late == 0 else 100 if days_late <= 7 else 500
    return (0 if concession else 800) + late

def show(name, tests):
    b = measure(total_fee, tests)["branches"]
    print(f"{name}: branches {b[0]} of {b[1]}")
    for days, concession in tests:
        c1, c2 = days >= 1, not concession               # the elif's two conditions
        reaches = days < 8                               # the elif is only evaluated below 8 days
        conds = f"c1={'T' if c1 else 'F'} c2={'T' if c2 else 'F'}" if reaches else "elif not reached"
        got, want = total_fee(days, concession), specified(days, concession)
        print(f"   ({days:>2}, {str(concession):<5}) {conds:<16} expected {want:>4},"
              f" got {got:>4} {'pass' if got == want else 'FAIL'}")

show("branch coverage", [(10, True), (4, False), (0, False)])
show("adding the test c2=F asks for", [(10, True), (4, False), (0, False), (4, True)])
branch coverage: branches 6 of 6
   (10, True ) elif not reached expected  500, got  500 pass
   ( 4, False) c1=T c2=T        expected  900, got  900 pass
   ( 0, False) c1=F c2=T        expected  800, got  800 pass
adding the test c2=F asks for: branches 6 of 6
   (10, True ) elif not reached expected  500, got  500 pass
   ( 4, False) c1=T c2=T        expected  900, got  900 pass
   ( 0, False) c1=F c2=T        expected  800, got  800 pass
   ( 4, True ) c1=T c2=F        expected  100, got    0 FAIL
munotes.in346

Branch Testing and Branch Coverage

The first three tests take all 6 branch outcomes, 100 per cent branch coverage, and every one passes. Look at the condition columns: the elif was true once (c1 and c2 both true) and false once, but it was false because c1 was false, the student was on time. c2 was never false: no test reached the elif with a concession. So the part of the decision that is wrong, and not concession, never had a chance to make a difference.

The fourth test is the one the conditions ask for: a student with a concession, 4 days late, so that c2 is false while c1 is true. The fee rule says Rs 100, the late fee alone. The code says Rs 0, because its condition quietly excuses a concession student from the late fee as well.

Beyond branch coverage, in outline

ISO/IEC/IEEE 29119-4 defines three criteria that look inside decisions. For the elif above:

CriterionWhat it requires (ISO/IEC/IEEE 29119-4)Tests for the elif
Branch condition testing"exercising Boolean values of the conditions within decisions and the decision outcomes"Each of c1 and c2 true once and false once, and the decision both ways: the four tests above do it
Branch condition combination testing"exercising combinations of Boolean values of conditions within a decision"Every combination of c1 and c2: four, adding one on time with a concession
MC/DC testing"demonstrating that a single Boolean condition within a decision can independently affect the outcome of the decision"Pairs of tests that differ in one condition and change the outcome: (4, False) against (0, False) shows c1, and (4, False) against (4, True) shows c2; three tests

Branch condition combination testing doubles its tests with each condition added to a decision; MC/DC asks only for pairs that show each condition's effect, which here took three tests instead of four. The ISTQB syllabus leaves such criteria out of the Foundation level, noting that "There are more rigorous white-box test techniques that are used in some safety-critical, mission-critical, or high-integrity environments to achieve more thorough code coverage". For MU's examination, knowing what each asks, and why branch coverage stops short, is enough.

Branch coverage in numbers

  • Compute the coverage. Branch coverage = branch outcomes taken ÷ total branch outcomes × 100. Two decisions give four outcomes; tests taking three of them give 75 per cent.
  • Find the fewest tests for 100 per cent. Each decision needs a test that makes it true and one that makes it false; tests can share decisions. late_fee needs two, since one input can make both decisions false and another both true.
  • Compare with statement coverage. A test set at 100 per cent branch coverage is always at 100 per cent statement coverage; the reverse fails whenever an if has no else.
munotes.in347

Branch Testing and Branch Coverage

What it does not mean

100 per cent branch coverage does not mean every path was taken. It means each outcome of each decision was taken at least once, not every combination of outcomes.

It does not mean every condition was tested. A decision can go both ways while one of its conditions never changes, as c2 did.

Branch coverage and decision coverage are not different ideas. They count the same outcomes (decision coverage) or all edges including straight-line ones (branch coverage in ISTQB's sense), and they agree at 100 per cent.

Branch testing does not choose expected results. The Rs 0 for a form on time came from the fee rule, and so did the Rs 100 that exposed the concession defect.

Quick revision

  • Branch (ISO/IEC/IEEE 24765): alternative sets of statements selected for execution; branch testing: "each outcome of each decision point".
  • ISTQB: a branch is a transfer of control between two nodes of the control flow graph, unconditional or conditional; coverage = branches exercised ÷ total branches.
  • Branch coverage subsumes statement coverage, not vice versa: the False side of an if without else has no statement.
  • Worked example 1: late_fee at 2 of 4 branches passed; the test for the missing False outcomes (0 days) crashed; 2 tests give 4 of 4.
  • Limits: defects needing a specific path, and defects inside compound conditions.
  • Worked example 2: 3 tests gave 6 of 6 branches and passed; the condition not concession was never false; the test that made it false (4 days late, concession) found Rs 0 charged where Rs 100 was due.
  • Beyond: branch condition, branch condition combination and MC/DC testing (ISO/IEC/IEEE 29119-4).

Test yourself

1. Define branch testing and branch coverage. Branch testing is a white-box technique that designs tests to exercise each outcome of each decision in the code, the branches of its control flow. Branch coverage is the number of branches exercised by the tests divided by the total number of branches, expressed as a percentage.

2. Why does branch coverage subsume statement coverage, but not the reverse? Every statement lies on some branch, so exercising every branch executes every statement. But a branch may contain no statement, such as the false side of an if without an else, so a test set can execute every statement without taking every branch.

3. Give a function and a test set with 100 per cent statement coverage but not 100 per cent branch coverage, and show the defect branch testing finds. late_fee assigns the fee only inside two ifs. The single test of 10 days late executes all five statements but takes only the true outcomes. Branch testing adds a test making both decisions false, 0 days late, and the function crashes because no fee was assigned.

munotes.in348

Branch Testing and Branch Coverage

4. What kinds of defect can 100 per cent branch coverage miss? Defects that need a particular combination of branches along a path, since each branch need be taken only once; defects inside a compound condition, since the decision can go both ways without each condition taking both values; and data-dependent defects, such as a division by zero.

5. Distinguish branch condition testing, branch condition combination testing and MC/DC. Branch condition testing requires each condition within a decision to take both values and the decision to take both outcomes. Branch condition combination testing requires every combination of condition values within a decision. MC/DC requires showing that each condition can independently change the decision's outcome, by pairs of tests differing only in that condition.

6. For the decision days_late >= 1 and not concession, give tests that satisfy MC/DC. Three tests: 4 days late without a concession (true), 0 days without a concession (false: only the first condition changed) and 4 days with a concession (false: only the second condition changed).

Contents This chapter on its own page

munotes.in349

Chapter Sixty-One

Experience-Based Testing

Syllabus topic Module 2, "Experience-based"

In one line

Experience-based testing designs tests from what testers, developers and users know about where software fails, rather than from a specification or the code; it finds what the formal techniques cannot see, it is only as good as the tester's knowledge, and it works best alongside the formal techniques, never instead of them.

In the wording a student can write in an examination: experience-based testing is a "class of test case design techniques based on using the experience of testers to generate test cases" (ISO/IEC/IEEE 29119-4). In the ISTQB syllabus's words, these techniques "effectively use the knowledge and experience of testers for the design and implementation of test cases. The effectiveness of these test techniques depends heavily on the tester's skills." They "can detect defects that may be missed using the black-box test techniques and white-box test techniques. Hence, experience-based test techniques are complementary to the black-box test techniques and white-box test techniques." MU names three: error guessing, exploratory testing and checklist-based testing.

The third family

Chapter Eight, on the three families of test design techniques, placed experience-based testing beside the black-box and white-box families. The first two each have a document to work from: the specification, or the code. The third works from something that is in no document: knowledge of how software goes wrong.

The 2018 edition of the ISTQB syllabus described the family's common ground plainly: its test conditions, test cases and test data "are derived from a test basis that may include knowledge and experience of testers, developers, users and other stakeholders", and that knowledge "includes expected use of the software, its environment, likely defects, and the distribution of those defects". The current syllabus lists where a tester's knowledge comes from, in its section on error guessing:

  • "How the application has worked in the past"
  • "The types of errors the developers tend to make and the types of defects that result from these errors"
  • "The types of failures that have occurred in other, similar applications"

Worked example: a rule with nothing to partition

ExamReg prints the student's name on the hall ticket, and its specification says only this: the student's name as on the admission record, at most 40 characters. That is a thin specification, and a formal technique can take from it only what it says. Equivalence partitioning and boundary value analysis find one ordered input, the length, and test it: empty, 1 character, 40, 41, and a typical name.

An experienced tester reads the same sentence and thinks of the names that actually arrive on forms: a surname with an apostrophe, initials with dots, a hyphenated first name, a field of nothing but spaces, a name typed with a space in front, a name written in Devanagari. None of these is in the specification. Each is a question about what as on the admission record means, and some of them have obvious answers.

munotes.in350

Experience-Based Testing

import re

def valid_name(name):                           # the version under test
    """ExamReg: may this name be printed on the hall ticket?"""
    return re.fullmatch(r"[A-Za-z ]{1,40}", name) is not None

# from the specification: "the student's name as on the admission record, at most 40 characters"
formal = [("", False), ("A", True), ("A" * 40, True), ("A" * 41, False), ("Priya Sharma", True)]

# from experience of names on real forms; "ask" marks a question the specification must answer
experience = [("Priya D'Souza", True),           # an apostrophe
              ("S. R. Iyer", True),              # initials with dots
              ("Anne-Marie Fernandes", True),    # a hyphen
              ("   ", False),                    # nothing but spaces
              (" Priya Sharma", "ask"),          # a leading space: store it, or trim it?
              ("अमित पाटील", "ask")]              # Devanagari: is the admission record in English only?

for name, tests in [("formal (length partitions and boundaries)", formal),
                    ("experience-based", experience)]:
    failed = [(n, want) for n, want in tests if want != "ask" and valid_name(n) != want]
    asked = [n for n, want in tests if want == "ask"]
    print(f"{name}: {len(tests)} tests, {len(failed)} failed, {len(asked)} questions")
    for n, want in failed:
        print(f"   FAIL {n!r}: expected {'accepted' if want else 'refused'},"
              f" got {'accepted' if valid_name(n) else 'refused'}")
    for n in asked:
        print(f"   ASK  {n!r}: the code says {'accepted' if valid_name(n) else 'refused'};"
              f" the specification does not say")
formal (length partitions and boundaries): 5 tests, 0 failed, 0 questions
experience-based: 6 tests, 4 failed, 2 questions
   FAIL "Priya D'Souza": expected accepted, got refused
   FAIL 'S. R. Iyer': expected accepted, got refused
   FAIL 'Anne-Marie Fernandes': expected accepted, got refused
   FAIL '   ': expected refused, got accepted
   ASK  ' Priya Sharma': the code says accepted; the specification does not say
   ASK  'अमित पाटील': the code says refused; the specification does not say

The formal tests are right and they pass: the length rule is implemented correctly. The experience-based tests find four defects the formal ones could not have been aimed at. Three real names are refused, because the developer's pattern allows only letters and spaces; a name made of nothing but spaces is accepted, because the pattern allows spaces anywhere. And two tests end not in a verdict but in a question for the person who owns the requirement: should a leading space be trimmed or stored, and may a name be printed in Devanagari at all? The specification cannot answer either. Asking is part of the result.

That is the family's value in one example. The formal techniques test what the specification says, and here it said almost nothing. The tester's knowledge filled the gap, found defects, and improved the specification.

munotes.in351

Experience-Based Testing

When experience-based testing is the right tool

The ISTQB syllabus names the situations in which exploratory testing, the most flexible of the three, is useful: "when there are few or inadequate specifications or there is significant time pressure on the testing", and "to complement other more formal test techniques". James Bach's list of where exploratory testing fits is longer; among its entries:

  • "You need to provide rapid feedback on a new product or feature."
  • "You need to learn the product quickly."
  • "You have already tested using scripts, and seek to diversify the testing."
  • "You want to find the single most important bug in the shortest time."
  • "You want to investigate and isolate a particular defect."

And in general, he writes, it is called for "in any situation where it's not obvious what the next test should be, or when you want to go beyond the obvious tests."

The worked example was the first of the syllabus's situations: an inadequate specification. It would have been worth doing even with a good one, as the complement the syllabus describes.

The three techniques compared

Error guessingExploratory testingChecklist-based testing
What guides the testerA list of likely errors, defects and failures (Chapter Sixty-Two, on error guessing)A charter, and what each test teaches about the next (Chapter Sixty-Three, on exploratory testing)A list of test conditions, often phrased as questions (Chapter Sixty-Four, on checklist-based testing)
Tests designedBefore execution, from the listDuring execution, one from the lastBefore or during, one or more per item
DocumentationThe list and the tests derived from itSession notes, bugs and issuesThe checklist, with each item's result
RepeatabilityModerate: the list can be reusedLow: the next session explores differentlyModerate to high, depending on how detailed the items are
How coverage is judgedWhich items of the list were triedWhich areas of the charter were exploredWhich items were checked

All three share the family's defining trait: the tests come from knowledge, not from a document. The ISTQB syllabus's description of checklists states the trade-off that runs through all of them: "If the checklists are high-level, some variability in the actual testing is likely to occur, resulting in potentially greater coverage but less repeatability."

Strengths and limits

Strengths. The family finds defects no specification describes, because it works from how software fails in practice. It works where specifications are thin or absent. It is quick to start: a skilled tester can find important defects in the first hour, before any formal test is designed. And it improves the specification, because a knowledgeable tester's tests raise questions that the specification's author did not think to answer.

munotes.in352

Experience-Based Testing

Limits. Its effectiveness "depends heavily on the tester's skills": a new tester has little experience to draw on. Its coverage is hard to measure, because there is no model of the software to count items against. Its tests may be hard to repeat. And it finds only the kinds of failure somebody has met or imagined. That is why it complements the formal techniques instead of replacing them.

What it does not mean

Experience-based testing is not unplanned testing. Each technique has its structure: a list of attacks, a charter and a session report, a checklist.

It is not only for senior testers. Lists and checklists carry one tester's experience to another, and a session charter gives a less experienced tester a mission.

It is not a substitute for the formal techniques. In the worked example the formal tests were still needed: they proved the length rule right, which the experience-based tests did not examine.

A question is not a failure. When the specification does not say what should happen, the right result of the test is a question to the requirement's owner.

Quick revision

  • Experience-based testing (ISO/IEC/IEEE 29119-4): test cases generated from "the experience of testers".
  • Depends "heavily on the tester's skills"; "complementary" to black-box and white-box techniques (ISTQB).
  • Knowledge comes from how the application has worked, the errors developers tend to make, and failures in similar applications (ISTQB).
  • Useful with "few or inadequate specifications" or "significant time pressure", and as a complement (ISTQB).
  • MU's three: error guessing, exploratory testing, checklist-based testing.
  • Worked example: 5 formal tests of the name rule passed; 6 experience-based tests found 4 defects (three real names refused, a blank name accepted) and raised 2 questions.

Test yourself

1. What is experience-based testing? How does it differ from the other two families? A family of test design techniques in which test cases are derived from the knowledge and experience of testers, developers and users about how software is used and how it fails. Black-box techniques derive tests from the specification and white-box techniques from the code; experience-based techniques from knowledge that neither document holds.

2. On what knowledge does experience-based testing draw? On how the application has worked in the past, the kinds of errors developers tend to make and the defects they cause, the failures seen in similar applications, the expected use of the software and its environment, and where defects have clustered before.

3. When is experience-based testing most useful? When specifications are few or inadequate, when there is significant time pressure, when rapid feedback or rapid learning is needed, and always as a complement to the formal techniques to go beyond the obvious tests.

munotes.in353

Experience-Based Testing

4. Name the three experience-based techniques in MU's syllabus and what guides each. Error guessing, guided by a list of likely errors, defects and failures; exploratory testing, guided by a charter and by what each test reveals; checklist-based testing, guided by a checklist of test conditions.

5. What are the limitations of experience-based testing? It depends heavily on the tester's skill and knowledge; its coverage is hard to measure; its tests can be hard to repeat; and it finds only kinds of failure that someone has met or can imagine.

6. In the worked example, why could the formal techniques not find the defects in valid_name? Because the specification gave them only one rule to work from, the length of at most 40 characters, which was implemented correctly. The defects lay in which characters a real name may contain, which the specification did not state and only the tester's knowledge of real names supplied.

Contents This chapter on its own page

munotes.in354

Chapter Sixty-Two

Error Guessing

Syllabus topic Module 2, "Experience-based: Error guessing"

In one line

Error guessing designs tests from a list of the mistakes programmers are known to make and the failures software is known to suffer, aimed deliberately at the places they are likely; done systematically it is called a fault attack, and its power comes entirely from the list.

In the wording a student can write in an examination: error guessing is a "test design technique in which test cases are derived on the basis of the tester's knowledge of past failures, or general knowledge of failure modes" (ISO/IEC/IEEE 29119-1). In the ISTQB syllabus's words it is "used to anticipate the occurrence of errors, defects, and failures, based on the tester's knowledge", including how the application has worked in the past, the errors developers tend to make, and the failures of similar applications. "Fault attacks are a way to implement error guessing. This test technique requires the tester to create or acquire a list of possible errors, defects and failures, and to design tests that will identify defects associated with the errors, expose the defects, or cause the failures."

Guessing, but not at random

The name undersells the technique. An experienced tester's guess that a form will mishandle an empty field is not a guess in the everyday sense: it is a prediction from having seen empty fields mishandled many times before. The ISTQB syllabus locates the knowledge in three places: "How the application has worked in the past", "The types of errors the developers tend to make and the types of defects that result from these errors", and "The types of failures that have occurred in other, similar applications".

It also sorts the errors into six kinds, which make a useful skeleton for any list: they "may be related to: input (e.g., correct input not accepted, parameters wrong or missing), output (e.g., wrong format, wrong result), logic (e.g., missing cases, wrong operator), computation (e.g., incorrect operand, wrong computation), interfaces (e.g., parameter mismatch, incompatible types), or data (e.g., incorrect initialization, wrong type)."

A starter list, from this book's own defects

A list is best built from failures that really happened. This book has run into a good many of them already, and each was a kind that recurs:

AttackWhat it catchesWhere this book met it
Empty or blank inputA blank value accepted, or a crash on nothingA hall-ticket name of only spaces (Chapter Sixty-One, on experience-based testing)
ZeroDivision by zero, a band that starts at 1No students appeared (Chapter Fifty-Nine, on statement testing)
Negative numbersValues that should be refused are computedNegative backlog papers (Chapter Thirty-Three, on unit testing)
Values that are not wholeFractions accepted where counts are meantOne and a half backlog papers (Chapter Fifty-Four, on equivalence partitioning)
The edge of every rangeA comparison one step out of placeA form exactly 7 days late (Chapter Fifty-Five, on boundary value analysis)
Special charactersReal text refused, or unsafe text acceptedAn apostrophe in a surname (Chapter Sixty-One, on experience-based testing)
Duplicates and repeatsSomething done twice that must happen onceA payment charged twice (Chapter Forty-Four, on recovery testing)
Wrong unitsA number passed in one unit and read in anotherPaise read as rupees (Chapter Thirty-Seven, on integration testing)
Dates and calendarsMonth lengths, leap years, the turn of the yearBelow
munotes.in355

Error Guessing

The list maps onto the syllabus's six kinds: most rows are input, the wrong units are interfaces, and the date row is computation. It will be wrong for some other application and incomplete for this one, which is exactly why a list is kept and extended rather than written once.

Worked example: a fault attack on the calendar

ExamReg counts the days a form is late: calendar days after the last date, and 0 for any form on or before it. Here is the developer's version, and a list of attacks aimed at the ways date arithmetic is known to go wrong. The expected results come from Python's own datetime.date, which knows the calendar and is independent of the code under test.

from datetime import date

def days_late(last_date, submitted):              # the version under test
    """Calendar days after the last date; 0 for a form on or before it."""
    days = (submitted.month - last_date.month) * 30 + (submitted.day - last_date.day)
    return max(days, 0)

def oracle(last_date, submitted):                 # the calendar, from Python's own date type
    return max((submitted - last_date).days, 0)

# the attack list: where date arithmetic is known to go wrong
attacks = [("on the last date itself",           date(2026, 10, 15), date(2026, 10, 15)),
           ("the day after",                     date(2026, 10, 15), date(2026, 10, 16)),
           ("before the last date",              date(2026, 10, 15), date(2026, 10, 1)),
           ("across the end of a 31-day month",  date(2026, 10, 31), date(2026, 11, 1)),
           ("across February, leap year",        date(2028, 2, 28),  date(2028, 3, 1)),
           ("across February, ordinary year",    date(2027, 2, 28),  date(2027, 3, 1)),
           ("across the end of the year",        date(2026, 12, 28), date(2027, 1, 2)),
           ("a year late",                       date(2026, 10, 15), date(2027, 10, 15))]

failed = 0
for name, last, sent in attacks:
    want, got = oracle(last, sent), days_late(last, sent)
    failed += want != got
    print(f"{name:<34} expected {want:>3}, got {got:>3}  {'pass' if want == got else 'FAIL'}")
print(f"{len(attacks)} attacks, {failed} failed")
on the last date itself            expected   0, got   0  pass
the day after                      expected   1, got   1  pass
before the last date               expected   0, got   0  pass
across the end of a 31-day month   expected   1, got   0  FAIL
across February, leap year         expected   2, got   3  FAIL
across February, ordinary year     expected   1, got   3  FAIL
across the end of the year         expected   5, got   0  FAIL
a year late                        expected 365, got   0  FAIL
8 attacks, 5 failed
munotes.in356

Error Guessing

The first three attacks are the ordinary cases, and they pass: a test designed from the fee rule alone, using dates in the middle of one month, would pass too. The other five fail, and every one of them is a failure date code is famous for, because the developer treated every month as 30 days long and ignored the year.

  • The end of a 31-day month. A form submitted on 1 November for a last date of 31 October is 1 day late; the code says 0. Every 31st of a month is a free day.
  • February, in both kinds of year. Across the end of February the code counts 3 days where there are 2 in a leap year and 1 in an ordinary one, so a student who was 1 day late is charged as if 3 days late.
  • The end of the year. A last date of 28 December and a form on 2 January is 5 days late. The month difference is negative (January minus December), the total is negative, and the code returns 0: the form counts as on time.
  • A year late. The same thing, at its worst: a form a whole year late counts as on time, and is accepted.

None of these came from a specification. The fee rule says calendar days after the last date and nothing about months or years. They came from knowing how programmers get dates wrong, and each took one line to try.

Building and keeping the list

The ISTQB syllabus says where lists come from: they "can be built based on experience, defect and failure data, or from common knowledge about why software fails". Three habits keep one useful.

  • Feed it from the defect records. Every defect that escaped the formal tests is a candidate entry, described as a kind: not the name D'Souza was refused but names with punctuation. ExamReg's release 2.0 defect records, which later chapters analyse, put input validation first among the defect types, 58 of 200, and a list should weight its attacks the same way.
  • Keep the entries specific enough to act on. Check the dates is a reminder; a last date of 31 October and a form on 1 November is a test.
  • Share it. A list written down turns one tester's experience into the whole team's, which is also the idea behind the checklists of Chapter Sixty-Four, on checklist-based testing.
munotes.in357

Error Guessing

Strengths and limits

Strengths. It is fast and cheap: each attack is a single test with an obvious target. It finds real defects early, often the ones users meet first, because the list is made of failures that really happen. And it needs no detailed specification.

Limits. It is only as good as the list and the tester's knowledge: a failure nobody has met or imagined is on no list. It gives no coverage measure beyond the list itself. And it is not systematic about the specification: it complements the formal techniques, which test every partition and boundary the specification defines, and does not replace them.

What it does not mean

Error guessing is not random testing. Random testing (Chapter Fifty-Three, on specification-based testing) chooses inputs by chance; error guessing chooses them by knowledge, aimed at a named kind of failure.

The guess is not the test. An attack names a kind of failure; the test still needs exact inputs and an expected result, here from the calendar.

A list is not finished. New kinds of failure are added as they are found, and entries that stop finding anything can be retired.

It is not only for input fields. Output formats, computations, interfaces between components and stored data all have their typical failures.

Quick revision

  • Error guessing (ISO/IEC/IEEE 29119-1): tests from "the tester's knowledge of past failures, or general knowledge of failure modes".
  • Knowledge (ISTQB): how the application worked in the past; the errors developers tend to make; failures in similar applications.
  • Six kinds (ISTQB): input, output, logic, computation, interfaces, data.
  • Fault attacks: a list of possible errors, defects and failures, and tests designed to expose each.
  • Starter list: empty, zero, negative, not whole, edges, special characters, duplicates, wrong units, dates.
  • Worked example: 8 date attacks on days_late, 5 failed: a 31-day month end, February in both kinds of year, the year end (5 days late counted as 0) and a year late (counted as 0).
  • Lists come from experience, defect and failure data, and common knowledge; keep them specific, shared and up to date.

Test yourself

1. What is error guessing? On what does the tester's guess depend? A test design technique in which tests are derived from the tester's knowledge of past failures and of how software typically fails. The knowledge comes from how the application has worked in the past, the kinds of errors its developers tend to make, and the failures of similar applications.

2. What is a fault attack? A systematic way of doing error guessing: the tester creates or acquires a list of possible errors, defects and failures, and designs tests that will expose each, cause the failure or identify the defect behind it.

munotes.in358

Error Guessing

3. List six categories of errors that an error guessing list can be organised by, with an example of each. Input, such as correct input not accepted; output, such as a wrong format; logic, such as a missing case; computation, such as a wrong operand; interfaces, such as a parameter mismatch; data, such as incorrect initialisation.

4. Write an attack list for a function that counts days between two dates. The same date; the next day; a date before the first; across the end of a 30-day and a 31-day month; across February in a leap year and in an ordinary year; across the end of a year; a whole year apart; and, where the dates can be far in the future, a century year that is not a leap year.

5. In the worked example, why did the ordinary tests pass while the attacks failed? Because the ordinary tests used dates within one month, where treating every month as 30 days and ignoring the year makes no difference. The defect shows only when the dates cross the end of a month of another length, the end of February, or the end of a year, and those are exactly the cases the attack list aims at.

6. What are the strengths and weaknesses of error guessing? It is quick and cheap, needs no detailed specification, and finds the defects users are likely to meet. It depends entirely on the tester's knowledge and the list, gives no coverage measure beyond the list, and is not systematic, so it complements the formal techniques rather than replacing them.

Contents This chapter on its own page

munotes.in359

Chapter Sixty-Three

Exploratory Testing

Syllabus topic Module 2, "Experience-based: ... Exploratory testing"

In one line

Exploratory testing is testing in which the tester designs each test while running the last one, using what the software has just shown to decide what to try next; organised into time-boxed sessions with a written mission, notes and a debrief, it becomes as manageable and reportable as scripted testing, and often finds the most important problems first.

In the wording a student can write in an examination: in James Bach's definition, "Exploratory testing is simultaneous learning, test design, and test execution." In other words, "any testing to the extent that the tester actively controls the design of the tests as those tests are performed and uses information gained while testing to design new and better tests." ISO/IEC/IEEE 29119-2 calls it a "type of unscripted experience-based testing in which the tester spontaneously designs and executes tests" from existing knowledge, earlier exploration and heuristic "rules of thumb". The ISTQB syllabus: "tests are simultaneously designed, executed, and evaluated while the tester learns about the test object." It is often structured by session-based testing: a test charter states the mission; the session is a time box; notes are kept on a session sheet; and a debriefing follows.

Learning, designing and executing at once

In scripted testing the three activities happen in order and usually by different people: someone designs test cases from the specification, someone else runs them later, and nobody changes a test because of what an earlier one revealed. In exploratory testing they happen together, in one head, and each result changes the next test.

Bach explains it with a jigsaw puzzle. Every glance at a piece is a test: does this piece connect to that one? As the picture forms, the solver changes tactics, collecting the pieces of one colour, working on the border, stepping back when things get disorganised. Designing every move in advance, before seeing the pieces, would be absurd. In his words, "the puzzle changes the puzzling".

He also treats it as a matter of degree, not of kind: it "is found on a continuum between pure scripted testing, prescribed completely in advance in every detail, and pure exploratory testing where every test idea emerges in the moment of test execution." A tester following a script who notices something odd and investigates it is exploring. Bach calls the pure end, with no script or outside direction at all, freestyle exploratory testing.

What the tester brings

Exploratory testing puts the design of tests in the middle of their execution, so everything depends on the tester. Bach lists what distinguishes an excellent explorer from a novice:

  • Test design. "An exploratory tester is first and foremost a test designer." Anyone can stumble on a test; the explorer crafts tests "that systematically explore the product".
  • Careful observation. A scripted tester need only watch what the script says to watch; the explorer "must watch for anything unusual or mysterious", and must "distinguish obervation from inference" (his spelling).
  • Critical thinking: explaining one's own logic and looking for errors in it.
  • Diverse ideas, helped by heuristics, which Bach describes as "mental devices such as guidelines, generic checklists, mnemonics, or rules of thumb". The attack list of Chapter Sixty-Two, on error guessing, is one such device.
  • Rich resources: tools, information, test data, and colleagues to draw on.
munotes.in360

Exploratory Testing

The ISTQB syllabus agrees on the conditions for success: exploratory testing "will be more effective if the tester is experienced, has domain knowledge and has a high degree of essential skills, like analytical skills, curiosity and creativeness".

Managing it: charters, sessions and debriefs

Unscripted does not have to mean unmanaged. Jonathan Bach's session-based test management, introduced in 2000, gives exploratory testing a unit of work that can be planned, counted and reviewed: the session, "an uninterrupted block of reviewable, chartered test effort".

The charter. Each session has a mission. A charter "states the mission and perhaps some of the tactics to be used", and may be chosen by the tester or assigned by the test lead. Bach's example charters for a decision-analysis product include short ones such as "Check UI against Windows interface standards." For ExamReg, a charter might read: explore the fee payment page with interrupted journeys: back, refresh, a second tab.

The time box. In the 2000 article's team, "sessions last 90 minutes, give or take"; one closer to 45 minutes is a short session, and one closer to two hours a long one. "Uninterrupted" means "no significant interruptions, no email, meetings, chatting or telephone calls."

The session sheet. Each session produces a report, the session sheet, in a fixed tagged format so that a program can read it: the charter and the areas tested, the tester, the start time, a task breakdown, the bugs found, and the issues raised. The task breakdown divides the session into three kinds of work, "test design and execution, bug investigation and reporting, and session setup", with the tester's estimate of each, and records the share of time spent "on charter" against "on opportunity", where "Opportunity testing is any testing that doesn't fit the charter of the session." Bugs are "concerns about the quality of the product"; issues are "questions or problems that relate to the test process or the project at large".

The debrief. "Each session is debriefed." The test lead's "primary objective in the debriefing is to understand and accept the session report. Another objective is to provide feedback and coaching to the tester."

munotes.in361

Exploratory Testing

The ISTQB syllabus describes the same arrangement: exploratory testing "is performed within a defined time box. The tester uses a test charter containing test objectives to guide the testing. The test session is usually followed by a debriefing". It adds a point about coverage: in this approach "test objectives may be treated as high-level test conditions. Coverage items are identified and exercised during the test session."

Worked example: four sessions on ExamReg

Four exploratory sessions were run on ExamReg before release, two by each of two testers, each with its own charter. Their session sheets, in the article's tagged format:

CHARTER
Explore the fee payment page with interrupted journeys: back, refresh, a second tab.
#AREAS
Page | Fee payment
START
2026-10-12 10:00
TESTER
Neha
#DURATION
normal
#TEST DESIGN AND EXECUTION
55
#BUG INVESTIGATION AND REPORTING
35
#SESSION SETUP
10
#CHARTER VS. OPPORTUNITY
90/10
BUGS
#BUG 1
Refreshing the confirmation page sends a second payment request.
#BUG 2
Back from the gateway shows the fee as unpaid although the payment went through.
ISSUES
#ISSUE 1
What should the page show if the gateway has not answered after 60 seconds?
====
CHARTER
Explore the fee payment page on a slow mobile connection, as a student in a hostel would.
#AREAS
Page | Fee payment
Platform | Android
START
2026-10-12 14:00
TESTER
Arjun
#DURATION
short
#TEST DESIGN AND EXECUTION
70
#BUG INVESTIGATION AND REPORTING
20
#SESSION SETUP
10
#CHARTER VS. OPPORTUNITY
100/0
BUGS
#BUG 3
The Pay button can be pressed again while the first press is still loading.
ISSUES
====
CHARTER
Explore the hall ticket download after payment, including a payment made minutes before the deadline.
#AREAS
Page | Hall ticket
Page | Fee payment
START
2026-10-13 10:00
TESTER
Neha
#DURATION
long
#TEST DESIGN AND EXECUTION
60
#BUG INVESTIGATION AND REPORTING
15
#SESSION SETUP
25
#CHARTER VS. OPPORTUNITY
70/30
BUGS
#BUG 4
A hall ticket downloaded before the payment is confirmed shows no seat number.
ISSUES
#ISSUE 2
The test database had no student with a concession; setup took longer than planned.
====
CHARTER
Explore the exam form's paper selection with unusual combinations of regular and backlog papers.
#AREAS
Page | Exam form
START
2026-10-13 14:00
TESTER
Arjun
#DURATION
normal
#TEST DESIGN AND EXECUTION
80
#BUG INVESTIGATION AND REPORTING
10
#SESSION SETUP
10
#CHARTER VS. OPPORTUNITY
100/0
BUGS
ISSUES

The session-based team read their sheets with a tool that "breaks them down into their basic elements, normalizes them, and summarizes them into tables and metrics". The program below does the same for these four: it reads each sheet, turns the duration into minutes, and totals the task breakdown weighted by session length.

import re

MINUTES = {"short": 45, "normal": 90, "long": 120}      # Bach: about 90 minutes; short, long

def parse(sheet):
    """One session sheet as {section: [lines]}; each #BUG and #ISSUE heading is one entry."""
    sections, tag = {}, None
    for line in sheet.strip().splitlines():
        if re.fullmatch(r"#(BUG|ISSUE) \d+", line):            # "#BUG 3", not "#BUG INVESTIGATION"
            sections.setdefault(line.split()[0], []).append(line)
            tag = None                                   # the bug's text is not a section
        elif line.isupper() or line.startswith("#"):
            tag = line
            sections.setdefault(tag, [])
        elif tag:
            sections[tag].append(line)
    return sections

sheets = [parse(s) for s in open("sessions.txt").read().split("====")]
totals = {"test": 0, "bug": 0, "setup": 0, "charter": 0, "minutes": 0}
print(f"{'tester':<7}{'minutes':>8}{'test':>6}{'bug':>5}{'setup':>6}{'on charter':>11}{'bugs':>6}{'issues':>7}")
for s in sheets:
    minutes = MINUTES[s["#DURATION"][0]]
    test, bug, setup = (int(s[k][0]) for k in ("#TEST DESIGN AND EXECUTION",
                        "#BUG INVESTIGATION AND REPORTING", "#SESSION SETUP"))
    on = int(s["#CHARTER VS. OPPORTUNITY"][0].split("/")[0])
    bugs, issues = len(s.get("#BUG", [])), len(s.get("#ISSUE", []))
    print(f"{s['TESTER'][0]:<7}{minutes:>8}{test:>5}%{bug:>4}%{setup:>5}%{on:>10}%{bugs:>6}{issues:>7}")
    for key, share in (("test", test), ("bug", bug), ("setup", setup), ("charter", on)):
        totals[key] += minutes * share / 100
    totals["minutes"] += minutes

m = totals["minutes"]
print(f"{len(sheets)} sessions, {m} minutes: test design and execution {100 * totals['test'] / m:.0f}%,"
      f" bug investigation {100 * totals['bug'] / m:.0f}%, setup {100 * totals['setup'] / m:.0f}%,"
      f" on charter {100 * totals['charter'] / m:.0f}%")
areas = sorted({a.split("|")[1].strip() for s in sheets for a in s["#AREAS"]})
print("areas explored:", ", ".join(areas))
munotes.in362

Exploratory Testing

tester  minutes  test  bug setup on charter  bugs issues
Neha         90   55%  35%   10%        90%     2      1
Arjun        45   70%  20%   10%       100%     1      0
Neha        120   60%  15%   25%        70%     1      1
Arjun        90   80%  10%   10%       100%     0      0
4 sessions, 345 minutes: test design and execution 65%, bug investigation 20%, setup 15%, on charter 87%
areas explored: Android, Exam form, Fee payment, Hall ticket

Four sessions, 345 minutes of testing, and four bugs, two of them in the first session, which is typical of a charter aimed at a risky area: interrupted payment journeys are where a web application's state goes wrong. Neither bug was in any test case, and neither would have been: the specification says what happens when a payment succeeds, not what happens when the student presses Back in the middle of it.

The totals are what a test lead reads before the debriefs. About two-thirds of the time went on test design and execution, a fifth on investigating and reporting bugs, and the rest on setup. The third session is the one to ask about: 25 per cent setup and 30 per cent off charter, and its issue says why, because the test data had no student with a concession. That is a problem with the test environment, not the tester, and the debrief is where it gets fixed. The last session found nothing, which is also a result: the exam form's paper selection survived 90 minutes of determined attempts to break it, and the notes of what was tried are the evidence.

munotes.in363

Exploratory Testing

Where exploratory testing fits

The ISTQB syllabus names its natural territory: it "is useful when there are few or inadequate specifications or there is significant time pressure on the testing", and "to complement other more formal test techniques". It "can incorporate the use of other test techniques (e.g., equivalence partitioning)". Bach's own list adds, among others, providing rapid feedback on a new feature, learning a product quickly, diversifying testing after scripts have been run, and investigating a particular defect. He also notes that "most formal written test procedures were probably created through a process of some sort of exploratory testing."

Strengths and limits

Strengths. It finds the problems no one wrote a test for, because it follows the software instead of a plan. It adapts at once to what it finds, spending time where the problems are. It needs little preparation and gives fast feedback. And, run in sessions, it can be planned, counted and reported.

Limits. It depends on the tester's skill more than any other technique. It is hard to repeat exactly, since the next session will explore differently. Its coverage is judged by charters and notes, not by a count of items fixed in advance. And without the discipline of charters, notes and debriefs, it degenerates into the aimless clicking it is sometimes mistaken for.

What it does not mean

Exploratory testing is not ad hoc testing. It has a mission, a time box, notes and a review; freestyle exploring is one end of a continuum, not the whole of it.

It is not undocumented. In session-based test management every session produces a sheet that a third party can review.

It does not replace scripted testing. It complements it: scripted regression tests check what is known, exploration looks for what is not.

A session that finds no bugs is not wasted. It is evidence about an area, and its notes say what was tried.

Quick revision

  • Bach: "Exploratory testing is simultaneous learning, test design, and test execution"; a continuum from pure scripted to freestyle.
  • ISO/IEC/IEEE 29119-2: unscripted experience-based testing, tests designed and executed spontaneously, guided by heuristics.
  • Explorer's skills (Bach): test design, careful observation, critical thinking, diverse ideas (heuristics), rich resources.
  • Session-based test management (Jonathan Bach, 2000): a session is "an uninterrupted block of reviewable, chartered test effort", about 90 minutes; a charter gives the mission; a session sheet records charter, areas, tester, start, task breakdown (test design and execution, bug investigation and reporting, session setup; on charter against opportunity), bugs and issues; every session is debriefed.
  • Worked example: 4 sessions, 345 minutes, 4 bugs; 65 per cent test design and execution, 20 per cent bug investigation, 15 per cent setup, 87 per cent on charter.
  • Useful with few or inadequate specifications, under time pressure, and as a complement (ISTQB).
munotes.in364

Exploratory Testing

Test yourself

1. Define exploratory testing. How does it differ from scripted testing? Exploratory testing is simultaneous learning, test design and test execution: the tester designs each test while testing, using what earlier tests revealed. In scripted testing the tests are designed in advance, often by someone else, and executed as written, without the results changing the tests.

2. What is a test charter? Write one for ExamReg. A short statement of the mission of an exploratory session, what to test and what to look for, perhaps with tactics. For example: explore the fee payment page with interrupted journeys (back, refresh, a second tab) and report any payment made twice or lost.

3. Explain session-based test management. A way of managing exploratory testing in sessions: uninterrupted, time-boxed blocks of chartered testing of about 90 minutes. Each session produces a session sheet recording the charter, areas, tester, time, task breakdown, bugs and issues, and each is debriefed by the test lead, who accepts the report and coaches the tester. The sheets are totalled into metrics that show progress.

4. What is the task breakdown of a session sheet, and what can a test lead learn from it? The tester's estimate of the share of session time spent on test design and execution, on bug investigation and reporting, and on session setup, and of time on charter against on opportunity. It shows where testing time goes: much setup or much off-charter work points to problems such as missing test data, as in the worked example's third session.

5. When is exploratory testing most useful? When specifications are few or inadequate, when time is short, when fast feedback on a new feature is needed, when a product must be learnt quickly, after scripted tests to diversify the testing, and to investigate a particular defect or risk.

6. What skills does an exploratory tester need? Test design, careful observation that separates what is seen from what is inferred, critical thinking about one's own logic, a supply of diverse ideas helped by heuristics, and resources such as tools, data and colleagues; with domain knowledge, curiosity and creativity.

Contents This chapter on its own page

munotes.in365

Chapter Sixty-Four

Checklist-Based Testing

Syllabus topic Module 2, "Experience-based: ... Checklist-based testing"

In one line

Checklist-based testing tests a product against a list of questions drawn from experience of what matters and what goes wrong; it gives experience-based testing consistency and a record, and a checklist stays useful only if it is pruned, sharpened and extended as the defects it is meant to catch change.

In the wording a student can write in an examination: in the ISTQB syllabus's words, "In checklist-based testing, a tester designs, implements, and executes tests to cover test conditions from a checklist. Checklists can be built based on experience, knowledge about what is important for the user, or an understanding of why and how software fails." Items "are often phrased in the form of a question. It should be possible to check each item separately and directly." Checklists "should not contain items that can be checked automatically, items better suited as entry criteria, exit criteria, or items that are too general". They "should be regularly updated based on defect analysis", but "care should be taken to avoid letting the checklist become too long". The same idea applied to documents is checklist-based reviewing, a "review technique guided by a list of questions or required attributes" (ISO/IEC 20246).

Where checklists come from

A checklist is experience written down. The syllabus names three sources: experience, "knowledge about what is important for the user", and "an understanding of why and how software fails". In practice they come from:

  • Defect history. Every defect that escaped is a candidate question: if a double click once charged a student twice, does a double click act only once? belongs on the list.
  • Standards and models. The quality characteristics of ISO/IEC 25010 (Chapter Twenty, on the quality model today) make a checklist for non-functional testing: is it secure, is it compatible, can a first-time user interact with it? The syllabus notes that checklists support "various test types, including functional and non-functional testing", giving as an example usability heuristics.
  • The users. What matters to the people who use the product: for ExamReg, that a student on a phone can finish the form before the deadline.

The same list serves before the software runs. Chapter Twenty-Eight's review of ExamReg's fee requirement used a checklist; that is checklist-based reviewing, and a team often keeps one list for reviewing requirements and another for testing the running product.

What makes a good checklist item

The syllabus's guidance turns into four tests of an item.

  1. It is a question with a definite answer. Does the form keep what the student typed after an error? can be answered yes or no by trying it.
  2. It can be checked on its own, directly. It does not depend on another item's answer, or on a long investigation.
  3. It is not better done by a machine. Is the page's HTML valid? is a question a validator answers in a second, every build; putting it on a person's list wastes the person and delays the answer. It belongs in an automated check.
  4. It is not too general. Is the form easy to use? cannot be checked directly; it has to be broken into questions that can.
munotes.in366

Checklist-Based Testing

And the list must not be a disguise for something else: an item such as has the build passed its unit tests? is an entry criterion for testing, not a test.

A checklist for ExamReg's web forms

ExamReg's testers keep a checklist for every form on the portal, and they record, release by release, how many defects each item found. Before release 3.0 it held ten items:

IdQuestion
C1Does every required field say it is required before the form is sent?
C2Is each error shown beside its field, in words a first-year student understands?
C3Does the form keep what the student typed after an error?
C4Does a double click on Submit or Pay act only once?
C5Are dates shown and accepted in one format, DD-MM-YYYY?
C6Does the form accept names with an apostrophe, a dot or a hyphen?
C7Can the whole form be filled in with the keyboard alone?
C8Does every page load the analytics script?
C9Is the form easy to use?
C10Is the page's HTML valid?

C6 is new in release 3.0, added after the names Chapter Sixty-One's experience-based tests found refused.

Keeping a checklist alive

The syllabus explains why a checklist decays: "Some checklist entries may gradually become less effective over time because the developers will learn to avoid making the same errors. New entries may also need to be added to reflect newly found high severity defects. Therefore, checklists should be regularly updated based on defect analysis." And it warns against the opposite failure, a list that only grows.

Worked example: the release 3.0 review

The program applies the syllabus's guidance to ExamReg's checklist and its history. An automatable item moves to an automated check; a too-general item is rewritten; an item that has found nothing in the last three releases is retired; everything else stays. Two new questions come from the high-severity defects of release 3.0 that no item covered: the payment sent twice by Refresh, found in Chapter Sixty-Three's exploratory sessions, and the days late miscounted across the end of a month and a year, found by Chapter Sixty-Two's error guessing.

RELEASES = ["1.0", "2.0", "2.1", "3.0"]

# (id, question, defects the item found in each release (None: not yet on the list), note)
checklist = [
    ("C1", "Does every required field say it is required before the form is sent?", [3, 1, 0, 0], ""),
    ("C2", "Is each error shown beside its field, in words a first-year student understands?",
                                                                                     [2, 2, 1, 1], ""),
    ("C3", "Does the form keep what the student typed after an error?",              [1, 0, 0, 0], ""),
    ("C4", "Does a double click on Submit or Pay act only once?",                    [0, 1, 1, 2], ""),
    ("C5", "Are dates shown and accepted in one format, DD-MM-YYYY?",                [2, 0, 0, 0], ""),
    ("C6", "Does the form accept names with an apostrophe, a dot or a hyphen?",      [None, None, None, 3], ""),
    ("C7", "Can the whole form be filled in with the keyboard alone?",               [0, 1, 0, 1], ""),
    ("C8", "Does every page load the analytics script?",                             [0, 0, 1, 0], "automatable"),
    ("C9", "Is the form easy to use?",                                               [1, 0, 1, 0], "too general"),
    ("C10", "Is the page's HTML valid?",                                             [1, 1, 0, 0], "automatable"),
]
# high-severity defects of release 3.0 that no item covered: each becomes a new question
new_items = ["Does pressing Back or Refresh after paying never send a second payment?",
             "Are days late counted correctly across the end of a month and of a year?"]

def verdict(found, note, quiet=3):
    if note == "automatable":
        return "move to an automated check"
    if note == "too general":
        return "rewrite as specific questions"
    recent = [n for n in found[-quiet:] if n is not None]
    if len(recent) == quiet and sum(recent) == 0:
        return f"retire: nothing found in the last {quiet} releases"
    return "keep"

counts = {}
for cid, question, found, note in checklist:
    v = verdict(found, note)
    counts[v.split(":")[0]] = counts.get(v.split(":")[0], 0) + 1
    shown = " ".join("-" if n is None else str(n) for n in found)
    print(f"{cid:<4} {shown:<8} {v}")
for q in new_items:
    print(f"new  from a 3.0 high-severity defect: {q}")

kept = counts.get("keep", 0)
print(f"{len(checklist)} items reviewed: " + ", ".join(f"{n} {v}" for v, n in counts.items()))
print(f"the working checklist: {kept} kept + {len(new_items)} new = {kept + len(new_items)} items")
munotes.in367

Checklist-Based Testing

C1   3 1 0 0  keep
C2   2 2 1 1  keep
C3   1 0 0 0  retire: nothing found in the last 3 releases
C4   0 1 1 2  keep
C5   2 0 0 0  retire: nothing found in the last 3 releases
C6   - - - 3  keep
C7   0 1 0 1  keep
C8   0 0 1 0  move to an automated check
C9   1 0 1 0  rewrite as specific questions
C10  1 1 0 0  move to an automated check
new  from a 3.0 high-severity defect: Does pressing Back or Refresh after paying never send a second payment?
new  from a 3.0 high-severity defect: Are days late counted correctly across the end of a month and of a year?
10 items reviewed: 5 keep, 2 retire, 2 move to an automated check, 1 rewrite as specific questions
the working checklist: 5 kept + 2 new = 7 items
munotes.in368

Checklist-Based Testing

Read each verdict against its history.

  • C3 and C5 are retired. Each found defects early and nothing in the last three releases: the developers learnt to keep a form's contents after an error and to use one date format, exactly the decay the syllabus describes. Retiring them is not saying the questions were wrong; it is making room. If a defect of either kind reappears, the item comes back.
  • C1 stays, although it found nothing in the last two releases, because the rule looks back three releases and C1 found a defect in 2.0. Where to draw that line is the team's judgement; the program only applies it consistently.
  • C4 stays and matters more each release: it found 0, 1, 1 and then 2 defects. A double-click defect costs a student money, so it would stay even if its count fell.
  • C8 and C10 move to automated checks. Whether every page loads its analytics script, and whether the HTML is valid, are questions a program answers on every build; a person checking them by hand, a few times a release, is both slower and less reliable.
  • C9 is rewritten. Is the form easy to use? found defects, but a tester cannot check it directly; it becomes several specific questions (can a first-time user find the Pay button without help? does every button say what it does?).
  • Two new items, each from a high-severity defect that no item would have caught.

The working checklist ends with 7 questions (plus whatever specific questions replace C9): shorter than it began, and aimed at the failures that are happening now.

High-level or detailed

A checklist can say check the error messages or it can say is each error shown beside its field, in words a first-year student understands?. The syllabus names the trade-off: "If the checklists are high-level, some variability in the actual testing is likely to occur, resulting in potentially greater coverage but less repeatability." A detailed item is tested the same way by every tester; a high-level one is interpreted, and different testers look at different things. And where detailed test cases do not exist, "checklist-based testing can provide guidelines and some degree of consistency for the testing."

Strengths and limits

Strengths. A checklist carries experience from one tester to the next, so a new tester starts with the team's knowledge. It gives experience-based testing consistency and a record: each item was checked, with a result. It is quick to apply and easy to review. And it sits comfortably with the other techniques: an item such as C6 can be tested with equivalence partitions of names.

munotes.in369

Checklist-Based Testing

Limits. It finds only what it asks about. It decays if nobody maintains it, and it bloats if nobody prunes it. A high-level item gives variable testing; a detailed one gives narrow testing. And ticking an item is not the same as testing it well: yes to C7 after pressing Tab twice is not a check of the whole form.

What it does not mean

A checklist is not a test case. An item names a test condition; the tester still decides how to check it, with what data, and what counts as a pass.

A longer checklist is not a better one. Items that no longer find defects, that a machine could check, or that cannot be checked directly make the list slower to use and easier to skim.

Retiring an item is not forgetting it. The defect records keep it, and it returns if its kind of defect does.

Checklists are not only for testing software that runs. The same idea guides reviews of requirements, designs and code.

Quick revision

  • Checklist-based testing (ISTQB): tests designed, implemented and executed to cover test conditions from a checklist.
  • Sources: experience, what matters to users, why and how software fails; defect history, standards and models, users.
  • Good items: questions, each checkable separately and directly; not automatable, not entry or exit criteria, not too general.
  • For functional and non-functional testing alike; checklist-based reviewing (ISO/IEC 20246) applies it to documents.
  • Keep it alive: entries lose effectiveness as developers learn; add entries for new high-severity defects; do not let it grow too long.
  • Worked example: 10 items reviewed: 5 kept, 2 retired (nothing in 3 releases), 2 automated, 1 rewritten; 2 added; 7 working items.
  • High-level checklists: more coverage, less repeatability.

Test yourself

1. What is checklist-based testing? Where do checklists come from? A technique in which the tester designs and executes tests to cover the test conditions on a checklist. Checklists are built from experience, from knowledge of what matters to users, and from an understanding of why and how software fails: in practice, from defect history, standards and quality models, and the users' priorities.

2. What should a checklist not contain, and why? Items that can be checked automatically, because a program checks them faster and on every build; items that are really entry or exit criteria, because they are conditions for testing, not tests; and items that are too general, because they cannot be checked directly.

3. Write five checklist items for testing a college's online examination form. Does every required field say it is required before the form is sent? Is each error shown beside its field in plain words? Does the form keep what the student typed after an error? Does a double click on Submit or Pay act only once? Does the form accept names with an apostrophe, a dot or a hyphen?

munotes.in370

Checklist-Based Testing

4. Why must a checklist be updated, and how? Because items lose their effectiveness as developers learn to avoid the errors they target, while new kinds of high-severity defects appear. It is updated from defect analysis: items that have stopped finding defects are retired or reviewed, new items are added for escaped high-severity defects, and the list is kept short.

5. What is the trade-off between high-level and detailed checklists? High-level items leave the tester to interpret them, so testing varies between testers, potentially covering more but with less repeatability; detailed items are tested the same way each time, more repeatably but more narrowly.

6. Compare checklist-based testing with exploratory testing. Both are experience-based. A checklist fixes in advance which conditions are checked, giving consistency and a record of each item; exploratory testing decides what to test during the session, guided by a charter and by what each test reveals, giving flexibility and discovery at the cost of repeatability.

Contents This chapter on its own page

munotes.in371

Chapter Sixty-Five

What a Software Metric Is, and Why Measure

Syllabus topic Module 2, "Software Metrics: Concept of software metrics and their importance"

In one line

A software metric turns an attribute of a product, a process or a project into a value on a defined scale, so that questions about cost, progress and quality can be answered with evidence instead of impressions; the value means something only together with how it was measured, on what scale, and for what purpose.

In the wording a student can write in an examination: a measure is a "variable to which a value is assigned as the result of measurement", and measurement is a "set of operations having the object of determining a value of a measure" (ISO/IEC/IEEE 15939, ISO/IEC 25000). A metric is a "defined measurement method and the measurement scale" (ISO/IEC 14102). A base measure is "defined in terms of an attribute and the method for quantifying it"; a derived measure is "defined as a function of two or more values of base measures"; an indicator is a "measure that provides an estimate or evaluation of specified attributes derived from a model with respect to defined information needs" (ISO/IEC/IEEE 15939). Metrics are taken of products, processes and resources (Basili, Caldiera and Rombach), and their values lie on nominal, ordinal, interval or ratio scales (ISO/IEC/IEEE 24765). Measurement matters because, in the GQM paper's words, it "is a mechanism for creating a corporate memory and an aid in answering a variety of questions associated with the enactment of any software process."

From an attribute to a decision

The vocabulary is precise because each word names one step in turning an observation into a decision. ISO/IEC/IEEE 15939, the standard for the measurement process, gives the chain:

TermDefinition (ISO/IEC/IEEE 15939)ExamReg example
Entity"object that is to be characterized by measuring its attributes"Release 2.0 of ExamReg
Attribute"inherent property or characteristic of an entity that can be distinguished quantitatively or qualitatively by human or automated means" (ISO/IEC 25000)Its size; its defects
Measurement method"logical sequence of operations, described generically, used in quantifying an attribute with respect to a specified scale"Count lines of code by a stated rule; count defects in the tracker
Base measure"measure defined in terms of an attribute and the method for quantifying it"20 KLOC; 200 defects
Measurement function"algorithm or calculation performed to combine two or more base measures"Defects divided by KLOC
Derived measure"measure that is defined as a function of two or more values of base measures"10 defects per KLOC
Indicator"measure that provides an estimate or evaluation of specified attributes derived from a model with respect to defined information needs"Defect density against the college's target
Information need"insight necessary to manage objectives, goals, risks, and problems"Is release 2.0 good enough to go live?
Decision criteria"thresholds, targets, or patterns used to determine the need for action or further investigation"Release only below a set density
munotes.in372

What a Software Metric Is, and Why Measure

Two older names map onto the first pair. A base measure is what many textbooks call a direct measure, taken straight from the entity, such as a count of defects; a derived measure is an indirect one, computed from others, such as defect density. CMMI gives examples of each: among base measures, "Estimates and actual measures of effort and cost (e.g., number of person hours)" and "Quality measures (e.g., number of defects by severity)"; among derived measures, "Defect density", "Test or verification coverage" and "Reliability measures (e.g., mean time to failure)".

What is measured: products, processes and resources

The Goal Question Metric paper, the subject of the next chapter, names three kinds of object a software organisation measures:

  • Products: "Artifacts, deliverables and documents that are produced during the system life cycle; E.g., specifications, designs, programs, test suites." Size, complexity and defect density are product metrics.
  • Processes: "Software related activities normally associated with time; E.g., specifying, designing, testing, interviewing." Review effectiveness, the time a defect waits to be fixed and defect removal efficiency are process metrics.
  • Resources: "Items used by processes in order to produce their outputs; E.g., personnel, hardware, software, office space." Staff hours and cost are resource metrics.

Measures kept to plan and track a single project, its effort, its cost against its budget, its schedule against its plan, are often grouped as project metrics; CMMI's examples include "Earned value" and "Schedule performance index". The same data can serve more than one purpose: the defects found in release 2.0 describe the product, and the stage at which each was found describes the process.

The paper adds a second distinction. Data are objective "If they depend only on the object that is being measured and not on the viewpoint from which they are taken; E.g., number of versions of a document, staff hours spent on a task, size of a program", and subjective "If they depend on both the object that is being measured and the viewpoint from which they are taken; E.g., readability of a text, level of user satisfaction."

Scales, and what they permit

A scale is an "ordered set of values, continuous or discrete, or a set of categories to which the attribute is mapped" (ISO/IEC/IEEE 15939). ISO/IEC/IEEE 24765 defines four, and they decide which calculations mean anything:

ScaleDefinition (ISO/IEC/IEEE 24765)ExamReg exampleMeaningful
Nominal"measurement values are categorical"Defect type: input validation, logic, interfaceCounts, the most frequent (mode)
Ordinal"measurement values are rankings"Severity: cosmetic, minor, major, criticalOrder, the middle value (median)
Interval"equal distances corresponding to equal quantities of the attribute"The calendar date a defect was foundDifferences, averages
Ratioequal distances, and "the value of zero corresponds to none of the attribute"Lines of code, number of defects, hoursAll of the above, and ratios: twice as many
munotes.in373

What a Software Metric Is, and Why Measure

The table is not pedantry. Severity is recorded in ExamReg's tracker as a ranking, and a ranking says that major is worse than minor, not by how much. Turning the rankings into numbers and averaging them is a very common mistake, and the program below shows what it costs.

Worked example: an average that changes its mind

Two releases of ExamReg each had 200 defects, graded by severity. The program computes an average severity for each, first coding the levels 1, 2, 3 and 4, then 1, 2, 5 and 6. Both codings keep the order, so for an ordinal scale both are equally valid. It then computes the medians, and finally the statistics that suit release 2.0's other measures.

from statistics import median_low

LEVELS = ["cosmetic", "minor", "major", "critical"]               # an ORDINAL scale: order only
releases = {"release 1.0": {"cosmetic": 68, "minor": 60, "major": 70, "critical": 2},
            "release 2.0": {"cosmetic": 48, "minor": 98, "major": 46, "critical": 8}}

def values(counts):
    """One entry per defect, as its position on the scale: 0 = cosmetic ... 3 = critical."""
    return [i for i, level in enumerate(LEVELS) for _ in range(counts[level])]

# two codings that both respect the order, so both are equally "correct" for an ordinal scale
codings = {"codes 1, 2, 3, 4": [1, 2, 3, 4], "codes 1, 2, 5, 6": [1, 2, 5, 6]}
for name, code in codings.items():
    means = {r: sum(code[v] for v in values(c)) / sum(c.values()) for r, c in releases.items()}
    worse = max(means, key=means.get)
    print(f"mean severity with {name}: " + ", ".join(f"{r} {m:.2f}" for r, m in means.items())
          + f"  -> {worse} looks worse")
for r, c in releases.items():
    print(f"median severity of {r}: {LEVELS[median_low(values(c))]}")

# release 2.0 on the other scales (FINDINGS 5.2): type is NOMINAL, size and counts are RATIO
types = {"input validation": 58, "logic and computation": 44, "interface": 30, "user interface": 24,
         "data and database": 18, "documentation": 12, "performance": 8, "security": 6}
print("most frequent defect type (the mode, a nominal statistic):", max(types, key=types.get))
print(f"defect density (a ratio of ratio measures): {200 / 20:.1f} defects per KLOC")
mean severity with codes 1, 2, 3, 4: release 1.0 2.03, release 2.0 2.07  -> release 2.0 looks worse
mean severity with codes 1, 2, 5, 6: release 1.0 2.75, release 2.0 2.61  -> release 1.0 looks worse
median severity of release 1.0: minor
median severity of release 2.0: minor
most frequent defect type (the mode, a nominal statistic): input validation
defect density (a ratio of ratio measures): 10.0 defects per KLOC
munotes.in374

What a Software Metric Is, and Why Measure

With one coding, release 2.0 has the higher average severity; with the other, release 1.0 does. Nothing about the defects changed, only the arbitrary numbers written against the words, so the average severity was never measuring the defects at all. The medians are the same under both codings, minor for each release, because a median uses only the order, which is all an ordinal scale has. To compare the releases' severity honestly, a team compares the counts at each level (release 1.0 has 72 major or critical defects against release 2.0's 54), or the median.

The last two lines use statistics that fit their scales: the most frequent type for a nominal measure, and a ratio for ratio measures, where 10 defects per KLOC and twice as many defects are both meaningful.

Why measure: the importance of metrics

The GQM paper gives the uses of measurement, each with a question it answers: "It helps support project planning (e.g., How much will a new project cost?); it allows us to determine the strengths and weaknesses of the current processes and products (e.g., What is the frequency of certain types of errors?); it provides a rationale for adopting/refining techniques (e.g., What is the impact of the technique XX on the productivity of the projects?); it allows us to evaluate the quality of specific processes and products (e.g., What is the defect density in a specific system after deployment?)." And during a project, measurement helps "to assess its progress, to take corrective action based on this assessment, and to evaluate the impact of such action."

For a tester the importance is concrete. Without measurement, is ExamReg ready to release? is answered by whoever argues loudest. With it, the answer rests on evidence: the defects still open by severity, the coverage reached, the rate at which new failures are appearing, each compared with criteria agreed before the argument started. The chapters that follow build those measures: size (Chapter Sixty-Seven, on lines of code and function points), quality and test metrics (Chapter Sixty-Eight), complexity (Chapters Seventy to Seventy-Two), and defects (Chapter Seventy-Nine, on defect metrics).

The same paper also warns what measurement cannot do on its own. To be effective it must be "Focused on specific goals", and "A bottom-up approach will not work because there are many observable characteristics in software (e.g., time, number of defects, complexity, lines of code, severity of failures, effort, productivity, defect density), but which metrics one uses and how one interprets them it is not clear without the appropriate models and goals to define the context." Collecting everything that can be counted, and hoping meaning appears, is how organisations drown in numbers. The next chapter's method starts from the goal instead.

munotes.in375

What a Software Metric Is, and Why Measure

What it does not mean

A metric is not just a number. It is a number with a defined method and scale; the same 200 defects means different things under different counting rules.

Measuring is not the same as improving. A measure is useful when it answers a question someone needs answered and leads to a decision.

Numbers on an ordinal scale are not quantities. Codes for severity or priority can be ranked, not averaged.

More metrics are not better metrics. Measurement must be focused on goals, or it produces data nobody can interpret.

Quick revision

  • Measure: a variable assigned a value by measurement; measurement: the operations that determine it (ISO/IEC/IEEE 15939). Metric: "defined measurement method and the measurement scale" (ISO/IEC 14102).
  • Chain (ISO/IEC/IEEE 15939): entity, attribute, measurement method, base measure (direct), measurement function, derived measure (indirect), indicator, information need, decision criteria.
  • Objects (GQM): products, processes, resources; project metrics track effort, cost and schedule. Data are objective or subjective.
  • Scales (ISO/IEC/IEEE 24765): nominal (categories), ordinal (rankings), interval (equal distances), ratio (equal distances and a true zero).
  • Worked example: the mean of severity codes reversed between two equally valid codings (2.03 against 2.07, then 2.75 against 2.61); the median, minor, did not.
  • Why measure (GQM): planning, strengths and weaknesses, choosing techniques, evaluating quality, tracking progress; measurement must be focused on goals.

Test yourself

1. Distinguish a measure, a measurement and a metric. A measure is a variable to which a value is assigned by measurement; measurement is the set of operations that determine that value; a metric is a defined measurement method together with its scale, so that the measurement is repeatable and its value interpretable.

2. Distinguish base and derived measures, with examples. A base measure is defined by an attribute and a method of quantifying it, taken directly, such as the number of defects found (200) or the size in KLOC (20). A derived measure is a function of two or more base measures, such as defect density, defects per KLOC (10).

3. What are product, process and resource metrics? Give an example of each. Product metrics measure artefacts such as specifications, programs and test suites (defect density); process metrics measure activities such as reviewing and testing (defect removal efficiency); resource metrics measure what processes use, such as people and hardware (staff hours). Project metrics track a project's effort, cost and schedule.

4. Name the four measurement scales and the statistics each permits. Nominal, categories only: counts and the mode. Ordinal, rankings: also the median. Interval, equal distances: also differences and the mean. Ratio, equal distances with a true zero: also ratios such as twice as many.

munotes.in376

What a Software Metric Is, and Why Measure

5. Why is the average severity of a set of defects not a meaningful metric? Because severity is ordinal: its codes record only order, and any order-preserving coding is equally valid, yet different codings give different averages and can reverse a comparison, as in the worked example. The median, or the count at each level, uses only the order and does not change.

6. Why do software organisations measure? To plan and estimate, to find the strengths and weaknesses of their processes and products, to decide which techniques to adopt, to evaluate the quality of products and processes, and to track a project's progress and the effect of corrective actions, all with evidence instead of impressions, provided the measures are focused on specific goals.

Contents This chapter on its own page

munotes.in377

Chapter Sixty-Six

Developing Metrics: Goal, Question, Metric

Syllabus topic Module 2, "Software Metrics: ... Developing and utilizing different types of metrics"

In one line

The goal, question, metric approach develops metrics from the top down: state a goal precisely, ask the questions that would show whether it is being met, and only then choose the measures that answer those questions; every number collected is tied to a question someone needs answered.

In the wording a student can write in an examination: the Goal Question Metric (GQM) approach, from Basili, Caldiera and Rombach, "is based upon the assumption that for an organization to measure in a purposeful way it must first specify the goals for itself and its projects, then it must trace those goals to the data that are intended to define those goals operationally, and finally provide a framework for interpreting the data with respect to the stated goals." Its model has three levels: the conceptual level (goal), in which "A goal is defined for an object, for a variety of reasons, with respect to various models of quality, from various points of view, relative to a particular environment"; the operational level (question), in which "A set of questions is used to characterize the way the assessment/achievement of a specific goal is going to be performed"; and the quantitative level (metric), in which "A set of data is associated with every question in order to answer it in a quantitative way." A goal has a purpose, an issue, an object and a viewpoint.

Why start from the goal

Chapter Sixty-Five, on what a software metric is, ended on the paper's warning that measurement must be "Focused on specific goals", and that "A bottom-up approach will not work". Software offers endless things to count, and a count with no question behind it cannot be interpreted. The paper's conclusion is the method's first rule: "measurement must be defined in a top-down fashion. It must be focused, based on goals and models."

The goal: purpose, issue, object, viewpoint

"A GQM model is a hierarchical structure (Figure 1) starting with a goal (specifying purpose of measurement, object to be measured, issue to be measured, and viewpoint from which the measure is taken)." The paper's own example states a goal in all four parts:

CoordinateThe paper's example
PurposeImprove
Issuethe timeliness of
Object (process)change request processing
Viewpointfrom the project manager's viewpoint

Read as a sentence: improve the timeliness of change request processing from the project manager's viewpoint. Each part does work. The purpose says what the measurement is for (to improve, to evaluate, to predict); the issue says which quality matters (timeliness, reliability, effectiveness); the object says what is measured, a product, a process or a resource; and the viewpoint says whose question it is, because a manager, a tester and a user want different things from the same object.

munotes.in378

Developing Metrics: Goal, Question, Metric

The paper names where each coordinate comes from. The issue and purpose come from "the policy and the strategy of the organization"; the object from "the description of the process and products of the organization"; and the viewpoint from "the model of the organization", with a check that the goals chosen are relevant to the people whose viewpoint they take.

The questions: three groups

"From the specification of each goal we can derive meaningful questions that characterize that goal in a quantifiable way. In general, we will ask at least three groups of questions":

  1. "How can we characterize the object (product, process, or resource) with respect to the overall goal of the specific GQM model?"
  2. "How can we characterize the attributes of the object that are relevant with respect to the issue of the specific GQM model?"
  3. "How do we evaluate the characteristics of the object that are relevant with respect to the issue of the specific GQM model?"

The first group describes where things stand, the second breaks the issue into its parts, and the third judges the result against what is wanted.

The metrics, and how they are chosen

Each question is refined into metrics, "some of them objective such as the one in the example, some of them subjective. The same metric can be used in order to answer different questions under the same goal." In choosing them, the paper weighs three factors:

  • "Amount and quality of the existing data: we will try to maximize the use of existing data sources if they are available and reliable";
  • "Maturity of the objects of measurement: we will apply objective measures to more mature measurement objects, and we will use more subjective evaluations when we deal with informal or unstable objects";
  • "Learning process: GQM models need always refinement and adaptation, therefore the measures we define must help us in evaluating not only the object of measurement but also the reliability of the model used to evaluate it."

In the paper's example, the question what is the current change request processing speed? is answered by the average cycle time, its standard deviation and the percentage of cases outside an upper limit; and is the performance of the process improving? by the current average cycle time as a percentage of a baseline average, and by a subjective rating of the manager's satisfaction.

Worked example: a GQM model for ExamReg's testing

ExamReg's test manager wants to know whether the testing before release 2.0 did its job. As a GQM goal: evaluate the effectiveness of ExamReg's defect detection before release from the test manager's viewpoint. One question from each group follows, and each metric is chosen from data that already exists, the defect tracker's record of where each of release 2.0's 200 defects was found, as the paper advises. The college has set a target of finding at least 90 per cent of defects before release; ISO/IEC/IEEE 15939 would call it a decision criterion, one of the "thresholds, targets, or patterns used to determine the need for action or further investigation".

munotes.in379

Developing Metrics: Goal, Question, Metric

# ExamReg release 2.0: where each of its 200 defects was found (FINDINGS 5.2)
found_by = {"requirements review": 16, "design review": 24, "code review": 34,
            "unit testing": 44, "integration testing": 32, "system testing": 28,
            "acceptance testing": 10, "after release": 12}
TARGET = 90                                   # decision criterion: % found before release

total = sum(found_by.values())
before = total - found_by["after release"]
reviews = sum(n for a, n in found_by.items() if a.endswith("review"))
tests = before - reviews
dre = 100 * before / total

goal = {"purpose": "Evaluate", "issue": "the effectiveness of",
        "object (process)": "ExamReg's defect detection before release",
        "viewpoint": "from the test manager's viewpoint"}
model = [
    ("Q1 What share of all the defects is found before release?",
     [("defects found before release", f"{before} of {total}"),
      ("defect removal efficiency", f"{dre:.1f}%")]),
    ("Q2 Which activities find the defects?",
     [("found by reviews", f"{reviews} ({100 * reviews / total:.0f}%)"),
      ("found by tests", f"{tests} ({100 * tests / total:.0f}%)"),
      ("the activity that finds most", max(found_by, key=found_by.get))]),
    ("Q3 Is the effectiveness satisfactory?",
     [("defect removal efficiency against the target", f"{dre:.1f}% against {TARGET}%: "
       + ("met" if dre >= TARGET else "NOT met")),
      ("defects that reached students", f"{found_by['after release']}"),
      ("the test manager's rating", "subjective: collected at the release review")]),
]

print("GOAL")
for coordinate, value in goal.items():
    print(f"   {coordinate:<17} {value}")
for question, metrics in model:
    print(question)
    for name, value in metrics:
        print(f"   {name:<45} {value}")
GOAL
   purpose           Evaluate
   issue             the effectiveness of
   object (process)  ExamReg's defect detection before release
   viewpoint         from the test manager's viewpoint
Q1 What share of all the defects is found before release?
   defects found before release                  188 of 200
   defect removal efficiency                     94.0%
Q2 Which activities find the defects?
   found by reviews                              74 (37%)
   found by tests                                114 (57%)
   the activity that finds most                  unit testing
Q3 Is the effectiveness satisfactory?
   defect removal efficiency against the target  94.0% against 90%: met
   defects that reached students                 12
   the test manager's rating                     subjective: collected at the release review

The model reads from the top. The goal says what is wanted and for whom; each question narrows it; each metric answers a question with a value that can be checked.

  • Q1 characterises the object. Of 200 defects, 188 were found before release: a defect removal efficiency of 94 per cent, a derived measure that Chapter Seventy-Nine, on defect metrics, defines fully.
  • Q2 breaks the issue into parts. Reviews found 74 defects and tests 114; unit testing found more than any other single activity. That answers which activities and points to where improvement effort would pay.
  • Q3 evaluates. Against the college's criterion of 90 per cent, 94 is met. The same defect removal efficiency answers Q1 and Q3, exactly as the paper allows: one metric, two questions. The 12 defects that reached students are the other side of the same number. And the test manager's own judgement sits beside the figures as a subjective metric, collected rather than computed.
munotes.in380

Developing Metrics: Goal, Question, Metric

What the model does not contain is equally telling. Lines of code written per day, the number of test cases, hours spent: all are countable, none answers any of the three questions, so none is collected for this goal.

The GQM process

The paper sets out the order of work: identify "a set of quality and/or productivity goals"; "From those goals and based upon models of the object of measurement, we derive questions"; specify "the measures that need to be collected in order to answer those questions, and to track the conformance of products and processes to the goals"; and then "develop the data collection mechanisms, including validation and analysis mechanisms." After the data is collected and interpreted, the model itself is revisited, because "GQM models need always refinement and adaptation". For MU's phrase developing and utilizing different types of metrics, the first steps are the developing and the last are the utilizing.

What makes a metric useful

Pulling the paper's advice and the measurement standard's vocabulary together, a metric earns its place when it:

  1. Answers a question that serves a goal. If no question needs it, it is not collected.
  2. Has a defined measurement method and a known scale (Chapter Sixty-Five, on what a software metric is), so that two people measuring the same thing get the same value and the right statistics are used.
  3. Is objective where the object is mature enough, and openly subjective where it is not.
  4. Uses existing, reliable data where it can, so that collecting it costs little and interrupts no one.
  5. Can be interpreted: it comes with decision criteria, a target, a threshold or a baseline, that say what value calls for action.
  6. Helps refine the model, by showing whether the questions and metrics are the right ones.

This list is the chapter's own synthesis; each item is traced to the paper or the standard above.

What it does not mean

GQM does not prescribe metrics. It is a way of deriving them; two organisations with different goals get different metrics from it.

A goal is not a slogan. Improve quality has no object, issue or viewpoint; evaluate the effectiveness of defect detection before release from the test manager's viewpoint has all four.

munotes.in381

Developing Metrics: Goal, Question, Metric

Subjective metrics are not forbidden. The paper includes them, and uses them for objects that are informal or unstable.

The model is not fixed once built. Its questions and metrics are refined as the organisation learns what it needs to know.

Quick revision

  • GQM (Basili, Caldiera and Rombach 1994): specify goals, trace them to data, provide a framework for interpreting the data; top-down.
  • Three levels: conceptual (goal), operational (question), quantitative (metric).
  • A goal has a purpose, an issue, an object (product, process or resource) and a viewpoint; the paper's example: improve the timeliness of change request processing from the project manager's viewpoint.
  • Three groups of questions: characterise the object; characterise its attributes relevant to the issue; evaluate them.
  • Choosing metrics: existing data; objective for mature objects, subjective for unstable ones; the model must be refined.
  • Worked example: 188 of 200 defects found before release, defect removal efficiency 94 per cent against a 90 per cent target, met; reviews 74, tests 114; one metric answered two questions.

Test yourself

1. What is the Goal Question Metric approach? A top-down method of developing measurement: an organisation first specifies its goals, then derives the questions that characterise whether each goal is being achieved, then defines the metrics that answer each question quantitatively, and provides a framework for interpreting the data against the goals.

2. Describe the three levels of a GQM model. The conceptual level is the goal, defined for an object with a purpose, a quality issue and a viewpoint. The operational level is the set of questions that characterise how the goal's achievement will be assessed. The quantitative level is the set of metrics, objective or subjective, associated with each question to answer it.

3. State the four parts of a GQM goal with an example. Purpose, issue, object and viewpoint: for example, evaluate (purpose) the effectiveness of (issue) ExamReg's defect detection before release (object, a process) from the test manager's viewpoint (viewpoint).

4. Build a GQM model for evaluating the effectiveness of testing, with one question from each group. Goal: evaluate the effectiveness of defect detection before release from the test manager's viewpoint. Q1: what share of defects is found before release? Metric: defect removal efficiency. Q2: which activities find the defects? Metrics: defects found by reviews and by each test level. Q3: is the effectiveness satisfactory? Metrics: defect removal efficiency against a target, defects found after release, and the test manager's rating.

5. What factors does GQM consider in choosing metrics? The amount and quality of existing data, which should be used where reliable; the maturity of the object of measurement, with objective measures for mature objects and subjective evaluations for informal or unstable ones; and the learning process, since the model must be refined and the measures should help evaluate the model itself.

munotes.in382

Developing Metrics: Goal, Question, Metric

6. Why is a bottom-up approach to measurement said not to work? Because software has many measurable characteristics, and without goals and models it is not clear which metrics to use or how to interpret them; data collected without a question behind it cannot be turned into a decision.

Contents This chapter on its own page

munotes.in383

Chapter Sixty-Seven

Size Metrics: Lines of Code and Function Points

Syllabus topic Module 2, "Software Metrics: ... different types of metrics"

In one line

Size is the measure most other software measures are divided by, and there are two families of it: lines of code, which count what programmers wrote and mean nothing until the counting rule is stated, and function points, which count what the software does for its users and can be counted before any code exists.

In the wording a student can write in an examination: SLOC, "source lines of code, the number of lines of programming language code in a program before compilation" (ISO/IEC 20968), is counted as physical source lines or logical source statements (Park, SEI 1992), and only a written counting rule makes the number meaningful. A function point is a "unit of measure for functional size", and functional size is the "size of the software derived by quantifying the functional user requirements" (ISO/IEC 20926 and related standards). Function point analysis counts five kinds of component: external inputs (EI), external outputs (EO), external inquiries (EQ), internal logical files (ILF) and external interface files (EIF); rates each low, average or high from its data element types (DET), record element types (RET) or file types referenced (FTR); adds their weights to give the unadjusted function points (UFP); and multiplies by a value adjustment factor, VAF = (TDI × 0.01) + 0.65, where TDI totals the ratings of 14 general system characteristics.

Why size

Park's report opens on why size matters: "Size measures have direct application to the planning, tracking, and estimating of software projects. They are used also to compute productivities, to normalize quality indicators, and to derive measures for memory utilization and test coverage." Normalising is the use testers meet first. Release 2.0 of ExamReg had 200 defects. Whether that is many or few depends on how big release 2.0 is, and a defect density, defects per thousand lines or per function point, is only as trustworthy as the size it divides by.

Lines of code, and the counting problem

Counting lines sounds like the one measure nobody could get wrong. Park's report, written for the SEI's Software Metrics Definition Working Group, found otherwise: "reported values for software size are often confusing and easily misinterpreted. This usually happens because neither the conveyors nor the receivers of the information know what the measurements include or whether the measures have been applied with any consistency." Its example: "reports like 'Our software activity produced 163,000 source code instructions on that job' can easily be misunderstood by a factor of three or more."

The report separates two measures:

  • Physical source lines: lines of the source file, counted by a stated rule about which kinds of line count.
  • Logical source statements: the statements of the language, however many lines each occupies.
munotes.in384

Size Metrics: Lines of Code and Function Points

Its answer to the ambiguity is a definition checklist, on which an organisation ticks, attribute by attribute, what its count includes and excludes. Its basic definition of physical source lines includes executable lines, declarations and compiler directives, and excludes comments (on their own lines or beside code), banners, blank comments and blank lines; and "When a line or statement contains more than one type, classify it as the type with the highest precedence", so a line of code with a comment after it counts as code.

Worked example 1: one module, five sizes

Here is a small module of ExamReg, with a docstring, comments, blank lines and one statement spread over several lines:

"""ExamReg: fees for the examination form (release 2.0).

The fee rules are the exam cell's; the numbers are rupees.
"""

FORM_FEE = 800          # waived with a concession
BACKLOG_FEE = 150       # for each backlog paper

LATE_FEES = {
    "on time": 0,
    "1 to 7 days": 100,
    "8 to 15 days": 500,
}


def late_band(days_late):
    # which row of LATE_FEES applies
    if days_late == 0:
        return "on time"
    if days_late <= 7:
        return "1 to 7 days"
    return "8 to 15 days"


def total_fee(days_late, backlog_papers, concession):
    """Form fee (waived with a concession), backlog fees and the late fee."""
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    fee = (0 if concession else FORM_FEE) + BACKLOG_FEE * backlog_papers
    return fee + LATE_FEES[late_band(days_late)]

The program counts it by five rules, each of which someone could honestly call "lines of code".

import ast

source = open("examreg_fees.py").read()
lines = source.splitlines()
tree = ast.parse(source)

docstring_lines = set()                     # lines inside a module or function docstring
for node in [tree] + [n for n in ast.walk(tree) if isinstance(n, ast.FunctionDef)]:
    first = node.body[0]
    if isinstance(first, ast.Expr) and isinstance(first.value, ast.Constant):
        docstring_lines.update(range(first.lineno, first.end_lineno + 1))

every = set(range(1, len(lines) + 1))
nonblank = {n for n in every if lines[n - 1].strip()}
comment = {n for n in nonblank if lines[n - 1].strip().startswith("#")}
statements = [n for n in ast.walk(tree) if isinstance(n, ast.stmt)
              and not (isinstance(n, ast.Expr) and isinstance(n.value, ast.Constant))]

counts = {
    "every line in the file": len(every),
    "nonblank lines": len(nonblank),
    "nonblank, noncomment lines (Park's basic SLOC)": len(nonblank - comment),
    "the same, docstrings treated as comments": len(nonblank - comment - docstring_lines),
    "logical statements (Python statements)": len(statements),
}
for rule, n in counts.items():
    print(f"{rule:<48} {n:>3}")
print(f"largest count / smallest: {max(counts.values()) / min(counts.values()):.1f}")
every line in the file                            30
nonblank lines                                    23
nonblank, noncomment lines (Park's basic SLOC)    22
the same, docstrings treated as comments          18
logical statements (Python statements)            14
largest count / smallest: 2.1

One file of thirty lines has five sizes, from 30 down to 14, a spread of more than two to one on a module chosen to be ordinary. Each rule is defensible, and a report that says only 30 lines of code or 14 lines of code tells the reader nothing about which was used. The row for docstrings shows why the checklist must be filled in for each language: Park's checklist has a row for comments, but a Python docstring is a string the language executes, which the reader thinks of as a comment. Which way to count it is a decision, and it has to be written down.

munotes.in385

Size Metrics: Lines of Code and Function Points

Park's recommendation, after weighing the two measures, is to "adopt physical source lines of code (SLOC) as one of their first measures of software size". Among his reasons: "It is easier and cheaper to build automated counters for physical lines than it is for logical statements", "Rules for determining when logical source statements begin and end are complex and different for every source language", and "Most historical data is in terms of physical source lines."

Function points

Lines of code have two weaknesses no counting rule can cure: they depend on the language (the same fee rule is a different number of lines in Python, Java and COBOL), and they exist only after the code is written, which is too late for estimating the project. Function points measure something else: what the software does for its users. They were defined in 1979 by Allan J. Albrecht at IBM, and the method in use is the International Function Point Users Group's, standardised as ISO/IEC 20926:2009.

Function point analysis is a "method for measuring functional size" (ISO/IEC 20926). It identifies five kinds of component, defined in that standard:

ComponentDefinition (ISO/IEC 20926)ExamReg examples
External input (EI)"elementary process that processes data or control information sent from outside the boundary"Submit the exam form; pay the fee
External output (EO)"elementary process that sends data or control information outside the application's boundary and includes additional processing logic beyond that of an external inquiry"The hall ticket; the payment receipt
External inquiry (EQ)"elementary process that sends data or control information outside the boundary"View registration status
Internal logical file (ILF)"user-recognizable group of logically related data or control information maintained within the boundary of the application being measured"Student records; exam registrations; payments
External interface file (EIF)a group of data "which is referenced by the application being measured, but which is maintained within the boundary of another application"The college's admission records

An elementary process is the "smallest unit of activity that is meaningful to the user". Each component is then rated by how much data it handles. OMG's Automated Function Points standard gives the definitions, taken from ISO/IEC 20926: a data element type (DET) is "a unique user recognizable, non-repeated attribute that is part of an ILF, EIF, EI or EO"; a record element type (RET) is a "user recognizable sub-group of data element types within a data function"; and a file type referenced (FTR) is a "data function (IFL or EIF) read and/or maintained by a transactional function" (the standard's own typing of ILF).

munotes.in386

Size Metrics: Lines of Code and Function Points

Files are rated from their DETs and RETs, transactions from their DETs and FTRs, each rating low, average or high, and each rating carries a weight:

ComponentLowAverageHigh
Internal logical file71015
External interface file5710
External input346
External output457
External inquiry346

The first four rows are OMG's; the inquiry row is Garmus's, from IFPUG's practice. OMG's standard counts no inquiries separately, for a reason worth knowing: "since automated counting tools cannot distinguish between External Inquiries and External Outputs, all External Inquiries will be included in and counted as External Outputs."

The sum of the weights is the unadjusted function point count. The IFPUG method then adjusts it for the system's general characteristics. Garmus lists the fourteen general system characteristics, from data communication to facilitate change, and gives the rule: "Each characteristic is given a value", the values are summed to the total degree of influence (TDI), and the factor is (TDI * .01) + .65 = VAF, as his slide writes it. The adjusted function points are the unadjusted count "multiplied by the Value Adjustment Factor".

Worked example 2: counting ExamReg

A counter has identified ExamReg's functions and counted each one's DETs and its RETs or FTRs. The program rates them by OMG's rules, adds the two inquiries as the counter rated them (IFPUG's rules for inquiries are in its counting manual, which is not on this book's shelf), applies the adjustment, and uses the result to express release 2.0's defects per function point.

LEVEL = {1: "low", 2: "average", 3: "high"}
WEIGHT = {"ILF": (7, 10, 15), "EIF": (5, 7, 10), "EI": (3, 4, 6), "EO": (4, 5, 7), "EQ": (3, 4, 6)}

def band(n, low_max, mid_max):             # 1, 2 or 3 as OMG AFP 1.0 rescales a count
    return 1 if n <= low_max else 2 if n <= mid_max else 3

def complexity(kind, det, other):
    """OMG AFP 1.0: data functions from DETs and RETs, transactions from DETs and FTRs."""
    if kind in ("ILF", "EIF"):
        total = band(det, 19, 50) + band(other, 1, 5)
    elif kind == "EI":
        total = band(det, 4, 15) + band(other, 1, 2)
    else:                                  # EO
        total = band(det, 5, 19) + band(other, 1, 3)
    return 1 if total <= 3 else 2 if total == 4 else 3

# ExamReg's functions as a counter identified them: (kind, name, DETs, RETs or FTRs)
functions = [("ILF", "Student record", 18, 2), ("ILF", "Exam registration", 24, 3),
             ("ILF", "Payment", 9, 1),
             ("EIF", "College admission records", 22, 1), ("EIF", "University paper catalogue", 8, 1),
             ("EI", "Submit the exam form", 16, 3), ("EI", "Pay the fee", 6, 2),
             ("EI", "Update contact details", 5, 1),
             ("EO", "Hall ticket", 14, 3), ("EO", "Payment receipt", 8, 2),
             ("EO", "Registrations by paper report", 21, 3)]
inquiries = [("View registration status", 1), ("Search the paper catalogue", 1)]   # rated by the counter

ufp = 0
for kind, name, det, other in functions:
    level = complexity(kind, det, other)
    ufp += WEIGHT[kind][level - 1]
    print(f"{kind:<4}{name:<32} DET {det:>2}  {'RET' if kind in ('ILF', 'EIF') else 'FTR'} {other}"
          f"  {LEVEL[level]:<8} {WEIGHT[kind][level - 1]:>2}")
for name, level in inquiries:
    ufp += WEIGHT["EQ"][level - 1]
    print(f"EQ  {name:<32} (rated by the counter)  {LEVEL[level]:<8} {WEIGHT['EQ'][level - 1]:>2}")

gsc = {"Data communication": 4, "Distributed data or processing": 1, "Performance objectives": 3,
       "Heavily used configuration": 2, "Transaction rate": 4, "On-line data entry": 5,
       "End-user efficiency": 4, "On-line update": 4, "Complex processing": 2, "Reusability": 1,
       "Conversion and install ease": 1, "Operational ease": 3, "Multiple-site use": 2,
       "Facilitate change": 3}
tdi = sum(gsc.values())
vaf = tdi * 0.01 + 0.65                    # Garmus 2006: (TDI * .01) + .65 = VAF
print(f"unadjusted function points {ufp}; {len(gsc)} characteristics, TDI {tdi};"
      f" VAF {vaf:.2f}; adjusted function points {ufp * vaf:.1f}")
print(f"release 2.0's 200 defects: {200 / (ufp * vaf):.2f} per function point")
munotes.in387

Size Metrics: Lines of Code and Function Points

ILF Student record                   DET 18  RET 2  low       7
ILF Exam registration                DET 24  RET 3  average  10
ILF Payment                          DET  9  RET 1  low       7
EIF College admission records        DET 22  RET 1  low       5
EIF University paper catalogue       DET  8  RET 1  low       5
EI  Submit the exam form             DET 16  FTR 3  high      6
EI  Pay the fee                      DET  6  FTR 2  average   4
EI  Update contact details           DET  5  FTR 1  low       3
EO  Hall ticket                      DET 14  FTR 3  average   5
EO  Payment receipt                  DET  8  FTR 2  average   5
EO  Registrations by paper report    DET 21  FTR 3  high      7
EQ  View registration status         (rated by the counter)  low       3
EQ  Search the paper catalogue       (rated by the counter)  low       3
unadjusted function points 70; 14 characteristics, TDI 39; VAF 1.04; adjusted function points 72.8
release 2.0's 200 defects: 2.75 per function point

Follow two rows through the rules. The exam registration file has 24 DETs, which OMG's rule places in its middle band (20 to 50), and 3 RETs, also its middle band (2 to 5); the two bands add to 4, which is average, weight 10. Submitting the exam form has 16 DETs, above the external input's middle band of 5 to 15, and 3 FTRs, above its band of exactly 2; the bands add to 6, which is high, weight 6.

munotes.in388

Size Metrics: Lines of Code and Function Points

The thirteen weights add to 70 unadjusted function points. The fourteen characteristics were rated to a total degree of influence of 39, so VAF = 39 × 0.01 + 0.65 = 1.04, and ExamReg is 72.8 adjusted function points. Divided into release 2.0's 200 defects, that is 2.75 defects per function point, a density that can be compared with another system's whatever language each was written in.

Lines of code or function points

Lines of codeFunction points
MeasuresWhat the programmers wroteWhat the software does for its users
AvailableOnly once the code existsFrom the requirements, before design
Depends on the languageYesNo
CountingEasy to automate, once the rule is written downNeeds trained counters and judgement; OMG's standard automates a version of it
Main riskAn unstated counting ruleInconsistent counters; unsuited to some kinds of software

Park weighed function points in 1992 and saw their strengths, "they are language-independent, often solution-independent, and usually computable early in development life cycles, even before specific product designs are available", and three reasons for caution, one of them that "Automated function point counters do not yet exist." That reason has since been overtaken: OMG's Automated Function Points standard is a specification for counting function points from source code by tool. It illustrates this book's rule that a source's claims have a date.

What it does not mean

A line count is not a size until its rule is stated. The same module gave five honest counts from 14 to 30.

Function points do not measure effort or quality. They measure functional size; effort and defects are divided by them.

Adjusted is not always used. The adjustment is a step of the IFPUG method; OMG's automated count and many comparisons use unadjusted function points.

Neither measure is the size. Each measures one attribute; which one to use depends on the question, as the last chapter's goal, question, metric method would ask.

Quick revision

  • Size normalises: defects, effort and cost are divided by it (Park).
  • SLOC (ISO/IEC 20968); Park 1992: physical source lines and logical source statements; a definition checklist states what is included; basic definition excludes comments and blank lines, and a line is classified by its highest-precedence type.
  • Worked example 1: one 30-line module measured 30, 23, 22, 18 and 14 by five rules.
  • Function points: functional size from the user's view (Albrecht, IBM, 1979; ISO/IEC 20926); EI, EO, EQ, ILF, EIF; complexity from DET with RET (files) or FTR (transactions).
  • Weights: ILF 7, 10, 15; EIF 5, 7, 10; EI 3, 4, 6; EO 4, 5, 7; EQ 3, 4, 6.
  • VAF = (TDI × 0.01) + 0.65 over 14 general system characteristics; adjusted = unadjusted × VAF.
  • Worked example 2: ExamReg 70 unadjusted, TDI 39, VAF 1.04, 72.8 adjusted; 2.75 defects per function point.
munotes.in389

Size Metrics: Lines of Code and Function Points

Test yourself

1. Why is "lines of code" an ambiguous measure? How is the ambiguity removed? Because a line count depends on whether blank lines, comments, declarations and continuation lines are counted, and whether physical lines or logical statements are meant; reports of size can be misunderstood by a factor of three or more. It is removed by a written definition, such as the SEI's definition checklist, stating exactly what the count includes and excludes.

2. Distinguish physical source lines from logical source statements. Physical source lines are lines of the source file counted by a stated rule, for example nonblank, noncomment lines. Logical source statements are the language's statements, however many lines each occupies, so one statement written over five lines counts once.

3. Name and define the five components of function point analysis. External input: an elementary process that processes data or control information coming from outside the boundary. External output: one that sends data outside the boundary with processing beyond an inquiry. External inquiry: one that sends data outside the boundary without that extra processing. Internal logical file: a user-recognisable group of related data maintained inside the application. External interface file: such a group referenced by the application but maintained by another.

4. A system has 3 low ILFs, 1 average EIF, 4 average EIs, 2 high EOs and 3 low EQs, and its 14 characteristics total 42. Compute its adjusted function points. Unadjusted: 3 × 7 + 1 × 7 + 4 × 4 + 2 × 7 + 3 × 3 = 67. VAF = 42 × 0.01 + 0.65 = 1.07. Adjusted: 67 × 1.07 = 71.69.

5. Give two advantages of function points over lines of code, and one disadvantage. They do not depend on the programming language, and they can be counted from the requirements before any code exists, so they help estimation. But counting needs trained counters and judgement, and it suits some kinds of software, such as business applications, better than others.

6. What was Park's recommendation on size measures, and why? To start with physical source lines of code, because counters for them are easier and cheaper to build, the rules for logical statements are complex and differ for every language, physical counts are easier to interpret and compare, and most historical data is in physical lines.

Contents This chapter on its own page

munotes.in390

Chapter Sixty-Eight

Quality, Process and Test Metrics

Syllabus topic Module 2, "Software Metrics: ... utilizing different types of metrics"

In one line

Metrics about software fall into three groups: product metrics say how good the thing is, process metrics say how well it was made and tested, and test metrics say how far the testing has got and what it has found; each is useful only when it answers a question and is read alongside the others.

In the wording a student can write in an examination: in the ISTQB syllabus's words, "Test metrics are gathered to show progress against the planned test schedule and budget, the current quality of the test object, and the effectiveness of the test activities with respect to the test objectives or an iteration goal." The common kinds are project progress metrics, test progress metrics, product quality metrics, defect metrics, risk metrics, coverage metrics and cost metrics. Among product quality metrics, defect density is the "number of defects per unit of product size", and the mean time to repair (sometimes called the mean time to change) is "the mean time the maintenance team requires to implement a change and restore the system to working order" (ISO/IEC/IEEE 24765). Among process metrics, the defects each activity finds and the effort it takes show which activities work best.

Three questions, three groups

Chapter Sixty-Five, on what a software metric is, divided the objects of measurement into products, processes and resources. For a tester the same division becomes three questions:

  • How good is the product? Product quality metrics: defects in it, how it performs, how reliable it is, how quickly it can be changed.
  • How well was it made and tested? Process metrics: where defects were found, at what cost in effort, how many escaped.
  • Where has the testing got to? Test metrics: tests run and passed, coverage reached, defects found and fixed, risks remaining.

The groups overlap on purpose. Defects found after release describe the product (they are in it) and the process (it let them through).

The ISTQB list of test metrics

The syllabus gives seven kinds, each with examples:

KindThe syllabus's examplesWhere this book meets it
Project progress"task completion, resource usage, test effort"Effort by activity, below
Test progresstest cases implemented, environment readiness, "number of test cases run/not run, passed/failed, test execution time"Chapter Ten's report of 160 of 240 run, on test reporting
Product quality"availability, response time, mean time to failure"Chapter Forty-Five's response times under load testing; Chapter Ninety's reliability
Defect"number and priorities of defects found/fixed, defect density, defect detection percentage"Below, and Chapter Seventy-Nine, on defect metrics
Risk"residual risk level"The risks still open at release
Coverage"requirements coverage, code coverage"Chapter Forty-One's traceability, on validation testing; Chapters Fifty-Eight to Sixty on code coverage
Cost"cost of testing, organizational cost of quality"Chapter One Hundred One, on the cost of quality
munotes.in391

Quality, Process and Test Metrics

CMMI's list of commonly used derived measures overlaps it: "Defect density", "Peer review coverage", "Test or verification coverage" and "Reliability measures (e.g., mean time to failure)".

Product quality metrics

Defect density divides the defects by the size of the product, so that a large system and a small one can be compared. It depends entirely on both counts being defined: which defects (all found, or only those after release; which severities) and which size (lines by what rule, or function points). Delivered defect density counts only the defects that reached users, and is the one that describes what users got.

Mean time to repair, in its maintenance sense, measures how quickly the product can be changed: from a problem being reported to a working fix in production. It belongs to maintainability, a product characteristic, although the team's process shows in it too.

The other product quality metrics the syllabus names, availability, response time and mean time to failure, are measured by running the system: Chapter Forty-Five measured ExamReg's response times under load, and Chapter Ninety defines reliability and computes them.

Process metrics

A process metric says how an activity performed. For defect detection the basic ones are:

  • Defects found by each activity and the share of all defects that each found, which Chapter Sixty-Six used in its goal, question, metric model.
  • Effort spent in each activity, in person-hours.
  • Defects found per person-hour, the activity's efficiency at finding defects.
  • Defect removal efficiency, the share of all defects found before release, which Chapter Seventy-Nine defines fully.

Worked example: release 2.0's product and process metrics

The program computes release 2.0's product metrics, by both measures of size, and its process metrics by activity. The effort figures and change times come from the project's time records.

from statistics import mean, median

KLOC, FUNCTION_POINTS = 20, 72.8          # release 2.0's size: FINDINGS 5.2, and Chapter 67's count
found_by = {"requirements review": 16, "design review": 24, "code review": 34,
            "unit testing": 44, "integration testing": 32, "system testing": 28,
            "acceptance testing": 10, "after release": 12}            # FINDINGS 5.2
hours = {"requirements review": 12, "design review": 20, "code review": 30, "unit testing": 80,
         "integration testing": 64, "system testing": 112, "acceptance testing": 40}  # FINDINGS 5.2.2
change_hours = [4, 6, 3, 30, 8, 5, 72, 10, 6, 4, 12, 20]  # the 12 fixes after release (FINDINGS 5.2.2)

total, delivered = sum(found_by.values()), found_by["after release"]
print("PRODUCT")
print(f"   defect density            {total / KLOC:.1f} per KLOC, {total / FUNCTION_POINTS:.2f} per function point")
print(f"   delivered defect density  {delivered / KLOC:.2f} per KLOC, {delivered / FUNCTION_POINTS:.2f} per function point")
print(f"   mean time to change       {mean(change_hours):.1f} hours (median {median(change_hours):.1f},"
      f" longest {max(change_hours)})")

print("PROCESS")
for activity, h in hours.items():
    print(f"   {activity:<21} {found_by[activity]:>3} defects in {h:>3} hours:"
          f" {found_by[activity] / h:.2f} per hour")
for kind in ("review", "testing"):
    d = sum(n for a, n in found_by.items() if a.endswith(kind))
    h = sum(n for a, n in hours.items() if a.endswith(kind))
    print(f"   all {kind + 's' if kind == 'review' else kind:<17} {d:>3} defects in {h:>3} hours: {d / h:.2f} per hour")
munotes.in392

Quality, Process and Test Metrics

PRODUCT
   defect density            10.0 per KLOC, 2.75 per function point
   delivered defect density  0.60 per KLOC, 0.16 per function point
   mean time to change       15.0 hours (median 7.0, longest 72)
PROCESS
   requirements review    16 defects in  12 hours: 1.33 per hour
   design review          24 defects in  20 hours: 1.20 per hour
   code review            34 defects in  30 hours: 1.13 per hour
   unit testing           44 defects in  80 hours: 0.55 per hour
   integration testing    32 defects in  64 hours: 0.50 per hour
   system testing         28 defects in 112 hours: 0.25 per hour
   acceptance testing     10 defects in  40 hours: 0.25 per hour
   all reviews            74 defects in  62 hours: 1.19 per hour
   all testing           114 defects in 296 hours: 0.39 per hour

The product. Release 2.0 had 10 defects per KLOC over its whole life, and 0.6 per KLOC reached students; by function points, 2.75 and 0.16. The two sizes give different numbers for the same defects, which is why a density is never reported without its size measure. The mean time to change is 15 hours, but the median is 7: one change took 72 hours and pulls the mean up. Reporting both, and the longest, tells the reader that most changes are quick and one was not, which a mean alone hides.

The process. Each review found more than one defect per person-hour; unit and integration testing found about half a defect per hour; system and acceptance testing a quarter. Taken together, the reviews found 1.19 defects per hour against 0.39 for testing, about three times as many. That is a real finding about this project, and a reason to look at where review effort could be increased, which Chapters Ninety-Six to Ninety-Eight, on reviews, examine with published evidence.

It is also easy to misread. Later activities find fewer defects per hour partly because earlier ones have already removed the easy ones, and partly because they look for different kinds of defect: a system test finds failures of the whole portal that no review of one document could have seen. Defects per hour compares how the activities performed on this project; it does not say that any activity could be dropped.

Reading metrics well

A few habits keep metrics honest.

  • Pair a rate with its volume. 1.33 defects per hour sounds best of all, but the requirements review ran for 12 hours; the largest number of defects came from unit testing.
  • Report the distribution, not only the mean, when a few values are extreme, as the change times were.
  • Watch trends, not single values. A density this release means most beside the same density last release, measured the same way.
  • Compare against decision criteria agreed in advance, the thresholds and targets of ISO/IEC/IEEE 15939 (Chapter Sixty-Six's goal, question, metric example set 90 per cent of defects found before release).
munotes.in393

Quality, Process and Test Metrics

What it does not mean

A test metric is not a quality metric. The number of tests passed says how far testing has got; the product's quality is in its defects, performance and reliability.

A low defect density is not proof of quality. It may mean the product is good, or that testing found little; it needs the process metrics beside it.

Defects per hour does not rank activities for removal. Each activity finds kinds of defect the others miss.

A mean is not always the typical value. With a skewed distribution, the median says more.

Quick revision

  • Test metrics show "progress against the planned test schedule and budget, the current quality of the test object, and the effectiveness of the test activities" (ISTQB).
  • ISTQB's seven kinds: project progress, test progress, product quality, defect, risk, coverage, cost.
  • Defect density: "number of defects per unit of product size"; mean time to repair in the maintenance sense: the mean time to implement a change and restore working order (ISO/IEC/IEEE 24765).
  • Process metrics: defects by activity, effort by activity, defects per person-hour, defect removal efficiency.
  • Worked example: 10.0 defects per KLOC (2.75 per function point), 0.60 delivered per KLOC (0.16 per function point); mean time to change 15.0 hours, median 7.0; reviews 1.19 defects per hour against testing 0.39.
  • Read metrics with their volumes, distributions, trends and decision criteria.

Test yourself

1. Why are test metrics gathered? Name the kinds ISTQB lists. To show progress against the planned test schedule and budget, the current quality of the test object, and the effectiveness of the test activities against their objectives. The kinds are project progress, test progress, product quality, defect, risk, coverage and cost metrics.

2. Distinguish product, process and test metrics, with an example of each. Product metrics describe the software itself, such as defect density or response time. Process metrics describe how it was developed and tested, such as defects found per person-hour of review. Test metrics describe the state of the testing, such as test cases run against planned, or requirements coverage.

3. A system of 50 KLOC had 150 defects, 9 of them found after release. Compute its defect density and delivered defect density. Defect density: 150 defects in 50 KLOC, 3.0 per KLOC. Delivered defect density: 9 in 50 KLOC, 0.18 per KLOC.

munotes.in394

Quality, Process and Test Metrics

4. What is mean time to change, and why report the median with it? The mean time needed to implement a change and restore the system to working order, a measure of maintainability. The median is reported because a few long changes can pull the mean far above the typical case, as a single 72-hour change did in the worked example (mean 15.0 hours, median 7.0).

5. In the worked example, reviews found 1.19 defects per hour and testing 0.39. What can and cannot be concluded? It can be concluded that on this project reviewing found defects more efficiently per hour of effort, a reason to consider more review. It cannot be concluded that testing could be reduced or dropped, because later activities find kinds of defect reviews cannot see, and they find fewer per hour partly because earlier activities removed the easier ones.

6. Why should a defect density always be reported with its size measure? Because the same defects give different densities by different sizes (10.0 per KLOC and 2.75 per function point for the same release), and the density is comparable only with others measured the same way.

Contents This chapter on its own page

munotes.in395

Chapter Sixty-Nine

Object-Oriented Metrics: The CK Suite

Syllabus topic Module 2, "Software Metrics: ... different types of metrics"

In one line

The CK suite measures an object-oriented design class by class: how much each class does (WMC), how deep it sits in the inheritance tree (DIT), how many classes inherit from it (NOC), how many other classes it depends on (CBO), how many methods a message to it can set running (RFC), and how little its methods have in common (LCOM); each value points a reviewer and a tester at the classes that deserve the most attention.

In the wording a student can write in an examination: Chidamber and Kemerer's suite of six metrics for object-oriented design is: Weighted Methods per Class (WMC), the sum of the complexities of the methods defined in a class ("If all method complexities are considered to be unity, then WMC = n, the number of methods"); Depth of Inheritance Tree (DIT), the depth of the class in the inheritance tree; Number of Children (NOC), the "number of immediate sub-classes subordinated to a class in the class hierarchy"; Coupling Between Object classes (CBO), "a count of the number of other classes to which it is coupled"; Response For a Class (RFC), the size of the response set, "a set of methods that can potentially be executed in response to a message received by an object of that class"; and Lack of Cohesion in Methods (LCOM), the number of pairs of methods that share no instance variable minus the number of pairs that share one, or zero if that is negative.

Why object-oriented designs needed their own metrics

Lines of code, function points and McCabe's complexity measure functions and programs. An object-oriented design is made of classes, and its quality lies in things those measures cannot see: how responsibilities are divided among classes, how deep the inheritance runs, how tangled the classes are with one another. Chidamber and Kemerer set out to supply measures for these, observing that "The need for such metrics is particularly acute when an organization is adopting a new technology for which established practices have yet to be developed." They grounded the metrics in a theory (Bunge's ontology, following Wand and Weber), evaluated them against a set of measurement principles, and built "An automated data collection tool" to collect them "at two field sites".

The six metrics

WMC, Weighted Methods per Class. For a class with methods M1 to Mn of complexities c1 to cn, WMC is their sum. The paper deliberately leaves complexity open ("This is left as an implementation decision"), and with every method counted as 1, WMC is simply the number of methods. Its first viewpoint: "The number of methods and the complexity of methods involved is a predictor of how much time and effort is required to develop and maintain the class."

munotes.in396

Object-Oriented Metrics: The CK Suite

DIT, Depth of Inheritance Tree. How far the class is from the root of its hierarchy; with multiple inheritance, "the maximum length from the node to the root of the tree". "DIT is a measure of how many ancestor classes can potentially affect this class." A deep class inherits much, "making it more complex to predict its behavior", and deeper trees mean "greater design complexity"; but "the greater the potential reuse of inherited methods".

NOC, Number of Children. The immediate subclasses. More children mean more reuse, but also "the greater the likelihood of improper abstraction of the parent class", and for the tester: "If a class has a large number of children, it may require more testing of the methods in that class."

CBO, Coupling Between Object classes. "Two classes are coupled when methods declared in one class use methods or instance variables defined by the other class." The viewpoints are all warnings: "Excessive coupling between object classes is detrimental to modular design and prevents reuse", and "The higher the inter-object class coupling, the more rigorous the testing needs to be."

RFC, Response For a Class. The response set is the class's own methods together with every method they call: the set {M} of the class's methods joined with every {Ri}, where {Ri} is the set of methods called by method i, counted "only up to the first level of nesting of method calls". RFC is the size of that set. "If a large number of methods can be invoked in response to a message, the testing and debugging of the class becomes more complicated", and "A worst case value for possible responses will assist in appropriate allocation of testing time."

LCOM, Lack of Cohesion in Methods. For each method, take the set of instance variables it uses. Count the pairs of methods whose sets have nothing in common (P) and the pairs that share at least one (Q). LCOM is P minus Q if that is positive, and 0 otherwise. The paper's example: methods using {a, b, c, d, e}, {a, b, e} and {x, y, z} give two disjoint pairs and one sharing pair, so "LCOM is the (number of null intersections - number of non-empty intersections), which in this case is 1." A high LCOM is a design finding: "Lack of cohesion implies classes should probably be split into two or more sub-classes."

Worked example: measuring ExamReg's form classes

ExamReg's forms are classes: a base Form, an ExamForm for the examination form, a BacklogForm for students registering only backlog papers, and a ContactForm for updating a phone number, with two helpers, Catalogue and FeeRule.

munotes.in397

Object-Oriented Metrics: The CK Suite

class Form:
    def __init__(self, student):
        self.student = student
        self.errors = []

    def validate(self):
        if not self.student.roll_no:
            self.errors.append("no roll number")
        return not self.errors

    def summary(self):
        return f"{self.student.name}: {len(self.errors)} problem(s)"


class ExamForm(Form):
    def __init__(self, student, papers):
        super().__init__(student)
        self.papers = papers

    def validate(self):
        ok = super().validate()
        for code in self.papers:
            if not Catalogue.exists(code):
                self.errors.append(f"unknown paper {code}")
        return ok and not self.errors

    def backlog_count(self):
        return sum(1 for code in self.papers if code.startswith("B"))

    def fee(self, days_late):
        return FeeRule.total(days_late, self.backlog_count(), self.student.concession)

    def allot_seat(self, centre, seat):
        self.centre, self.seat = centre, seat

    def seat_label(self):
        return f"{self.centre}-{self.seat}"


class BacklogForm(ExamForm):
    def validate(self):
        return super().validate() and self.backlog_count() > 0


class ContactForm(Form):
    def __init__(self, student, phone):
        super().__init__(student)
        self.phone = phone

    def validate(self):
        if len(self.phone) != 10:
            self.errors.append("phone number must have 10 digits")
        return super().validate()


class Catalogue:
    PAPERS = {"USCS501", "USCS502", "BUSCS401"}

    @staticmethod
    def exists(code):
        return code in Catalogue.PAPERS


class FeeRule:
    @staticmethod
    def total(days_late, backlog_papers, concession):
        late = 0 if days_late == 0 else 100 if days_late <= 7 else 500
        return (0 if concession else 800) + 150 * backlog_papers + late

The paper defines its metrics for any object-oriented language, and applying them to Python needs a few counting rules, which the program states in its comments: DIT counts from ExamReg's own root class (every Python class ultimately inherits from object, which would add the same 1 to every class); CBO counts the classes named inside a class's methods, in either direction, and leaves inheritance to DIT and NOC; RFC counts calls written as self.m(), super().m() or Class.m(), not calls on lists and strings; and WMC counts each method as 1. The program checks its LCOM against the paper's own example before measuring anything.

import ast
from itertools import combinations

def lcom(ivar_sets):
    """Chidamber and Kemerer's LCOM: pairs of methods sharing no instance variable (P)
    minus pairs sharing one (Q), or 0 if that is negative."""
    if not any(ivar_sets):
        return 0
    p = sum(1 for a, b in combinations(ivar_sets, 2) if not a & b)
    q = sum(1 for a, b in combinations(ivar_sets, 2) if a & b)
    return max(p - q, 0)

print("the paper's own example, {a,b,c,d,e} {a,b,e} {x,y,z}: LCOM =",
      lcom([set("abcde"), set("abe"), set("xyz")]))

tree = ast.parse(open("examreg_forms.py").read())
classes = {c.name: c for c in tree.body if isinstance(c, ast.ClassDef)}
parent = {name: next((b.id for b in c.bases if b.id in classes), None) for name, c in classes.items()}
methods = {name: [f for f in c.body if isinstance(f, ast.FunctionDef)] for name, c in classes.items()}

def ancestry(name):                          # the class, then its parent, then its parent ...
    while name:
        yield name
        name = parent[name]

def is_self_attribute(node):
    return isinstance(node, ast.Attribute) and isinstance(node.value, ast.Name) and node.value.id == "self"

def instance_variables(name):
    """Every self.x assigned in the class or in one of its ancestors."""
    found = set()
    for c in ancestry(name):
        for n in ast.walk(classes[c]):
            for target in getattr(n, "targets", []) + [getattr(n, "target", None)]:
                for t in (target.elts if isinstance(target, ast.Tuple) else [target]):
                    if is_self_attribute(t):
                        found.add(t.attr)
    return found

def response_set(name):
    """The class's own methods, plus every method they call as self.m(), super().m() or Class.m()."""
    def where(cls, method):                  # the class in which a called method is defined
        return next(c for c in ancestry(cls) if method in {f.name for f in methods[c]})
    rs = {f"{name}.{f.name}" for f in methods[name]}
    for f in methods[name]:
        for call in ast.walk(f):
            if not (isinstance(call, ast.Call) and isinstance(call.func, ast.Attribute)):
                continue
            target, method = call.func.value, call.func.attr
            if isinstance(target, ast.Name) and target.id == "self":
                rs.add(f"{where(name, method)}.{method}")
            elif isinstance(target, ast.Call) and getattr(target.func, "id", "") == "super":
                rs.add(f"{where(parent[name], method)}.{method}")
            elif isinstance(target, ast.Name) and target.id in classes:
                rs.add(f"{target.id}.{method}")
    return rs

# classes named inside a class's methods (inheritance is measured by DIT and NOC, not counted again)
named = {name: {n.id for f in methods[name] for n in ast.walk(f) if isinstance(n, ast.Name)}
               & (classes.keys() - {name}) for name in classes}

print(f"{'class':<12}{'WMC':>5}{'DIT':>5}{'NOC':>5}{'CBO':>5}{'RFC':>5}{'LCOM':>6}")
for name in classes:
    ivars = instance_variables(name)
    used = [{n.attr for n in ast.walk(f) if is_self_attribute(n) and n.attr in ivars} for f in methods[name]]
    wmc = len(methods[name])                 # every method's complexity taken as 1
    dit = len(list(ancestry(name))) - 1
    noc = sum(1 for c in classes if parent[c] == name)
    cbo = sum(1 for other in classes if other != name and (other in named[name] or name in named[other]))
    print(f"{name:<12}{wmc:>5}{dit:>5}{noc:>5}{cbo:>5}{len(response_set(name)):>5}{lcom(used):>6}")
munotes.in398

Object-Oriented Metrics: The CK Suite

the paper's own example, {a,b,c,d,e} {a,b,e} {x,y,z}: LCOM = 1
class         WMC  DIT  NOC  CBO  RFC  LCOM
Form            3    0    2    0    3     0
ExamForm        6    1    1    2   10     7
BacklogForm     1    2    0    0    3     0
ContactForm     2    1    0    0    4     0
Catalogue       1    0    0    1    1     0
FeeRule         1    0    0    1    1     0

The first line proves the LCOM function agrees with the paper. Then the table, read as a reviewer and a tester would read it:

  • ExamForm stands out on every column that matters. It has the most methods (WMC 6), is coupled to both helpers (CBO 2), and a message to it can set ten different methods running (RFC 10). By the paper's viewpoints it will take the most effort to maintain and the most care to test.
  • Its LCOM of 7 is a design finding. Of its six methods, four work on the papers, the errors and the student; two, allot_seat and seat_label, work only on the centre and the seat, and share nothing with the rest. Fifteen pairs of methods, eleven sharing nothing, four sharing something: 11 - 4 = 7. Allotting a seat is a different responsibility from checking an exam form, and the metric suggests what a reviewer would: move it to a class of its own, a SeatAllotment.
  • Form's NOC of 2 is a testing finding. Two classes inherit its validate, and both call it through super(); a defect in it breaks all three forms.
  • BacklogForm's DIT of 2 is the deepest: to know what its validate does, a reader must read three classes. Its RFC of 3 is small, but two of the three methods it can run belong to its ancestors.
  • Form's LCOM of 0 and ContactForm's of 0 say those classes hang together: every pair of methods shares an instance variable.
munotes.in399

Object-Oriented Metrics: The CK Suite

Using the suite

The metrics are useful as a way of choosing where to look, not as scores. A tester reads them to direct effort: classes with high RFC and CBO need the most integration testing and the most careful test doubles (Chapter Thirty-Four, on unit testing techniques); classes with high NOC need their inherited methods tested in every subclass that uses them. A reviewer reads them to question the design: a high LCOM asks whether the class should be split, and a high CBO whether the classes should know so much about each other.

Two cautions come from the paper itself. WMC's complexity weighting is "left as an implementation decision", so two tools' WMC values are comparable only with the same weighting. And an LCOM of 0 "does not imply maximal cohesiveness, since within the set of classes with LCOM = 0, some may be more cohesive than others." Every value, as Chapter Sixty-Five insisted of every metric, means something only with its counting rule, which is why this chapter states its own.

What it does not mean

A high value is not a defect. It is a reason to look: a large class may be large for good reasons.

Deeper inheritance is not simply worse. The paper gives both sides: harder to predict, but more reuse.

LCOM 0 is not perfect cohesion. It means the sharing pairs are at least as many as the disjoint ones.

The suite does not replace testing. It points testing at the classes where defects are most likely to be expensive.

Quick revision

  • CK suite (Chidamber and Kemerer, IEEE TSE 1994): six metrics for object-oriented design, measured per class.
  • WMC: sum of method complexities; with complexities of 1, the number of methods.
  • DIT: depth in the inheritance tree; NOC: number of immediate subclasses.
  • CBO: number of other classes coupled to it (one uses the other's methods or instance variables).
  • RFC: size of the response set, the class's methods plus the methods they call (first level).
  • LCOM: disjoint method pairs minus sharing pairs, or 0; the paper's example gives 1.
  • Worked example: ExamForm WMC 6, CBO 2, RFC 10, LCOM 7 (11 disjoint pairs, 4 sharing): split out seat allotment; Form NOC 2; BacklogForm DIT 2.
  • Use: to direct testing and review effort; state the counting rules.
munotes.in400

Object-Oriented Metrics: The CK Suite

Test yourself

1. Name and define the six CK metrics. WMC, the sum of the complexities of a class's methods (the number of methods if each counts 1); DIT, the depth of the class in the inheritance tree; NOC, the number of its immediate subclasses; CBO, the number of other classes it is coupled to; RFC, the number of methods that can be executed in response to a message to it, its own methods plus those they call; LCOM, the number of pairs of its methods sharing no instance variable minus the number sharing one, or zero.

2. Compute LCOM for a class whose three methods use the instance variables {a, b}, {b, c} and {d}. Pairs: the first two share b (one sharing pair); the first and third, and the second and third, share nothing (two disjoint pairs). LCOM = 2 - 1 = 1.

3. What do high values of CBO and RFC mean for a tester? A class coupled to many others, or whose methods call many methods, is harder to test and debug: its behaviour depends on more code, it needs more test doubles in unit testing and more integration tests, and it deserves more of the testing time.

4. What does a high LCOM suggest, and what did it suggest in the worked example? That the class's methods fall into groups that do not share data, so the class probably has more than one responsibility and should be split. In ExamForm, two methods used only the centre and seat, so seat allotment could become a class of its own.

5. Give one advantage and one disadvantage of a deep inheritance tree, in the CK paper's terms. A class deep in the tree can reuse many inherited methods; but it inherits so much that its behaviour is harder to predict, and deeper trees mean more classes and methods are involved, so the design is more complex.

6. Why must the counting rules for CK metrics be stated? Because the paper leaves some choices open, such as the complexity weighting in WMC, and applying the metrics to a particular language needs more (whether the language's root class counts in DIT, which calls enter RFC, whether inheritance counts as coupling), so two tools can give different values for the same class.

Contents This chapter on its own page

munotes.in401

Chapter Seventy

Cyclomatic Complexity

Syllabus topic Module 2, "Software Metrics: ... Complexity metrics"

In one line

Cyclomatic complexity counts the independent paths through a module's control flow graph, V(G) = E - N + 2, which for code with only binary decisions is simply the number of decisions plus one; it tells a developer when a module has grown too tangled to understand and a tester how many tests it takes to cover its logic.

In the wording a student can write in an examination: in NIST Special Publication 500-235, "Cyclomatic complexity is defined for each module to be e - n + 2, where e and n are the number of edges and nodes in the control flow graph, respectively." (McCabe's general form is V(G) = e - n + 2p, where p is the number of connected components, 1 for a single module.) It can be computed three ways: from the graph, E - N + 2; from the decisions, "If all decisions are binary and there are p binary decision predicates, v(G) = p + 1"; and, for a graph drawn without crossing edges, as the number of regions, including the outside. McCabe proposed an upper limit of 10 for a module. The measure gives "the number of independent paths" through the graph, which is the number of tests basis path testing needs.

The problem McCabe set out to solve

McCabe opened his 1976 paper with a question about modules: how to divide a system into pieces that are "both testable and maintainable". The practice of the day limited a module's physical size, and he showed why that was not enough: imagine a 50-line program made of 25 consecutive IF THEN constructs. "Such a program could have as many as 33.5 million distinct control paths", only a few of which would ever be tested. Size and difficulty are different things.

He therefore measured the number of paths, but not all of them, since "Any program with a backward branch potentially has an infinite number of paths." Instead he counted basic paths, the independent paths that, "when taken in combination", generate every other path. Their number is the cyclomatic number of the program's control flow graph, a quantity from graph theory.

Three ways to compute V(G)

From the graph. Count the edges E and the nodes N of the control flow graph (Chapter Fifty-Eight, on structural testing, drew one) and compute E - N + 2. NIST 500-235 explains the + 2: the cyclomatic number of graph theory is e - n + 1 for a strongly connected graph; a program's graph becomes strongly connected when a "virtual edge" is added from its exit back to its entry, and adding one for that edge gives e - n + 2.

munotes.in402

Cyclomatic Complexity

From the decisions. In a graph where every decision has two ways out, each decision adds exactly one edge more than it adds nodes, so V(G) is the number of decisions plus one. NIST 500-235 puts it simply: "Starting with one and adding the number of such nodes yields the complexity." A decision with more ways out, such as a case statement, adds one less than its number of branches.

From the regions. Drawn on paper without edges crossing, a graph divides the page into regions, and by Euler's formula for planar graphs their number is E - N + 2. So "counting the regions gives a quick visual method for determining complexity", provided the outside of the drawing is counted as one region.

Here is check_form's graph again, from Chapter Fifty-Eight's structural testing:

The control flow graph of check_form: nine nodes, four of them decisions, and twelve edges

Figure 70.1 check_form's control flow graph: 12 edges, 9 nodes, 4 decisions, 5 regions

All three methods agree on it. There are 12 edges and 9 nodes, so V(G) = 12 - 9 + 2 = 5. There are 4 decisions (the shaded nodes), so V(G) = 4 + 1 = 5. And the drawing has 5 regions: the triangle 2, 3, 4; the two loops closed by the back edges into node 4 (one between 4, 5 and the False edge back from 5, one around 5, 6 and the edge back from 6); the triangle 7, 8, 9; and the outside.

What counts as a decision in source code

Counting decisions straight from the code is the method used in practice, but "constructs that appear similar in different languages may not have identical control flow semantics, so caution is advised", NIST 500-235 warns. Its rules:

  • An if statement, a while loop and the like are binary decisions: each adds one. A for loop decides, each time round, whether to go round again, so it adds one too.
  • Compound conditions. McCabe noted that a condition such as c1 AND c2 is two decisions in disguise, since without the AND it would be written as two nested ifs, and "it has been found to be more convenient to count conditions instead of predicates when calculating complexity". NIST 500-235 makes it precise: "Boolean operators add either one or nothing to complexity, depending on whether they have short-circuit evaluation semantics". Python's and and or are short-circuit (the right side is evaluated only if it can change the result), so each adds one.
  • A conditional expression, a if c else b, is a decision inside an expression, and adds one.

Worked example: measuring ExamReg's functions

The file holds four of ExamReg's functions: late_band and total_fee from the fee module, check_form from Chapter Fifty-Eight's structural testing, and hall_ticket_status, which decides what a student's hall ticket page shows.

munotes.in403

Cyclomatic Complexity

def late_band(days_late):
    if days_late == 0:
        return "on time"
    if days_late <= 7:
        return "1 to 7 days"
    return "8 to 15 days"


def total_fee(days_late, backlog_papers, concession):
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    fee = (0 if concession else 800) + 150 * backlog_papers
    return fee + {"on time": 0, "1 to 7 days": 100, "8 to 15 days": 500}[late_band(days_late)]


def check_form(papers, days_late):
    problems = []
    if not papers:
        problems.append("no papers selected")
    for code in papers:
        if len(code) != 7:
            problems.append(f"bad paper code {code}")
    if days_late > 15:
        problems.append("more than 15 days late")
    return problems


def hall_ticket_status(student, form, paid, today, last_date):
    if not student.active:
        return "blocked: not an active student"
    if student.dues > 0 and not student.concession:
        return "blocked: college dues unpaid"
    if not form.submitted:
        return "blocked: no exam form"
    if not paid:
        if today <= last_date:
            return "waiting: fee not paid yet"
        return "blocked: fee not paid"
    if not form.approved and not form.late_approval:
        return "blocked: form not approved by the college"
    for paper in form.papers:
        if paper.clash or (paper.backlog and not paper.eligible):
            return f"blocked: paper {paper.code}"
    if student.photo is None or student.signature is None:
        return "blocked: photo or signature missing"
    return "issued"

The program counts each function's decisions by NIST 500-235's rules, computes V(G) as decisions plus one, checks check_form a second way from the edge list of its graph, and checks McCabe's 33.5 million.

import ast

def decisions(function):
    """Binary decisions, counted as NIST 500-235 counts them from source code."""
    count = 0
    for node in ast.walk(function):
        if isinstance(node, (ast.If, ast.For, ast.While, ast.IfExp, ast.ExceptHandler)):
            count += 1                       # a statement or expression with two ways out
        elif isinstance(node, ast.BoolOp):
            count += len(node.values) - 1    # each short-circuit "and" or "or" adds one
        elif isinstance(node, ast.comprehension):
            count += 1 + len(node.ifs)       # a comprehension's loop, and each filter
    return count

tree = ast.parse(open("examreg_code.py").read())
print(f"{'function':<20}{'decisions':>10}{'V(G)':>6}")
for f in tree.body:
    v = decisions(f) + 1
    print(f"{f.name:<20}{decisions(f):>10}{v:>6}" + ("   over 10: split it" if v > 10 else ""))

# check_form's control flow graph as Chapter 58 drew it: 9 nodes, and these edges
edges = [(1, 2), (2, 3), (2, 4), (3, 4), (4, 5), (4, 7), (5, 6), (5, 4), (6, 4), (7, 8), (7, 9), (8, 9)]
nodes = {n for edge in edges for n in edge}
print(f"check_form from its graph: {len(edges)} edges - {len(nodes)} nodes + 2 = {len(edges) - len(nodes) + 2}")

# McCabe's warning about size: 25 consecutive IF THEN constructs, each taken or not
print(f"paths through 25 IF THEN constructs in a row: {2 ** 25:,}; their V(G): {25 + 1}")
function             decisions  V(G)
late_band                    2     3
total_fee                    3     4
check_form                   4     5
hall_ticket_status          14    15   over 10: split it
check_form from its graph: 12 edges - 9 nodes + 2 = 5
paths through 25 IF THEN constructs in a row: 33,554,432; their V(G): 26
munotes.in404

Cyclomatic Complexity

The small functions. late_band has two decisions and V(G) 3: three independent paths, one for each band. total_fee has 3 decisions although it has only one if: the or in its first condition adds one, and the conditional expression that waives the form fee adds another. That is the counting rule doing its job, since each is a place where the code can go two ways and a test can go wrong. check_form comes out at 5 by counting decisions and at 5 again from its graph.

The function over the limit. hall_ticket_status looks orderly: a list of checks, one after another. But it has 14 decisions: nine ifs, a for loop, and four ands and ors hidden in its conditions. V(G) = 15, over McCabe's limit, and the program says so. Fifteen independent paths is fifteen tests to cover its logic, and a reader must hold all fifteen routes in mind to change it safely. McCabe's projects had a rule for this case: when a module's complexity went over 10, its programmers had to recognise the subfunctions inside it and make them modules of their own, or else rewrite it. Here the checks fall into natural groups (the student's standing, the form and its fee, the papers, the documents), and each group can become a function of its own, each well under 10.

The size warning. Twenty-five IF THEN constructs in a row, each either taken or not, give 33,554,432 paths (2 to the power 25): McCabe's 33.5 million. Their V(G) is only 26. That is the point of the measure: it counts the independent paths a tester can actually cover, not the combinations of them.

The limit of 10, and its exceptions

McCabe reported that his projects' "particular upper bound" was 10. NIST 500-235 records the debate since: "The original limit of 10 as proposed by McCabe has significant supporting evidence, but limits as high as 15 have been used successfully as well." Limits over 10, it says, "should be reserved for projects that have several operational advantages over typical projects", such as experienced staff and code walkthroughs; "an organization can pick a complexity limit greater than 10, but only if it is sure it knows what it is doing and is willing to devote the additional testing effort required by more complex modules."

The one exception McCabe allowed was a module made of a single large case statement, many independent cases following one selection. NIST 500-235 tells how that exception was abused: one developer reported a "modified" complexity that divided by the number of case branches, and "could take a module with complexity 90 and reduce it to 'modified' complexity 10 simply by adding a ten-branch multiway decision statement to it that did nothing." A metric that can be lowered without making the code any simpler is no longer measuring anything.

munotes.in405

Cyclomatic Complexity

What V(G) tells a tester

Cyclomatic complexity is the number of independent paths, so it is the number of tests in a basis set: tests that between them exercise every decision outcome independently. NIST 500-235 states it for multiway decisions too: "Cyclomatic complexity gives the number of tests, which for a multiway decision statement is the number of decision branches." Chapter Seventy-Two, on basis path testing, builds such a set.

It also says where to spend effort. The more complex a module, the more tests it needs and, in NIST 500-235's words, overly complex modules "are more prone to error, are harder to understand, are harder to test, and are harder to modify."

What it does not mean

V(G) is not the number of paths. With a loop there may be infinitely many; V(G) counts the independent ones, which combine to form the rest.

It is not a measure of size. A long straight-line function has V(G) 1; a short one full of conditions can be over 10.

It does not measure everything that makes code hard. Deep nesting, obscure names and complicated data all make code harder to understand without changing V(G); the next chapter's measures look at some of them.

Two tools may disagree. Whether and, or, conditional expressions and case branches count is a rule, and the rule must be stated with the number.

Quick revision

  • V(G) = E - N + 2 (NIST 500-235; McCabe's general form E - N + 2P, P connected components).
  • V(G) = decisions + 1 when every decision is binary; a multiway decision adds its branches minus one.
  • V(G) = regions of a planar drawing of the graph, counting the outside.
  • Short-circuit and and or add one each; so does a conditional expression.
  • McCabe (1976): 25 IF THEN constructs give 33.5 million paths (2 to the power 25) but V(G) 26; limit 10, "reasonable, but not magical"; NIST: up to 15 with advantages.
  • Worked example: late_band 3, total_fee 4, check_form 5 (and 12 - 9 + 2 = 5 from its graph), hall_ticket_status 15, over the limit.
  • V(G) is the number of tests in a basis set of independent paths.

Test yourself

1. Define cyclomatic complexity and give three ways to compute it. A measure of the number of independent paths through a module's control flow graph. It is E - N + 2, from the numbers of edges and nodes; the number of binary decisions plus one; and, for a graph drawn without crossings, the number of regions including the outside.

munotes.in406

Cyclomatic Complexity

2. A control flow graph has 11 edges and 9 nodes. What is its cyclomatic complexity? 11 - 9 + 2 = 4.

3. Compute V(G) for check_form and show that the three methods agree. It has 12 edges and 9 nodes, so 12 - 9 + 2 = 5; it has 4 binary decisions (the empty-form check, the loop, the code-length check and the lateness check), so 4 + 1 = 5; and its drawn graph has 5 regions: two triangles, two loops into the loop node, and the outside.

4. How do compound conditions affect cyclomatic complexity? A condition with a short-circuit and or or is several decisions, because the second part is evaluated only in some cases; each such operator adds one to the complexity, so if a and b adds two, not one.

5. Why did McCabe propose a limit of 10, and how should a module over the limit be treated? Because modules with high complexity are harder to understand, test and modify; he found 10 a reasonable upper bound for keeping modules testable. A module over the limit should be split into subfunctions or rewritten, except, in his rule, a module consisting of one large case statement.

6. Why is cyclomatic complexity more useful than lines of code for judging testability? Because testability depends on the number of paths through the logic, not on length: a 50-line module of 25 consecutive ifs has 33.5 million paths, while a longer straight-line module has one. Cyclomatic complexity counts the independent paths, which is also the number of tests needed to cover them.

Contents This chapter on its own page

munotes.in407

Chapter Seventy-One

Halstead's Measures and Other Complexity Metrics

Syllabus topic Module 2, "Software Metrics: ... Complexity metrics"

In one line

Halstead's measures treat a program as a stream of operators and operands and compute from their counts its length, vocabulary, volume, difficulty and the effort to write it; they are easy to count and useful for comparing programs, but the theory behind them is weak and the numbers depend on counting rules that must be stated; structural measures such as fan-in, fan-out and nesting depth look at how the pieces connect instead.

In the wording a student can write in an examination: in Halstead's software science, as Shen, Conte and Dunsmore summarise it, "A computer program is considered in Software Science to be a series of tokens which can be classified as either 'operators' or 'operands'." From η1 (unique operators), η2 (unique operands), N1 (total occurrences of operators) and N2 (total occurrences of operands) come the length N = N1 + N2, the vocabulary η = η1 + η2, the volume V = N × log2 η (in bits), the estimated length N^ = η1 log2 η1 + η2 log2 η2, the difficulty D = (η1 / 2) × (N2 / η2), the effort E = D × V (in "elementary mental discriminations") and the time T = E / S seconds, with S, the Stroud number, "normally set to 18". The Purdue review found the theory's foundations weak ("the very base of Software Science (counting operators and operands) is shaky due to ambiguities concerning what should be counted"), but some measures, notably the difficulty as a measure of error-proneness, supported by data.

Programs as operators and operands

Maurice Halstead's Elements of Software Science (1977) proposed that a program could be measured like a text, by counting its words. "Generally, any symbol or keyword in a program that specifies an algorithmic action is considered an operator, and a symbol used to represent data is considered an operand. Most punctuation marks are also considered as operators." In if days_late <= 7: return "1 to 7 days", the operators are if, <=, : and return; the operands are days_late, 7 and the string.

Four counts are taken:

SymbolCount
η1the number of unique operators
η2the number of unique operands
N1the total occurrences of operators
N2the total occurrences of operands

The measures

Everything else is computed from those four, and the review gives each formula:

  • Length, N = N1 + N2: the total number of tokens.
  • Vocabulary, η = η1 + η2: the number of different tokens.
  • Volume, V = N × log2 η. Its unit is the bit: "It is the actual size in a computer if a uniform binary encoding for the Vocabulary is used." The review also reads it as "the number of mental comparisons needed to write a program of Length N".
  • Program level and difficulty. A program written in the most compact form possible has the potential volume V; any program's level is L = V / V, between 0 and 1, and "The inverse of the Program Level is termed the Difficulty", D = 1 / L. Since V* is usually unknown, the level is estimated from the counts as (2 / η1) × (η2 / N2), so the difficulty estimate is D = (η1 / 2) × (N2 / η2). "An intuitive argument for this formula is that programming difficulty increases if additional operators are introduced ... and if an operand is used repetitively".
  • Effort, E = V / L = D × V, in "elementary mental discriminations".
  • Time, T = E / S seconds, where S is the Stroud number: a psychologist, John Stroud, had suggested the mind makes a limited number of elementary discriminations per second, between 5 and 20, and "S is normally set to 18".
  • The length equation, N^ = η1 log2 η1 + η2 log2 η2: Halstead's claim that a well-structured program's length can be predicted from its vocabulary alone.
munotes.in408

Halstead's Measures and Other Complexity Metrics

Worked example 1: counting three ExamReg functions

Halstead's measures stand or fall by the counting rules, so this chapter states its own for Python, and the program applies them: keywords and punctuation are operators, a pair of brackets counts once (at the opening bracket), a name followed by an opening bracket (a function being defined or called) is an operator, and every other name, number and string is an operand, including True, False and None.

def late_band(days_late):
    if days_late == 0:
        return "on time"
    if days_late <= 7:
        return "1 to 7 days"
    return "8 to 15 days"


def total_fee(days_late, backlog_papers, concession):
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    fee = (0 if concession else 800) + 150 * backlog_papers
    return fee + {"on time": 0, "1 to 7 days": 100, "8 to 15 days": 500}[late_band(days_late)]


def days_late(last_date, submitted):
    return max((submitted - last_date).days, 0)


def fee_due(last_date, submitted, backlog_papers, concession):
    return total_fee(days_late(last_date, submitted), backlog_papers, concession)


def check_form(papers, days):
    problems = []
    if not papers:
        problems.append("no papers selected")
    for code in papers:
        if len(code) != 7:
            problems.append("bad paper code " + code)
    if days > 15:
        problems.append("more than 15 days late")
    return problems
import ast, io, keyword, math, tokenize

SKIP = {tokenize.NEWLINE, tokenize.NL, tokenize.INDENT, tokenize.DEDENT, tokenize.COMMENT,
        tokenize.ENDMARKER}

def operators_and_operands(source):
    """The chapter's counting rules: keywords and punctuation are operators (a bracket pair
    counts once, at its opening bracket), a called name is an operator, and every other
    name, number and string is an operand; True, False and None are operands."""
    tokens = [t for t in tokenize.generate_tokens(io.StringIO(source).readline) if t.type not in SKIP]
    operators, operands = [], []
    for i, t in enumerate(tokens):
        after = tokens[i + 1].string if i + 1 < len(tokens) else ""
        if t.type == tokenize.OP:
            if t.string not in ")]}":
                operators.append(t.string)
        elif t.type == tokenize.NAME and keyword.iskeyword(t.string) and t.string not in ("True", "False", "None"):
            operators.append(t.string)
        elif t.type == tokenize.NAME and after == "(":
            operators.append(t.string + "()")
        else:
            operands.append(t.string)
    return operators, operands

module = open("examreg_module.py").read()
functions = {f.name: ast.get_source_segment(module, f) for f in ast.parse(module).body}

ops, opnds = operators_and_operands(functions["late_band"])
print("late_band's unique operators:", " ".join(sorted(set(ops))))
print("late_band's unique operands: ", " ".join(sorted(set(opnds))))

print(f"{'function':<12}{'n1':>4}{'n2':>4}{'N1':>4}{'N2':>4}{'N':>5}{'N est':>7}{'V':>7}{'D':>6}{'E':>8}{'T (s)':>7}")
for name in ("late_band", "total_fee", "check_form"):
    ops, opnds = operators_and_operands(functions[name])
    n1, n2, N1, N2 = len(set(ops)), len(set(opnds)), len(ops), len(opnds)
    N = N1 + N2                                          # length
    N_est = n1 * math.log2(n1) + n2 * math.log2(n2)      # the length equation
    V = N * math.log2(n1 + n2)                           # volume, in bits
    D = (n1 / 2) * (N2 / n2)                             # difficulty: 1 / the level estimate
    E = D * V                                            # effort
    print(f"{name:<12}{n1:>4}{n2:>4}{N1:>4}{N2:>4}{N:>5}{N_est:>7.1f}{V:>7.1f}{D:>6.1f}{E:>8.0f}{E / 18:>7.0f}")
munotes.in409

Halstead's Measures and Other Complexity Metrics

late_band's unique operators: ( : <= == def if late_band() return
late_band's unique operands:  "1 to 7 days" "8 to 15 days" "on time" 0 7 days_late
function      n1  n2  N1  N2    N  N est      V     D       E  T (s)
late_band      8   6  13   8   21   39.5   80.0   5.3     426     24
total_fee     19  14  31  22   53  134.0  267.4  14.9    3991    222
check_form    18   9  32  18   50  103.6  237.7  18.0    4279    238

The first two lines show the rules at work on late_band: eight distinct operators, among them def, if, return, the two comparisons, the colon and the function's own name, and six distinct operands, its parameter, two numbers and three strings. The columns n1 and n2 in the table are the η1 and η2 of the formulas.

Read the table as Halstead's theory asks. late_band is the smallest by every measure: volume 80 bits, difficulty 5.3. total_fee and check_form are close in volume (267 and 238 bits), but check_form has the higher difficulty, 18.0 against 14.9, because it reuses its operands more (its problems list is named five times): the formula counts repetition as difficulty. The effort column multiplies the two, and the time column divides the effort by 18, predicting that check_form took about four minutes of concentrated work to write. Whether it did is exactly what the critique doubts.

And look at the length equation. It predicts 39.5 tokens for late_band, which has 21; 134.0 for total_fee, which has 53. The review observed that the equation "appears to work best in the range of N between 2000 and 4000", far larger than any of these functions.

munotes.in410

Halstead's Measures and Other Complexity Metrics

The Purdue critique

Shen, Conte and Dunsmore's review, from the Software Metrics Research Group at Purdue, is the standard account of what is and is not known about software science. Its findings:

  • The counting is ambiguous. "Ambiguities, both theoretical and practical, in the classification and treatment of some operators and operands may lead to substantially different values of some Software Science metrics." Halstead counted each GO TO label as a different operator for each label, but every IF as the same operator. Tools resolve the ambiguities "using some convenient strategy", which is why this chapter wrote its own rules down.
  • The derivations rest on unverifiable assumptions. The review found that "several implied assumptions are made for which there seem to be no theoretical justifications".
  • The time equation is suspect. "Among psychologists there is no general acceptance of Stroud's hypothesis that the mind is capable of making a constant number (S) of mental discriminations per second. As a theoretical concept, the Time equation must therefore be considered suspect."
  • The early experiments were weak: the samples were too small ("Many of Halstead's conclusions were based on sample sizes less than 10") and the programs were small.

But its conclusion is not a dismissal. "Published data does seem to sustain the usefulness of D (the so-called Difficulty metric) as a measure of error-proneness", and "Results also suggest that the Software Science E is a better effort measure than most others being used." In practice, it concluded, "the 'real world' use of Software Science measures in their current state must be done very carefully."

Other complexity measures: fan-in, fan-out and nesting

ISO/IEC/IEEE 24765 defines complexity as the "degree to which a system's design or code is difficult to understand because of numerous components or relationships among components". McCabe's measure looks at the paths inside a module and Halstead's at its tokens; the relationships between modules need measures of their own.

  • Fan-out: the number of other modules a module calls. A module with high fan-out depends on many others, and changes in any of them can break it. CMMI's examples of process objectives include "Keep design complexity (fan-out rate) below a specified threshold".
  • Fan-in: the number of modules that call a given module. High fan-in is a sign of reuse, and also of risk: a defect in it reaches every caller.
  • Nesting depth: how deeply compound statements are placed inside one another. Nesting is "embedding one construct inside another" (ISO/IEC/IEEE 24765); every level is one more condition a reader must hold in mind.

These are the chapter's operational definitions, computed on the module's own call graph by the program below.

munotes.in411

Halstead's Measures and Other Complexity Metrics

Worked example 2: the module's call graph

import ast

tree = ast.parse(open("examreg_module.py").read())
functions = {f.name: f for f in tree.body if isinstance(f, ast.FunctionDef)}

# the call graph: which of the module's own functions each function calls
calls = {name: {n.func.id for n in ast.walk(f) if isinstance(n, ast.Call)
                and isinstance(n.func, ast.Name) and n.func.id in functions}
         for name, f in functions.items()}

COMPOUND = (ast.If, ast.For, ast.While, ast.With, ast.Try)

def nesting(node, depth=0):
    """The deepest nesting of compound statements inside node (an elif counts at its if's level)."""
    deepest = depth
    for child in ast.iter_child_nodes(node):
        elif_ = isinstance(node, ast.If) and node.orelse == [child] and isinstance(child, ast.If)
        inner = depth + 1 if isinstance(child, COMPOUND) and not elif_ else depth
        deepest = max(deepest, nesting(child, inner))
    return deepest

print(f"{'function':<12}{'fan-in':>7}{'fan-out':>8}{'nesting':>8}  calls")
for name, f in functions.items():
    fan_in = sum(1 for caller in functions if name in calls[caller])
    print(f"{name:<12}{fan_in:>7}{len(calls[name]):>8}{nesting(f):>8}  {', '.join(sorted(calls[name])) or '-'}")
function     fan-in fan-out nesting  calls
late_band         1       0       1  -
total_fee         1       1       1  late_band
days_late         1       0       0  -
fee_due           0       2       0  days_late, total_fee
check_form        0       0       2  -

fee_due has the highest fan-out: it calls days_late and total_fee, and through total_fee it depends on late_band too, so a test of it is an integration test of all three (Chapter Thirty-Eight's top-down and bottom-up orders are orders over exactly this graph). late_band, total_fee and days_late each have a fan-in of 1: each is used in one place, so a change to one can break one caller. check_form has the deepest nesting, 2, an if inside a for; the others never go deeper than one level. None of these values is alarming, which is the point of measuring: they show that the module's complexity sits in total_fee's conditions and check_form's loop, not in its structure.

What it does not mean

Halstead's time is not a schedule. The critique found the time equation suspect in theory; T is at most a way of comparing programs.

Volume is not lines of code. It counts tokens and weights them by the vocabulary, so two programs with the same number of lines can differ widely.

The numbers are not portable without the counting rules. Two tools that classify brackets, calls or declarations differently give different values for the same code.

High fan-in is not bad in itself. A widely used module is often a well-designed one; it simply deserves more testing, because its defects spread.

Quick revision

  • Tokens are operators (keywords, punctuation, actions) or operands (data): η1, η2 unique; N1, N2 total.
  • Length N = N1 + N2; vocabulary η = η1 + η2; volume V = N × log2 η bits; estimated length N^ = η1 log2 η1 + η2 log2 η2.
  • Level L = V* / V; difficulty D = 1 / L, estimated as (η1 / 2) × (N2 / η2); effort E = D × V; time T = E / 18 seconds (Stroud number).
  • Worked example: late_band V 80.0, D 5.3; total_fee V 267.4, D 14.9; check_form V 237.7, D 18.0; the length equation overestimated every one.
  • Purdue critique (Shen, Conte and Dunsmore 1981): ambiguous counting, unjustified assumptions, a suspect time equation, small samples; but D useful for error-proneness and E a good effort measure; use "very carefully".
  • Fan-out: modules called; fan-in: modules calling; nesting depth: compound statements inside one another.
munotes.in412

Halstead's Measures and Other Complexity Metrics

Test yourself

1. Define Halstead's four basic counts and the measures derived from them. η1 and η2 are the numbers of unique operators and unique operands; N1 and N2 their total occurrences. Length N = N1 + N2; vocabulary η = η1 + η2; volume V = N × log2 η; difficulty D = (η1 / 2) × (N2 / η2); effort E = D × V; time T = E / 18 seconds.

2. A program has 10 unique operators, 16 unique operands, 60 operator occurrences and 48 operand occurrences. Compute its length, vocabulary, volume, difficulty and effort. Length N = 60 + 48 = 108. Vocabulary η = 10 + 16 = 26. Volume V = 108 × log2 26, about 108 × 4.70, which is about 507.6 bits. Difficulty D = (10 / 2) × (48 / 16) = 15. Effort E = 15 × 507.6, about 7614.

3. What is the Stroud number, and why is the time equation doubted? The number of elementary mental discriminations the mind is supposed to make per second, taken as 18, which turns Halstead's effort into seconds. It is doubted because psychologists do not generally accept that the mind makes a constant number of such discriminations per second.

4. Summarise the Purdue critique of software science. The counting of operators and operands is ambiguous and changes the values; several derivations rest on unjustified assumptions; the time equation is suspect; and the early experiments used too few and too small programs. Yet the difficulty measure seems useful for predicting error-proneness and the effort measure compares well with others, so the measures can be used, very carefully.

5. Define fan-in and fan-out, and say what a high value of each suggests. Fan-out is the number of modules a module calls; a high value means it depends on many others and is sensitive to their changes. Fan-in is the number of modules that call it; a high value means it is widely reused, and a defect in it affects many callers, so it deserves thorough testing.

munotes.in413

Halstead's Measures and Other Complexity Metrics

6. Why must the counting rules for Halstead's measures be stated? Because whether brackets, calls, declarations and constants count as operators or operands changes the counts, and so every measure computed from them; values from two tools are comparable only if both used the same rules.

Contents This chapter on its own page

munotes.in414

Chapter Seventy-Two

Why Complexity Matters to Testing: Basis Path Testing

Syllabus topic Module 2, "Software Metrics: ... Complexity metrics, and their significance in testing"

In one line

A module's cyclomatic complexity is the number of independent paths through it, and basis path testing tests exactly that many paths, chosen so that each varies one decision from the paths before it; because every decision is then tested independently, defects that hide when decisions go the same way cannot hide, and the number of tests grows with the complexity, which is where the defects are.

In the wording a student can write in an examination: NIST 500-235 states the structured testing criterion: "Test a basis set of paths through the control flow graph of each module. This means that any additional path through the module's control flow graph can be expressed as a linear combination of paths that have been tested." "Basis path testing, another name for structured testing, is the requirement that a basis set of paths should be tested." A basis set has exactly V(G) paths. The baseline method generates one: "start with a baseline path, then vary exactly one decision outcome to generate each successive path until all decision outcomes have been varied, at which time a basis will have been generated." Structured testing "subsumes branch and statement coverage testing", and its number of tests is proportional to complexity: "the minimum number of tests required to satisfy the structured testing criterion is exactly the cyclomatic complexity."

Complexity's significance in testing

MU's syllabus asks for complexity metrics "and their significance in testing". NIST 500-235 gives three answers, and this chapter demonstrates each.

  1. It says how many tests. V(G) is the number of independent paths, so it is the minimum number of tests that covers the module's logic in the structured sense. "Statement and branch coverage testing do not even come close to sharing this property. All statements and branches of an arbitrarily complex module can be covered with just one test".
  2. It says where to test hardest. "Given the correlation between complexity and errors, it makes sense to concentrate testing effort on the most complex and therefore error-prone software." A module of complexity 15 gets three times the basis tests of one of complexity 5.
  3. It finds defects that coverage hides. "Another strength of structured testing is that, for the precise mathematical interpretation of 'independent' as 'linearly independent,' structured testing guarantees that all decision outcomes are tested independently." The worked example shows what that buys.

Paths as vectors

The word independent is meant literally. Number the edges of a control flow graph, and write a path as a list of 1s and 0s, one for each edge, saying whether the path takes it. Paths are then vectors, and one path is a combination of others when its vector is their sum and difference. A set of paths is independent when none of them is a combination of the others, and it is a basis when every path through the graph is a combination of paths in the set. McCabe's result, restated in NIST 500-235, is that a basis always has V(G) paths: fewer cannot generate every path, and any more are combinations of the rest.

munotes.in415

Why Complexity Matters to Testing: Basis Path Testing

The baseline method

A basis can be found by trial and error, but NIST 500-235 gives a method. "The first step is to pick a functional 'baseline' path through the program that represents a legitimate function and not just an error exit", in the tester's judgement "the most important path to test". Then, one decision at a time, flip a decision's outcome on a path already chosen and follow the program on to its exit, until every decision outcome has been taken. Each new path adds exactly one edge the earlier paths had not taken, which is why the paths come out independent, and why there are exactly V(G) of them.

Worked example: two defects that cancel

Here is a version of ExamReg's fee rule for forms up to 7 days late: the form fee of Rs 800, waived by a concession, plus a Rs 100 late fee for any form after the last date. The version under test has two defects: its concession removes Rs 700 instead of Rs 800, and its late fee adds nothing instead of Rs 100. When a student has a concession and is late, the two mistakes cancel: 800 - 700 + 0 = 100, which happens to be right.

The function has two decisions, so V(G) = 3, and four paths through it, named by the outcomes of its two decisions: TT, TF, FT and FF. The program writes each path as a vector over the graph's seven edges, measures independence by the rank of those vectors, and compares a branch coverage test set with a basis.

from fractions import Fraction
from itertools import combinations

def total_fee(days_late, concession):                 # the version under test (days_late 0 to 7)
    fee = 800
    if concession:                                    # decision d1
        fee = fee - 700
    if days_late > 0:                                 # decision d2
        fee = fee + 0
    return fee

def specified(days_late, concession):                 # the fee rule for 0 to 7 days late
    return (0 if concession else 800) + (100 if days_late > 0 else 0)

# the control flow graph's edges, and the edges each path through it takes
EDGES = ["entry-d1", "d1 true", "after d1", "d1 false", "d2 true", "after d2", "d2 false"]
PATHS = {"TT": ["entry-d1", "d1 true", "after d1", "d2 true", "after d2"],
         "TF": ["entry-d1", "d1 true", "after d1", "d2 false"],
         "FT": ["entry-d1", "d1 false", "d2 true", "after d2"],
         "FF": ["entry-d1", "d1 false", "d2 false"]}
INPUT = {"TT": (3, True), "TF": (0, True), "FT": (3, False), "FF": (0, False)}

def vector(path):
    return [Fraction(int(e in PATHS[path])) for e in EDGES]

def rank(rows):
    """Gaussian elimination: the number of linearly independent rows."""
    rows, r = [row[:] for row in rows], 0
    for col in range(len(EDGES)):
        pivot = next((i for i in range(r, len(rows)) if rows[i][col] != 0), None)
        if pivot is None:
            continue
        rows[r], rows[pivot] = rows[pivot], rows[r]
        for i in range(len(rows)):
            if i != r and rows[i][col] != 0:
                factor = rows[i][col] / rows[r][col]
                rows[i] = [a - factor * b for a, b in zip(rows[i], rows[r])]
        r += 1
    return r

def run(name, paths):
    print(f"{name}: paths {' '.join(paths)}, rank {rank([vector(p) for p in paths])}")
    for p in paths:
        args = INPUT[p]
        got, want = total_fee(*args), specified(*args)
        print(f"   {p} {str(args):<11} expected {want:>3}, got {got:>3}  {'pass' if got == want else 'FAIL'}")

print(f"V(G) = {len(EDGES)} edges - 6 nodes + 2 = {len(EDGES) - 6 + 2}")
run("branch coverage, two tests", ["TT", "FF"])
run("basis from the baseline FT", ["FT", "TT", "FF"])
combo = [a + b - c for a, b, c in zip(vector("TT"), vector("FF"), vector("FT"))]
print("TT + FF - FT equals TF:", combo == vector("TF"))

bases = [c for c in combinations(PATHS, 3) if rank([vector(p) for p in c]) == 3]
finding = [c for c in bases if any(total_fee(*INPUT[p]) != specified(*INPUT[p]) for p in c)]
print(f"every basis of 3 paths: {len(bases)}; bases that find the defect: {len(finding)}")
munotes.in416

Why Complexity Matters to Testing: Basis Path Testing

V(G) = 7 edges - 6 nodes + 2 = 3
branch coverage, two tests: paths TT FF, rank 2
   TT (3, True)   expected 100, got 100  pass
   FF (0, False)  expected 800, got 800  pass
basis from the baseline FT: paths FT TT FF, rank 3
   FT (3, False)  expected 900, got 800  FAIL
   TT (3, True)   expected 100, got 100  pass
   FF (0, False)  expected 800, got 800  pass
TT + FF - FT equals TF: True
every basis of 3 paths: 4; bases that find the defect: 4

Branch coverage passes. Two tests, a late student with a concession (TT) and an on-time student without one (FF), make each decision go both ways: 4 of 4 branch outcomes, 100 per cent branch coverage. Both pass, because in TT the two defects cancel and in FF neither is reached. The rank line says why this is not enough: the two paths have rank 2, one short of V(G).

The basis fails. The tester takes as the baseline an ordinary late form: FT, 3 days late, no concession. Flipping the first decision gives TT; flipping the second gives FF. Three paths, rank 3: a basis. And the baseline itself fails: the student should pay 800 + 100 = 900 rupees and is charged 800, because the late fee adds nothing.

munotes.in417

Why Complexity Matters to Testing: Basis Path Testing

Why the basis could not miss it. The line TT + FF - FT equals TF shows the fourth path is a combination of the three tested ones, which is what makes them a basis. The last line checks every possible set of three paths: all 4 are bases, and all 4 find the defect. The reason is the one NIST 500-235 gives for its own example: testing the two decisions independently requires at least one path on which they go different ways, TF or FT, and on those paths the two defects no longer cancel. In NIST 500-235's words, structured testing "does not allow interactions between decision outcomes during testing to hide errors". Branch coverage does allow it: its two tests varied both decisions together.

The steps, for a module of any size

  1. Draw the control flow graph and compute V(G), by any of the three methods of Chapter Seventy, on cyclomatic complexity.
  2. Choose a baseline path: a typical, important function of the module, not an error exit.
  3. Flip one decision at a time: for each decision on a chosen path whose other outcome has not yet been taken, take that outcome and follow the program to the exit, keeping the other decisions as typical as possible.
  4. Stop at V(G) paths, when every outcome of every decision has been taken.
  5. Find inputs that drive each path, write expected results from the specification, and run.

Some paths cannot be driven by any input, for instance when a module makes the same decision twice; NIST 500-235 treats that case separately, by changing the code to remove the dependency or by settling for the largest number of independent paths that can be exercised.

Where the effort goes

Complexity turns into a test budget. Chapter Seventy's measurements of ExamReg put late_band at 3, check_form at 5 and hall_ticket_status at 15: a basis for the last needs fifteen tests, five times as many as the first. That is not a penalty on the complex module; it is the honest cost of its logic, and a reason, before testing starts, to split it, since three modules of complexity 5 are easier to test well than one of 15.

What it does not mean

A basis is not all paths. It is V(G) paths from which every other path can be built; with loops, all paths are infinitely many.

Basis path testing does not choose expected results. The 900 rupees that exposed the defect came from the fee rule.

It is not the same as weak structured testing. Running V(G) different paths that happen to cover all branches is not enough; NIST 500-235 calls that weak structured testing, and it does not guarantee independence.

munotes.in418

Why Complexity Matters to Testing: Basis Path Testing

Any basis is not equally good. Every basis catches defects like the one above; the baseline method adds the tester's judgement by starting from the most important path.

Quick revision

  • Structured (basis path) testing (NIST 500-235): "Test a basis set of paths through the control flow graph of each module".
  • A basis has V(G) linearly independent paths; every other path is a linear combination of them.
  • Baseline method: start from the most important functional path; flip exactly one decision outcome at a time until every outcome has been taken.
  • Significance in testing: V(G) is the minimum number of tests; test effort should follow complexity; decisions are tested independently, so interactions cannot hide defects.
  • Worked example: two cancelling defects; branch coverage (TT, FF, rank 2) passed; the basis (FT, TT, FF, rank 3) failed at FT, 800 charged for 900; all 4 possible bases find the defect.
  • Structured testing subsumes branch and statement coverage.

Test yourself

1. State the structured testing criterion. What is a basis set of paths? Test a basis set of paths through each module's control flow graph. A basis set is a set of linearly independent paths from which every path through the graph can be formed as a linear combination; it always has V(G) paths.

2. Describe the baseline method for generating basis paths. Choose a baseline path that represents a typical, important function of the module. Then repeatedly take a path already chosen, flip the outcome of one decision whose other outcome has not yet been exercised, and follow the program to its exit. Each new path adds one new decision outcome, and after V(G) paths every outcome has been taken and the paths form a basis.

3. Why does a module with V(G) = 5 need at least five tests under basis path testing, when branch coverage might need two? Because five linearly independent paths are needed to form a basis, and each test follows one path. Branch coverage only requires each decision outcome to be taken once, which a few long paths can do together without the decisions ever being varied independently.

4. In the worked example, why did branch coverage miss the defects while every basis found them? Branch coverage used two tests in which both decisions went the same way, and on those paths the two defects cancelled out or were not reached. Any basis of three paths must include a path on which the decisions go different ways, and on such a path the defects do not cancel, so the wrong fee appears.

munotes.in419

Why Complexity Matters to Testing: Basis Path Testing

5. Derive the basis paths for check_form, whose complexity is 5. Baseline: a valid form with one correctly coded paper, on time (papers present, loop entered once, code valid, loop exits, not late). Flip the empty-form decision: no papers (the loop is then not entered). Flip the code-length decision: one badly coded paper. Flip the loop decision: two papers, so the loop repeats. Flip the lateness decision: the baseline form more than 15 days late. Five paths, with expected results from the specification.

6. What is the significance of cyclomatic complexity in testing? It gives the minimum number of tests for basis path testing; it points testing effort at the most complex, most error-prone modules; and testing a basis set guarantees each decision is tested independently, so defects that depend on how decisions combine cannot hide as they can under branch coverage.

Contents This chapter on its own page

munotes.in420

Chapter Seventy-Three

What a Defect Is

Syllabus topic Module 2, "Defect Management: Definition of defects"

In one line

A defect is an imperfection in a work product, whether code, a requirement, a design, a document or a test, that makes it fail its requirements and must be repaired or replaced; it is one possible outcome of investigating a reported problem, and it carries a set of attributes (where it was found, how bad it is, where it lies, what changed to fix it) that make defects countable and manageable.

In the wording a student can write in an examination: a defect is an "imperfection or deficiency in a work product where that work product does not meet its requirements or specifications and needs to be either repaired or replaced" (ISO/IEC 23531). It is found by investigating an anomaly, "anything observed in the documentation or operation of a system that deviates from expectations" (IEEE 1012-2024), reported as an incident, an "anomalous or unexpected event or set of events at any time during the life cycle of a project, product, service, or system" (ISO/IEC/IEEE 12207:2026). Not every reported problem is a defect: in Florac's words, "some of the problems reported are not software failures or software defects, and may not even be software product problems." Each problem record carries attributes, "mutually exclusive" in value, such as its finding activity, criticality, problem type, uniqueness, urgency and the work product the defect was found in.

From something odd to a confirmed defect

Chapter Two, on errors, faults and failures, set out the chain from a person's mistake to a defect in a work product to a failure in operation. Defect management runs the chain the other way: it starts from something someone noticed and works back to whether there is a defect at all, and where.

StageThe wordDefinition
Something is noticedAnomaly"anything observed in the documentation or operation of a system that deviates from expectations based on previously verified system, software, or hardware products or reference documents" (IEEE 1012-2024)
It is recordedIncident, and its incident reportan "anomalous or unexpected event or set of events at any time during the life cycle"; the report is "documentation of the occurrence, nature, and status of an incident" (ISO/IEC/IEEE 29119-2)
It is investigatedProblema "difficulty, uncertainty, or otherwise realised and undesirable condition or situation that is investigated and can receive corrective action" (ISO/IEC/IEEE 12207:2026); in issue management, the "root cause of one or more incidents" (ISO/IEC 23531)
A cause is confirmed in a work productDefectan "imperfection or deficiency in a work product where that work product does not meet its requirements or specifications and needs to be either repaired or replaced" (ISO/IEC 23531)

The last clause of the definition matters: a defect is something that "needs to be either repaired or replaced". Deciding that is a judgement, made during the investigation, and the answer may be no.

munotes.in421

What a Defect Is

Not every reported problem is a defect

Florac's report for the SEI's Software Metrics Definition Working Group is plain about it: "Problems are unsatisfactory encounters with the software; consequently, some of the problems reported are not software failures or software defects, and may not even be software product problems." It sorts every problem into one of three subtypes, and each subtype into values:

  • Software defects, by the work product that contains them: a requirements defect, "A mistake made in the definition or specification of the customer needs for a software product"; a design defect; a code defect; a document defect; a test case defect, where "A mistake in the test case causes the software product to give an unexpected result"; and a defect in another work product, such as a test tool or a configuration library.
  • Other problems, with no evidence of a software defect: a hardware problem, an operating system problem, a user mistake, an operations mistake, or a new requirement or enhancement outside the agreed requirements.
  • Undetermined: not repeatable or cause unknown, or not yet evaluated.

A separate attribute, uniqueness, marks a report as original or duplicate; a duplicate is one where "The problem or defect has been previously discovered."

Two consequences follow for a tester. A failed test is not automatically a product defect: the test case itself may be wrong, which is Florac's test case defect and the false positive of Chapter Two, on errors, faults and failures. And a defect need not be in code: requirements, designs and documents are work products, and their defects count.

The attributes a defect carries

Florac's framework asks "who, what, why, when, where, and how" of every problem, and turns the answers into attributes with defined values: "In developing the attribute values, we are careful to ensure they are mutually exclusive; that is, any given problem may have one, and only one, value for each attribute." The main ones, filled in for DR-311, the major defect ExamReg's acceptance testing found (Chapter Forty-Two, on acceptance testing):

AttributeThe question it answers (Florac)DR-311
Identification"What software product or software work product is involved?"ExamReg 2.0, the hall ticket
Finding activity"What activity discovered the problem or defect?"User acceptance testing
Finding mode"How was the problem or defect found?"By running the software
Criticality"How critical or severe is the problem or defect?"Major
Problem status"What work needs to be done to dispose of the problem?"Open, accepted under waiver W-07
Problem type"What is the nature of the problem? If a defect, what kind?"Software defect: design
Uniqueness"What is the similarity to previous problems or defects?"Original
Urgency"What urgency or priority has been assigned?"Fix in release 2.1
Environment"Where was the problem discovered?"A basic Android phone
Timingwhen reported, discovered and correctedReported during UAT; not yet corrected
Originator"Who reported the problem?"A student in the acceptance test group
Defects found in"What software artifacts caused or contain the defect?"The hall ticket's PDF generator
Changes made to"What software artifacts were changed to correct the defect?"None yet
munotes.in422

What a Defect Is

Two of the attributes look alike and are not. Criticality "is a measure of the disruption a problem gives users when they encounter it", and is normally given by whoever reported it. Urgency "is the degree of importance that the evaluation, resolution, and closure of a problem is given by the organization charged with executing the problem management process", assigned by the team that fixes it. They are the severity and the priority that Chapter Seventy-Six, on writing a defect report, keeps apart.

Worked example: how many defects did release 2.0 have?

Release 2.0 of ExamReg drew 250 problem reports, from reviewers, testers and, after release, students. Each was investigated and classified by problem type and uniqueness. The program counts them under Florac's subtypes, and then answers one simple question under four definitions.

# ExamReg release 2.0's problem reports, classified by Florac's attributes (FINDINGS 5.2.3):
# (problem type, uniqueness) -> number of reports
reports = {("requirements defect", "original"): 28, ("design defect", "original"): 40,
           ("code defect", "original"): 120, ("operational document defect", "original"): 12,
           ("test case defect", "original"): 9, ("code defect", "duplicate"): 18,
           ("user mistake", "original"): 7, ("operations mistake", "original"): 3,
           ("new requirement/enhancement", "original"): 6,
           ("not repeatable/cause unknown", "value not identified"): 5,
           ("value not identified", "value not identified"): 2}

PRODUCT = {"requirements defect", "design defect", "code defect", "operational document defect"}
SOFTWARE = PRODUCT | {"test case defect", "other work product defect"}
UNDETERMINED = {"not repeatable/cause unknown", "value not identified"}

def subtype(problem_type):
    return ("software defect" if problem_type in SOFTWARE
            else "undetermined" if problem_type in UNDETERMINED else "other problem")

def count(rule):
    return sum(n for key, n in reports.items() if rule(*key))

print("by Florac's subtypes:")
for s in ("software defect", "other problem", "undetermined"):
    print(f"   {s:<16} {count(lambda t, u: subtype(t) == s):>4}")

print("'how many defects did release 2.0 have?', by four definitions:")
definitions = {
    "every problem report": lambda t, u: True,
    "software defects, originals and duplicates": lambda t, u: t in SOFTWARE,
    "original software defects (product and tests)": lambda t, u: t in SOFTWARE and u == "original",
    "original defects in the product itself": lambda t, u: t in PRODUCT and u == "original",
}
for name, rule in definitions.items():
    print(f"   {name:<47} {count(rule):>4}")

print("the product's defects by the work product they were in:")
for t in sorted(PRODUCT, key=lambda t: -reports[(t, "original")]):
    print(f"   {t:<30} {reports[(t, 'original')]:>4}")
munotes.in423

What a Defect Is

by Florac's subtypes:
   software defect   227
   other problem      16
   undetermined        7
'how many defects did release 2.0 have?', by four definitions:
   every problem report                             250
   software defects, originals and duplicates       227
   original software defects (product and tests)    209
   original defects in the product itself           200
the product's defects by the work product they were in:
   code defect                     120
   design defect                    40
   requirements defect              28
   operational document defect      12

Four honest answers, from 250 down to 200, to the same question. A report that says release 2.0 had 250 defects counts reports, including duplicates, users' mistakes and enhancement requests; one that says 200 counts original defects in the product, which is the figure this book has used since Chapter Twenty-Four, on quality control and quality assurance, and the one every defect density in Module 2 divides. The difference is not a matter of accuracy but of definition, exactly the problem Park found with lines of code (Chapter Sixty-Seven, on size metrics). Florac's report exists to fix it the same way: with attributes whose values are exclusive, and a checklist that states which values a count includes.

The last lines show where the product's defects were: 120 in code, but 80 in requirements, designs and documents together. Two defects in five were not in the code at all, which is why reviews of requirements and designs (Chapters Ninety-Six to Ninety-Eight) are part of finding defects.

What it does not mean

A problem report is not a defect. It is a claim that something is wrong, which investigation confirms or rejects.

A failed test is not always a product defect. The test case may be wrong, and Florac counts that as a test case defect.

A defect is not only in code. Requirements, designs, documents and tests are work products with defects of their own.

Criticality is not urgency. One is the disruption to users, reported by the originator; the other is the order of work, set by the team.

Quick revision

  • Defect: "imperfection or deficiency in a work product where that work product does not meet its requirements or specifications and needs to be either repaired or replaced" (ISO/IEC 23531).
  • Anomaly (IEEE 1012-2024), incident and problem (ISO/IEC/IEEE 12207:2026), incident report (ISO/IEC/IEEE 29119-2): the stages before a defect is confirmed.
  • Florac (SEI 1992): problems are "unsatisfactory encounters with the software"; subtypes software defect (requirements, design, code, document, test case, other work product), other problem (hardware, operating system, user mistake, operations mistake, enhancement), undetermined; uniqueness: original or duplicate.
  • Attributes, one value each: identification, finding activity, finding mode, criticality, status, type, uniqueness, urgency, environment, timing, originator, defects found in, changes made to.
  • Criticality (disruption to users) is not urgency (order of work).
  • Worked example: 250 reports; 227 software defects, 16 other problems, 7 undetermined; "how many defects" is 250, 227, 209 or 200 by definition; 80 of the 200 were outside code.
munotes.in424

What a Defect Is

Test yourself

1. Define a defect. How does it differ from an anomaly and an incident? A defect is an imperfection or deficiency in a work product that does not meet its requirements or specifications and needs to be repaired or replaced. An anomaly is anything observed that deviates from expectations, and an incident is the unexpected event recorded; both are what is noticed and reported, before investigation shows whether a defect is behind them.

2. Why are some problem reports not defects? Give four examples. Because a problem is any unsatisfactory encounter with the software, and investigation may find the cause elsewhere. Examples: a duplicate of a defect already reported; a user's mistake in using the software; an operations mistake such as a wrongly set server clock; a request for a new feature outside the requirements; a problem that cannot be repeated.

3. In which work products can defects be found? Give an example of each from ExamReg. Requirements (a fee rule that does not say what happens on the last date itself), design (a hall ticket generated in a way too slow for a basic phone), code (a late fee charged at the wrong boundary), documents (a user guide that gives the wrong last date), and test cases (a test that expects the wrong fee).

4. List the attributes a defect record carries. Identification of the product; finding activity; finding mode; criticality; problem status; problem type; uniqueness; urgency; environment; timing (reported, discovered, corrected); originator; the work product the defect was found in; the work products changed to correct it; related changes; projected availability and release of the fix.

5. Distinguish criticality from urgency. Criticality measures how much the problem disrupts users, and is normally given by whoever reported it. Urgency is the priority the problem management organisation gives to evaluating, fixing and closing it, and decides the order of work.

6. Why can two reports of "the number of defects" for the same release differ, and how is that prevented? Because they may count different things: all problem reports, all software defects including duplicates and test case defects, or only original defects in the product. It is prevented by recording defects with attributes whose values are mutually exclusive and by stating, as a checklist, which values the count includes.

Contents This chapter on its own page

munotes.in425

Chapter Seventy-Four

The Defect Life Cycle

Syllabus topic Module 2, "Defect Management: Definition of defects and their lifecycle"

In one line

A reported defect moves through a fixed set of states, from new through assigned, fixed and verified to closed, with side exits for reports that are rejected, duplicated or deferred and a loop back for fixes that fail; each move has a rule about who may make it, and a defect tracker enforces the rules so that no defect is closed unverified or lost.

In the wording a student can write in an examination: the defect life cycle is the set of states a defect report passes through and the transitions allowed between them. The ISTQB syllabus says the defect management process "includes a workflow for handling individual defects or anomalies from their discovery to their closure", comprising "activities to log the reported anomalies, analyze and classify them, decide on a suitable response such as to fix or keep it as it is and finally to close the defect report", and lists typical statuses: "open, deferred, duplicate, waiting to be fixed, awaiting confirmation testing, re-opened, closed, rejected". Florac's generic model has two statuses, Open (with the sub-states recognized, evaluated and resolved) and Closed, each entered when defined criteria are met.

Why a defect needs a life cycle

A defect report is work in progress for several people: a tester who found the failure, someone who decides whether and when to fix it, a developer who fixes it, a tester who checks the fix. Without agreed states, a report can sit unread, be fixed and never retested, or be closed by the person who wants it to go away. Florac puts the point in terms of measurement: "Because problem status is dependent on the amount of information known about the problem, it is important that we define and understand the possible states and the criteria used to change from one state to another".

Florac's generic statuses

The SEI framework keeps the minimum and lets each organisation add detail:

  • Open: "the problem is recognized and some level of investigation and action will be undertaken to resolve it". Its three generic sub-states follow the work:
  • Recognized: "Sufficient data has been collected to permit an evaluation of the problem to be made."
  • Evaluated: enough investigation has been done "to at least determine the problem type".
  • Resolved: "sufficient information is available to satisfy the rules for resolution", which may include "a designed and tested change, change verification with the originator, change control approval, or a causal analysis of the problem".
  • Closed: "the investigation is complete and the action required to resolve the problem has been proposed, accepted, and completed to the satisfaction of all concerned." And: "In some cases, a problem report will be recognized as invalid as part of the recognition process and be closed immediately."
munotes.in426

The Defect Life Cycle

Two rules come with the states. "A problem cannot be in more than one state at any point in time", and since states change, "one must always specify a time when measuring the problem status".

ExamReg's life cycle

ExamReg's team uses nine states, which fill in Florac's generic ones and match the ISTQB syllabus's typical statuses:

ExamReg's defect life cycle: new, assigned, fixed, verified and closed, with rejected, duplicate, deferred and reopened

Figure 74.1 ExamReg's defect life cycle: eleven moves; shaded states are closed

StateMeaningFlorac's status
NewReported with enough detail to investigateOpen: recognized
AssignedAccepted as a defect; a developer owns itOpen: evaluated
DeferredAccepted as a defect; the fix is postponed to a later releaseOpen: evaluated
ReopenedA fix failed, or a closed defect has come backOpen: evaluated
FixedA change has been madeOpen: resolved
VerifiedA tester has confirmed the fix, and the regression tests passOpen: resolved
ClosedFinished, to everyone's satisfactionClosed
RejectedNot a defect: a user's mistake, the software working as specified, a wrong testClosed
DuplicateThe same defect is already reportedClosed

The moves between them, and who may make each:

FromToWhoWhen
NewAssignedTriageIt is a defect, to be fixed now
NewRejectedTriageIt is not a defect
NewDuplicateTriageIt is already reported
NewDeferredTriageIt is a defect, to be fixed in a later release
DeferredAssignedTriageThe later release begins
AssignedFixedDeveloperThe change is made
FixedVerifiedTesterThe confirmation test passes, and so do the regression tests
FixedReopenedTesterThe confirmation test fails
VerifiedClosedTesterNothing is outstanding
ClosedReopenedTesterThe failure comes back
ReopenedAssignedTriageSomeone owns it again

Two moves are missing on purpose. There is no move from Assigned or Fixed straight to Closed, so no defect can be closed without a tester's check of the fix. And the developer cannot move a defect to Verified: whoever fixed it does not decide that it is fixed. Triage is the decision taken on each new report, often by a defect review board; Chapter Seventy-Seven, on tracking defects to closure, describes it.

Worked example: a tracker that enforces the rules

A defect tracker is a state machine, the kind Chapter Fifty-Seven tested, and its main job is to refuse the moves the life cycle forbids. Bugzilla's documentation describes its own workflow in exactly those terms: its configuration page shows every status twice, as the starting status and as the target, and "If the checkbox is checked, then the transition from the left to the top status is legal; if it's unchecked, that transition is forbidden."

The program writes ExamReg's life cycle as such a table, with the role allowed to make each move. It walks DR-205, a rounding defect in the fee page, through a life that includes a failed fix, and then tries three moves the rules forbid.

munotes.in427

The Defect Life Cycle

# ExamReg's defect life cycle: (from, to) -> the role allowed to make the move
MOVES = {("new", "assigned"): "triage",     ("new", "rejected"): "triage",
         ("new", "duplicate"): "triage",    ("new", "deferred"): "triage",
         ("deferred", "assigned"): "triage",
         ("assigned", "fixed"): "developer",
         ("fixed", "verified"): "tester",   ("fixed", "reopened"): "tester",
         ("verified", "closed"): "tester",  ("closed", "reopened"): "tester",
         ("reopened", "assigned"): "triage"}
# each state against Florac's generic statuses (SEI 1992): Open (recognized, evaluated, resolved), Closed
FLORAC = {"new": "open: recognized", "assigned": "open: evaluated", "deferred": "open: evaluated",
          "reopened": "open: evaluated", "fixed": "open: resolved", "verified": "open: resolved",
          "closed": "closed", "rejected": "closed", "duplicate": "closed"}

class Defect:
    def __init__(self, ident):
        self.ident, self.state, self.history = ident, "new", ["new"]

    def move(self, to, role):
        allowed = MOVES.get((self.state, to))
        if allowed is None:
            raise ValueError(f"{self.state} -> {to} is not a move in the life cycle")
        if allowed != role:
            raise PermissionError(f"{self.state} -> {to} is for the {allowed}, not the {role}")
        self.state = to
        self.history.append(to)

d = Defect("DR-205")                       # the fee page rounds a concession fee wrongly
for to, role in [("assigned", "triage"), ("fixed", "developer"), ("reopened", "tester"),
                 ("assigned", "triage"), ("fixed", "developer"), ("verified", "tester"),
                 ("closed", "tester")]:
    d.move(to, role)
print(d.ident, " -> ".join(d.history))
print("   in Florac's terms:", " -> ".join(FLORAC[s] for s in d.history))

for state, to, role in [("assigned", "closed", "developer"), ("fixed", "verified", "developer"),
                        ("duplicate", "assigned", "triage")]:
    e = Defect("DR-3xx")
    e.state = state
    try:
        e.move(to, role)
    except (ValueError, PermissionError) as refusal:
        print(f"refused ({type(refusal).__name__}): {refusal}")
DR-205 new -> assigned -> fixed -> reopened -> assigned -> fixed -> verified -> closed
   in Florac's terms: open: recognized -> open: evaluated -> open: resolved -> open: evaluated -> open: evaluated -> open: resolved -> open: resolved -> closed
refused (ValueError): assigned -> closed is not a move in the life cycle
refused (PermissionError): fixed -> verified is for the tester, not the developer
refused (ValueError): duplicate -> assigned is not a move in the life cycle

DR-205's history is typical of a real defect's: the first fix failed its confirmation test, the report went back to triage and a second fix passed. In Florac's terms it was open for its whole life until the last step, moving back from resolved to evaluated when the first fix failed, which is why counting defects by status always needs a date.

The three refusals are the tracker doing its job. A developer cannot close an assigned defect, because Closed is reachable only through Verified. A developer cannot verify a fix, because that move belongs to a tester. And a duplicate cannot be assigned, because its work belongs to the original report. The two kinds of refusal differ: the first and third moves do not exist in the life cycle at all, while the second exists but belongs to someone else.

munotes.in428

The Defect Life Cycle

Reading the life cycle for testing

Because the life cycle is a state machine, it can be tested like one: every valid move made, every invalid one attempted, as Chapter Fifty-Seven, on state transition testing, did for ExamReg's login. And its states are measurement points. Florac notes that "Measurements of the numbers of problem reports in the various problem states, shown with respect to time, will inform the manager of the rate the project is progressing towards a goal"; how many are open, how long they have been open, and how many were reopened are the tracking and defect metrics of Chapters Seventy-Seven and Seventy-Nine.

What it does not mean

There is no single standard life cycle. Tools and teams name their states differently; the ISTQB syllabus lists typical statuses, and Florac gives only Open and Closed with three generic sub-states.

Fixed is not closed. A fix is closed only after a tester has confirmed it, and the regression tests show it broke nothing.

Rejected is not ignored. A rejected report records why it is not a defect, and its reporter can challenge the decision.

Deferred is not forgotten. A deferred defect is still open, and it is counted as open when the release ships.

Quick revision

  • Workflow from discovery to closure: log, analyse and classify, decide on a response, close (ISTQB); typical statuses: open, deferred, duplicate, waiting to be fixed, awaiting confirmation testing, re-opened, closed, rejected.
  • Florac (SEI 1992): Open (recognized, evaluated, resolved) and Closed; one state at a time; state counts need a date; an invalid report may be closed at recognition.
  • ExamReg: New, Assigned, Deferred, Reopened, Fixed, Verified, Closed, Rejected, Duplicate; eleven moves, each with its role.
  • No move from Assigned or Fixed to Closed; the developer cannot verify.
  • Worked example: DR-205 went new, assigned, fixed, reopened, assigned, fixed, verified, closed; three forbidden moves refused.
  • The life cycle is a state machine: test it like one, and measure the counts in each state over time.

Test yourself

1. What is the defect life cycle? Draw a typical one. The set of states a defect report passes through from its discovery to its closure, and the transitions allowed between them. A typical one: New, then Assigned, Fixed, Verified and Closed, with New also leading to Rejected, Duplicate or Deferred, Deferred leading back to Assigned, and Fixed or Closed leading to Reopened, which leads back to Assigned.

2. Explain Florac's generic problem statuses. Open, meaning the problem is recognised and will be investigated and acted on, with the sub-states Recognized (enough data to evaluate), Evaluated (investigated enough to know the problem type) and Resolved (the rules for resolution satisfied); and Closed, meaning the investigation is complete and the action has been completed to everyone's satisfaction.

munotes.in429

The Defect Life Cycle

3. Why should a developer not be able to move a defect to Verified or Closed? Because verification is an independent check that the fix works and broke nothing else; if the person who made the fix could declare it verified, a fix that fails would be closed unchecked.

4. Distinguish rejected, duplicate and deferred. A rejected report is not a defect, for example a user's mistake or the software behaving as specified; a duplicate reports a defect already reported, and its work belongs to the original; a deferred report is a real defect whose fix has been postponed to a later release, and it stays open.

5. When is a defect reopened? When the confirmation test of its fix fails, or when a failure it caused appears again after the defect was closed.

6. How can a defect tracker's life cycle itself be tested? As a state machine: by making every valid transition, attempting every invalid one, and checking that the tracker refuses moves that do not exist and moves attempted by a role that may not make them.

Contents This chapter on its own page

munotes.in430

Chapter Seventy-Five

The Defect Management Process

Syllabus topic Module 2, "Defect Management: ... Defect management process"

In one line

Defect management is the process that takes every reported problem from the moment it is noticed to a decision and, for real defects, through a fix and its verification to closure, and then asks why the defects happened so that fewer happen next time; its stages are prevention, discovery, recording, triage, resolution, verification, closure and improvement.

In the wording a student can write in an examination: the ISTQB syllabus describes the defect management process as including "a workflow for handling individual defects or anomalies from their discovery to their closure and rules for their classification. The workflow typically comprises activities to log the reported anomalies, analyze and classify them, decide on a suitable response such as to fix or keep it as it is and finally to close the defect report." It warns that "the reported anomalies may turn out to be real defects or something else (e.g., false-positive result, change request)". In testing, ISO/IEC/IEEE 29119-2 names the test incident reporting process, the "dynamic test process for reporting incidents requiring further action that were identified during the test execution process to the relevant stakeholders". The process closes with causal analysis, "analysis of a defect to determine its cause" (ISO/IEC/IEEE 24765), so that defects are prevented as well as removed.

What the process is for

The ISTQB syllabus gives three objectives for defect reports, and they are the process's objectives too:

  • "Provide those responsible for handling and resolving reported defects with sufficient information to resolve the issue"
  • "Provide a means of tracking the quality of the work product"
  • "Provide ideas for improvement of the development and test process"

The first is about each defect, the second about the product, the third about the process that made it. A process that fixes defects but never looks at them together has met only the first.

The stages

1. Prevention. The cheapest defect is the one never made. CMMI's Causal Analysis and Resolution states it plainly: "Reliance on detecting defects and problems after they have been introduced is not cost effective. It is more effective to prevent defects and problems by integrating Causal Analysis and Resolution activities into each phase of the project." Prevention draws on the last stage: what caused last release's defects is what the checklists, reviews and coding standards of this release guard against.

2. Discovery. Defects are found by reviews, static analysis, tests at every level, and users. ISO/IEC/IEEE 29119-3 calls what testing finds a test incident, an "event occurring during the execution of a test that requires investigation". The word investigation is the point: an incident is not yet a defect.

3. Recording. Each incident is reported, with enough information to reproduce and investigate it; Chapter Seventy-Six, on writing a defect report, sets out what that means. "Anomalies may be reported during any phase of the SDLC", the ISTQB syllabus notes, and defects found by static testing are best handled the same way.

munotes.in431

The Defect Management Process

4. Triage. Every new report gets a decision. Is it a defect? Is it new or a duplicate? Should it be fixed now, deferred, or does it describe a new requirement, which belongs to a change control board, a "formally chartered group responsible for reviewing, evaluating, approving, delaying, or rejecting changes to a project" (the PMBOK Guide)? This is where the reports that are "something else" leave the defect stream.

5. Resolution. A developer finds the defect behind the failure (Chapter Forty-Seven's debugging) and changes the work product.

6. Verification. A tester runs the confirmation test, the original failing test on the fixed version, and regression tests around the change (Chapter Thirty-Two, on test levels and test types). A fix that fails is reopened.

7. Closure. The report is closed when the fix is verified, or when it is rejected, found to be a duplicate, or turned into a change request, with the reason recorded.

8. Improvement. The defects are studied together: which kinds, from where, found when, and why. The causes found lead to corrective action, "action to eliminate the cause of a nonconformity and to prevent recurrence", and preventive action, "action to eliminate the cause of a potential nonconformity" (ISO/IEC 19770-1). Chapter Eighty, on using defect data to improve the process, takes this stage further.

Florac's generic problem management system draws the middle of the same process as four activities, collection, evaluation, resolution and closure, which his problem statuses (Chapter Seventy-Four, on the defect life cycle) follow: a report is Recognized once collected, Evaluated once its type is known, Resolved once the rules for resolution are met, and then Closed.

Worked example: release 2.0 through the process

Release 2.0 drew 250 incident reports. ExamReg's triage rules decide what happens to each kind, and the program routes all 250 through them, then follows the accepted defects through resolution, verification and closure, and finally picks the target for causal analysis.

from collections import Counter

# release 2.0's 250 incident reports by Florac's problem type and uniqueness (FINDINGS 5.2.3)
reports = {("requirements defect", "original"): 28, ("design defect", "original"): 40,
           ("code defect", "original"): 120, ("operational document defect", "original"): 12,
           ("test case defect", "original"): 9, ("code defect", "duplicate"): 18,
           ("user mistake", "original"): 7, ("operations mistake", "original"): 3,
           ("new requirement/enhancement", "original"): 6,
           ("not repeatable/cause unknown", "value not identified"): 5,
           ("value not identified", "value not identified"): 2}

def triage(problem_type, uniqueness):
    """ExamReg's triage rules: the decision taken on each new report."""
    if uniqueness == "duplicate":
        return "duplicate: linked to the original report, closed"
    if problem_type in ("user mistake", "operations mistake"):
        return "rejected: not a defect, reason recorded"
    if problem_type == "new requirement/enhancement":
        return "change request: sent to the change control board"
    if problem_type == "test case defect":
        return "test case defect: returned to the test team"
    if problem_type in ("not repeatable/cause unknown", "value not identified"):
        return "returned to the reporter for more information"
    return "accepted as a product defect"

decisions = Counter()
for (problem_type, uniqueness), n in reports.items():
    decisions[triage(problem_type, uniqueness)] += n
print(f"TRIAGE of {sum(decisions.values())} reports")
for decision, n in decisions.most_common():
    print(f"   {n:>3}  {decision}")

accepted = decisions["accepted as a product defect"]
deferred = ["DR-311 (major, waiver W-07)", "DR-318 (minor)"]          # FINDINGS 5.2.4
fixed, reopened = accepted - len(deferred), 14
print(f"RESOLUTION of {accepted} product defects: {fixed} fixed, {len(deferred)} deferred to 2.1:"
      f" {', '.join(deferred)}")
print(f"VERIFICATION of {fixed} fixes: {fixed - reopened} passed the first confirmation test,"
      f" {reopened} reopened and fixed again")
print(f"CLOSED by the end of the count: {fixed} product defects; still open, deferred: {len(deferred)}")

# the loop back into the process: the commonest defect type goes to causal analysis (FINDINGS 5.2)
types = {"input validation": 58, "logic and computation": 44, "interface": 30, "user interface": 24,
         "data and database": 18, "documentation": 12, "performance": 8, "security": 6}
top = max(types, key=types.get)
print(f"IMPROVEMENT: {top} ({types[top]} of {sum(types.values())}) is selected for causal analysis")
munotes.in432

The Defect Management Process

TRIAGE of 250 reports
   200  accepted as a product defect
    18  duplicate: linked to the original report, closed
    10  rejected: not a defect, reason recorded
     9  test case defect: returned to the test team
     7  returned to the reporter for more information
     6  change request: sent to the change control board
RESOLUTION of 200 product defects: 198 fixed, 2 deferred to 2.1: DR-311 (major, waiver W-07), DR-318 (minor)
VERIFICATION of 198 fixes: 184 passed the first confirmation test, 14 reopened and fixed again
CLOSED by the end of the count: 198 product defects; still open, deferred: 2
IMPROVEMENT: input validation (58 of 200) is selected for causal analysis

Read the output as the process's record.

Triage is where one report in five leaves the defect stream. Of 250 reports, 200 were product defects; the other 50 were something else, and each went somewhere definite. Duplicates were linked to their originals, so that the original's fix closes them and the count of defects is not inflated. Users' and operators' mistakes were rejected with the reason recorded, so the reporter learns why. Test case defects went back to the test team, because the fix is to the test, not the product. Enhancement requests went to the change control board, because a new requirement is a decision about scope, not a repair. And unrepeatable reports went back to their reporters for more information, rather than being closed unexamined.

munotes.in433

The Defect Management Process

Resolution fixed 198 defects and deferred 2, the two defects the college accepted at acceptance testing (Chapter Forty-Two): DR-311 under its waiver, and the minor DR-318. Deferred defects stay open, and they appear in the release's report as known defects.

Verification caught 14 fixes that did not work the first time, one in fourteen, and sent them back; all passed at the second attempt. Without that step, 14 defects would have been closed while still present.

Improvement closes the loop: input validation, the commonest defect type at 58 of 200, is selected for causal analysis, the first step of the prevention that starts the next release's process.

Who does what

On ExamReg's project the stages are divided like this; other teams divide them differently, but the verification stays independent of the fix.

StageDone by, at ExamReg
PreventionThe whole team, through reviews, standards and checklists
Discovery and recordingReviewers, testers, users
TriageA defect review board: test lead, development lead, product owner
ResolutionThe developer who owns the work product
VerificationA tester, not the person who made the fix
ClosureThe tester, or the triage board for reports that are not defects
ImprovementThe team together, led by quality assurance

What it does not mean

Defect management is not only fixing. Triage, verification and improvement are stages of the process; fixing is one of eight.

Every report does not become a defect. A fifth of release 2.0's reports were duplicates, mistakes, test errors, requests or unrepeatable, and each needed its own decision.

A fix is not the end. The confirmation and regression tests decide whether a defect is really gone.

Improvement is not optional. Without it, the process removes this release's defects and lets the next release's in by the same causes.

Quick revision

  • Defect management process (ISTQB): a workflow from discovery to closure and rules for classification; log, analyse and classify, decide a response, close.
  • Report objectives (ISTQB): enough information to resolve; tracking product quality; ideas for improving the process.
  • Test incident (ISO/IEC/IEEE 29119-3): an event in test execution requiring investigation; the test incident reporting process (29119-2) reports incidents needing action to stakeholders.
  • Stages: prevention, discovery, recording, triage, resolution, verification, closure, improvement.
  • Triage decisions: accept, reject, duplicate, defer, change request, return for information.
  • Causal analysis (ISO/IEC/IEEE 24765); corrective and preventive action (ISO/IEC 19770-1); CMMI: prevention beats detection.
  • Worked example: 250 reports; 200 product defects; 198 fixed, 2 deferred; 14 fixes reopened; input validation (58) to causal analysis.

Test yourself

1. What is defect management? List its stages. The process of handling reported problems from discovery to closure, and of using what they show to improve the process. Its stages are prevention, discovery, recording, triage, resolution, verification, closure and improvement.

munotes.in434

The Defect Management Process

2. What are the objectives of a defect report? To give those who handle and resolve the defect enough information to resolve it; to provide a means of tracking the quality of the work product; and to provide ideas for improving the development and test process.

3. What happens at triage? Give the possible decisions. Each new report is examined and a decision is taken: accept it as a defect to fix now, defer the fix to a later release, reject it as not a defect with the reason recorded, link it as a duplicate, send it to the change control board as a new requirement, or return it to the reporter for more information.

4. Why is verification a separate stage, and who performs it? Because a fix can fail or break something else; the confirmation test on the fixed version and regression tests around the change show whether the defect is really gone. It is performed by a tester, not by the developer who made the fix.

5. How does defect management prevent defects, not only remove them? By analysing the defects of a release together to find their causes, and taking corrective and preventive actions, such as new checklist items, review focus, coding standards or training, so that the same kinds of defect are not introduced again.

6. In the worked example, what became of the 250 reports? 200 were accepted as product defects, of which 198 were fixed (14 after a failed first fix) and 2 deferred to the next release; 18 were duplicates, 10 were rejected as users' or operators' mistakes, 9 were test case defects returned to the test team, 7 went back to their reporters for more information, and 6 became change requests.

Contents This chapter on its own page

munotes.in435

Chapter Seventy-Six

Writing a Defect Report: Severity and Priority

Syllabus topic Module 2, "Defect Management: ... including defect reporting"

In one line

A defect report must let someone who was not there reproduce the failure, understand why it is wrong, and decide what to do about it; that takes a precise title, exact steps from a known starting point, the expected and the actual result, the environment, and two separate judgements: how bad the failure is (severity) and how soon it must be fixed (priority).

In the wording a student can write in an examination: a defect report "logged during dynamic testing typically includes" (ISTQB v4.0.1) a unique identifier; a title "with a short summary of the anomaly being reported"; the date, the issuing organisation and the author with their role; the test object and test environment; the context, such as the test case being run; a "Description of the failure to enable reproduction and resolution including the test steps that detected the anomaly, and any relevant test logs, database dumps, screenshots, or recordings"; the "Expected results and actual results"; the severity, the "degree of impact" on stakeholders or requirements; the priority to fix; the status; and references, such as to the test case. Severity is the "relative detrimental effect of a failure" (IEEE 982:2024); priority is the order in which it should be fixed. They are set by different people for different reasons, and they do not have to agree.

Who reads a defect report

The ISTQB syllabus gives a defect report three jobs (Chapter Seventy-Five, on the defect management process, listed them), and each has a reader. The developer needs to reproduce the failure and find its cause. The triage board needs to judge how serious it is and when to fix it. And whoever studies the release's defects later needs the report's attributes to classify it. A report written for the reporter's own memory ("Fee wrong!!") serves none of them.

The parts of a report, and how to write each

PartWhat the ISTQB syllabus asks forHow to write it well
Identifier"Unique identifier"Usually assigned by the tool
Title"a short summary of the anomaly being reported"Say what fails, where and when, so the report can be found and understood from its title alone
Date, authorthe date observed, the organisation, the author and their roleSo questions can be asked of the right person
Test object and environment"Identification of the test object and test environment"The build number, server, browser, operating system and device: the failure may happen on only one
Contextthe test case being run, the activity, the technique or data usedLinks the failure to its test, so the fix can be confirmed with it
Descriptiontest steps, logs, dumps, screenshots or recordings "to enable reproduction and resolution"Numbered steps from a stated start; only the steps needed; the evidence attached
Expected and actual"Expected results and actual results"Say where the expected result comes from: the requirement, the test case
Severity"degree of impact"From the agreed scale, judged from the failure's effect
Priority"Priority to fix"Set by whoever decides the order of work
Statusopen, deferred, duplicate, and so onKept by the tool as the life cycle moves (Chapter Seventy-Four)
References"e.g., to the test case"Test case, requirement, related reports
munotes.in436

Writing a Defect Report: Severity and Priority

Three habits make the difference between a report that is fixed and one that bounces back.

  • Make the steps the smallest that still fail. Chapter Forty-Seven, on debugging, reduced a failing case to its essentials; the reporter who does that first saves the developer hours.
  • Report what was seen, not what was guessed. The gateway rejects Rs 249.99 is an observation; the rounding code is broken is a guess, and may be wrong.
  • One failure per report. A report holding two failures is closed when one is fixed, and the other is lost.

Worked example: rewriting a report

A student helping with ExamReg's system test filed a report whose title was Fee wrong!!, whose only step was pay the fee, whose actual result was wrong amount, and whose severity was very high.

The tester who investigated it rewrote it as DR-205, the defect whose life cycle Chapter Seventy-Four followed. The program checks both versions against ISTQB's list of contents and three rules of this chapter's own: a title of at least six words that is not mostly words like wrong and error, more than one step, and a build number in the environment.

# the contents of a defect report logged in dynamic testing (ISTQB CTFL v4.0.1, section 5.5)
FIELDS = ["id", "title", "date", "author", "test object", "environment", "context",
          "steps", "expected", "actual", "severity", "priority", "status", "references"]
SEVERITY = ["critical", "major", "minor", "cosmetic"]          # ExamReg's scale
PRIORITY = ["P1", "P2", "P3", "P4"]
VAGUE_WORDS = {"error", "bug", "problem", "issue", "wrong", "broken", "fails", "not working"}

def check(report):
    """What stops a developer acting on this report."""
    problems = [f"missing: {f}" for f in FIELDS if not report.get(f)]
    title = report.get("title", "")
    words = [w.strip("!.?,").lower() for w in title.split()]
    if title and (len(words) < 6 or sum(w in VAGUE_WORDS for w in words) > len(words) / 3):
        problems.append("title: too short or too vague to find the report again")
    if report.get("steps") and len(report["steps"]) < 2:
        problems.append("steps: one step is rarely enough to reproduce a failure")
    if report.get("severity") and report["severity"] not in SEVERITY:
        problems.append(f"severity: {report['severity']!r} is not on the scale {SEVERITY}")
    if report.get("environment") and "build" not in report["environment"]:
        problems.append("environment: no build number, so nobody knows which version failed")
    return problems

first = {"title": "Fee wrong!!", "author": "Rohan", "steps": ["pay the fee"],
         "actual": "wrong amount", "severity": "very high"}
rewritten = {
    "id": "DR-205",
    "title": "Fee page shows Rs 249.99 instead of Rs 250 for a concession student 3 days late",
    "date": "2026-10-14", "author": "Neha, tester",
    "test object": "ExamReg 2.0, fee payment page",
    "environment": "build 2.0.7 on the test server; Chrome 140 on Windows 11",
    "context": "system test cycle 1, test case TC-FEE-07 (decision table rule R7)",
    "steps": ["log in as test student 2026CS014, who has a concession",
              "fill the exam form with 1 backlog paper, submitted 3 days after the last date",
              "open Pay the fee"],
    "expected": "Rs 250: Rs 150 for the backlog paper and Rs 100 late fee; no form fee",
    "actual": "Rs 249.99, and the payment gateway rejects the amount",
    "severity": "major", "priority": "P1", "status": "new",
    "references": "TC-FEE-07; requirement FEE-2; screenshot fee-page-DR205.png",
}
for name, report in [("the first report", first), ("the rewritten report", rewritten)]:
    problems = check(report)
    print(f"{name}: {len(problems)} problem(s)")
    for p in problems:
        print("   " + p)
munotes.in437

Writing a Defect Report: Severity and Priority

the first report: 12 problem(s)
   missing: id
   missing: date
   missing: test object
   missing: environment
   missing: context
   missing: expected
   missing: priority
   missing: status
   missing: references
   title: too short or too vague to find the report again
   steps: one step is rarely enough to reproduce a failure
   severity: 'very high' is not on the scale ['critical', 'major', 'minor', 'cosmetic']
the rewritten report: 0 problem(s)

The first report has twelve problems, and each is a question the developer would have to come back with. Which build? Which student, which papers, how late? What amount was expected, and why? Very high on what scale? The rewrite answers all of them before they are asked. Its title alone says what fails (the amount shown), where (the fee page), for whom (a concession student 3 days late) and by how much (a paisa). Its steps start from a named test student, and its expected result shows its arithmetic and its source, the fee rule the decision table of Chapter Fifty-Six tested. And its context links it to the test case that will confirm the fix.

A program can check a report's form, as this one does; it cannot check its truth. That the steps really reproduce the failure is something only the reporter can make sure of, by following them once more before filing.

Severity and priority

The two fields look alike and answer different questions.

Severity is about the failure. IEEE 982:2024 defines it as the "relative detrimental effect of a failure". Florac calls the same attribute criticality, "a measure of the disruption a problem gives users when they encounter it", normally given by whoever reported it. Bugzilla's Severity field "indicates how severe the problem is", ranging "from blocker ('application unusable') to trivial ('minor cosmetic issue')". ExamReg's scale has four levels: critical, major, minor, cosmetic.

munotes.in438

Writing a Defect Report: Severity and Priority

Priority is about the work. Florac's urgency "determines the order in which problems are evaluated, resolved, and closed", and it is assigned by the organisation that fixes problems. Bugzilla: "The Priority field is used to prioritize bugs, either by the assignee, or someone else with authority to direct their time such as a project manager."

Usually a severe failure is also urgent, but not always, and the exceptions are why both fields exist. ExamReg's four combinations:

High priorityLow priority
High severityA payment charged twice when a student presses Pay again (Chapter Forty-Four, on recovery testing): money is taken wrongly, and it happens every registration season. Fix now.The yearly admin report crashes when asked for more than 10,000 rows: a crash, but it is run once a year, after the results, and a filtered report is a workaround. Fix before it is next needed.
Low severityThe college's name is misspelt on the login page, days before registration opens: nothing is broken, but every student sees it. Fix before registration.DR-318, the hall ticket's font slightly small when printed (Chapter Forty-Two, on acceptance testing): noticed, harmless, deferred to 2.1.

The top-right and bottom-left boxes are the ones a student should be able to explain in an examination: severity is judged from the failure, priority from the failure, the deadline, the number of people affected, the workaround and the cost of fixing now.

What it does not mean

A defect report is not an accusation. It describes a failure and its evidence; it does not name whose mistake it was.

Severity is not priority. A critical failure in a rarely used function can wait; a cosmetic flaw on the first page every student sees cannot.

More detail is not always better. The steps should be the fewest that reproduce the failure; everything else goes in attachments.

A report checker does not make a report true. It finds missing parts; only reproducing the steps shows the report is right.

Quick revision

  • ISTQB v4.0.1 contents: identifier; title; date, organisation, author and role; test object and environment; context; description with steps and evidence; expected and actual results; severity; priority; status; references.
  • Good reports: a searchable title (what, where, when), minimal numbered steps from a stated start, observations not guesses, one failure per report, evidence attached.
  • Severity: "relative detrimental effect of a failure" (IEEE 982:2024); Florac's criticality, from the reporter.
  • Priority: the order of fixing; Florac's urgency, from the organisation that fixes.
  • Four combinations: high and high (double charge), high severity and low priority (yearly report crash), low severity and high priority (misspelt name on the login page), low and low (small font).
  • Worked example: the first report had 12 problems; the rewritten DR-205 had none.
munotes.in439

Writing a Defect Report: Severity and Priority

Test yourself

1. List the contents of a defect report. A unique identifier; a title summarising the anomaly; the date, organisation and author with role; the test object and test environment; the context, such as the test case run; a description with the steps to reproduce and evidence such as logs and screenshots; the expected and actual results; the severity; the priority; the status; and references such as the test case.

2. Distinguish severity from priority, and say who sets each. Severity is the degree of harm the failure causes, judged from its effect on users and requirements, and is usually given by the reporter. Priority is how soon the defect should be fixed relative to others, set by whoever directs the fixing work, such as a project manager or triage board, taking into account deadlines, workarounds and cost.

3. Give an example of a high-severity, low-priority defect and of a low-severity, high-priority one. High severity, low priority: a crash in an annual report that is rarely run and has a workaround. Low severity, high priority: a misspelt college name on the login page just before registration opens.

4. Rewrite the report titled Fee wrong!!, with the single step pay the fee and the actual result wrong amount. DR-205: Fee page shows Rs 249.99 instead of Rs 250 for a concession student 3 days late. Environment: build 2.0.7, Chrome 140 on Windows 11. Steps: log in as test student 2026CS014 (concession); fill the exam form with one backlog paper, submitted 3 days after the last date; open Pay the fee. Expected: Rs 250 (Rs 150 backlog fee and Rs 100 late fee, no form fee). Actual: Rs 249.99, and the gateway rejects it. Severity major, priority P1; test case TC-FEE-07.

5. Why should a defect report contain only one failure? Because the report is resolved and closed as a unit; if it holds two failures, fixing one may close the report while the other remains unfixed and forgotten.

6. Why must the expected result state its source? Because the developer must know whether the tester's expectation is right; an expected result traced to a requirement or test case can be checked, while one without a source may itself be the mistake, a test case defect rather than a product defect.

Contents This chapter on its own page

munotes.in440

Chapter Seventy-Seven

Tracking Defects to Closure

Syllabus topic Module 2, "Defect Management: ... defect reporting and tracking"

In one line

Tracking follows every defect from report to closure and the whole population of defects week by week: how many arrive, how many close, how many stay open and for how long, and which requirements they belong to; the trends say when a product is nearly ready, and the ages say which defects are being neglected.

In the wording a student can write in an examination: defect tracking is the continuing record and review of each defect's status, from its report to its closure, and of the defects together. Florac's framework records the date of each status change, and "These attributes are used to determine status, problem age, and problem arrival rate. This information is also of primary importance in product readiness models to determine when the product is ready for acceptance testing or delivery." The main tracking measures are the arrival rate (new defects per period), the closure rate (defects closed per period), the open count and its trend, and the age of open defects. A defect review board meets regularly to triage new reports and act on old ones. And through traceability, which the ISTQB syllabus asks to be maintained "between the test basis elements, testware associated with these elements (e.g., test conditions, risks, test cases), test results, and defects", each defect is tied to the requirement and test it concerns.

Tracking one defect, and tracking them all

Chapter Seventy-Four, on the defect life cycle, followed one defect, DR-205, through its states. Tracking asks a second kind of question, about all of them at once. Florac puts it in terms of the process: "The rate of problem arrival and the time it takes to process a problem report through closure addresses the efficacy of the problem management process." A single defect can be handled perfectly while the population grows out of control; only the counts over time show it.

The defect review board

Somebody has to look at the numbers and at the reports behind them. On ExamReg's project that is the defect review board: the test lead, the development lead and the exam cell's product owner, meeting twice a week during testing. It does two jobs. It triages every new report, taking the decisions of Chapter Seventy-Five, on the defect management process: accept, reject, duplicate, defer, or send to the change control board, which is "responsible for reviewing, evaluating, approving, delaying, or rejecting changes to a project" (the PMBOK Guide). And it reviews the open list, oldest and most severe first, asking of each why it is still open and what would close it.

Arrival, closure and the open count

Plot the defects reported each week and the defects closed each week, and a test cycle usually shows a shape. Early on, testers find defects faster than developers fix them, and the open count rises. As the product stabilises, fewer new defects arrive, the fixes catch up, and the open count falls. The week in which closures first overtake arrivals is the turning point; the open count's approach to zero, or to a small set of known defects, is what Florac's "product readiness models" look for.

munotes.in441

Tracking Defects to Closure

Worked example: release 2.0's ten test weeks

The program tracks release 2.0's 114 test-phase defects over its ten test weeks, then takes the list of defects open at the end of week 6, the midpoint of testing, and analyses it as the review board did that week.

from collections import Counter
from itertools import accumulate

# release 2.0's ten test weeks (FINDINGS 5.2.5): defects reported and closed each week
new    = [10, 16, 18, 17, 14, 12, 10, 8, 5, 4]
closed = [4, 10, 14, 16, 16, 15, 13, 11, 8, 5]
open_at_end = [a - c for a, c in zip(accumulate(new), accumulate(closed))]

print("week  new  closed  open")
for week, (n, c, o) in enumerate(zip(new, closed, open_at_end), 1):
    print(f"{week:>4} {n:>4} {c:>7} {o:>5}  {'#' * o}")
turn = next(w for w, (n, c) in enumerate(zip(new, closed), 1) if c > n)
print(f"closures first exceed new reports in week {turn}; open at the end: {open_at_end[-1]}")

# the defects open at the end of week 6: (id, severity, requirement area, days open)
week6 = [("DR-221", "major", "FEE", 3), ("DR-224", "minor", "FORM", 5),
         ("DR-226", "cosmetic", "HALLTICKET", 5), ("DR-218", "minor", "FEE", 8),
         ("DR-213", "major", "FORM", 9), ("DR-209", "minor", "LOGIN", 12),
         ("DR-207", "cosmetic", "FORM", 13), ("DR-204", "minor", "FEE", 15),
         ("DR-199", "major", "FEE", 20), ("DR-195", "minor", "FORM", 26),
         ("DR-190", "cosmetic", "HALLTICKET", 31), ("DR-183", "minor", "FEE", 40)]
bands = [("0 to 7 days", 0, 7), ("8 to 14 days", 8, 14), ("15 to 30 days", 15, 30), ("over 30 days", 31, 999)]
print(f"week 6: {len(week6)} open, by age:")
for name, low, high in bands:
    ids = [d for d, s, r, age in week6 if low <= age <= high]
    print(f"   {name:<14} {len(ids):>2}  {' '.join(ids)}")
print("   by requirement area:", ", ".join(f"{r} {n}" for r, n in Counter(r for _, _, r, _ in week6).most_common()))
old_major = [d for d, s, r, age in week6 if age > 14 and s in ("critical", "major")]
print("   for the review board, major and open over 14 days:", " ".join(old_major))
week  new  closed  open
   1   10       4     6  ######
   2   16      10    12  ############
   3   18      14    16  ################
   4   17      16    17  #################
   5   14      16    15  ###############
   6   12      15    12  ############
   7   10      13     9  #########
   8    8      11     6  ######
   9    5       8     3  ###
  10    4       5     2  ##
closures first exceed new reports in week 5; open at the end: 2
week 6: 12 open, by age:
   0 to 7 days     3  DR-221 DR-224 DR-226
   8 to 14 days    4  DR-218 DR-213 DR-209 DR-207
   15 to 30 days   3  DR-204 DR-199 DR-195
   over 30 days    2  DR-190 DR-183
   by requirement area: FEE 5, FORM 4, HALLTICKET 2, LOGIN 1
   for the review board, major and open over 14 days: DR-199
munotes.in442

Tracking Defects to Closure

The trend. The open count climbs for four weeks, to 17, while testers find defects faster than they are fixed, and the bars show the shape at a glance. In week 5 closures (16) overtake new reports (14) for the first time, and from then on the open count falls every week. By week 10 new reports are down to 4 a week and only 2 defects remain open: DR-311 and DR-318, the two the college accepted as known defects at acceptance testing. A falling arrival rate is only good news if testing has not slowed down; the review board checks it against the test progress figures (Chapter Ten, on test reporting) before reading it as stability.

The snapshot. At the end of week 6, twelve defects were open. Seven were under two weeks old, which is the normal flow. Five had been open more than 14 days, and two more than a month; the review board's rule is to ask about every defect over 14 days old, and to act at once on any that is major or critical. That picks out DR-199, a major defect in the fee area, 20 days old. The others old enough to ask about are minor or cosmetic, which may be why they waited, and the board decides for each whether to fix it now or defer it deliberately rather than by neglect.

The requirement view. Five of the twelve belonged to the fee requirements and four to the exam form. Traced to requirements, open defects say which parts of the product are not yet ready, which the ISTQB syllabus names as one use of traceability: it "can be used to evaluate the level of residual risk in a test object". The fee area was the one to watch.

What a tracking report contains

A weekly tracking report for a defect review board answers five questions, each with a number:

  1. How many new defects, and how many closed, this period, and the trend over recent periods?
  2. How many are open, by severity and by status (new, assigned, fixed and awaiting verification, reopened, deferred)?
  3. How old are the open ones, and which exceed the agreed age for their severity?
  4. Which requirements, modules or areas have open defects?
  5. What was decided about each defect that needed a decision?
munotes.in443

Tracking Defects to Closure

What it does not mean

Tracking is not reporting once. It is a continuing record, read as a trend; one week's numbers say little.

A falling arrival rate does not prove quality. It may mean the product is stabilising, or that testing has slowed; it is read beside the test progress figures.

An old defect is not always a neglected one. It may be deferred by decision; the point of ageing is to make every such decision explicit.

Zero open defects is not the goal. Known, accepted defects with agreed workarounds, like DR-311 and DR-318, may be released; unknown ones are the danger.

Quick revision

  • Florac: status dates give "status, problem age, and problem arrival rate", and are "of primary importance in product readiness models".
  • Measures: arrival rate, closure rate, open count and its trend, age of open defects.
  • Defect review board: triages new reports and reviews the open list, oldest and most severe first.
  • Traceability (ISTQB): between test basis, testware, results and defects; open defects by requirement show residual risk.
  • Worked example: open count peaked at 17 in week 4; closures overtook arrivals in week 5; 2 open at the end (DR-311, DR-318); week 6: 12 open, 5 over 14 days, one major (DR-199) escalated; fee area 5, form 4.

Test yourself

1. What does defect tracking involve? Keeping a continuing record of every defect's status from report to closure, with the date of each change, and reviewing the defects together over time: how many arrive and close each period, how many are open, how old they are, and which requirements they affect, so that neglected defects and trends are seen and acted on.

2. What are the arrival rate and the closure rate, and what does their crossing mean? The arrival rate is the number of new defects reported per period; the closure rate the number closed per period. When closures first exceed arrivals, the open count starts to fall, a sign that the product is stabilising, provided testing effort has not dropped.

3. Why track the age of open defects? Because a defect can sit open without anyone deciding anything about it. Age shows which defects have waited longest; with a rule such as reviewing every defect open more than two weeks, each old defect is either fixed or deferred by an explicit decision.

4. What is a defect review board, and what does it do? A group, typically the test lead, development lead and product owner, that meets regularly during testing to triage every new defect report and to review the open defects, deciding for each whether to fix, defer, reject or escalate it.

munotes.in444

Tracking Defects to Closure

5. How does traceability help in tracking defects? By linking each defect to the requirement and test case it concerns, it shows which requirements still have open defects, and so which parts of the product are not ready and where residual risk remains; and it tells the tester which test confirms a fix.

6. In the worked example, which defect did the board escalate at week 6, and why? DR-199, because it was the only major defect that had been open more than 14 days; the other old defects were minor or cosmetic and were considered for fixing or deliberate deferral.

Contents This chapter on its own page

munotes.in445

Chapter Seventy-Eight

A Defect Tracker at Work: Bugzilla

Syllabus topic Practical, "Bug Tracking and Defect Lifecycle Management"

In one line

Bugzilla is an open-source defect tracker: each bug is a record with a summary, product and component, severity and priority, an owner and a status, and the tool moves it through a workflow that an administrator can configure, refusing moves the workflow forbids, and turns the collection of bugs into tables and charts for tracking.

In the wording a student can write in an examination: in Bugzilla, a bug's Status and Resolution "define exactly what state the bug is in", running "from not even being confirmed as a bug, through to being fixed and the fix confirmed by Quality Assurance" (its user guide). The default workflow has the statuses UNCONFIRMED, CONFIRMED, IN_PROGRESS, RESOLVED and VERIFIED, and a resolved bug takes one of the resolutions FIXED, DUPLICATE, WONTFIX, WORKSFORME and INVALID. Its Severity field "indicates how severe the problem is", from blocker to trivial, and its Priority field "is used to prioritize bugs", with default values "P1 to P5". The workflow "can be customized". Bugzilla's reports are "a view of the current state of the bug database", as tables or graphs, and its charts "a view of the state of the bug database over time".

The practical's exercise

MU's practical syllabus sets the task in one line: "Log software defects using Bugzilla. Assign severity and priority levels. Track defect lifecycle stages and generate defect summary reports." Its course outcome adds the aim: to "Manage and track software defects using Bugzilla and prepare structured defect lifecycle reports." Everything in the three chapters before this one (the life cycle, the report, the tracking) is what the tool is for; this chapter shows how Bugzilla embodies it.

A bug's fields

Bugzilla's user guide describes the fields of the screen that shows a bug. The main ones:

FieldWhat the user guide says
Summary"A one-sentence summary of the problem, displayed in the header next to the bug number."
Status and Resolution"These define exactly what state the bug is in"
Product and Component"Bugs are divided up by Product and Component, with a Product having one or more Components in it."
Version"used to indicate the version(s) affected by the bug report"
Hardware"These indicate the computing environment where the bug was found."
Priority"used to prioritize bugs, either by the assignee, or someone else with authority to direct their time such as a project manager"; the defaults are P1 to P5
Severity"indicates how severe the problem is", from blocker ("application unusable") to trivial ("minor cosmetic issue"); it can also mark "an enhancement request"
Target Milestone"A future version by which the bug is to be fixed."
Assigned To"The person responsible for fixing the bug."
munotes.in446

A Defect Tracker at Work: Bugzilla

Compare the list with the ISTQB contents of a defect report in Chapter Seventy-Six, on writing a defect report: the summary is the title, product, component, version and hardware identify the test object and environment, and the severity and priority are the same two judgements. The steps, expected and actual results go into the bug's description and comments.

The values in most of these lists belong to each installation. The administration guide: "Legal values for the operating system, platform, bug priority and severity, and custom fields of type Drop Down and Multiple-Selection Box (see Custom Fields), as well as the list of valid bug statuses and resolutions, can be customized from the same interface." ExamReg's installation uses the college's own four severities, critical, major, minor and cosmetic, in place of the defaults.

The default workflow

The user guide says that "The life cycle of a bug, also known as workflow, is customizable to match the needs of your organization", and draws the default as a diagram. Read off that diagram:

FromToThe diagram's label
(new bug)UNCONFIRMED"Bug is filed by a non-empowered user in a product where the UNCONFIRMED state is enabled"
(new bug)CONFIRMED(a bug filed directly as confirmed)
UNCONFIRMEDCONFIRMED"Bug determined to be present"
CONFIRMEDIN_PROGRESS"Developer is working on the bug"
IN_PROGRESSCONFIRMED"Developer stops work on bug"
IN_PROGRESSRESOLVED"Fix checked in"
RESOLVEDVERIFIED"QA verifies that the solution works"
RESOLVEDCONFIRMED"QA not satisfied with the solution"
VERIFIEDCONFIRMED"Fix turns out to be wrong"
UNCONFIRMEDRESOLVED"Bug is not fixable (e.g because it is invalid)"
VERIFIEDUNCONFIRMED"Bug is reopened, was never confirmed"

The diagram also draws, without labels, moves from CONFIRMED to RESOLVED, from UNCONFIRMED to IN_PROGRESS and from RESOLVED back to UNCONFIRMED.

Set it beside ExamReg's own life cycle in Chapter Seventy-Four, on the defect life cycle, and the ideas are the same with other names. UNCONFIRMED is a new report not yet triaged; CONFIRMED is an accepted defect; IN_PROGRESS is being fixed; RESOLVED covers both the fixed state and the three triage exits, with the resolution saying which (FIXED, or DUPLICATE, WONTFIX, WORKSFORME, INVALID); VERIFIED is the tester's confirmation; and the arrows back to CONFIRMED are reopening. There is no DEFERRED status by default: a deferred bug stays open, with a later Target Milestone.

Bugzilla enforces its workflow as a table of legal moves. The administration guide describes the configuration page: every status appears "first on the left for the starting status, and on the top for the target status in the transition. If the checkbox is checked, then the transition from the left to the top status is legal; if it's unchecked, that transition is forbidden."

munotes.in447

A Defect Tracker at Work: Bugzilla

Worked example: six ExamReg bugs through the workflow

The program models the default workflow exactly as the diagram draws it. It enforces the transitions and one more rule from the diagram: RESOLVED is the status that carries a resolution. It logs six ExamReg bugs, moves them as the practical asks, tries two forbidden moves, and produces two defect summary reports.

from collections import Counter

# Bugzilla's default workflow, as the diagram in its user guide draws it
WORKFLOW = {"UNCONFIRMED": {"CONFIRMED", "IN_PROGRESS", "RESOLVED"},
            "CONFIRMED":   {"IN_PROGRESS", "RESOLVED"},
            "IN_PROGRESS": {"CONFIRMED", "RESOLVED"},
            "RESOLVED":    {"CONFIRMED", "VERIFIED", "UNCONFIRMED"},
            "VERIFIED":    {"CONFIRMED", "UNCONFIRMED"}}
RESOLUTIONS = {"FIXED", "DUPLICATE", "WONTFIX", "WORKSFORME", "INVALID"}
OPEN = {"UNCONFIRMED", "CONFIRMED", "IN_PROGRESS"}

class Bug:
    def __init__(self, number, summary, component, severity, priority, status="CONFIRMED"):
        self.number, self.summary, self.component = number, summary, component
        self.severity, self.priority, self.status, self.resolution = severity, priority, status, ""

    def change(self, status, resolution=""):
        if status not in WORKFLOW[self.status]:
            raise ValueError(f"bug {self.number}: {self.status} -> {status} is not in the workflow")
        if status == "RESOLVED" and resolution not in RESOLUTIONS:
            raise ValueError(f"bug {self.number}: RESOLVED needs one of {sorted(RESOLUTIONS)}")
        if status != "RESOLVED" and resolution:
            raise ValueError(f"bug {self.number}: only RESOLVED takes a resolution")
        self.status = status
        if status == "RESOLVED":
            self.resolution = resolution
        elif status in OPEN:
            self.resolution = ""                 # reopening clears the old resolution
        # VERIFIED keeps the resolution the bug was resolved with

bugs = [Bug(1, "Payment charged twice when Pay is pressed again", "Fee Payment", "critical", "P1"),
        Bug(2, "Fee shows Rs 249.99 for a concession student 3 days late", "Fee Payment", "major", "P1"),
        Bug(3, "Hall ticket takes 4.2 s to open on a basic phone", "Hall Ticket", "major", "P2"),
        Bug(4, "Login accepts a roll number of only spaces", "Login", "minor", "P2"),
        Bug(5, "Exam form rejects the name D'Souza", "Exam Form", "major", "P1"),
        Bug(6, "Hall ticket font small", "Hall Ticket", "cosmetic", "P4", status="UNCONFIRMED")]
b = {bug.number: bug for bug in bugs}

b[1].change("IN_PROGRESS"); b[1].change("RESOLVED", "FIXED"); b[1].change("VERIFIED")
b[2].change("IN_PROGRESS"); b[2].change("RESOLVED", "FIXED")
b[2].change("CONFIRMED")                           # QA not satisfied with the solution
b[2].change("IN_PROGRESS"); b[2].change("RESOLVED", "FIXED"); b[2].change("VERIFIED")
b[4].change("RESOLVED", "FIXED")
b[5].change("RESOLVED", "DUPLICATE")               # the same defect as an earlier report
b[6].change("CONFIRMED")                           # bug determined to be present
try:
    b[3].change("VERIFIED")
except ValueError as refusal:
    print("refused:", refusal)
try:
    b[4].change("VERIFIED", "WONTFIX")
except ValueError as refusal:
    print("refused:", refusal)

print(f"{'bug':>3}  {'status':<12}{'resolution':<11}{'severity':<9}{'pri':<4}summary")
for bug in bugs:
    print(f"{bug.number:>3}  {bug.status:<12}{bug.resolution:<11}{bug.severity:<9}{bug.priority:<4}{bug.summary}")

print("summary:", ", ".join(f"{s} {n}" for s, n in sorted(Counter(bug.status for bug in bugs).items())),
      "| resolutions:", ", ".join(f"{r} {n}" for r, n in sorted(Counter(x.resolution for x in bugs if x.resolution).items())))

# a tabular report, as Bugzilla draws one: severity against component, for the open bugs
open_bugs = [bug for bug in bugs if bug.status in OPEN]
components = sorted({bug.component for bug in bugs})
print(f"open bugs, severity by component: {len(open_bugs)} open")
print(f"{'':<10}" + "".join(f"{c:>13}" for c in components))
for sev in ("critical", "major", "minor", "cosmetic"):
    row = Counter(bug.component for bug in open_bugs if bug.severity == sev)
    print(f"{sev:<10}" + "".join(f"{row[c]:>13}" for c in components))
munotes.in448

A Defect Tracker at Work: Bugzilla

refused: bug 3: CONFIRMED -> VERIFIED is not in the workflow
refused: bug 4: only RESOLVED takes a resolution
bug  status      resolution severity pri summary
  1  VERIFIED    FIXED      critical P1  Payment charged twice when Pay is pressed again
  2  VERIFIED    FIXED      major    P1  Fee shows Rs 249.99 for a concession student 3 days late
  3  CONFIRMED              major    P2  Hall ticket takes 4.2 s to open on a basic phone
  4  RESOLVED    FIXED      minor    P2  Login accepts a roll number of only spaces
  5  RESOLVED    DUPLICATE  major    P1  Exam form rejects the name D'Souza
  6  CONFIRMED              cosmetic P4  Hall ticket font small
summary: CONFIRMED 2, RESOLVED 2, VERIFIED 2 | resolutions: DUPLICATE 1, FIXED 3
open bugs, severity by component: 2 open
              Exam Form  Fee Payment  Hall Ticket        Login
critical              0            0            0            0
major                 0            0            1            0
minor                 0            0            0            0
cosmetic              0            0            1            0

The lives of the bugs. Bug 1, the double charge, went straight through: in progress, resolved as fixed, verified. Bug 2, the rounding defect Chapter Seventy-Four followed as DR-205, went through twice: its first fix was resolved, QA was "not satisfied with the solution", and the bug went back to CONFIRMED, then round again to VERIFIED. Bug 4 is fixed but not yet verified, which is where a tracker makes waiting work visible. Bug 5 was resolved as a DUPLICATE, so its work belongs to the earlier report. Bug 6 was filed by a student as UNCONFIRMED and triaged to CONFIRMED. Bug 3, the slow hall ticket, is confirmed and waiting, which in a real installation would carry a Target Milestone of 2.1.

The refusals. A confirmed bug cannot jump to VERIFIED, because the diagram has no such arrow: nothing can be verified that was never resolved. And a bug cannot be given a resolution on the way to VERIFIED, because only RESOLVED takes one.

The reports. The summary line is the practical's "defect summary report" in its simplest form: two bugs each confirmed, resolved and verified; three fixes and a duplicate. The table is the kind of report the user guide describes, a set of bugs chosen by a search and plotted by one attribute against another; its example is to "plot their severity against their component to see which component had had the largest number of bad bugs reported against it." Here the only open bugs are both in the Hall Ticket component, one major and one cosmetic: that is where the release's remaining risk sits.

Reports and charts in Bugzilla

Bugzilla offers "two more ways of viewing sets of bugs" besides the list: reports, which "give different views of the current state of the database", and charts, which "plot the changes in particular sets of bugs over time". A report can be shown as a table, as CSV, or as a bar, line or pie chart. A chart is made of data sets, each a saved search counted over time; Bugzilla can "Sum a number of data sets (e.g. you could Sum data sets representing RESOLVED, VERIFIED and CLOSED in a particular product to get a data set representing all the resolved bugs in that product.)". The arrival and closure trends of Chapter Seventy-Seven, on tracking defects to closure, are exactly such charts.

munotes.in449

A Defect Tracker at Work: Bugzilla

What it does not mean

Bugzilla's statuses are not the only right ones. They are a default; each installation can change them, and ExamReg's own life cycle uses other names for the same ideas.

RESOLVED does not mean fixed. A bug resolved as DUPLICATE, WONTFIX, WORKSFORME or INVALID is resolved without a fix; the resolution says which.

A tracker does not manage defects by itself. It enforces the workflow and keeps the record; triage, fixing and verification are still people's decisions.

The program is not Bugzilla. It models the documented workflow so that its rules can be run and checked here; the practical drives the real tool.

Quick revision

  • Bugzilla fields: summary, status and resolution, product and component, version, hardware, priority (P1 to P5), severity (blocker to trivial), target milestone, assigned to.
  • Default statuses: UNCONFIRMED, CONFIRMED, IN_PROGRESS, RESOLVED, VERIFIED; resolutions: FIXED, DUPLICATE, WONTFIX, WORKSFORME, INVALID.
  • Main path: CONFIRMED, IN_PROGRESS ("Developer is working on the bug"), RESOLVED ("Fix checked in"), VERIFIED ("QA verifies that the solution works"); back to CONFIRMED when QA is not satisfied or the fix turns out to be wrong.
  • The workflow and the legal values of priority, severity, statuses and resolutions are customisable; the workflow is a table of legal transitions.
  • Reports: the current state, as tables or graphs; charts: data sets over time.
  • Worked example: six bugs; two forbidden moves refused; summary 2 confirmed, 2 resolved, 2 verified; open bugs both in Hall Ticket.

Test yourself

1. Describe the default life cycle of a bug in Bugzilla. A bug filed by a user who cannot confirm bugs starts UNCONFIRMED; once it is determined to be present it becomes CONFIRMED (bugs filed by empowered users start there). When a developer works on it, it is IN_PROGRESS; when the fix is checked in, it is RESOLVED with the resolution FIXED; when QA verifies the fix, it is VERIFIED. If QA is not satisfied, or the fix turns out to be wrong, it returns to CONFIRMED. A bug can also be RESOLVED without a fix, as DUPLICATE, WONTFIX, WORKSFORME or INVALID.

munotes.in450

A Defect Tracker at Work: Bugzilla

2. What is the difference between a status and a resolution in Bugzilla? The status says where the bug is in the workflow (unconfirmed, confirmed, in progress, resolved, verified); the resolution, given when the bug is resolved, says how it was disposed of: fixed, duplicate, will not fix, works for me, or invalid.

3. List the main fields of a Bugzilla bug. Summary; status and resolution; product and component; version; hardware (platform and operating system); priority; severity; target milestone; the person it is assigned to; and a QA contact, with the description and comments holding the steps and results.

4. How does Bugzilla let an organisation adapt it? An administrator can customise the workflow, choosing which transitions between statuses are legal on a table of starting and target statuses, and can customise the legal values of fields such as priority, severity, operating system and platform, and the lists of statuses and resolutions.

5. What reports does Bugzilla provide, and how would you produce a defect summary report for the practical? Tabular and graphical reports of the current state of the bug database, and charts of chosen sets of bugs over time. A defect summary report counts the bugs by status and resolution, and tabulates open bugs by severity against component, to show where the remaining defects are.

6. How does Bugzilla's workflow correspond to the generic defect life cycle? UNCONFIRMED is a new, untriaged report; CONFIRMED is an accepted defect; IN_PROGRESS is being fixed; RESOLVED with FIXED is fixed, and RESOLVED with DUPLICATE, WONTFIX, WORKSFORME or INVALID corresponds to the triage exits; VERIFIED is the tester's confirmation; moves back to CONFIRMED are reopening. Deferral is expressed by a later target milestone rather than a status.

Contents This chapter on its own page

munotes.in451

Chapter Seventy-Nine

Defect Metrics

Syllabus topic Module 2, "Defect Management: ... Metrics related to defects"

In one line

Defect metrics turn a release's defect records into answers: how dense the defects are and where, how many were caught before users met them and by which activity, how many slipped past the review meant to catch them, how long they survived, how many fixes failed, and how much triage effort went on reports that were not defects.

In the wording a student can write in an examination: the main metrics related to defects are defect density, the "number of defects per unit of product size" (ISO/IEC/IEEE 24765); defect removal efficiency (DRE), the defects found before release as a percentage of all defects found, before and after, which the ISTQB syllabus lists among defect metrics as the "defect detection percentage"; the effectiveness of each stage, which Fagan defined for inspections as "Errors found by an inspection" divided by "Total errors in the product before inspection"; leakage, the defects that escape the activity meant to catch them; defect age, how long or how many stages a defect survives; the reopen rate, the fixes that fail their confirmation test; the share of reports that are not defects; and a severity index, a weighted count whose weights are a policy, not a measurement.

One dataset, many questions

Chapter Sixty-Eight, on quality, process and test metrics, computed release 2.0's overall defect density and its defects found per hour of each activity. This chapter goes further into the defects themselves, and it needs one more table: for each of the 200 defects, the work product it was in and the stage that found it.

OriginReq reviewDesign reviewCode reviewUnitIntegrationSystemAcceptanceAfter releaseTotal
Requirements16311123128
Design02153441240
Code002840262033120
Documents0000123612
Total1624344432281012200

The row totals are the problem types of Chapter Seventy-Three, on what a defect is; the column totals are the activities that found them, used since Chapter Twenty-Four, on quality control and quality assurance. The table adds only the link between them.

Removal efficiency, overall and stage by stage

Defect removal efficiency asks the question every release is judged by: of all the defects it contained, what share did the team catch before users did? It can only be computed after release, once users have had time to find what testing missed, and it is always provisional: a defect found next year lowers it.

Stage effectiveness asks the same question of each activity. Fagan defined it for inspections: "Error detection efficiency" is "Errors found by an inspection" divided by the "Total errors in the product before inspection". The denominator is the hard part: the errors present at a stage are those already in the work products it can see, minus those earlier stages removed, and it is known only in hindsight, when later stages have found the rest.

munotes.in452

Defect Metrics

Worked example 1: efficiency, leakage and age

STAGES = ["req review", "design review", "code review", "unit", "integration",
          "system", "acceptance", "after release"]
# where release 2.0's defects came from (rows) and which stage found them (columns): FINDINGS 5.2.6
FOUND = {"requirements": [16, 3, 1, 1, 1, 2, 3, 1],
         "design":       [0, 21, 5, 3, 4, 4, 1, 2],
         "code":         [0, 0, 28, 40, 26, 20, 3, 3],
         "documents":    [0, 0, 0, 0, 1, 2, 3, 6]}
BORN = {"requirements": 0, "design": 1, "code": 2, "documents": 2}   # the first stage that can see them

total = sum(sum(row) for row in FOUND.values())
delivered = sum(row[-1] for row in FOUND.values())
print(f"defect removal efficiency: {total - delivered} of {total} found before release"
      f" = {100 * (total - delivered) / total:.1f}%")

print(f"{'stage':<14}{'present':>8}{'found':>6}{'effectiveness':>15}")
found_so_far = 0
for i, stage in enumerate(STAGES[:-1]):
    present = sum(sum(row) for o, row in FOUND.items() if BORN[o] <= i) - found_so_far
    found = sum(row[i] for row in FOUND.values())
    print(f"{stage:<14}{present:>8}{found:>6}{100 * found / present:>14.1f}%")
    found_so_far += found

print("leakage: defects that escaped the review of the work product they were in")
for origin, stage in (("requirements", 0), ("design", 1), ("code", 2)):
    row = FOUND[origin]
    print(f"   {origin:<13} {sum(row) - row[stage]:>3} of {sum(row):>3} escaped {STAGES[stage]}:"
          f" {100 * (sum(row) - row[stage]) / sum(row):.1f}%")

print("age: how many stages each kind of defect survived before it was found, on average")
for origin, row in FOUND.items():
    stages_survived = sum(n * (i - BORN[origin]) for i, n in enumerate(row))
    print(f"   {origin:<13} {stages_survived / sum(row):.2f}")
defect removal efficiency: 188 of 200 found before release = 94.0%
stage          present found  effectiveness
req review          28    16          57.1%
design review       52    24          46.2%
code review        160    34          21.2%
unit               126    44          34.9%
integration         82    32          39.0%
system              50    28          56.0%
acceptance          22    10          45.5%
leakage: defects that escaped the review of the work product they were in
   requirements   12 of  28 escaped req review: 42.9%
   design         19 of  40 escaped design review: 47.5%
   code           92 of 120 escaped code review: 76.7%
age: how many stages each kind of defect survived before it was found, on average
   requirements  1.68
   design        1.40
   code          1.49
   documents     4.17

Overall. 188 of 200 defects were found before release: a removal efficiency of 94.0 per cent, the figure Chapter Sixty-Six's goal, question, metric model judged against its target of 90.

Stage by stage. The requirements review found 16 of the 28 defects the requirements held, 57.1 per cent. The design review saw the 12 requirements defects that escaped plus the 40 design defects, 52 in all, and caught 24. The code review is the weakest link at 21.2 per cent: 160 defects were present when it ran, most of them in 120 freshly written code units, and it found 34. The four test levels (34.9, 39.0, 56.0 and 45.5 per cent) are exactly Chapter Thirty-One's figures, on a strategic approach to software testing, computed there from the same data; system testing was the most effective single stage of the whole release.

munotes.in453

Defect Metrics

Leakage. Leakage looks at the same data from each work product's side. Of the 28 requirements defects, 12 escaped the requirements review, 42.9 per cent; of the design defects, 47.5 per cent escaped the design review; of the code defects, 76.7 per cent escaped the code review. Every defect that leaks is found later, if at all, by an activity that costs more per defect (Chapter Ninety-Six, on why reviews pay, has the published evidence).

Age. Defects in the documentation survived more than four stages on average, far longer than any other kind: no review looked at the user documents, so they waited for testers and users to trip over them. That single number is a finding with an obvious remedy.

Worked example 2: density by module, reopening, rejection and severity

# release 2.0 (FINDINGS 5.2, 5.2.3, 5.2.4, 5.2.6)
by_module = {"Exam form": (74, 5.5), "Fee payment": (58, 4.0), "Admin reports": (20, 3.5),
             "Hall ticket": (18, 2.5), "Login": (12, 1.5), "Profile": (10, 1.5),
             "Notifications": (8, 1.5)}                         # (defects, KLOC)
print("defect density by module, highest first:")
for module, (d, kloc) in sorted(by_module.items(), key=lambda m: -m[1][0] / m[1][1]):
    print(f"   {module:<14} {d:>3} defects in {kloc:>3} KLOC: {d / kloc:>5.1f} per KLOC")

reports, product, duplicates, rejected = 250, 200, 18, 10
fixes, reopened = 198, 14
print(f"reopen rate: {reopened} of {fixes} fixes = {100 * reopened / fixes:.1f}%")
print(f"reports that were not product defects: {reports - product} of {reports}"
      f" = {100 * (reports - product) / reports:.1f}%"
      f" (duplicates {duplicates}, rejected as mistakes {rejected})")

severity = {"critical": 8, "major": 46, "minor": 98, "cosmetic": 48}
weights = {"critical": 10, "major": 5, "minor": 2, "cosmetic": 1}        # the team's policy
index = sum(severity[s] * weights[s] for s in severity) / sum(severity.values())
print(f"severity index with weights 10/5/2/1: {index:.2f} per defect")
print("   the counts it summarises:", ", ".join(f"{s} {n}" for s, n in severity.items()))
defect density by module, highest first:
   Fee payment     58 defects in 4.0 KLOC:  14.5 per KLOC
   Exam form       74 defects in 5.5 KLOC:  13.5 per KLOC
   Login           12 defects in 1.5 KLOC:   8.0 per KLOC
   Hall ticket     18 defects in 2.5 KLOC:   7.2 per KLOC
   Profile         10 defects in 1.5 KLOC:   6.7 per KLOC
   Admin reports   20 defects in 3.5 KLOC:   5.7 per KLOC
   Notifications    8 defects in 1.5 KLOC:   5.3 per KLOC
reopen rate: 14 of 198 fixes = 7.1%
reports that were not product defects: 50 of 250 = 20.0% (duplicates 18, rejected as mistakes 10)
severity index with weights 10/5/2/1: 2.77 per defect
   the counts it summarises: critical 8, major 46, minor 98, cosmetic 48
munotes.in454

Defect Metrics

Density by module. The exam form had the most defects, 74, but it is also the largest module; per KLOC, the fee payment module is worse, 14.5 against 13.5. Counts say where most of the defects are; densities say where the code is weakest. Both point at the same two modules, which between them held two-thirds of the release's defects.

Reopen rate. 14 of 198 fixes failed their confirmation test, 7.1 per cent. Each was a defect the developer believed fixed; the rate measures how well fixes are tested before they are handed back, and a rising rate is an early warning.

Reports that were not defects. One report in five was a duplicate, a user's or operator's mistake, a test error, an enhancement request or unrepeatable (Chapter Seventy-Five, on the defect management process). The figure measures the load on triage, and a high duplicate share says testers cannot easily search the existing reports.

Severity index. Weighted 10, 5, 2 and 1, release 2.0's defects score 2.77 per defect. The number is only as meaningful as the weights: Chapter Sixty-Five, on what a software metric is, showed that averaging ranked values can reverse a comparison when the codes change, and a severity index is such an average. It is useful as long as its weights are a stated policy (here, the team's rough view of relative cost) and used unchanged from release to release; it is never a measurement of severity, and the counts it summarises are always reported beside it.

What each metric is for

MetricThe question it answersWatch out for
Defect densityHow defective is the product, or each part of it?The size measure and the counting rule for defects
Removal efficiencyWhat share of defects did we catch before users?It needs time after release, and it only ever falls
Stage effectivenessWhich activities catch the defects present when they run?The denominator is known only in hindsight
LeakageWhich reviews let their own kind of defect through?Small counts per work product
AgeHow long do defects survive?Measured in stages or in days: say which
Reopen rateHow often does a fix fail?Count reopenings of the same defect once or each time
Reports not defectsHow much triage effort goes on non-defects?Duplicates, mistakes and requests mean different things
Severity indexA single weighted figure for severityIt depends on the weights; report the counts too
munotes.in455

Defect Metrics

What it does not mean

A high removal efficiency does not mean few defects. It means few escaped; a release full of defects can have a high efficiency if testing was thorough.

A module with many defects is not necessarily the worst. Density, not count, compares modules of different sizes.

Stage effectiveness is not a score for the people. It measures an activity on one release, with a denominator that later stages supplied.

A severity index is not an average severity. It is a weighted count under a policy, and the counts behind it must be shown.

Quick revision

  • Defect density: defects per unit of size (ISO/IEC/IEEE 24765); by module, it shows the weakest code.
  • Defect removal efficiency: defects found before release over all defects; release 2.0: 188 of 200, 94.0 per cent.
  • Stage effectiveness (Fagan's error detection efficiency): errors found by a stage over errors present before it; release 2.0: requirements review 57.1, design review 46.2, code review 21.2, unit 34.9, integration 39.0, system 56.0, acceptance 45.5 per cent.
  • Leakage: requirements 42.9, design 47.5, code 76.7 per cent escaped their own review.
  • Age in stages: documents 4.17, far above the rest.
  • Density: fee payment 14.5 and exam form 13.5 defects per KLOC; reopen rate 7.1 per cent; not defects 20 per cent of reports; severity index 2.77 with weights 10, 5, 2, 1.

Test yourself

1. Define defect removal efficiency and compute it for release 2.0. The defects found before release as a percentage of all defects found, before and after release. Release 2.0: 188 found before release of 200 in all, 94.0 per cent.

2. What is the effectiveness of a stage, and why is its denominator hard to know? The defects a stage finds divided by the defects present in the product when it runs, as Fagan defined error detection efficiency for inspections. The defects present include those no stage has yet found, which are known only later, when subsequent stages and users find them.

3. A module of 2 KLOC has 30 defects and one of 6 KLOC has 48. Which is more defective, and why? The first: 30 defects in 2 KLOC is 15 per KLOC, against 48 in 6 KLOC, 8 per KLOC. The second has more defects only because it is larger.

4. What is defect leakage? Give release 2.0's figures. The share of a work product's defects that escape the activity meant to catch them. In release 2.0, 12 of 28 requirements defects escaped the requirements review (42.9 per cent), 19 of 40 design defects escaped the design review (47.5 per cent), and 92 of 120 code defects escaped the code review (76.7 per cent).

munotes.in456

Defect Metrics

5. What does the reopen rate measure, and what was release 2.0's? The share of fixes that fail their confirmation test and are reopened, a measure of the quality of fixes and of the testing done before they are handed back. Release 2.0: 14 of 198, 7.1 per cent.

6. Why must a severity index be reported with the counts it summarises? Because it is a weighted sum of ranked categories, and its value depends on the weights chosen, which are a policy, not a measurement; the counts at each severity level are the data, and show what the index hides.

Contents This chapter on its own page

munotes.in457

Chapter Eighty

Using Defect Data to Improve the Process

Syllabus topic Module 2, "Defect Management: ... their utilization for process improvement"

In one line

Defect data improves the process when a team selects a pattern worth acting on, digs from the symptoms to a cause the process can change, changes the process, and then measures whether the defects of that kind actually fell, against the kinds it did not act on; the loop is CMMI's causal analysis and resolution, and its tools include root cause analysis, the five whys and orthogonal defect classification.

In the wording a student can write in an examination: "The purpose of Causal Analysis and Resolution (CAR) is to identify causes of selected outcomes and take action to improve process performance" (CMMI for Development v1.3). Its two goals are to determine causes (select outcomes for analysis; analyse their causes) and to address causes (implement action proposals; evaluate the effect of the actions; record the causal analysis data). Causal analysis is the "analysis of a defect to determine its cause", and a root cause is the "source of a defect such that if it is removed, the defect is decreased or removed" (ISO/IEC/IEEE 24765). The five whys ask why repeatedly, because, in Sakichi Toyoda's words as ASQ quotes them, "by repeating why five times, the nature of the problem as well as its solution becomes clear." Orthogonal defect classification reads the distribution of defect types across the life cycle as a signature of the process.

From counting to changing

Chapters Seventy-Three to Seventy-Nine recorded, classified, tracked and measured release 2.0's defects. None of that improves the next release by itself. The ISTQB syllabus's third objective for defect reports is to "Provide ideas for improvement of the development and test process" (Chapter Seventy-Five, on the defect management process), and CMMI's Causal Analysis and Resolution is the discipline that turns the ideas into changes. Its introductory notes give the reason: "Reliance on detecting defects and problems after they have been introduced is not cost effective. It is more effective to prevent defects and problems by integrating Causal Analysis and Resolution activities into each phase of the project."

The loop, as CMMI sets it out

CMMI practiceWhat it means for release 2.0's defects
Select outcomes for analysisChoose a pattern worth the effort: "Since it is impractical to perform causal analysis on all outcomes, targets are selected by tradeoffs on estimated investments and estimated returns"
Analyze causesDig from the pattern to causes in the process: root cause analysis, the five whys, a cause-effect diagram (Chapter One Hundred Five)
Implement action proposalsChange the process: a standard, a shared component, a checklist item, a review, training
Evaluate the effect of implemented actionsMeasure the same pattern in the next release, and compare
Record causal analysis dataKeep what was found and done, so other projects learn from it
munotes.in458

Using Defect Data to Improve the Process

The first step is the one most often skipped. Release 2.0's defects offer many patterns: the fee and exam form modules' high density (Chapter Seventy-Nine, on defect metrics), the code review's low effectiveness, the documents' long survival. Input validation is chosen here because it is the largest defect type, 58 of 200, so one successful action on it removes more defects than an action on any other type; Chapter One Hundred Four, on Pareto diagrams, sets out the same reasoning for all the types at once.

Root causes and the five whys

A root cause is not the first cause found. The exam form accepted a negative number of backlog papers is a symptom; the developer forgot the check is a cause, but not one a process can change: people forget. The root cause is the reason the process let the forgetting through. ASQ describes the five whys as "a questioning process designed to drill down into the details of a problem or a solution and peel away the layers of symptoms", and notes that it "may take less or more than five times to reach the root cause". On release 2.0's input validation defects:

  1. Why were 58 defects input validation? Forms accepted values the rules forbid: negative days, blank names, fractional backlog counts.
  2. Why did the forms accept them? Each form's developer wrote that form's checks, and several missed some.
  3. Why did each developer write their own? ExamReg had no shared validation for its common field types.
  4. Why was there none? The design left validation to each screen; no one owned it.
  5. Why did the design leave it? The requirements listed the valid values but never said how invalid ones must be handled; the ruling on negative days (Chapter Eight) arrived late, after the code.

The last answer is a cause the process can change, and so is the fourth. The team's actions follow from them: a shared validation module for ExamReg's field types; a requirements review checklist item, does every input say what happens to an invalid value? (Chapter Sixty-Four, on checklist-based testing); and unit test templates built from equivalence partitions and boundaries (Chapters Fifty-Four and Fifty-Five, on equivalence partitioning and boundary value analysis) for every input.

Orthogonal defect classification: reading the process from the defects

Chillarege and his colleagues at IBM proposed a way to get process feedback from defects without a separate study for each question. ODC classifies each defect by the kind of fix it needed (the types listed in Chapter Two, on errors, faults and failures), and chooses the types so that each "can be associated with a few specific phases in the process". Then the mix of types found at each stage is a signature: "the defect type distribution changes with time, and the distribution provides an indication of where the development is, logically." Function defects, for instance, should be found early; "the bar corresponding to function defects should be diminishing through the process."

munotes.in459

Using Defect Data to Improve the Process

The payoff is a diagnosis. "If at system test the profile of the distribution looks like it should be in unit test or integration test, then the distribution indicates that the product is prematurely in system test." And "When a departure in the process is identified by a deviation in the distribution curve, the offending defect type also points to the part of the process that is probably responsible for this departure." Release 2.0's input validation defects, most of them checking defects in ODC's terms, pointing at the design and coding of screens, are exactly the kind of signal ODC formalises.

Worked example: did the action work?

Release 2.1, the next release, was built with the three actions in place. It added or changed 8 KLOC and had 69 defects. Because the two releases differ in size, the program compares defects per KLOC by type; and because many things change between releases, it compares the targeted type with all the others, which the actions did not aim at and which therefore act as a control.

# defects by type: release 2.0 (FINDINGS 5.2, 20 KLOC) and release 2.1 (FINDINGS 5.3, 8 KLOC)
r20 = {"input validation": 58, "logic and computation": 44, "interface": 30, "user interface": 24,
       "data and database": 18, "documentation": 12, "performance": 8, "security": 6}
r21 = {"input validation": 9, "logic and computation": 20, "interface": 13, "user interface": 10,
       "data and database": 7, "documentation": 5, "performance": 3, "security": 2}
KLOC_20, KLOC_21 = 20, 8
TARGET = "input validation"                   # the type causal analysis acted on

print(f"{'defect type':<22}{'2.0 per KLOC':>13}{'2.1 per KLOC':>13}{'change':>8}")
for t in r20:
    before, after = r20[t] / KLOC_20, r21[t] / KLOC_21
    mark = "   <- the action's target" if t == TARGET else ""
    print(f"{t:<22}{before:>13.2f}{after:>13.2f}{100 * (after - before) / before:>7.0f}%{mark}")

def rate(release, kloc, types):
    return sum(release[t] for t in types) / kloc

others = [t for t in r20 if t != TARGET]
print(f"target type:  {rate(r20, KLOC_20, [TARGET]):.2f} -> {rate(r21, KLOC_21, [TARGET]):.2f} per KLOC")
print(f"all others:   {rate(r20, KLOC_20, others):.2f} -> {rate(r21, KLOC_21, others):.2f} per KLOC")
print(f"share of all defects: {100 * r20[TARGET] / sum(r20.values()):.0f}% -> "
      f"{100 * r21[TARGET] / sum(r21.values()):.0f}%")
defect type            2.0 per KLOC 2.1 per KLOC  change
input validation               2.90         1.12    -61%   <- the action's target
logic and computation          2.20         2.50     14%
interface                      1.50         1.62      8%
user interface                 1.20         1.25      4%
data and database              0.90         0.88     -3%
documentation                  0.60         0.62      4%
performance                    0.40         0.38     -6%
security                       0.30         0.25    -17%
target type:  2.90 -> 1.12 per KLOC
all others:   7.10 -> 7.50 per KLOC
share of all defects: 29% -> 13%
munotes.in460

Using Defect Data to Improve the Process

Input validation defects fell from 2.90 to 1.12 per KLOC, a drop of 61 per cent, and from 29 to 13 per cent of all defects. The other types together did not fall: 7.10 per KLOC before, 7.50 after. That contrast is the evidence. Had every type fallen by a similar amount, the drop in input validation could have come from anything that changed between the releases (a smaller, simpler release, a more careful team); because only the targeted type fell, the actions are the likeliest explanation.

The practice CMMI calls "Evaluate the Effect of Implemented Actions" asks for exactly this, and the evaluation is also a caution. One release is one measurement, and 9 defects is a small number; the team keeps watching the rate over the next releases before calling the problem solved, and the rise in logic defects, 14 per cent, is itself a candidate for the next causal analysis.

What it does not mean

Improvement is not fixing more defects. It is changing the process so that fewer of them are made or more are caught early.

A root cause is not a person. The developer forgot is where the five whys start, not where they stop; the answer must be something the process can change.

One release is not proof. An effect is confirmed over several releases, and against the defect types that were not targeted.

ODC does not replace judgement. A deviation in the type distribution points at a part of the process; people still have to find out what went wrong there.

Quick revision

  • CMMI Causal Analysis and Resolution: "to identify causes of selected outcomes and take action to improve process performance"; select outcomes, analyse causes, implement actions, evaluate the effect, record the data.
  • Root cause (ISO/IEC/IEEE 24765): a source whose removal decreases or removes the defect; causal analysis: analysis of a defect to determine its cause.
  • Five whys (ASQ): ask why repeatedly, from the symptom to a cause the process can change; "It may take less or more than five times".
  • ODC (Chillarege 1992): defect types associated with process phases; the type distribution by stage is a process signature; a deviation points at the responsible part of the process.
  • Worked example: input validation 58 of 200 in release 2.0; three actions; release 2.1: 2.90 to 1.12 per KLOC (a 61 per cent drop) while other types rose from 7.10 to 7.50; share 29 to 13 per cent.

Test yourself

1. What is causal analysis and resolution? List its practices. A process, in CMMI, for identifying the causes of selected outcomes, such as a class of defects, and acting on them to improve the process. Its practices are selecting outcomes for analysis, analysing their causes, implementing action proposals, evaluating the effect of the actions, and recording the causal analysis data.

munotes.in461

Using Defect Data to Improve the Process

2. Apply the five whys to a class of defects. For input validation defects: forms accepted invalid values; because each developer wrote their own checks and some missed them; because there was no shared validation; because the design left validation to each screen; because the requirements never said how invalid values must be handled. The last two are root causes the process can change: shared validation, and a requirements review item on invalid input.

3. Why is "the developer made a mistake" not a root cause? Because people will always make mistakes; a cause is useful only if changing it prevents the defect, so the analysis continues until it reaches something in the process, such as missing requirements, a missing shared component or a missing review item, that can be changed.

4. How does orthogonal defect classification give feedback on the process? It classifies each defect by the kind of fix it needed, with types associated with particular phases of development. The distribution of types found at each stage is a signature of the process; if it departs from the expected pattern, for example function defects still common at system test, the product may be prematurely in that stage, and the defect type points at the part of the process responsible.

5. How should the effect of a process change be evaluated? By measuring the targeted class of defects in the following release, normalised by size, and comparing it with the classes that were not targeted, which act as a control; a drop only in the targeted class supports the change as the cause. The evaluation continues over several releases.

6. In the worked example, why was the unchanged rate of the other defect types important? Because it rules out explanations that would have lowered every type, such as a simpler release or a more careful team; since only input validation defects fell, the three actions aimed at them are the likeliest cause of the fall.

Contents This chapter on its own page

munotes.in462

Chapter Eighty-One

Quality Concepts: Variation, Design and Conformance

Syllabus topic Module 2, "Software Quality Assurance: Understanding quality concepts"

In one line

Every process varies, and quality work starts by telling the variation built into a process from the variation that has a cause of its own; it asks two separate questions of every product, whether the design is right and whether the product conforms to it; it holds the answers within limits by quality control; it counts what all of this costs; and the decisions that change a process belong to management.

In the wording a student can write in an examination: variation is the difference between one output of a process and the next. A common cause is a "source of variation of a process that exists because of normal and expected interactions among components of a process" (ISO/IEC/IEEE 24765:2017); a special cause is a "source of variation that is not inherent in the system, is not predictable, and is intermittent" (ISO/IEC/IEEE 24765c:2014). Quality of design is how well the characteristics chosen for a product (its requirements, its design, its grade) suit the people who will use it; quality of conformance is how closely the product as built matches that design. Quality control measures the product, compares it with what the design requires and acts on the difference. The cost of quality is the "life-cycle costs associated with assuring that a product or service conforms to requirements, plus failure costs from non-conformance to requirements" (ISO/IEC/IEEE 24774:2021). Changing the system that produces the variation is management's part.

Where these concepts come from

Module 2's quality assurance topics begin with ideas that were worked out in factories long before software existed. ASQ's history of quality dates the turning point: "Walter Shewhart began to focus on controlling processes in the mid-1920s, making quality relevant not only for the finished product but for the processes that created it." Shewhart saw that the data a process yields can be analysed "to see whether a process is stable and in control, or if it is being affected by special causes that should be fixed." Chapters Eighty-Two and Eighty-Three, on the quality movement, follow the people who carried that idea from Shewhart to total quality management. This chapter sets out the concepts they share, in software's terms.

Variation: no two runs are alike

Run a program twice on the same input and it gives the same answer; run a process twice and it does not give the same result. Release 2.0's testers found 16 defects in their second week and 18 in their third (Chapter Seventy-Seven, on tracking defects to closure). Two requests for the same fee page were served in 380 and 410 milliseconds (Chapter Forty-Five, on load testing). Variation is the name for these differences, and every process has it, including the processes that build and test software.

munotes.in463

Quality Concepts: Variation, Design and Conformance

The standards sort its sources into two kinds.

  • A common cause is a "source of variation of a process that exists because of normal and expected interactions among components of a process". Every request for ExamReg's fee page crosses the same network, waits behind the same server's other work and reads the same database. Small differences in each of these, from one request to the next, add up to differences in response time, and no single one of them can be named as the cause of a particular page's time.
  • A special cause is a "source of variation that is not inherent in the system, is not predictable, and is intermittent". A disk that fails, a setting changed by mistake, a release with a new defect: something happens that is not part of how the process normally works, and it can be found.

ASQ's page on the control chart puts the pair briefly: variation comes from "special causes (non-routine events) or common causes (built into the process)". A process whose variation comes from common causes alone is a stable process, one "from which all special causes of process variation have been removed and prevented from recurring, so that only common causes of process variation of the process remain" (ISO/IEC/IEEE 24765:2017). The same ASQ page says what stability buys: it shows "whether the process variation is consistent (in control) or is unpredictable (out of control, affected by special causes of variation)". Stable means predictable. It does not mean good: a stable process can be predictably too slow.

Worked example 1: the fee page's forty requests

Chapter Forty-Five, on load testing, read the results of forty requests for ExamReg's fee page. The program below takes the same forty, in the order they were sent, and looks at their variation: the spread of the pages that were served, drawn as a histogram with one # for each page, and the requests that stand apart from it.

from statistics import mean, stdev

# Chapter 45's forty requests for ExamReg's fee page, in order: (elapsed ms, HTTP response code)
runs = [(380, 200), (410, 200), (417, 200), (463, 200), (454, 200), (516, 200), (491, 200),
        (569, 200), (528, 200), (622, 200), (565, 200), (675, 200), (602, 200), (728, 200),
        (639, 200), (431, 200), (676, 200), (484, 200), (413, 200), (537, 200), (450, 200),
        (590, 200), (487, 200), (643, 200), (524, 200), (696, 200), (561, 200), (749, 200),
        (598, 200), (452, 200), (635, 200), (505, 200), (672, 200), (558, 200), (409, 200),
        (120, 503), (446, 200), (664, 200), (95, 503), (717, 200)]

served = [ms for ms, code in runs if code == 200]
print(f"{len(served)} pages served: {min(served)} to {max(served)} ms,"
      f" mean {mean(served):.0f}, standard deviation {stdev(served):.0f}")
for low in range(350, 750, 50):
    n = sum(low <= ms < low + 50 for ms in served)
    print(f"   {low}-{low + 49} ms  {'#' * n}")

for i, (ms, code) in enumerate(runs, 1):
    if code != 200:
        print(f"request {i}: no fee page; response code {code} after {ms} ms")
munotes.in464

Quality Concepts: Variation, Design and Conformance

38 pages served: 380 to 749 ms, mean 551, standard deviation 104
   350-399 ms  #
   400-449 ms  ######
   450-499 ms  #######
   500-549 ms  #####
   550-599 ms  ######
   600-649 ms  #####
   650-699 ms  #####
   700-749 ms  ###
request 36: no fee page; response code 503 after 120 ms
request 39: no fee page; response code 503 after 95 ms

The served pages. Thirty-eight pages were served, in times from 380 to 749 milliseconds around a mean of 551. They form one spread with no gaps in it. Nothing in the picture marks any of them out, and nothing in the data says why request 28 took 749 ms while request 1 took 380. Every one of them passed through the same network, server and database, and variation that comes from shared, normal sources like these is common-cause variation by the definition above.

The two failures. Requests 36 and 39 are different in kind. They got no fee page at all: they came back with response code 503, in 120 and 95 ms, faster than any page that was served. They are not the slow end of the spread; they are something that happened, twice in forty requests, and not on the other thirty-eight. That is what the definition of a special cause describes: not inherent in the system, not predictable, intermittent. The test's own load does not explain them: Chapter Forty-Five, on load testing, found that never more than two of its requests were in flight at once. So the tester reports them as a defect, and someone finds the cause.

What the histogram cannot show. A histogram throws away the order in which the pages were served. Whether the thirty-eight served pages are all common-cause variation, or hide a special cause of their own, such as a slow start or a drift over time, depends on that order. The tools that keep the order are the control chart and the run chart, which Chapters Ninety-One and One Hundred Seven, on statistical process control and run charts, apply to ExamReg's data.

Two kinds of cause, two responses

The two kinds of cause need different responses. ASQ lists among the uses of a control chart "determining whether your quality improvement project should aim to prevent specific problems or to make fundamental changes to the process".

  • A special cause is found and removed where it happened. The 503 responses have a cause that can be traced, fixed and prevented from recurring, and the people who run and maintain the portal can do it.
  • Common-cause variation is reduced only by changing the process. No single served page has a cause of its own to remove. To make pages like request 28 rarer, the system that serves every page has to change: a faster server, a cached fee table, a lighter page. Each of these is a decision about design and money, not a repair.
munotes.in465

Quality Concepts: Variation, Design and Conformance

Confusing the two wastes effort either way. Investigating request 28 as if it had a cause of its own spends time on a difference the process produces all the time, and finds nothing a fix would remove. Filing the 503 responses under the portal is sometimes slow leaves a defect in the product.

The Juran Institute's account of Joseph Juran's trilogy draws a similar line, from the side of cost. It describes a sporadic spike in failures, which "resulted from some unplanned event such as a power failure, process breakdown, or human error", and a chronic level of waste, which "goes on and on until the organization decides to find its root causes and remove it". Against the spike, the people doing the work restore the usual level. Against the chronic waste they cannot act alone: "Under conventional responsibility patterns, the operating forces are unable to get rid of the defects or waste." Chapter Eighty-Two, on Shewhart, Deming and Juran, returns to the trilogy.

Quality of design and quality of conformance

Chapter Twenty-Two, on what quality means, gave ISO's definition of quality and two classic short ones: Juran's "fitness for use" and Philip Crosby's "conformance to requirements". ASQ's glossary sets out two technical meanings side by side: quality is "1) the characteristics of a product or service that bear on its ability to satisfy stated or implied needs; 2) a product or service free of deficiencies."

The two meanings ask two different questions of any product, and this book gives each a name for what it judges.

  • Quality of design judges the plan: are the characteristics chosen for the product, its requirements, its design and its grade, the right ones for the people who will use it? A fee page designed for desktop browsers alone would have a low quality of design for students who register from their phones, however well it was built.
  • Quality of conformance judges the build: does the product as made match its plan? ExamReg's fee rules say 1 to 7 days late costs Rs 100; a version coded so that the band ends at 6 days, and charges Rs 500 at exactly 7, has a conformance defect, the one Chapter Fifty-Five, on boundary value analysis, found.
munotes.in466

Quality Concepts: Variation, Design and Conformance

A product needs both. A perfect build of the wrong design fails its users as surely as a faulty build of the right one. The two questions are ones this book has asked before: conformance is what verification checks, that "you built it right", and design is what validation checks, that "you built the right thing" (CMMI's plain words, Chapter Twenty-Six, on verification and validation). ASQ's page on the cost of quality speaks in the same terms when it describes failures that occur "when the results of work fail to reach design quality standards".

Worked example 2: which checks find which kind of defect

Release 2.0's 200 defects were recorded with the work product each was in and the activity that found it: the table Chapter Seventy-Nine, on defect metrics, worked from. A defect in the requirements or the design is a flaw in the plan, a failure of quality of design. A defect in the code is a departure from the plan, a failure of conformance. (The twelve defects in the user documents are left out: a document can be wrong in either way.) The program asks which checks found each kind.

# release 2.0's defects: the work product each was in, and the activity that found it (FINDINGS 5.2.6)
found_by = ["requirements review", "design review", "code review", "unit testing",
            "integration testing", "system testing", "acceptance testing", "after release"]
origin = {"requirements": [16, 3, 1, 1, 1, 2, 3, 1],
          "design":       [0, 21, 5, 3, 4, 4, 1, 2],
          "code":         [0, 0, 28, 40, 26, 20, 3, 3]}
checks = {"reviews": ["requirements review", "design review", "code review"],
          "tests built from the plan": ["unit testing", "integration testing", "system testing"],
          "the users": ["acceptance testing", "after release"]}

for name, rows in [("flaws in the plan (requirements, design)", ["requirements", "design"]),
                   ("departures from the plan (code)", ["code"])]:
    total = sum(sum(origin[r]) for r in rows)
    print(f"{name}: {total} defects, found by")
    for check, activities in checks.items():
        n = sum(origin[r][found_by.index(a)] for r in rows for a in activities)
        print(f"   {check:<26} {n:>3}  ({n / total:.0%})")
flaws in the plan (requirements, design): 68 defects, found by
   reviews                     46  (68%)
   tests built from the plan   15  (22%)
   the users                    7  (10%)
departures from the plan (code): 120 defects, found by
   reviews                     28  (23%)
   tests built from the plan   86  (72%)
   the users                    6  (5%)

The two kinds were found by different checks. Unit, integration and system tests, which are built from the requirements and the design, found 86 of the 120 departures from the plan, 72 per cent, but only 15 of the 68 flaws in it, 22 per cent. That is what their construction predicts: a test derived from a wrong requirement expects the wrong result, and passes when the code faithfully does the wrong thing. The plan's flaws were found mostly by reviews, 46 of the 68. And twice the share of them got as far as the users, at acceptance testing or after release: 10 per cent, against 5 per cent of the code's defects.

munotes.in467

Quality Concepts: Variation, Design and Conformance

The lesson is the one the two names carry. Conformance can be checked against the plan; the plan itself has to be checked against the users' needs, by reviews of the requirements and the design and by validation with the people who will use the product. Chapter Ninety-Six, on why reviews pay, prices the difference between finding a flaw in the plan early and finding it late.

Quality control: holding the variation within limits

Quality control is where variation and conformance meet. ExamReg's fee page will never be served in the same time twice, and its code will never be written without defects; what matters is whether the variation stays within what the design allows, and whether the departures are caught. ASQ's glossary gives one definition of quality control as "the operational techniques and activities used to fulfill requirements for quality"; an older definition in SEVOCAB describes the loop: "monitoring service performance or product quality, recording results, and recommending necessary changes" (ISO/IEC/IEEE 24765c:2014). Chapter Twenty-Four, on quality control and quality assurance, set out the current definitions and where testing fits.

For software the loop has three steps, each already met in this book.

  1. Measure a characteristic of the product: response times from a load test, defects from a review, results from a test run.
  2. Compare it with what the design requires: the fee page within 2 seconds, every fee as the rules say.
  3. Act on the difference: report the defect, fix it and confirm the fix.

ASQ's glossary adds a warning about the word itself. Control can mean "an evaluation to indicate needed corrective responses", "the act of guiding" or "the state of a process in which the variability is attributable to a constant system of chance causes". The first is the loop above; the last is the stable process of the previous sections. A project can run the loop faithfully, catching and correcting every departure it finds, while its process is not in control in the last sense.

The cost of quality, in outline

Everything in this chapter costs money: the reviews that judge the plan, the tests that check conformance, the fixes, and the failures that get through to the users. The cost of quality counts it. ASQ describes it as a way for an organisation to determine "the extent to which its resources are used for activities that prevent poor quality, that appraise the quality of the organization's products or services, and that result from internal and external failures", and so in four categories.

munotes.in468

Quality Concepts: Variation, Design and Conformance

CategoryWhat it pays forOn ExamReg
PreventionAvoiding defects: quality planning, requirements, trainingWriting the review checklist; training developers in the fee rules
AppraisalFinding defects: checking products against their specifications, auditsReviews of requirements, design and code; every test level
Internal failureDefects found before the customer has the product: rework, failure analysisFixing and retesting the defects found before release
External failureDefects found by the customer: repairs, complaintsFixing the defects students found, and answering their complaints

The standard definition puts the same total in two parts: the costs of "assuring that a product or service conforms to requirements", plus the "failure costs from non-conformance to requirements". Crosby, as ASQ reports, called the measure the "price of nonconformance" and "argued that organizations choose to pay for poor quality". Chapters One Hundred One and One Hundred Two, on the cost of quality and on using quality costs for decision making, draw up ExamReg's statement and use it to decide.

Management's part

Three of this chapter's findings end at the same place. Common-cause variation is reduced only by changing the system. Flaws in the plan are caught by reviews that someone has to schedule and staff. And the cost of quality depends on how the budget is divided between prevention, appraisal and failure. Each is a decision about the process, its people and its money, and the people doing the work cannot take it on their own.

ISO's quality management principles make leadership the second of seven: "Leaders at all levels establish unity of purpose and direction and create conditions in which people are engaged in achieving the organization's quality objectives." W. Edwards Deming gave the reason in the tenth of his fourteen points for management, which asks managers to stop exhorting the work force to zero defects, because "the bulk of the causes of low quality and low productivity belong to the system and thus lie beyond the power of the work force." The Juran Institute's account of the trilogy says what the operating forces can do without such leadership: "What they can do is to carry out control", which is holding the chronic level where it is. Moving it needs more: "Breakthrough requires special methods and leadership support to attain significant changes and results." ASQ's page on total quality management lists among its common difficulties "Insufficient resources or lack of sustained commitment of those resources" and "Management's failure to recognize and/or reward achievements", both of them management's to put right.

On ExamReg the division is plain. The maintenance team can find and remove the cause of the 503 responses. A faster fee page for every student needs the college to pay for a better server or the software house to schedule a redesign. The review checklist of Chapter Eighty, on using defect data to improve the process, needed someone with authority over the project's plan to give reviews the time. Chapter Ninety-Four, on the ISO 9000 family and its seven principles, sets out the rest.

munotes.in469

Quality Concepts: Variation, Design and Conformance

What it does not mean

Variation is not a defect. Every process varies. A defect is a departure from what the design allows, and a special cause is a reason to look, not a verdict.

Stable is not good. A stable process is predictable; it can predictably fall short of its requirements, and then only a change to the process helps.

Conformance is not the whole of quality. A product that conforms perfectly to a poor design is still poor. Quality of design and quality of conformance are judged separately, by different checks.

Quality control is not only testing. Testing is "a major form of quality control", in the ISTQB syllabus's words, and not the only one; a review of a design is quality control too.

Management's part is not a slogan. It is a set of decisions: time for reviews, money for a better system, and changes to how the work is done.

Quick revision

  • Variation: every process varies. Common cause (ISO/IEC/IEEE 24765): "normal and expected interactions among components of a process"; special cause: "not inherent in the system, is not predictable, and is intermittent". ASQ: "special causes (non-routine events) or common causes (built into the process)".
  • Stable process: only common causes remain; stable means predictable, not good.
  • Responses: a special cause is found and removed; common-cause variation is reduced only by changing the process.
  • Quality of design: are the chosen requirements, design and grade right for the users? Checked by reviews of the plan and by validation. Quality of conformance: does the product match its design? Checked by verification, mostly by tests built from the plan.
  • Quality control: measure, compare with what the design requires, act on the difference.
  • Cost of quality: prevention, appraisal, internal failure, external failure (ASQ); the costs of assuring conformance plus the failure costs of non-conformance (ISO/IEC/IEEE 24774:2021).
  • Management: owns the decisions that change the system; ISO's leadership principle.
  • Worked examples: 38 fee pages served in 380 to 749 ms, and two requests that got no page, with the marks of special causes; release 2.0's 68 flaws in the plan were found 68 per cent by reviews and 22 per cent by tests, its 120 departures from the plan 72 per cent by tests.
munotes.in470

Quality Concepts: Variation, Design and Conformance

Test yourself

1. What is variation, and what are its two kinds of cause? Variation is the difference between one output of a process and the next; every process has it. Common causes are built into the process: the normal interactions of its parts, present all the time, none of which can be singled out as the cause of a particular result. Special causes are not part of the process: unpredictable, intermittent events, such as a failed disk or a faulty change, which can be found and removed.

2. Why does it matter which kind of cause lies behind a variation? Because the responses differ. A special cause is found and removed where it happened. Common-cause variation has no single cause to remove; it is reduced only by changing the process, which is usually a management decision. Treating one as the other either wastes effort on routine differences or leaves a real problem in place.

3. Distinguish quality of design from quality of conformance, with an example. Quality of design is how well the chosen requirements, design and grade suit the users' needs; quality of conformance is how closely the product as built matches that design. A fee page designed for desktop browsers alone has a low quality of design for students on phones, however well it is built; a correctly specified fee page whose code ends the Rs 100 late fee band a day early has a conformance defect.

4. Why do tests find conformance defects more easily than design defects? Because tests are derived from the requirements and the design; a test built from a wrong requirement expects the wrong result and passes. In release 2.0, tests built from the plan found 72 per cent of the code's defects but only 22 per cent of the requirements and design defects, most of which were found by reviews.

5. What is quality control? Describe its loop. The activities that check a product against what its design requires and act on the differences. The loop: measure a characteristic of the product, compare it with the requirement, and act on any difference by reporting, fixing and confirming the fix.

6. What is management's part in quality? Taking the decisions the people doing the work cannot take alone: changing the system that produces common-cause variation, giving reviews the time and people to catch flaws in the plan, and dividing spending between prevention, appraisal and failure. ISO's leadership principle puts it as establishing unity of purpose and direction and creating the conditions in which people achieve the quality objectives.

Contents This chapter on its own page

munotes.in471

Chapter Eighty-Two

The Quality Movement: Shewhart, Deming and Juran

Syllabus topic Module 2, "Software Quality Assurance: Understanding quality concepts and the Quality Movement"

In one line

The quality movement moved quality from inspecting finished products to managing the processes that make them: Walter Shewhart gave it the control chart; W. Edwards Deming gave it the Plan-Do-Study-Act cycle and fourteen points for management, which place most causes of poor quality in the system; Joseph Juran gave it fitness for use and the trilogy of quality planning, control and improvement; and Japan, where Deming and Juran taught in the 1950s, showed what the ideas could do.

In the wording a student can write in an examination: Walter Shewhart proposed the control chart in the 1920s: a centre line at the process mean, with control limits a set number of standard deviations either side (three, by the accepted standard), which show whether a process is stable. W. Edwards Deming taught Shewhart's methods in Japan in the 1950s. His fourteen points tell management to improve the system constantly, to cease dependence on inspection, to drive out fear and to drop slogans, quotas and merit ratings, because "the bulk of the causes of low quality and low productivity belong to the system". His PDSA cycle (Plan, Do, Study, Act) is "a systematic process for gaining valuable learning and knowledge for the continual improvement of a product, process, or service" (the Deming Institute). Joseph Juran defined quality as fitness for use and managed it through a trilogy: "quality planning (developing the products and processes required to meet customer needs), quality control (meeting product and process goals) and quality improvement (achieving unprecedented levels of performance)" (ASQ's glossary).

Before the movement: quality by inspection

ASQ's history of quality begins with the craft guilds of medieval Europe, whose inspection committees "enforced the rules by marking flawless goods with a special mark or symbol". A master who sold to local customers inspected his own goods before sale. The factory system split the crafts into specialised tasks, and quality came to rest on the skill of labourers "supplemented by audits and/or inspections". Late in the nineteenth century, Frederick Taylor's methods raised productivity and lowered quality, and "To remedy the quality decline, factory managers created inspection departments to keep defective products from reaching customers."

The Second World War stretched inspection to its limit. The U.S. armed forces inspected almost every unit produced, a practice that "required huge inspection forces", and they turned to sampling inspection, with sampling tables published in a military standard, Mil-Std-105. They also sponsored training courses in "Walter Shewhart's statistical quality control (SQC) techniques".

Every one of these methods found defects after they were made. A software team that relies on a test phase at the end is in the same position, which is the line Chapter Twenty-Four, on quality control and quality assurance, drew between checking the product and assuring the process. The quality movement is the story of attention moving from the product to the process.

munotes.in472

The Quality Movement: Shewhart, Deming and Juran

Walter Shewhart: control the process, not just the product

Shewhart worked at Bell Laboratories; the Deming Institute calls him "Walter Shewhart of the famous Bell Laboratories in New York". ASQ dates his turn to processes to the mid-1920s: he saw that industrial processes yield data, and that the data could show "whether a process is stable and in control, or if it is being affected by special causes that should be fixed". Chapter Eighty-One, on quality concepts, set out his two kinds of cause. The control chart is his tool for telling them apart.

The NIST/SEMATECH e-Handbook sets out the general model he proposed. Take a statistic that measures some quality characteristic, such as the average of a sample, with a known mean and standard deviation. The chart's centre line is that mean. The upper control limit is the mean plus k standard deviations, and the lower control limit the mean minus k. With k set to 3, the handbook says, "we speak of 3-sigma control charts", and three "has become an accepted standard in industry". Values are plotted in the order they occur; values inside the limits, with no pattern, show a process in statistical control, and a value outside is a signal to look for a special cause.

Shewhart's model carries two cautions. The handbook separates the chart's limits from a product's specification: "Control Limits are used to determine if the process is in a state of statistical control", while "Specification Limits are used to determine if the product will function in the intended fashion." A process can be in control and still make a product that fails its specification, the stable but slow process of Chapter Eighty-One, on quality concepts. And a process is not to be judged in control too soon: the handbook quotes his rule of thumb that it should first give at least twenty-five samples of four, taken under the same essential conditions, that are in control. His book, Economic Control of Quality of Manufactured Product, appeared in 1931. Chapter Ninety-One, on statistical process control, builds his chart for software data.

W. Edwards Deming: the system belongs to management

ASQ's history introduces Deming as "a statistician with the U.S. Department of Agriculture and Census Bureau" who "became a proponent of Shewhart's SQC methods and later became a leader of the quality movement in both Japan and the United States." In the 1950s, invited by JUSE, the Union of Japanese Scientists and Engineers, he went back to Japan "to teach methods for statistical analysis and control of quality to Japanese engineers and executives"; ASQ's page on total quality management calls this the origin of TQM.

munotes.in473

The Quality Movement: Shewhart, Deming and Juran

The fourteen points. Deming's best-known statement is his fourteen points for management, first presented in his book Out of the Crisis (1986). The Deming Institute keeps a condensed wording of them. The table gives each as a short label, and beside it one reading of what the point asks of a software team; that reading is this book's, not Deming's.

PointIn shortFor a software team, one reading
1Constancy of purpose toward improving product and serviceQuality goals that survive the next deadline
2Adopt the new philosophy; management learns its responsibilities and leads the changeManagers learn what quality work needs, not only testers
3Cease dependence on inspection; build quality into the product in the first placeClear requirements, reviews and unit tests, not a test phase at the end, as the way to quality
4Stop awarding business on price alone; minimise total costChoose a supplier or a library for its total cost, its defects included
5Improve the system of production and service constantly and foreverCausal analysis after every release
6Institute training on the jobTrain testers and developers in the techniques they use
7Institute leadership: supervision should help people do a better jobA lead removes obstacles instead of counting output
8Drive out fearPeople report defects, and their own mistakes, without fear of blame
9Break down barriers between departmentsDevelopers, testers and users work as one team
10Eliminate slogans, exhortations and targets for the work forceNo zero defects poster in place of a better process
11Eliminate quotas and management by numerical goals; substitute leadershipNo quota of test cases a day or defects a week
12Remove barriers to pride of workmanship, including the annual merit ratingDo not rank people by the defects they find or make
13A vigorous programme of education and self-improvementTime to learn new tools and methods
14Put everybody to work on the transformationQuality is every role's job, not the SQA group's alone

The tenth point gives the reason behind several of the others. Such exhortations, in the Institute's wording, "only create adversarial relationships, as the bulk of the causes of low quality and low productivity belong to the system and thus lie beyond the power of the work force." That sentence joins Deming to Shewhart: common-cause variation belongs to the system, and the system belongs to management. Deming later presented the points as following "naturally as application of the System of Profound Knowledge", his name for the theory behind them (The New Economics, as the Institute quotes it).

munotes.in474

The Quality Movement: Shewhart, Deming and Juran

Worked example: ranking people in a stable system

The twelfth point asks for an end to the annual merit rating, and the tenth point's reason explains why. If most of the variation in results comes from the system, a ranking of the people inside it is largely a ranking of chance. The program shows what that means with a model, not data: eight developers who are identical by construction, each making 100 changes a month in the same system, where every change has the same 5 per cent chance of carrying a defect. It prints a year of monthly defect counts and asks what a manager who ranked the developers would conclude.

import random

rng = random.Random(82)
developers = [f"dev {i}" for i in range(1, 9)]
# a model, not data: every developer works in the same system, making 100 changes a month,
# and every change has the same 5 per cent chance of carrying a defect
defects = {d: [sum(rng.random() < 0.05 for _ in range(100)) for month in range(12)]
           for d in developers}

print("month  " + "".join(f"{m:>3}" for m in range(1, 13)) + "   year")
for d in developers:
    print(f"{d:<7}" + "".join(f"{n:>3}" for n in defects[d]) + f"{sum(defects[d]):>7}")

blamed = set()
for m in range(12):
    most = max(defects[d][m] for d in developers)
    blamed |= {d for d in developers if defects[d][m] == most}
totals = [sum(defects[d]) for d in developers]
print(f"developers with the month's most defects at least once: {len(blamed)} of {len(developers)}")
print(f"yearly totals from {min(totals)} to {max(totals)}; the model expects {12 * 100 * 0.05:.0f} from each")
month    1  2  3  4  5  6  7  8  9 10 11 12   year
dev 1    4  3  6  0  6  4  1  9  2  8  7  2     52
dev 2    4  5  8  6  6  7  8  7  7  9  5  5     77
dev 3    4  7  7  9 11  7  6  1  8  4  7  2     73
dev 4    2  7  6  2  4  3  4  5  5  2  5  9     54
dev 5    5  8  2  6  5  5  4  6  5  3  5  4     58
dev 6    7  5  4  6  3  6  5  3  8  6  2  5     60
dev 7    3  5  3  5 10 10  4  6  5  6  8  5     70
dev 8    4  2  4  5  6  5  5  4  6  6  4  6     57
developers with the month's most defects at least once: 7 of 8
yearly totals from 52 to 77; the model expects 60 from each

Every row comes from the same process, and yet the table supports every conclusion a merit system draws. Seven of the eight developers had the month's most defects at least once. Dev 3 had the worst single month, 11 defects in month 5. Dev 1, whose year was the best at 52, had the most defects of anyone in month 8. A yearly rating would call dev 2, at 77, the weakest and dev 1 the strongest, though the model expects exactly the same from both. The spread from 52 to 77 is common-cause variation, and nothing in it is about the people.

munotes.in475

The Quality Movement: Shewhart, Deming and Juran

The model does not show that people never differ. It shows that a difference has to be larger than the variation the system produces on its own before it says anything about a person, and deciding which differences are that large is Shewhart's question, answered by his chart (Chapter Ninety-One, on statistical process control). It also shows why Deming put the burden on management: the only way to lower all eight rows is to lower the 5 per cent, which is a property of the system, not of any developer.

PDSA. Deming's cycle for learning and improvement came from Shewhart; the Deming Institute says it "was first introduced to Dr. Deming by his mentor, Walter Shewhart". Its four steps are Plan (a goal, a theory, a prediction of the result and measures of success), Do (carry out the plan, often first on a small scale), Study (compare what happened with what was predicted) and Act (use what was learned: adjust the goal or the method, rethink the theory, or widen the change). Deming "emphasized the PDSA Cycle, not the PDCA Cycle". In the Institute's account, Check is about whether a change succeeded or failed, while Deming's focus was on predicting the results, studying the actual ones and comparing them "to possibly revise the theory". Chapter Ninety-Nine, on PDCA and kaizen, puts the cycle to work.

Joseph Juran: fitness for use and the trilogy

Juran's short definition of quality, "fitness for use", which Chapter Twenty-Two, on what quality means, compared with the others, puts the user's purpose first. His way of managing quality is the trilogy, which the Juran Institute calls "a universal way of thinking about managing for quality leadership". It has three processes.

  • Quality planning, which the Institute also calls quality by design: creating the product and the processes that will make it, to meet customers' needs.
  • Quality control: holding performance at the planned level during operations, and restoring it when something goes wrong.
  • Quality improvement: raising performance to levels not reached before. The Institute defines breakthrough as "the organized creation of beneficial change and the attainment of unprecedented levels of performance."

The Institute's description of the trilogy diagram puts time across and the cost of poor quality up. Planning comes first. Once operations begin, some work fails and has to be redone: in its example "more than 20 percent of the work", a chronic level of waste that goes on until someone decides to remove its root causes. A sporadic spike takes the failures above 40 per cent, and control brings them back to the chronic level. Improvement then drives the chronic level "far below the original level".

munotes.in476

The Quality Movement: Shewhart, Deming and Juran

Cost of poor quality over time: planning; then a chronic level of about 20 per cent with a sporadic spike above 40 that control brings back; then improvement to a new, lower chronic level

Figure 82.1 The Juran trilogy: control removes the sporadic spike, and only improvement lowers the chronic level

Read beside Chapter Eighty-One, on quality concepts, the diagram tells the same story from the side of cost. The sporadic spike behaves like a special cause, and control deals with it. The chronic waste is built in, and only improvement, with the leadership support the Institute says breakthrough requires, lowers it. For a software project, one reading is that planning is the requirements, the design and the quality plan (Chapter Eighty-Seven, on the quality assurance plan); control is the reviews, tests and defect tracking during development (Chapters Seventy-Three to Seventy-Nine, on defects); and improvement is causal analysis between releases (Chapter Eighty, on using defect data to improve the process).

The Pareto principle. Improvement cannot attack everything at once. The Juran Institute ties the trilogy's improvement to the Pareto principle: an organisation that succeeds "By attaining just a few vital breakthroughs year after year" can outperform its competitors. ASQ states the principle as the 80/20 rule: "80% of outcomes result from 20% of causes". For software, Boehm and Basili (2001) found that "About 80 percent of the defects come from 20 percent of the modules, and about half the modules are defect free", with the share across studies "between 60 and 90 percent". The ratio is a tendency, not a law: Chapter Four, on the seven principles of testing, found two of ExamReg's seven modules holding 66 per cent of the defects. Chapter One Hundred Four, on Pareto diagrams, draws the chart.

Japan after the war

ASQ's history tells the turn. After the Second World War, Japanese manufacturers converted from military goods to civilian goods for trade, and "At first, Japan had a widely held reputation for shoddy exports". They welcomed foreign lecturers, Deming and Juran among them. Juran "predicted the quality of Japanese goods would overtake the quality of goods produced in the United States by the mid-1970s", because of the pace at which Japan was improving. The approach Japan built was the total quality approach: "Rather than relying purely on product inspection, Japanese manufacturers focused on improving all organizational processes through the people who used them." As a result, in ASQ's words, "Japan was able to produce higher-quality exports at lower prices". American manufacturers first answered on price and with import restrictions. The answer that worked, total quality management, is the subject of Chapter Eighty-Three, on Feigenbaum, Ishikawa, Crosby and TQM.

munotes.in477

The Quality Movement: Shewhart, Deming and Juran

The three compared

ShewhartDemingJuran
Known forThe control chart (1920s)The fourteen points; PDSAThe trilogy; fitness for use
Central ideaTell common from special causes by a process's own dataMost causes of poor quality belong to the system, which management ownsPlan quality in, hold it by control, raise it by breakthrough
What it asks of a teamMeasure the process, and react only to signalsImprove the system; drop slogans, quotas and ratingsPlan, control and improve as three separate jobs
A book CMMI citesEconomic Control of Quality of Manufactured Product (1931)Out of the Crisis (1986)Juran on Planning for Quality (1988)

What it does not mean

Ceasing dependence on inspection is not stopping testing. Deming's third point asks for quality built into the product "in the first place". Testing still finds what gets through; it is not how quality is made.

Control limits are not specification limits. Control limits come from the process's own data and say whether it is stable; a specification says whether the product will do its job.

The 80/20 rule is not a law. It describes a tendency, which Boehm and Basili put between 60 and 90 per cent of defects in 20 per cent of modules.

PDSA is not PDCA with a new letter. Study compares the results with a prediction and revises the theory; Check, in the Institute's account, asks whether the plan succeeded.

Rejecting merit ratings is not ignoring performance. It means judging a result against the variation the system produces, not against other people's chance results.

Quick revision

  • Before the movement: guild marks, factory inspection departments, wartime sampling inspection (Mil-Std-105): quality by finding defects after they were made.
  • Shewhart (Bell Laboratories, 1920s): the control chart; centre line at the mean, limits at k standard deviations, k = 3 the accepted standard; control limits are not specification limits; Economic Control of Quality of Manufactured Product (1931).
  • Deming: statistician; taught Shewhart's methods in Japan in the 1950s; fourteen points (constancy of purpose, cease dependence on inspection, improve the system constantly, training, leadership, drive out fear, break down barriers, no slogans, quotas or merit rating, education, everybody in the transformation); most causes "belong to the system"; PDSA, with Study rather than Check.
  • Juran: fitness for use; the trilogy of planning, control and improvement; a sporadic spike removed by control, chronic waste lowered only by improvement; breakthrough; the Pareto principle.
  • Pareto: "80% of outcomes result from 20% of causes" (ASQ); in software about 80 per cent of defects from 20 per cent of modules, 60 to 90 per cent across studies (Boehm and Basili).
  • Japan: from shoddy exports to higher quality at lower prices, by improving all processes through the people who use them.
  • Worked example (a model): eight identical developers; 7 of 8 had a month's most defects; yearly totals from 52 to 77, against 60 expected from each.
munotes.in478

The Quality Movement: Shewhart, Deming and Juran

Test yourself

1. How was quality achieved before Shewhart, and what did he change? By inspecting finished goods: the marks of guild inspection committees, the inspection departments of factories, and in the Second World War sampling inspection. Shewhart turned attention to the process that makes the goods, using the data it yields to see whether it is stable or affected by special causes, with the control chart as the tool.

2. Describe Shewhart's control chart. A plot, in the order the values occur, of a statistic that measures a quality characteristic, with a centre line at its mean and upper and lower control limits k standard deviations either side, k = 3 by the accepted standard. Values inside the limits without a pattern show a stable process; a value outside signals a special cause to be found. Control limits come from the process's data; specification limits come from the requirements.

3. State six of Deming's fourteen points and explain one for a software team. Create constancy of purpose; cease dependence on inspection; improve the system constantly; drive out fear; break down barriers between departments; eliminate slogans, targets and quotas. For a software team, ceasing dependence on inspection means building quality in through clear requirements, reviews and unit tests, instead of relying on a test phase at the end to find defects that have already been made.

4. Why did Deming oppose slogans, targets and merit ratings? Because most of the causes of low quality belong to the system, which the work force cannot change: exhortations only create adversarial relationships, and ratings reward chance. In the chapter's model, eight identical developers produced yearly totals from 52 to 77, and seven of them had a month with the most defects.

5. Explain Juran's trilogy with its diagram. Planning designs the product and its processes to meet customers' needs; control, during operations, holds performance at the planned level and removes sporadic spikes; improvement lowers the chronic level of waste by breakthrough. The diagram plots the cost of poor quality against time: a chronic level of about 20 per cent of work redone, a sporadic spike above 40 per cent that control brings back, and then improvement to a much lower chronic level.

6. What is the Pareto principle, and how does it apply to software? A few causes account for most of an effect, often stated as 80 per cent of outcomes from 20 per cent of causes. In software, about 80 per cent of defects come from 20 per cent of the modules (between 60 and 90 per cent across the studies Boehm and Basili report), so improvement and extra testing go first to the few modules or causes that account for most of the defects.

Contents This chapter on its own page

munotes.in479

Chapter Eighty-Three

The Quality Movement: Feigenbaum, Ishikawa, Crosby and TQM

Syllabus topic Module 2, "Software Quality Assurance: Understanding quality concepts and the Quality Movement"

In one line

The second half of the quality movement widened quality from the factory floor to the whole organisation: Armand Feigenbaum's total quality control, which first sorted quality costs into prevention, appraisal and internal and external failure; Kaoru Ishikawa's company-wide quality control, with quality circles and the cause-and-effect diagram; Philip Crosby's conformance to requirements and zero defects; and total quality management, which gathered these ideas into one management approach that Six Sigma, the ISO 9000 standards and, for software, the capability maturity models carried on.

In the wording a student can write in an examination: Feigenbaum's name and the term total quality control are, in ASQ's words, "virtually synonymous"; his "was the first text to characterize quality costs as the costs of prevention, appraisal, and internal and external failure." Ishikawa led Japan's company-wide quality control, whose hallmark is "broad involvement in quality, not only top to bottom within the organization, but also start to finish in the product life cycle"; he promoted quality circles, small groups of employees "(10 or fewer) and their supervisor" who study and improve their own work, and created the cause-and-effect (fishbone) diagram. Crosby's Four Absolutes of Quality Management are: "Quality means conformance to requirements, not goodness"; "Quality is achieved by prevention, not appraisal"; "Quality has a performance standard of Zero Defects, not acceptable quality levels"; and "Quality is measured by the Price of Nonconformance", not indexes. Total quality management (TQM) is "a management system for a customer-focused organization that engages all employees in continual improvement of the organization" (ASQ).

Armand Feigenbaum: total quality control

ASQ's page on its honorary member Armand V. Feigenbaum says that his name and the term total quality control are "virtually synonymous". His ideas are in his book Total Quality Control, "first published in 1951 under the title Quality Control: Principles, Practice, and Administration". He had been the manager of worldwide manufacturing operations and quality control at the General Electric Company, and he argued for quality as an international discipline: "The belief that quality travels under an exclusive foreign passport is a myth."

Two of his ideas run through the rest of this book. The first is in the name. Quality control that is total reaches every part of the organisation that affects quality, not only the inspectors at the end of the line; ASQ's page on total quality management calls his book "a forerunner for the present understanding of TQM". Software shows why the word matters. Release 2.0's 200 defects began in the requirements (28), the design (40), the code (120) and the user documents (12), as Chapter Seventy-Three, on what a defect is, classified them. They were made by analysts, designers, developers and writers, so quality control confined to testers running tests on the code cannot be total.

munotes.in480

The Quality Movement: Feigenbaum, Ishikawa, Crosby and TQM

The second idea is the cost of quality. According to ASQ, Feigenbaum's "was the first text to characterize quality costs as the costs of prevention, appraisal, and internal and external failure": the four categories that Chapter Eighty-One, on quality concepts, outlined, and that Chapters One Hundred One and One Hundred Two, on the cost of quality and on using quality costs for decision making, put to work.

Kaoru Ishikawa: company-wide quality control

ASQ's page on Kaoru Ishikawa begins with a parallel: "Ishikawa, like Japan as a whole, learned the basics of statistical quality control developed by Americans", and then went beyond it. "Perhaps Ishikawa's most important contribution has been his key role in the development of a specifically Japanese quality strategy", which the page names company-wide quality control (CWQC). Its hallmark is "broad involvement in quality, not only top to bottom within the organization, but also start to finish in the product life cycle." ASQ's timeline of TQM records that in 1968 "The Japanese named their approach to total quality" enterprise quality control, and that "Kaoru Ishikawa's synthesis of the philosophy contributed to Japan's ascendancy as a quality leader."

Quality circles carried the involvement to the bottom of the organisation. ASQ's glossary defines a quality circle as "A quality improvement or self-improvement study group composed of a small number of employees (10 or fewer) and their supervisor", and adds that circles "originated in Japan, where they are called quality control circles." Ishikawa, who directed the Quality Control Circle Headquarters at JUSE, the Union of Japanese Scientists and Engineers, "played a major role in the growth of quality circles", which spread to more than 50 countries. He never treated them as a substitute for management: ASQ notes that he "was always aware of the importance of top management support", developed quality control courses for executives from the late 1950s, and devised the audit for the Deming Prize, which requires the participation of a company's top executives. A software team's nearest equivalent, in this book's reading, is the retrospective, where the people who do the work study how to do it better; the Scrum Guide gives its purpose as "to plan ways to increase quality and effectiveness."

Tools for everyone. The cause-and-effect diagram, often called the Ishikawa diagram, "has provided a powerful tool that can easily be used by non-specialists to analyze and solve problems". ASQ also records that "The seven basic quality tools were first highlighted in Kaoru Ishikawa's classic book Guide to Quality Control." Chapters One Hundred Three to One Hundred Seven, from the basic tools to run charts, teach them. Ishikawa called standardization and quality control "two wheels of the same cart", but he stressed that standards must change and must rest on an analysis of what customers need.

munotes.in481

The Quality Movement: Feigenbaum, Ishikawa, Crosby and TQM

Philip Crosby: conformance, prevention and zero defects

ASQ records Philip Crosby as widely recognised for promoting zero defects and for defining quality as conformance to requirements. His career in quality began in 1952. According to the biography published by the firm he later founded, he was at Martin-Marietta from 1957 to 1965, and in 1964 the Department of the Army gave him its Distinguished Civilian Service Medal for developing the concept of Zero Defects; he was then ITT's corporate vice president of quality from 1965 to 1979. His book Quality Is Free (1979) "has been credited with playing a large part in beginning the quality revolution in the United States and Europe", and in 1979 he founded Philip Crosby Associates, "teaching management how to establish a preventive culture to get things done right the first time."

His ideas are summed up in his Four Absolutes of Quality Management:

  1. "Quality means conformance to requirements, not goodness." Quality is defined by the requirements, so it can be checked; Chapter Twenty-Two, on what quality means, set this definition beside Juran's fitness for use.
  2. "Quality is achieved by prevention, not appraisal." Finding defects after they are made is appraisal; quality comes from not making them.
  3. "Quality has a performance standard of Zero Defects, not acceptable quality levels." The standard is to get it right, not to agree in advance how much may be wrong.
  4. "Quality is measured by the Price of Nonconformance", not indexes: in money, as the cost of doing things wrong, the measure ASQ says he used "to raise awareness of the importance of quality."

Zero defects, Crosby and Deming. Deming's tenth point lists targets "asking for zero defects" among the things management should stop putting to the work force, while Crosby made zero defects a standard. Read side by side, the two are less opposed than they sound. Deming objected to asking workers for results whose causes "lie beyond the power of the work force". Crosby's firm taught management to build the preventive culture that makes the standard reachable. Both put the work of preventing defects on management.

Worked example: what an acceptable quality level accepts

Crosby's third absolute sets zero defects against acceptable quality levels, a term that came from acceptance sampling. The NIST/SEMATECH e-Handbook explains the method. Inspecting every item may be impossible or too costly, so a random sample is taken from a lot, and the lot is accepted or rejected on what the sample shows; acceptance sampling is "the middle of the road" approach between no inspection and 100% inspection. A single sampling plan (n, c) inspects n items and rejects the lot if the sample holds more than c defectives. If the lot is large, the number of defectives in the sample is approximately binomial, and the probability of accepting a lot whose fraction defective is p is the chance of c or fewer defectives in n. Plans are set up around an acceptable quality level (AQL), "the base line requirement for the quality of the producer's product", and designed so that the plan's operating characteristic (OC) curve "yields a high probability of acceptance at the AQL." The wartime tables became Mil. Std. 105A in 1950; its revision Mil. Std. 105D (1963) was adopted as ANSI Z1.4 in 1971 and, with minor changes, as ISO 2859 in 1974.

munotes.in482

The Quality Movement: Feigenbaum, Ishikawa, Crosby and TQM

The program computes the probability of acceptance for the handbook's own (52, 3) plan and compares it with the table the handbook prints. It then imagines ExamReg's exam cell using the same plan to approve each day's fees: it checks 52 of the day's 2,000 computed fees by hand, and approves the day's batch if no more than 3 are wrong.

from math import comb

def p_accept(n, c, p):
    """Probability that a single sampling plan (n, c) accepts a lot whose fraction defective is p."""
    return sum(comb(n, d) * p**d * (1 - p)**(n - d) for d in range(c + 1))

# the NIST/SEMATECH e-Handbook's (52, 3) plan, and the probabilities its table prints
handbook = [0.998, 0.980, 0.930, 0.845, 0.739, 0.620, 0.502, 0.394, 0.300, 0.223, 0.162, 0.115]
print("defective  computed  handbook")
for k, printed in enumerate(handbook, 1):
    computed = f"{p_accept(52, 3, k / 100):.3f}"
    note = "" if computed == f"{printed:.3f}" else "   <- differs"
    print(f"{k / 100:>9.2f}  {computed:>8}  {printed:>8.3f}{note}")

# the same plan used by ExamReg's exam cell on one day's 2,000 computed fees
day = 2000
for p in (0.01, 0.05, 0.10):
    accept = p_accept(52, 3, p)
    print(f"{p:.0%} wrong: {day * p:.0f} wrong fees; batch accepted with probability {accept:.3f};"
          f" wrong fees in accepted batches, on average {day * p * accept:.0f}")
defective  computed  handbook
     0.01     0.998     0.998
     0.02     0.980     0.980
     0.03     0.930     0.930
     0.04     0.846     0.845   <- differs
     0.05     0.738     0.739   <- differs
     0.06     0.620     0.620
     0.07     0.502     0.502
     0.08     0.394     0.394
     0.09     0.300     0.300
     0.10     0.223     0.223
     0.11     0.162     0.162
     0.12     0.115     0.115
1% wrong: 20 wrong fees; batch accepted with probability 0.998; wrong fees in accepted batches, on average 20
5% wrong: 100 wrong fees; batch accepted with probability 0.738; wrong fees in accepted batches, on average 74
10% wrong: 200 wrong fees; batch accepted with probability 0.223; wrong fees in accepted batches, on average 45

The table. The formula reproduces the handbook's table at ten of its twelve points. At 4 and 5 per cent defective the exact binomial sum gives 0.846 and 0.738, where the handbook prints 0.845 and 0.739; the sums were checked again with exact fractions, and the handbook does not say how its table was computed. A difference in the third decimal changes nothing that follows, but it is a reminder that a printed table is a claim like any other, and a program can check it.

munotes.in483

The Quality Movement: Feigenbaum, Ishikawa, Crosby and TQM

The exam cell. The plan does exactly what it was designed to do. A day with 1 per cent of fees wrong, 20 students charged the wrong amount, is approved with probability 0.998. A day with 5 per cent wrong is approved with probability 0.738, and on average 74 of its 100 wrong fees go out in approved batches. Even a day with 10 per cent wrong is approved more than one time in five.

That is Crosby's objection in numbers. An acceptable quality level is a decision, made in advance, that some wrong fees will reach students; the sampling plan decides only how many, and how often. The handbook is candid about its purpose: "the main purpose of acceptance sampling is to decide whether or not the lot is likely to be acceptable, not to estimate the quality of the lot." Crosby's route to zero defects was prevention: for ExamReg, a fee function built from clear rules and tested by the techniques of Chapters Fifty-Four to Fifty-Six, from equivalence partitioning to decision tables. And where checking is still wanted, the handbook's reasons for sampling (testing that destroys the item, or 100 per cent inspection that costs too much or takes too long) rarely hold for software's output: a second program can recompute all 2,000 fees in a moment.

Total quality management

The ideas of both halves of the movement came together as total quality management. ASQ describes TQM as "a management system for a customer-focused organization that engages all employees in continual improvement of the organization", and "an integrative system that uses strategy, data, and effective communications to integrate the quality discipline into the processes, products, services, and culture of the organization." Its methods, ASQ says, are found in the teachings of Crosby, Deming, Feigenbaum, Ishikawa and Juran. The term itself "began initially as a term coined by the Naval Air Systems Command to describe its Japanese-style management approach to quality improvement."

ASQ lists eight principles of TQM. The right-hand column is this book's reading of each for a software team.

TQM principleFor a software team
Customer focusedRequirements and acceptance tests start from what students and the exam cell need
Employee involvementDevelopers, testers and analysts all own quality, not the SQA group alone
Process approachDefects are traced to the process step that let them in
Integrated systemDevelopment, testing and support work to the same quality objectives
Strategic and systematic approachQuality goals are part of the project's plan, not added at the end
Continual improvementCausal analysis after every release, measured against the next
Fact-based decision makingDecisions rest on defect and test data, not impressions
CommunicationsDefect reports, test reports and reviews that everyone can read
munotes.in484

The Quality Movement: Feigenbaum, Ishikawa, Crosby and TQM

ASQ is frank about what goes wrong. Among the common difficulties it lists "Lack of cooperation and teamwork among different workgroups" and "Management's failure to recognize and/or reward achievements". And it records that TQM as a name "has fallen out of favor as international standards for quality management have been developed": its principles now live in quality management systems, in the ISO 9000 series, first published in 1987 (Chapter Ninety-Four, on the ISO 9000 family), and in award programmes such as the Malcolm Baldrige National Quality Award, established the same year.

Six Sigma. The movement's later arrival, in ASQ's words, is "Six Sigma, a methodology developed by Motorola to improve its business processes by minimizing defects", which "evolved into an organizational approach that achieved breakthroughs and significant bottom-line results." ASQ's TQM page adds that Motorola, one of the first Baldrige winners, showed that TQM and Six Sigma were compatible. Chapter Eighty-Nine, on statistical software quality assurance and Six Sigma, computes its measures.

Software joins the movement

The capability maturity models brought the movement to software, and CMMI for Development tells the lineage in its own introduction. Shewhart's principles of statistical quality control were refined by W. Edwards Deming, Phillip Crosby and Joseph Juran; from there, in CMMI's own words, "Watts Humphrey, Ron Radice, and others extended these principles further and began applying them to software in their work at IBM (International Business Machines) and the SEI". Humphrey's book, Managing the Software Process, appeared in 1989. The premise the SEI took over is the movement's own, restated for software: "the quality of a system or product is highly influenced by the quality of the process used to develop and maintain it." Chapter Eighty-Four, on background issues in software quality assurance, follows the software side of the story from 1968, and Chapter One Hundred, on Lean and CMMI, describes the maturity levels.

The movement at a glance

WhenWhat
Late 13th centuryGuilds set quality rules and mark flawless goods
Mid-1920sShewhart turns to controlling processes; the control chart
1931Shewhart, Economic Control of Quality of Manufactured Product
Second World WarSampling inspection replaces inspecting every unit; Mil. Std. 105A follows in 1950
1950sDeming and Juran teach in Japan
1951Feigenbaum's book, later titled Total Quality Control
From the late 1950sIshikawa's courses for executives; company-wide quality control and quality circles
1964Crosby honoured by the Department of the Army for the concept of Zero Defects
1968Japan names its approach enterprise quality control
1979Crosby, Quality Is Free; Philip Crosby Associates founded
1986Deming, Out of the Crisis
1987The ISO 9000 series first published; the Baldrige award established
1989Humphrey, Managing the Software Process
munotes.in485

The Quality Movement: Feigenbaum, Ishikawa, Crosby and TQM

What it does not mean

Total quality control is not a bigger quality department. It is quality made part of every function's work, from requirements to support.

Zero defects is not a slogan for the work force. As Crosby used it, it is a standard for the management system; as a poster addressed to workers, it is what Deming's tenth point rejects.

A sampling plan does not measure quality. In the handbook's words, it decides "whether or not the lot is likely to be acceptable, not to estimate the quality of the lot."

Quality circles do not replace management. Ishikawa paired them with top management support, and his Deming Prize audit required the participation of the top executives.

TQM is not a certificate. It is a management approach; the standard an organisation can be certified against is ISO 9001 (Chapter Ninety-Five, on ISO 9001 and certification).

Quick revision

  • Feigenbaum: total quality control (first published 1951 as Quality Control: Principles, Practice, and Administration); the first text to sort quality costs into prevention, appraisal, internal failure and external failure.
  • Ishikawa: company-wide quality control, "top to bottom" and "start to finish"; quality circles (10 or fewer employees and their supervisor); the cause-and-effect diagram; the seven basic tools in Guide to Quality Control; top management support; the Deming Prize audit.
  • Crosby: the Four Absolutes: conformance to requirements, not goodness; prevention, not appraisal; Zero Defects, not acceptable quality levels; measured by the Price of Nonconformance, not indexes. Quality Is Free (1979).
  • Acceptance sampling: a (n, c) plan accepts a lot when a sample of n has c or fewer defectives; the plan is built to accept lots at the AQL; its purpose is to decide on the lot, not to estimate its quality.
  • TQM (ASQ): a customer-focused management system engaging all employees in continual improvement; eight principles; the name gave way to quality management standards such as ISO 9000 (1987).
  • Six Sigma: developed by Motorola to minimise defects. Software: Humphrey, Radice and others at IBM and the SEI; the premise that process quality drives product quality.
  • Worked example: the handbook's (52, 3) plan, recomputed (two values differ in the third decimal); a day with 5 per cent of 2,000 fees wrong is approved with probability 0.738, sending 74 wrong fees out on average.
munotes.in486

The Quality Movement: Feigenbaum, Ishikawa, Crosby and TQM

Test yourself

1. What is total quality control, and what did Feigenbaum contribute to the cost of quality? Quality control extended from inspection to every part of the organisation that affects quality, across the whole life of the product; Feigenbaum's book, first published in 1951, set it out, and it was the first text to classify quality costs as prevention, appraisal, internal failure and external failure.

2. What are quality circles, and what was Ishikawa's part in them? Small study groups of employees, ten or fewer, with their supervisor, who meet to improve quality in their own work; they originated in Japan as quality control circles. Ishikawa directed JUSE's Quality Control Circle Headquarters, edited its books on circles and played a major role in their growth, while insisting on top management support.

3. State Crosby's Four Absolutes of Quality Management. Quality means conformance to requirements, not goodness; quality is achieved by prevention, not appraisal; the performance standard is Zero Defects, not acceptable quality levels; quality is measured by the Price of Nonconformance, not indexes.

4. Why did Crosby reject acceptable quality levels? Use acceptance sampling to explain. An acceptable quality level is set in advance, and a sampling plan built on it accepts lots at that level with high probability, so some defects are knowingly passed on. With the (52, 3) plan, a day with 5 per cent of ExamReg's 2,000 fees wrong would be approved with probability 0.738, sending 74 wrong fees out on average. Crosby's standard is zero defects, reached by prevention.

5. What is total quality management? Name its principles. A management system for a customer-focused organisation that engages all employees in continual improvement (ASQ). Its eight principles are customer focus, employee involvement, a process approach, an integrated system, a strategic and systematic approach, continual improvement, fact-based decision making, and communications.

6. How did the quality movement reach software? Through the capability maturity models: Watts Humphrey, Ron Radice and others at IBM and the SEI applied the principles of Shewhart, Deming, Crosby and Juran to software processes, on the premise that the quality of a product is highly influenced by the quality of the process used to develop and maintain it.

Contents This chapter on its own page

munotes.in487

Chapter Eighty-Four

Background Issues in Software Quality Assurance

Syllabus topic Module 2, "Software Quality Assurance: ... Background issues and challenges in SQA"

In one line

Software quality assurance grew out of the manufacturing quality movement once software became large, costly and critical enough to fail in public: the 1968 NATO conference named software engineering and debated a software crisis; its papers already set quality assurance apart from quality control as an independent function speaking for the user; military procurement and then the IEEE turned such ideas into standards; and responsibility for quality settled on everyone who builds software, with an assurance function kept objective to check it.

In the wording a student can write in an examination: the background issues of SQA are the conditions it grew out of and still answers. Software grew faster than the ability to build it well (the 1968 NATO conference, where "software engineering" was a name "deliberately chosen as being provocative"). Its failures could be "a matter of life and death". Quality, as Dijkstra said there, "can never be established afterwards", so it had to be built into the process. Quality assurance was seen from the start as a function distinct from quality control, done "by an independently reporting agency representing the interests of the eventual user" (Bemer's checklist, 1968). Military procurement produced early software quality models and standards (McCall's factors for the U.S. Air Force, 1977; DOD-STD-2167A, 1988), and the IEEE's standard for SQA processes, IEEE 730, is now in its 2026 edition. Responsibility for quality lies with everyone on a project, while an SQA function, kept objective, gives the assurance.

From the factory to software

Chapters Eighty-Two and Eighty-Three, on the quality movement, followed quality from inspection to process control and total quality management, and ended where CMMI's introduction does: with Watts Humphrey and others applying the movement's principles to software. The background issues of this chapter are the reasons software needed the movement's ideas in a form of its own. Chapter Twenty-Three, on quality in software development, gave the technical reasons: software does not wear out, its defects are design defects, and its complexity is invisible. This chapter gives the historical ones: the moment the software field admitted it had a quality problem, and what it decided to do about it.

1968: the NATO conference

The NATO Science Committee's study group on computer science proposed a working conference of about fifty experts from computer manufacturers, universities, software houses and computer users; it met at Garmisch, Germany, from 7 to 11 October 1968. Its report explains the title: the phrase software engineering was "deliberately chosen as being provocative", implying that software manufacture should rest on the kind of theoretical foundations and practical disciplines traditional in the established branches of engineering.

Growth. The report's section on software and society begins with three quotations that "indicate the rate of growth of software". Helms: "In Europe alone there are about 10,000 installed computers", a number "increasing at a rate of anywhere from 25 per cent to 50 per cent per year." And d'Agapeyeff: "In 1958 a European general purpose computer manufacturer often had less than 50 software programmers, now they probably number 1,000-2,000 people; what will be needed in 1978?" The program carries these figures forward at the rates they imply. It is arithmetic on the conference's own numbers, not a record of what happened.

munotes.in488

Background Issues in Software Quality Assurance

# figures quoted at the 1968 NATO conference (the report, section 2), carried ten years forward
computers = 10_000                          # installed in Europe in 1968 (Helms)
for rate in (0.25, 0.50):                   # "from 25 per cent to 50 per cent per year"
    print(f"computers growing {rate:.0%} a year: {computers * (1 + rate) ** 10:,.0f} by 1978")

# d'Agapeyeff: fewer than 50 programmers at a manufacturer in 1958, 1,000 to 2,000 in 1968
for now in (1_000, 2_000):
    factor = now / 50
    rate = factor ** (1 / 10) - 1
    print(f"50 to {now:,} programmers in ten years: {factor:.0f} times, {rate:.1%} a year;"
          f" at that rate, {now * factor:,.0f} by 1978")
computers growing 25% a year: 93,132 by 1978
computers growing 50% a year: 576,650 by 1978
50 to 1,000 programmers in ten years: 20 times, 34.9% a year; at that rate, 20,000 by 1978
50 to 2,000 programmers in ten years: 40 times, 44.6% a year; at that rate, 80,000 by 1978

Ten years at Helms's rates would take Europe's computers from 10,000 to somewhere between about 93,000 and 577,000. D'Agapeyeff's figures mean the programming staff of one manufacturer grew at least 20 to 40 times in a decade, between 35 and 45 per cent a year, since the starting figure was "less than 50"; another decade at those rates would need 20,000 to 80,000 people. No discipline, testing included, could train people and invent methods that fast, and the report says the growth "was viewed with more alarm than pride."

Failure. The alarm was about what failures could do. David and Fraser wrote: "Particularly alarming is the seemingly unavoidable fallibility of large software, since a malfunction in an advanced hardware-software system can be a matter of life and death". Dijkstra put the risk in one line: "the massive dissemination of error-loaded software is frightening." Kinslow described IBM's OS/360 and TSS/360 as "straight-through, start-to-finish, no-test-development, revolutions", and added: "I have never seen an engineer build a bridge of unprecedented span, with brand new materials, for a kind of traffic never seen before".

The software crisis. Some participants called the situation a software crisis, and the name itself was argued over. Kolence did not like "the use of the word 'crisis'"; for him the problem was that "certain classes of systems are placing demands on us which are beyond our capabilities", and "It is large systems that are encountering great difficulties." Hastings, running large installations, found users "reasonably satisfied". Dijkstra welcomed the debate itself: "the admission of shortcomings is the primary condition for improvement."

munotes.in489

Background Issues in Software Quality Assurance

Quality cannot be added afterwards

One remark at the conference states the principle this whole module rests on. Discussing whether design could be separated from production, Dijkstra said: "I am convinced that the quality of the product can never be established afterwards. Whether the correctness of a piece of software can be guaranteed or not depends greatly on the structure of the thing made." It is Deming's third point, to cease dependence on inspection (Chapter Eighty-Two, on Shewhart, Deming and Juran), arrived at independently for software. Testing at the end can find defects; it cannot put quality in. That is why quality assurance is about the process that makes the software, from its first requirement.

1968: quality assurance as a separate function

The conference's working papers include R.W. Bemer's Checklist for planning software system production, and its ninth section is headed Quality Assurance. Its questions, written as a manager's checklist, contain the core of SQA as later standards define it.

Bemer's question, 1968What it became
"Is the quality assurance function recognized to be different from implicit and continuous quality control during fabrication"QA and QC as distinct (Chapter Twenty-Four, on quality control and quality assurance)
"Is software quality assurance done by an independently reporting agency representing the interests of the eventual user?"SQA independence (Chapter Twenty-Five, on quality management and SQA)
"Is the product tested to ensure that it is the most useful for the customer in addition to matching functional specifications?"Validation as well as verification (Chapter Twenty-Six, on verification and validation)
"Are they defined and constructed concurrently with the software?" (of the QA test programs)Tests designed from the start, not after coding
"Is at least one person engaged in software quality assurance for every ten engaged in its fabrication?"Staffing SQA as a planned part of the project
"Is this test library applied upon issuance of each modification of the software system?"Regression testing (Chapter Thirty-Nine, on regression and smoke testing)

The ratio of one in ten is Bemer's question, not a rule any later standard sets; its point is that assurance needs people of its own, planned from the start.

Military and government origins

Much of the early formal work on software quality was paid for by governments that bought large software systems. McCall, Richards and Walters's Factors in Software Quality, the model of Chapter Nineteen, on McCall's quality factors, was a 1977 technical report (RADC-TR-77-369) written by General Electric for the Rome Air Development Center of the U.S. Air Force Systems Command. The U.S. Department of Defense set standards for how its contractors developed software: Boehm noted in 1988 that its standard on software management, DoD-Std-2167, "requires that developers produce and use risk management plans", and the SEI's 1992 measurement reports cite DOD-STD-2167A, Military Standard, Defense System Software Development (1988). Space agencies keep standards of their own: NASA's software assurance standard, NASA-STD-8739.8B, was approved in September 2022.

munotes.in490

Background Issues in Software Quality Assurance

The pattern explains a feature of SQA that surprises students: its vocabulary of plans, audits, reviews and records comes from buyers who could not inspect software themselves and needed evidence that it had been built properly. That is what assurance means: "grounds for justified confidence that a claim has been or will be achieved" (ISO/IEC/IEEE 15026-1:2025, Chapter Twenty-Four).

The IEEE and ISO standards

The civilian standards followed. IEEE 730, IEEE Standard for Software Quality Assurance Processes, establishes requirements "for initiating, planning, controlling, and executing" the SQA processes of a software development or maintenance project (IEEE SA's page). Its 2014 edition superseded 730-2002, and in 2026 it was itself superseded by IEEE 730-2026, published on 21 August 2026 and "harmonized with the software life cycle processes of ISO/IEC/IEEE 12207:2017". The definition of SQA that SEVOCAB still carries is the 2014 edition's: the "set of activities that define and assess the adequacy of software processes to provide evidence that establishes confidence that the software processes are appropriate for and produce software products of suitable quality for their intended purposes". NASA's standard adopts the same definition, and adds: "A key attribute of software assurance is the objectivity of the software assurance function with respect to the project."

Two further standards frame the rest of this module. ISO/IEC/IEEE 12207, now in its 2026 edition, places quality assurance among the life cycle processes (Chapter Twenty-Four gave its definition), and ISO 9001, now ISO 9001:2026, is the quality management standard an organisation can be certified against (Chapters Ninety-Four and Ninety-Five, on the ISO 9000 family and ISO 9001). A textbook written before 2026 will cite older editions of all of these; the definitions this book quotes are the current ones, and it says where an older one is used.

Who is responsible for quality?

The history gives two answers, and both are right.

Everyone. Quality cannot be added afterwards, so it is made by everyone whose work goes into the product. The ISTQB syllabus says that QA "is the responsibility of everyone on a project", and Deming's fourteenth point says "The transformation is everybody's job."

And an assurance function that is objective. Bemer wanted an agency that reports independently and speaks for the user; IEEE 730-2014 defines SQA independence as freedom "from technical, managerial, and financial influences"; NASA makes objectivity "a key attribute". The assurance function does not replace the others' responsibility. It checks that the process is being followed and says so to people who can act. NASA's standard also names the condition for any of it to work: "Project and SMA Management support of the software assurance function is essential".

munotes.in491

Background Issues in Software Quality Assurance

On ExamReg's project the division looks like this. It is the book's own, and Chapter Eighty-Six, on SQA activities, describes the SQA group's side in full.

WhoTheir part in quality
The exam cell (the customer)States the rules and needs; takes part in reviews and acceptance testing
AnalystsRequirements that are complete, testable and say what happens to invalid input
DevelopersCode that conforms to the design; unit tests; code reviews
TestersTest design and execution at every level; defect reports
The project managerPlans, schedules and resources for reviews and testing
The SQA groupAudits of process and product; reports noncompliance to management, independently of the project
The college and the software house's managementQuality policy and objectives; the money and time for all of the above

What it does not mean

The software crisis was not the whole field failing. Participants at the conference itself disagreed; Kolence placed the difficulty in large systems, and Buxton said most computers worked tolerably well. The concern was about scale and criticality.

Quality assurance is not the SQA group's job alone. The group gives independent assurance; the quality itself is made by everyone on the project.

Independence is not isolation. An SQA function reports separately so that its findings reach people who can act; it still works with the project throughout.

Standards are not the history's end. Each edition replaces the one before; IEEE 730 alone has had editions in 2002, 2014 and 2026.

Quick revision

  • NATO 1968 (Garmisch, October): "software engineering" chosen as provocative; growth of 25 to 50 per cent a year in computers (Helms) and 20 to 40 times in programmers in a decade (d'Agapeyeff); failures "a matter of life and death"; the software crisis debated.
  • Dijkstra: "the quality of the product can never be established afterwards"; "the admission of shortcomings is the primary condition for improvement".
  • Bemer's checklist, section 9: QA distinct from QC; independent, speaking for the user; usefulness as well as the specification; tests built concurrently; one in ten; the test library rerun on each modification.
  • Military and government: McCall's factors for the U.S. Air Force (1977); DOD-STD-2167A (1988); NASA-STD-8739.8B (2022).
  • IEEE 730: SQA processes; 2002, 2014, now 2026; SQA definition (2014 edition, through SEVOCAB).
  • Responsibility: everyone on the project makes quality; an objective SQA function assures it; management support is essential.
  • Program: the conference's figures carried forward: 93,132 to 576,650 computers; 20,000 to 80,000 programmers per manufacturer by 1978 at the same rates.
munotes.in492

Background Issues in Software Quality Assurance

Test yourself

1. What were the main background issues that gave rise to software quality assurance? The rapid growth of software in size and number, faster than people and methods could keep up; the seriousness of failures in large and critical systems; the conviction that quality cannot be added after the product is built; and buyers, especially governments, who needed evidence that software had been built properly. The 1968 NATO conference brought these together and named software engineering.

2. Why was the phrase software engineering chosen for the 1968 conference? It was deliberately provocative: it implied that software manufacture should be based on the theoretical foundations and practical disciplines traditional in the established branches of engineering, which at the time it was not.

3. What was the software crisis? Did everyone at the conference agree there was one? The name some participants gave to the gap between what large software systems were expected to do and what could be built reliably, on time and within cost. They did not agree: Kolence disliked the word and placed the difficulty in large systems; Hastings found users reasonably satisfied; others, such as David and Fraser, stressed that failures could be a matter of life and death.

4. What did Bemer's 1968 checklist say about quality assurance? That QA is distinct from the continuous quality control of production; that it should be done by an independently reporting agency representing the user; that the product should be tested for usefulness as well as against its specification; that QA tests should be built concurrently with the software; that there should be at least one QA person for every ten in production; and that the test library should be rerun on every modification.

5. Who is responsible for software quality? Everyone on the project: analysts, developers, testers, managers and the customer each make part of it, since quality cannot be added afterwards. An SQA function, objective and independent of the project, gives assurance that the process is followed, and management must support it.

6. Name a military and an IEEE standard in the history of software quality. DOD-STD-2167A, Military Standard, Defense System Software Development (1988), and IEEE 730, the IEEE Standard for Software Quality Assurance Processes, whose current edition is IEEE 730-2026.

Contents This chapter on its own page

munotes.in493

Chapter Eighty-Five

The Challenges in Software Quality Assurance

Syllabus topic Module 2, "Software Quality Assurance: ... Background issues and challenges in SQA"

In one line

Software quality assurance is hard for reasons that belong to software itself and reasons that belong to the way it is made: the product is complex, must fit what others designed, keeps changing and cannot be seen; its requirements move; its schedules squeeze the checks at the end; its quality is hard to measure; its people find bad news hard to give and hear; and more and more of it is code the team did not write.

In the wording a student can write in an examination: the challenges in SQA are the conditions that make it hard to give justified confidence in software quality. Brooks's essential difficulties, "complexity, conformity, changeability, and invisibility", mean that assurance can never be complete, depends on parts outside the team's control, is valid for one version only, and must rest on evidence rather than inspection by eye. To them projects add changing requirements, which reopen verification; schedule pressure, which cuts testing and reviews and, in the ISTQB syllabus's words, makes people err, since "Humans make errors for various reasons, such as time pressure"; measurement, since quality has many characteristics and every count depends on its definition; people and culture, since testers "are often the bearers of bad news" and assurance can be seen as policing; and third-party and reused code, which the team did not write and may not be able to change.

The essential difficulties, seen from assurance

Chapter Twenty-Three, on quality in software development, set out the four properties Frederick Brooks called "the inherent properties of this irreducible essence of modern software systems: complexity, conformity, changeability, and invisibility", and what each asks of testing. For assurance, whose job is to give justified confidence, each property sets a limit.

Brooks's propertyThe challenge for assuranceHow SQA answers it
ComplexityConfidence can never be complete: no test set covers every stateRisk-based choice of what to check; several kinds of check (reviews, static analysis, tests)
ConformityQuality depends on interfaces and rules the team did not designAssurance of requirements and interfaces, and testing at the seams
ChangeabilityEvery assurance is for one version; a change can undo itConfiguration control, change review and regression testing
InvisibilityNobody can inspect software by looking at itEvidence: reviewed documents, records, measures and test results

The last row explains why so much of SQA is paperwork. A welding inspector can look at a weld; nobody can look at ExamReg's fee function and see that it is right. What SQA can examine is the evidence the process left behind: a requirements review record, a test report, a defect log, a coverage figure. When the evidence is missing, so is the assurance.

munotes.in494

The Challenges in Software Quality Assurance

Changing requirements

Requirements change, and modern methods treat that as normal: the Agile Manifesto's principles say "Welcome changing requirements, even late in development." Each change is also a challenge to assurance. A requirement that changes after the code is written reopens the design, the code, the tests written against the old version and every result that depended on them. ExamReg met this in miniature when the exam cell ruled that a negative number of days must be refused: the ruling arrived after the fee code was written (Chapter Eighty, on using defect data to improve the process, traced a class of defects back to it).

The answers are traceability and discipline, not resistance to change. If every requirement is linked to its design, code and tests, a change shows at once which tests must be revised and rerun (Chapter Seventy-Seven, on tracking defects to closure, used the same links to see which requirements still had open defects). Changes go through a change control board that weighs their cost; and regression testing (Chapter Thirty-Nine, on regression and smoke testing) confirms that what was not meant to change did not.

Schedule pressure

Release dates rarely move, and the activities at the end of a project, system and acceptance testing, absorb every delay before them. Schedule pressure also causes defects in the first place: the ISTQB syllabus lists time pressure first among the reasons "Humans make errors". The program below measures what squeezing the test phase would have cost release 2.0. It takes the ten test weeks of Chapter Seventy-Seven, on tracking defects to closure, and asks, for each week the testing might have been stopped at, how many defects would have gone into the release: those that testing found only in the later weeks, and those found but still open.

from itertools import accumulate

# release 2.0's ten test weeks (FINDINGS 5.2.5): defects found and defects closed each week
found = [10, 16, 18, 17, 14, 12, 10, 8, 5, 4]
closed = [4, 10, 14, 16, 16, 15, 13, 11, 8, 5]
found_by, closed_by = list(accumulate(found)), list(accumulate(closed))

print(f"{'testing stops':>13} {'found':>9} {'found only in':>15} {'found but':>11} {'left in the':>12}")
print(f"{'after week':>13} {'so far':>9} {'later weeks':>15} {'still open':>11} {'release':>12}")
for week in range(4, 11):
    later = found_by[-1] - found_by[week - 1]
    still_open = found_by[week - 1] - closed_by[week - 1]
    print(f"{week:>13} {found_by[week - 1]:>9} {later:>15} {still_open:>11} {later + still_open:>12}")
testing stops     found   found only in   found but  left in the
   after week    so far     later weeks  still open      release
            4        61              53          17           70
            5        75              39          15           54
            6        87              27          12           39
            7        97              17           9           26
            8       105               9           6           15
            9       110               4           3            7
           10       114               0           2            2
munotes.in495

The Challenges in Software Quality Assurance

With the full ten weeks, release 2.0 shipped with 2 known, accepted defects. Had the release date cut testing to six weeks, the 27 defects that testing found only in weeks 7 to 10 would have stayed in the product unfound, and the 12 open at the end of week 6 would have shipped with them: 39 defects instead of 2. Cutting to four weeks would have left 70. Every week counts: dropping only the last week would have shipped 7 defects instead of 2, and each earlier week dropped adds more, from 8 to 16.

Two cautions keep the reading honest. Some defects found late may have been introduced by fixes made during testing, and would not have existed had testing stopped earlier; the data cannot separate them, so the counts are an upper bound. And the 12 defects that students found after release are in none of these columns: they escaped the full ten weeks, and no length of this testing would have caught them. What the table does show is the challenge: when the schedule decides the length of testing, it also decides how many defects are shipped, and SQA's task is to make that trade visible to the people who set the date, with numbers like these.

Measuring quality

Quality is hard to assure partly because it is hard to measure. It is not one quantity: ISO/IEC 25010:2023 describes product quality with nine characteristics (Chapter Twenty, on the quality model today), and a release can improve in one and decline in another. The counts themselves depend on definitions. Chapter Seventy-Three, on what a defect is, gave four honest answers to how many defects did release 2.0 have?: 250, 227, 209 or 200, depending on what was counted. And some measures invite misuse: Chapter Sixty-Five, on what a software metric is, showed that averaging severity codes can reverse a comparison, because the codes are ordinal.

The challenge for SQA is to measure without misleading: to define each measure before collecting it, as the goal-question-metric approach of Chapter Sixty-Six asks, to use measures that fit the scale of the data, and to read any single number, such as a defect count or a pass rate, beside the others.

People and culture

Assurance is done by people to the work of other people, and that makes it delicate. The ISTQB syllabus is plain about it: "Testers are often the bearers of bad news. It is a common human trait to blame the bearer of bad news." Test results "may be perceived as criticism of the product and of its author", and "Confirmation bias can make it difficult to accept information that disagrees with currently held beliefs." An SQA group that audits a project's process meets the same reaction more strongly, because its findings go to management.

munotes.in496

The Challenges in Software Quality Assurance

Three answers recur in the sources. Communicate constructively: the syllabus asks that "information about defects and failures should be communicated in a constructive way." Remove fear: Deming's eighth point, "Drive out fear", because people who fear blame hide defects instead of reporting them. And share the responsibility: in the whole team approach "everyone is responsible for quality", so a defect report is information for the team, not an accusation. Independence, which SQA needs for its findings to be trusted, has to be combined with this; Chapter Twenty-Five, on quality management and SQA, described its three kinds.

Third-party and reused code

Every product contains code its team did not write: at the least a language's libraries and an operating system, and usually frameworks, services and components carried over from earlier systems. Brooks's conformity applies to all of it, and assurance has less to work with, since the team may not see the code's history, its tests or even its source.

Two cases in this book show the two sides of the risk. Ariane 5's inertial reference system reused software from Ariane 4, where it had worked; on Ariane 5's faster trajectory it failed, and the Inquiry Board found that the reused software had not been checked against the new rocket's flight conditions (Chapter Three, on why software must be tested). And a library can stop being maintained while applications still depend on it: Apache's page for Log4j 1 records that it reached end of life on 5 August 2015, and that "Since Log4j 1 is no longer maintained none of the issues listed will be fixed" (Chapter Nine, on test execution).

The assurance answers follow from the challenge. Know what the product contains: an inventory of every third-party component and its version. Know each component's support status, and plan to replace what is no longer maintained. Re-verify reused code against the new system's requirements and conditions, not the old one's. And test at the seams, where the product meets what it did not build.

The challenges at a glance

ChallengeWhy it is hardSQA's main answers
Essential difficultiesComplex, conforming, changing, invisibleRisk-based checks; evidence; configuration control
Changing requirementsEach change reopens verificationTraceability; change control; regression testing
Schedule pressureTesting absorbs delays; people err under time pressureMake the trade visible with data; protect reviews early
MeasurementMany characteristics; counts depend on definitionsDefined measures; the right scale; several measures together
People and cultureBad news, blame, confirmation biasConstructive reporting; no fear; the whole team owns quality
Third-party and reused codeNot written, seen or changeable by the teamInventory; support status; re-verification; testing at the seams
munotes.in497

The Challenges in Software Quality Assurance

What it does not mean

The challenges are not reasons to give up assurance. They explain why assurance gives justified confidence, never certainty, and why it must be planned rather than improvised.

Changing requirements are not a defect of the customer. They are normal; the challenge is to change the product under control, with traceability and regression testing.

A shorter test phase does not save its cost. In release 2.0's data, every week cut leaves defects in the product, to be found later by users, when fixing them costs more (Chapter Ninety-Six, on why reviews pay, gives the evidence on the cost of a late defect).

Third-party code is not someone else's quality problem. Once shipped in the product, its defects are the product's defects.

Quick revision

  • Brooks's essential difficulties for assurance: complexity (confidence never complete), conformity (depends on others' interfaces), changeability (assurance is per version), invisibility (evidence, not inspection).
  • Changing requirements: "Welcome changing requirements, even late in development" (Agile Manifesto); answers: traceability, change control, regression testing.
  • Schedule pressure: time pressure causes errors (ISTQB); cutting release 2.0's testing from ten weeks to six would have shipped 39 defects instead of 2; to four, 70.
  • Measurement: nine characteristics (ISO/IEC 25010:2023); 250, 227, 209 or 200 defects by definition; ordinal data.
  • People and culture: testers are "the bearers of bad news"; confirmation bias; constructive communication; drive out fear; the whole team approach.
  • Third-party and reused code: Ariane 5's reused software; Log4j 1's end of life (2015); inventory, support status, re-verification, testing at the seams.

Test yourself

1. How do Brooks's four essential difficulties make software quality assurance hard? Complexity means no set of checks covers every state, so confidence is never complete; conformity means quality depends on interfaces and rules the team did not design; changeability means any assurance holds for one version only and a change can undo it; invisibility means software cannot be inspected by looking, so assurance must rest on evidence such as reviews, records, measures and test results.

2. Why are changing requirements a challenge for SQA, and how is it met? Each change after work has been done reopens the design, code, tests and results that depended on the old requirement. It is met with traceability from requirements to design, code and tests, so the impact of a change is known; with change control that weighs each change; and with regression testing to confirm that nothing else changed.

3. How does schedule pressure affect software quality? Use release 2.0's data. Testing at the end absorbs earlier delays, and people make more errors under time pressure. Had release 2.0's testing been cut from ten weeks to six, the 27 defects found only in weeks 7 to 10 and the 12 still open at the end of week 6 would have shipped, 39 defects instead of 2.

munotes.in498

The Challenges in Software Quality Assurance

4. Why is software quality hard to measure? Quality has many characteristics, nine in ISO/IEC 25010:2023, which can move in different directions; even simple counts depend on definitions, as release 2.0's defects numbered 250, 227, 209 or 200 by four definitions; and some data, such as severity codes, is ordinal and cannot be averaged meaningfully.

5. What people and cultural challenges does SQA face, and how are they addressed? Testers and SQA staff bring bad news, which people tend to blame on the bearer; results can be felt as criticism, and confirmation bias resists unwelcome information. They are addressed by communicating defects constructively, removing fear of blame so that people report problems, and sharing responsibility for quality across the whole team.

6. What is the challenge of third-party and reused code, and what can SQA do about it? The team did not write it and may not see its tests or source, it may stop being maintained, and it may be used in conditions it was not built for, as Ariane 5's reused software was. SQA keeps an inventory of components and versions, tracks their support status, re-verifies reused code against the new system's requirements and tests the interfaces where the product meets it.

Contents This chapter on its own page

munotes.in499

Chapter Eighty-Six

SQA Activities: What the SQA Group Does

Syllabus topic Module 2, "Software Quality Assurance: ... Activities and approaches in Software Quality Assurance"

In one line

The SQA group gives the people doing the work, and their managers, an objective view of whether the project is following its own processes: it helps set those processes up, evaluates processes and work products against them, records every case of noncompliance, sees each one resolved in the project or escalated to management, keeps the records and reports the trends.

In the wording a student can write in an examination: CMMI states the purpose of process and product quality assurance as "to provide staff and management with objective insight into processes and associated work products". Its activities are "Objectively evaluating performed processes and work products against applicable process descriptions, standards, and procedures", "Identifying and documenting noncompliance issues", "Providing feedback to project staff and managers on the results of quality assurance activities" and "Ensuring that noncompliance issues are addressed". The SQA group also takes part in planning from the start, keeps records of its work and analyses noncompliance trends. A noncompliance that the project cannot resolve is escalated "to an appropriate level of management for resolution".

What the SQA group checks, and what testing checks

The first thing to understand about SQA activities is what they are not. CMMI draws the line between quality assurance and verification in one sentence: "The practices in the Process and Product Quality Assurance process area ensure that planned processes are implemented, while the practices in the Verification process area ensure that specified requirements are satisfied."

A tester asks whether ExamReg charges the right fee. The SQA group asks whether the fee module was built the way the project said it would be: whether its requirements were reviewed with the exam cell present, whether its code was reviewed, whether its unit tests were run and recorded, whether its defects were closed only after a confirmation test. Both can look at the same work product, a test report for instance, but from different sides; CMMI advises projects to use the overlap "to minimize duplication of effort while taking care to maintain separate perspectives." Chapter Twenty-Four, on quality control and quality assurance, drew the same line between checking the product and assuring the process.

Objectivity: independence and criteria

An evaluation is worth only as much as its objectivity. CMMI defines to objectively evaluate as "To review activities and work products against criteria that minimize subjectivity and bias by the reviewer", and says how it is achieved: "Objectivity is achieved by both independence and the use of criteria."

Independence. "Traditionally, a quality assurance group that is independent of the project provides objectivity." IEEE 730-2014 names three kinds of independence, technical, managerial and financial (Chapter Twenty-Five, on quality management and SQA). CMMI also allows another arrangement: in an organisation "with an open, quality oriented culture", quality assurance can be embedded in the process and done by peers, which may be the most feasible approach for a small organisation. It sets conditions: everyone doing it is trained in quality assurance; those who evaluate a work product are separate from those who produced it; and "An independent reporting channel to the appropriate level of organizational management should be available so that noncompliance issues can be escalated as necessary."

munotes.in500

SQA Activities: What the SQA Group Does

Criteria. Every evaluation works from criteria stated in advance. CMMI's subpractice lists what they settle: what will be evaluated, when or how often, how the evaluation will be conducted, and who must be involved. For ExamReg these are checklists, one for each process, built from the project's own standards.

The activities

Taking part in planning

SQA does not start with the first audit. CMMI says that "Quality assurance should begin in the early phases of a project to establish plans, processes, standards, and procedures that will add value to the project", and that those who do quality assurance take part in setting them up "to ensure that they fit project needs and that they will be usable for performing quality assurance evaluations." The processes and work products to be evaluated are chosen at this stage too. The result is the software quality assurance plan, the subject of Chapter Eighty-Seven.

Evaluating processes

The SQA group checks that the processes the project performs match their descriptions, standards and procedures: that a requirements review was held as the procedure says, with the people it names; that builds are tagged in version control; that defects are retested before they are closed. CMMI lists ways to do it, from formal to informal:

  • "Formal audits by organizationally separate quality assurance organizations"
  • "Peer reviews, which can be performed at various levels of formality"
  • "In-depth review of work at the place it is performed (i.e., desk audits)"
  • "Distributed review and comment of work products"
  • Process checks built into the processes themselves, such as a fail-safe that stops a process done incorrectly (CMMI's example is Poka-Yoke).

An audit, in CMMI's glossary, is "An objective examination of a work product or set of work products against specific criteria (e.g., requirements)", and the term covers process compliance audits as well as configuration audits.

Evaluating work products

The SQA group also checks selected work products (a requirements specification, a test plan, a release package) against the standards and procedures that apply to them, choosing them by documented sampling criteria when not everything can be checked. CMMI's examples of when to evaluate them: "Before delivery to the customer", "During delivery to the customer", "Incrementally, when it is appropriate", "During unit testing", "During integration" and "When demonstrating an increment".

munotes.in501

SQA Activities: What the SQA Group Does

Handling noncompliance

A noncompliance issue is a problem "identified in evaluations that reflect a lack of adherence to applicable standards, process descriptions, or procedures." Every one is identified and recorded. CMMI then sets the order of resolution. First, resolve it in the project, with the people concerned. Its examples of ways to do that are "Fixing the noncompliance", "Changing the process descriptions, standards, or procedures that were violated" and "Obtaining a waiver to cover the noncompliance". Only if the project cannot resolve it, escalate: "When noncompliance issues cannot be resolved in the project, use established escalation mechanisms to ensure that the appropriate level of management can resolve the issue." Either way, "Track noncompliance issues to resolution."

The second way deserves attention. A noncompliance does not always mean the people were wrong: sometimes the process description is, and the right resolution is to change it.

Keeping records and reporting

The group keeps "records of quality assurance activities" in enough detail that "status and results are known": evaluation logs, quality assurance reports, the status of corrective actions and reports of quality trends. It makes sure the people concerned hear the results in time, and it periodically reviews open noncompliance issues and trends with the manager designated to act on them.

Learning

Every evaluation also looks for "lessons learned that could improve processes", and the group analyses its noncompliance issues "to see if there are quality trends that can be identified and addressed". This is where SQA feeds the improvement loop of Chapter Eighty, on using defect data to improve the process.

ActivityCMMI practiceWhat it produces
Take part in planningIntroductory notesThe SQA plan; processes, standards and checklists that can be evaluated
Evaluate processesSP 1.1 Objectively Evaluate ProcessesEvaluation reports; noncompliance reports
Evaluate work productsSP 1.2 Objectively Evaluate Work ProductsEvaluation reports; noncompliance reports
Resolve or escalate noncomplianceSP 2.1 Communicate and Resolve Noncompliance IssuesCorrective actions; escalations; quality trends
Keep recordsSP 2.2 Establish RecordsEvaluation logs; QA reports; status of corrective actions

Public standards phrase the same work as tasks. NASA's software assurance standard, for example, pairs each engineering requirement with assurance tasks that begin "Confirm that" or "Assess": to "Confirm that all plans, including security plans, are in place and have expected content", or to "Assess plans for compliance". The verbs are the point. The assurance function does not write the plans or the code; it confirms and assesses them.

Worked example: release 2.0's SQA audits

ExamReg's software house has a small SQA group, independent of the project team and reporting to the head of delivery. During release 2.0 it held eight audits, one for each process in the SQA plan, each against a checklist. The program reads the audit log and produces the measures the group reported to management.

munotes.in502

SQA Activities: What the SQA Group Does

from statistics import mean, median

# release 2.0's SQA audits, in the order held (FINDINGS 5.4): (audit, items checked)
audits = [("requirements review", 12), ("design review", 10), ("coding standard", 15),
          ("unit testing", 12), ("configuration management", 10), ("defect tracking", 10),
          ("test reporting", 8), ("release readiness", 12)]
# each noncompliance: (id, audit, how it was resolved, days from report to resolution)
issues = [("NC-01", "requirements review", "fixed", 5), ("NC-02", "requirements review", "fixed", 3),
          ("NC-03", "design review", "fixed", 1), ("NC-04", "coding standard", "waiver", 7),
          ("NC-05", "coding standard", "fixed", 4), ("NC-06", "coding standard", "fixed", 2),
          ("NC-07", "unit testing", "fixed", 6), ("NC-08", "unit testing", "process changed", 10),
          ("NC-09", "configuration management", "fixed", 1),
          ("NC-10", "configuration management", "escalated", 21),
          ("NC-11", "defect tracking", "fixed", 3), ("NC-12", "test reporting", "fixed", 2),
          ("NC-13", "release readiness", "fixed", 1)]

print(f"{'audit':<26}{'items':>6}{'noncompliant':>14}{'compliance':>12}")
for name, items in audits:
    nc = sum(1 for _, audit, _, _ in issues if audit == name)
    print(f"{name:<26}{items:>6}{nc:>14}{(items - nc) / items:>12.0%}")
total = sum(items for _, items in audits)
print(f"{'all audits':<26}{total:>6}{len(issues):>14}{(total - len(issues)) / total:>12.1%}")

ways = {"fixed": "fixed in the project", "process changed": "process description changed",
        "waiver": "waiver granted", "escalated": "escalated to management"}
for how, label in ways.items():
    ids = [i for i, _, resolved, _ in issues if resolved == how]
    print(f"{label:<28}{len(ids):>3}  {' '.join(ids)}")
days = [d for *_, d in issues]
print(f"days to resolve: mean {mean(days):.1f}, median {median(days)}, longest {max(days)}")
print("over 14 days, for the manager's review:", " ".join(i for i, *_, d in issues if d > 14))
audit                      items  noncompliant  compliance
requirements review           12             2         83%
design review                 10             1         90%
coding standard               15             3         80%
unit testing                  12             2         83%
configuration management      10             2         80%
defect tracking               10             1         90%
test reporting                 8             1         88%
release readiness             12             1         92%
all audits                    89            13       85.4%
fixed in the project         10  NC-01 NC-02 NC-03 NC-05 NC-06 NC-07 NC-09 NC-11 NC-12 NC-13
process description changed   1  NC-08
waiver granted                1  NC-04
escalated to management       1  NC-10
days to resolve: mean 5.1, median 3, longest 21
over 14 days, for the manager's review: NC-10

Compliance. Of 89 checklist items, 13 were found out of compliance, 85.4 per cent compliance overall. The coding standard and configuration management audits were the weakest at 80 per cent, and the release readiness audit the best at 92. The numbers are only as meaningful as the checklists: an audit of 8 items and one of 15 are not the same test, and a trend is read across releases, audit by audit, not across different audits.

Resolution. Ten noncompliances were fixed in the project, most within a few days, as CMMI expects. Three show the other paths:

munotes.in503

SQA Activities: What the SQA Group Does

  • NC-08, process description changed. The unit testing standard asked for a coverage report for every module on every build, and nobody read them. The right resolution was to change the standard, to one report per release, not to force the team to produce reports nobody used.
  • NC-04, waiver granted. The coding standard sets a limit of 10 on cyclomatic complexity, and hall_ticket_status measured 15 (Chapter Seventy, on cyclomatic complexity). Splitting it days before release would have risked new defects, so a waiver was granted with a condition: split it in release 2.1. A waiver is a recorded decision, not an oversight.
  • NC-10, escalated to management. The database scripts were kept outside version control. The project lead declined to move them before release; the SQA group could not resolve it in the project, so it used its independent reporting channel and escalated to the head of delivery, who ordered the move. It took 21 days, the only issue over the 14-day threshold at which the group reviews open issues with the manager.

What the measures are for. The mean time to resolve, 5.1 days, is pulled up by the one escalation; the median, 3 days, describes the typical issue better, as Chapter Sixty-Eight, on quality, process and test metrics, showed for change times. And none of these numbers measures ExamReg's quality directly. They measure whether the project did what its plan said, which is exactly what process and product quality assurance is for; the product's quality is measured by the defect metrics of Chapter Seventy-Nine.

What it does not mean

The SQA group is not the test team. Testing checks the product against its requirements; SQA checks that the planned processes, testing among them, were followed.

An audit is not a hunt for culprits. Its findings are about adherence to processes, and one of CMMI's own resolutions is to change the process when it is the thing that is wrong.

Escalation is not the first step. Noncompliance is resolved in the project wherever possible; escalation is for what the project cannot resolve.

A waiver is not an exception quietly made. It is a decision recorded with its reason and its conditions, and tracked like any other issue.

High compliance is not high quality. A project can follow a weak process faithfully; compliance measures adherence, and the process itself must be judged by its results.

Quick revision

  • PPQA purpose (CMMI): "to provide staff and management with objective insight into processes and associated work products".
  • Against verification: PPQA ensures "planned processes are implemented"; verification that "specified requirements are satisfied".
  • Objectivity: "by both independence and the use of criteria"; an independent group, or QA embedded in the process with training, separation from the work product's authors, and an independent reporting channel.
  • Activities: take part in planning; objectively evaluate processes (SP 1.1) and work products (SP 1.2); communicate and resolve noncompliance issues (SP 2.1): fix, change the process description, waiver, or escalate; establish records (SP 2.2); analyse trends and lessons learned.
  • Evaluation methods: formal audits, peer reviews, desk audits, distributed review, built-in process checks.
  • Worked example: 8 audits, 89 items, 13 noncompliances (85.4 per cent compliance); 10 fixed, 1 process changed, 1 waiver, 1 escalated (21 days); mean 5.1 days, median 3.
munotes.in504

SQA Activities: What the SQA Group Does

Test yourself

1. What is the purpose of software quality assurance as CMMI states it, and how does it differ from verification? To provide staff and management with objective insight into processes and associated work products. SQA ensures that the planned processes are implemented; verification ensures that specified requirements are satisfied. Both may examine the same work product, from different perspectives.

2. List the activities of an SQA group. Taking part in establishing the project's plans, processes, standards and procedures; objectively evaluating performed processes against their descriptions, standards and procedures; objectively evaluating selected work products; identifying, documenting and communicating noncompliance issues and ensuring they are resolved, escalating those the project cannot resolve; keeping records of quality assurance activities; and analysing trends and lessons learned to improve processes.

3. How is objectivity achieved in SQA evaluations? By independence and by the use of criteria: evaluations are done against stated criteria by people who did not produce the work product, traditionally a quality assurance group independent of the project. Where QA is embedded in the process, its people are trained, separate from the work's authors, and have an independent reporting channel for escalation.

4. What is a noncompliance issue, and how is it resolved? A problem found in an evaluation that reflects a lack of adherence to applicable standards, process descriptions or procedures. It is resolved in the project if possible, by fixing the noncompliance, changing the process description that was violated, or obtaining a waiver; if the project cannot resolve it, it is escalated to the designated level of management; in every case it is tracked to resolution.

5. Give four ways of performing objective evaluations. Formal audits by an organisationally separate quality assurance group; peer reviews at various levels of formality; desk audits, reviewing work where it is performed; distributed review and comment of work products; and process checks built into the processes themselves.

6. In the worked example, why was NC-10 escalated, and what do the resolution times show? The project lead declined to put the database scripts under version control before release, so the SQA group could not resolve the issue in the project and escalated it to the head of delivery, who ordered it; it took 21 days. Most issues were resolved in a few days: the median was 3 days, while the mean of 5.1 days was pulled up by the escalation.

Contents This chapter on its own page

munotes.in505

Chapter Eighty-Seven

The Software Quality Assurance Plan

Syllabus topic Module 2, "Software Quality Assurance: ... Activities and approaches in Software Quality Assurance"

In one line

A software quality assurance plan says, before the work starts, what quality assurance will do on a project and how: its purpose and scope, its activities and methods, the people, resources and reporting lines, the standards to be audited and when, how noncompliances are tracked and escalated, the measures to be reported, what the assurance products must contain to be accepted, and how the plan itself is changed.

In the wording a student can write in an examination: the software quality assurance plan (SQAP) is the document that defines the SQA activities for a project. IEEE 730 sets requirements "for initiating, planning, controlling, and executing" a project's SQA processes; its current edition is IEEE 730-2026. NASA's handbook describes the same document in public: the plan "provides insight into the methods, approaches, responsibilities, and processes for the assurance activities of all life cycle and mission phases". It covers the introduction (purpose, scope, overview), the activities and methods, the stakeholders, resources, roles and organisation, data management and acceptance criteria for its products, risk and safety, training and communication, the metrics, issue tracking, a glossary, the change procedure and the schedule; a topic that does not apply is marked not applicable rather than left out.

Why assurance needs a plan of its own

Chapter Eighty-Six, on SQA activities, showed what the SQA group does. Each of those activities needs decisions made in advance: which processes and work products will be evaluated, against which criteria, when, by whom, and to whom the findings go. CMMI's process area on quality assurance says those decisions come first: "Quality assurance should begin in the early phases of a project to establish plans, processes, standards, and procedures that will add value to the project". The SQA plan is where they are written down, agreed with the project and management, and made checkable.

A plan also protects the assurance function's objectivity. When the reporting line, the escalation path and the audits are agreed before any finding exists, nobody can argue later that an unwelcome audit was improvised.

The SQA plan is not the test plan. A test plan is a "detailed description of test objectives to be achieved and the means and schedule for achieving them, organized to coordinate testing activities" (ISO/IEC/IEEE 29119-2:2021), the document of Chapter Eleven, on writing a test plan. The SQA plan covers the assurance of every process, testing among them: one of its audits checks that the test plan is being followed.

What a plan contains

IEEE 730-2026 and its 2014 predecessor define a plan's content, but both are sold by the IEEE and neither is on disk for this book, so it does not reproduce their outline; a textbook may show one from an earlier edition. NASA's Software Engineering Handbook publishes an equivalent outline, the "minimum recommended content for a Software Assurance Plan (SAP)", and the table groups its topics.

munotes.in506

The Software Quality Assurance Plan

GroupTopics (NASA's minimum content)What the plan states
IntroductionPurpose; scope; overviewWhy the plan exists, what it covers, how it is organised
The workAssurance activities; assurance methodsPlanned audits and assessments, status reporting, analyses; how they are done (reviews attended, products and processes reviewed, tests witnessed, issues reported)
The peopleStakeholders; resources; roles and responsibilities; organisation and management; training; communicationWho is involved; the staff, tools and access SQA needs; who does what; to whom SQA reports; what SQA staff must learn; how findings travel
The productsData management; acceptance criteriaThe assurance products, where they are stored and for how long, and when each is accepted
RiskSafety-critical assessment (if needed); risk managementWhether the software is safety-critical; how risks SQA finds are handled
ControlRequirements mapping; metrics; issue tracking and reportingWhat will be checked against what; the measures and how they are reported; how problems are reported, tracked and resolved
HousekeepingAcronyms; glossary; change procedure and historyTerms; how the plan is changed and its history kept
TimeScheduleThe assurance activities and audits, aligned with the project's schedule and milestones

Two rules in the handbook matter as much as the list. "If a content section does not apply", the plan says so, marked not applicable, instead of silently leaving it out: an absent topic could be an oversight, and a marked one is a decision. And the schedule must align the audits "with the project schedule, milestones, and life cycle products", because an audit held after the milestone it should inform is too late to matter.

Two of NASA's entries belong to NASA's own rules: a classification of the software under NASA's procedural requirements, and a matrix mapping NASA's requirements to assurance tasks. A plan outside NASA replaces the first with its own statement of how critical the software is, and the second with the list of standards and procedures the project must follow and the audits that check each one. ExamReg's plan does exactly that.

Worked example: reviewing ExamReg's plan

ExamReg's SQA engineer wrote a first draft of the plan for release 2.0, one line per topic followed by the standards to be audited and the audit schedule. Reviewing a plan is itself quality assurance, and two checks are mechanical enough for a program: does the plan address every required topic, and does every standard the project must follow have an audit that checks it? The program reads draft 1, then draft 1 with the two changes the review led to.

munotes.in507

The Software Quality Assurance Plan

# ExamReg release 2.0: software quality assurance plan, draft 1 (one line per topic)
purpose: assure that release 2.0 follows its planned processes, and report the results objectively
scope: release 2.0 of ExamReg, all seven modules, from the requirements review to the release
overview: one line for each topic, then the standards to be audited and the audit schedule
activities: the process audits below; a status report every two weeks; a trend report at release
methods: audits against checklists; attending requirements and design reviews; desk audits of code
stakeholders: the exam cell; the project lead; the test lead; the head of delivery
resources: one SQA engineer, one day a week; read access to the repository and defect tracker
roles: SQA audits and reports; the project lead resolves noncompliances; escalations go up
organisation: the SQA engineer reports to the head of delivery, not to the project lead
data management: audit reports and the noncompliance log, in version control, kept three years
safety: not applicable; ExamReg controls no physical process and endangers no one
risk management: risks that SQA finds go on the project's risk register
training: the SQA engineer is trained in the coding standard and the review procedure
communication: findings to the project lead in two days; open issues reviewed fortnightly
metrics: compliance per audit; noncompliances by process; days to resolve; escalations
issue tracking: noncompliance log NC-nn; resolved in the project, or escalated after 14 days
glossary: SQA software quality assurance; NC noncompliance; V(G) cyclomatic complexity
change procedure: changes approved by the head of delivery and recorded with their date
schedule: each audit is held in the phase it checks, before that phase's exit review
standard: requirements template
standard: review procedure
standard: coding standard
standard: unit test standard
standard: configuration management procedure
standard: defect tracking procedure
standard: test reporting procedure
standard: release procedure
audit: requirements; requirements review; requirements template, review procedure
audit: design; design review; review procedure
audit: coding; coding standard; coding standard
audit: testing; unit testing; unit test standard
audit: testing; defect tracking; defect tracking procedure
audit: testing; test reporting; test reporting procedure
audit: release; release readiness; release procedure
acceptance criteria: an audit report is accepted when each finding has an NC number and an owner
audit: testing; configuration management; configuration management procedure
# the topics a plan must address: NASA's minimum content (SWEHB Topic 5.17), adapted for ExamReg
REQUIRED = ["purpose", "scope", "overview", "activities", "methods", "stakeholders", "resources",
            "roles", "organisation", "data management", "acceptance criteria", "safety",
            "risk management", "training", "communication", "metrics", "issue tracking",
            "glossary", "change procedure", "schedule"]

def read(*paths):
    plan, standards, audits = {}, [], []
    for path in paths:
        with open(path) as f:
            for line in f:
                if line.startswith("#") or not line.strip():
                    continue
                key, value = (part.strip() for part in line.split(":", 1))
                if key == "standard":
                    standards.append(value)
                elif key == "audit":
                    phase, name, checks = (part.strip() for part in value.split(";"))
                    audits.append((phase, name, [c.strip() for c in checks.split(",")]))
                else:
                    plan[key] = value
    return plan, standards, audits

for draft, paths in [("draft 1", ["examreg-sqa-plan.txt"]),
                     ("draft 2", ["examreg-sqa-plan.txt", "draft-2-changes.txt"])]:
    plan, standards, audits = read(*paths)
    missing = [topic for topic in REQUIRED if topic not in plan]
    audited = {s for _, _, checks in audits for s in checks}
    unaudited = [s for s in standards if s not in audited]
    print(f"{draft}: {len(audits)} audits for {len(standards)} standards")
    print(f"   topics missing: {', '.join(missing) or 'none'}")
    print(f"   standards with no audit: {', '.join(unaudited) or 'none'}")
munotes.in508

The Software Quality Assurance Plan

draft 1: 7 audits for 8 standards
   topics missing: acceptance criteria
   standards with no audit: configuration management procedure
draft 2: 8 audits for 8 standards
   topics missing: none
   standards with no audit: none

Draft 1 had two gaps. It never said when an assurance product counts as done, so a vague audit report could not be sent back; and the project's configuration management procedure, which governs how builds are tagged and what is under version control, had no audit at all. Neither gap is visible by reading the plan line by line, because nothing on the page is wrong; both are things that are absent.

Draft 2 adds an acceptance criterion for audit reports, that each finding has a noncompliance number and an owner, and a configuration management audit during testing. It addresses all twenty topics and audits all eight standards. The change mattered: the configuration management audit, when it was held, found the two noncompliances NC-09 and NC-10, one of which had to be escalated (Chapter Eighty-Six, on SQA activities). Without draft 2, nobody would have looked.

The program checks only what can be checked mechanically. Whether one day a week is enough SQA effort, whether 14 days is the right time before escalation, whether the checklists behind each audit are any good: those are judgements, for the head of delivery and the project lead to make when they approve the plan.

Reading ExamReg's plan against the outline

A few lines of the plan show what each topic asks for in practice.

  • Safety: not applicable, with a reason. ExamReg controls no physical process; the plan says so instead of omitting the topic.
  • Organisation. The SQA engineer reports to the head of delivery, not to the project lead whose work is audited: the managerial independence of Chapter Twenty-Five, on quality management and SQA, written into the plan.
  • Issue tracking. A numbered noncompliance log, resolution in the project first, escalation after 14 days: the rules of CMMI's practice for resolving noncompliance, turned into ExamReg's procedure.
  • Metrics. The four measures that the program of Chapter Eighty-Six, on SQA activities, computed, decided before the first audit, so the report at release uses definitions agreed at the start.
  • Change procedure. Draft 2 itself went through it: approved by the head of delivery and recorded with its date.
munotes.in509

The Software Quality Assurance Plan

What it does not mean

The SQA plan is not the test plan. The test plan organises testing; the SQA plan organises the assurance of all the processes, testing included.

A plan is not finished when it is written. It is reviewed, approved, and changed under its own change procedure as the project learns, as draft 2 shows.

Not applicable is not the same as missing. A topic that does not apply is marked so, with the reason; a topic left out is a gap.

Completeness is not adequacy. A plan can mention every topic and still assign too little effort or weak checklists; a program can find gaps, and only people can judge sufficiency.

Quick revision

  • SQAP (IEEE 730-2014, now IEEE 730-2026): the plan that defines a project's SQA activities; written early, as CMMI says QA "should begin in the early phases of a project".
  • Contents (NASA's minimum content, Topic 5.17): purpose, scope, overview; activities and methods; stakeholders, resources, roles, organisation, training, communication; data management and acceptance criteria; safety assessment and risk management; requirements mapping, metrics, issue tracking; acronyms, glossary, change procedure; schedule.
  • Rules: mark a topic that does not apply as not applicable; align audits with milestones.
  • Against the test plan: the test plan covers testing; the SQA plan covers assurance of all processes.
  • Worked example: draft 1 missed acceptance criteria and had no configuration management audit; draft 2 addressed all 20 topics and audited all 8 standards; the added audit found NC-09 and NC-10.

Test yourself

1. What is a software quality assurance plan, and why is it prepared at the start of a project? The document that defines what quality assurance will do on a project and how: activities, methods, responsibilities, reporting lines, the standards to be audited and the schedule. It is prepared early because evaluations need criteria and a schedule agreed in advance, because QA helps set up the processes it will later check, and because an agreed reporting and escalation path protects its objectivity.

2. List the main contents of an SQA plan. Introduction (purpose, scope, overview); the assurance activities and methods; stakeholders, resources, roles and responsibilities, organisation, training and communication; data management and acceptance criteria for the assurance products; safety assessment and risk management; the standards to be checked, the metrics and issue tracking; acronyms, glossary and the change procedure; and the schedule of assurance activities and audits.

munotes.in510

The Software Quality Assurance Plan

3. How does an SQA plan differ from a test plan? A test plan describes the objectives, means and schedule of testing for a test item. An SQA plan describes the assurance of all the project's processes and products, including testing, by audits and reviews against standards; one of its audits checks that the test plan is followed.

4. Why must a plan state "not applicable" rather than leave a topic out? Because an absent topic cannot be told apart from an oversight, while a topic marked not applicable, with a reason, records a decision that a reviewer can check.

5. What did the review of ExamReg's draft plan find, and why did it matter? The draft had no acceptance criteria for assurance products and no audit of the configuration management procedure. Draft 2 added both. The configuration management audit later found two noncompliances, one of which had to be escalated; without the change it would not have been held.

6. Which parts of a plan's review can be automated, and which cannot? Mechanical checks can: whether every required topic is present and whether every standard has an audit scheduled. Judgements cannot: whether the effort assigned is enough, whether the escalation time is right, and whether the checklists are good; those are for the people who approve the plan.

Contents This chapter on its own page

munotes.in511

Chapter Eighty-Eight

Approaches to SQA: Formal Methods, Cleanroom and Process Models

Syllabus topic Module 2, "Software Quality Assurance: ... Activities and approaches in Software Quality Assurance"

In one line

Software quality can be assured in four broad ways, which real projects combine: by proving the software correct against its specification (formal methods); by building it so that defects are prevented and certifying it with statistical, usage-based testing (cleanroom); by making the process capable of producing quality (process models such as CMMI and ISO 9001); and by examining the product itself with reviews, measurement and testing (the product approach).

In the wording a student can write in an examination: a proof of correctness is a "formal technique used to prove mathematically that a computer program satisfies its specified requirements" (ISO/IEC/IEEE 24765:2017), and it needs a formal specification, one "written in a formal notation, often for use in proof of correctness". Cleanroom software engineering is "a theory-based, team-oriented process for development and certification of high-reliability software systems under statistical quality control" (SEI, 1996). Its development teams verify correctness before any execution, and an independent certification team tests by random samples from a model of use. The process approach, taken by CMMI and ISO 9001, rests on the premise that "the quality of a system or product is highly influenced by the quality of the process used to develop and maintain it" (CMMI). The product approach evaluates the software itself against its quality requirements.

Why there is more than one approach

Testing alone cannot assure quality, and the reason was put plainly in 1972. Dijkstra, in his Turing lecture: "program testing can be a very effective way to show the presence of bugs, but is hopelessly inadequate for showing their absence." Chapter One, on what software testing is, met that sentence as a limit of testing. Dijkstra drew a conclusion from it: "The only effective way to raise the confidence level of a program significantly is to give a convincing proof of its correctness." He did not mean proving a finished program afterwards: "the programmer should let correctness proof and program grow hand in hand."

Each approach in this chapter answers the limit in its own way. Formal methods replace sampling with reasoning. Cleanroom combines reasoning during development with sampling that is statistically sound. Process models improve the way software is made, so that fewer defects are made. And the product approach keeps examining what was made, knowing that it can find defects but not prove their absence.

Formal methods and proof of correctness

A proof of correctness shows, by mathematical reasoning, that a program does what its specification says for every input it can receive, not only for the inputs a test tried. The standards give the related terms. Formal verification is the "activity proving or disproving the correctness of intended applications with respect to a formal specification or a property, using formal methods of mathematics" (ISO/IEC 23643:2020). The ISTQB syllabus counts formal methods, "model checking and proof of correctness", among the forms of quality control besides testing (Chapter Twenty-Four, on quality control and quality assurance).

munotes.in512

Approaches to SQA: Formal Methods, Cleanroom and Process Models

Three facts about proofs decide where they are used.

  1. A proof is relative to a specification. It shows that the program meets the specification, not that the specification is right; a wrong specification is proved faithfully.
  2. A proof holds only under its preconditions. It states what the program may assume about its inputs, and says nothing outside them.
  3. Proofs cost effort. The cleanroom reference model reserves written proofs for where they are worth it: "Written proofs of correctness based on function-theoretic techniques provide additional rigor if necessary for life-, mission- and enterprise-critical software."

Worked example 1: verifying ExamReg's late fee band

Cleanroom verifies a program structure by structure, not path by path. Its Correctness Theorem gives, for each basic control structure, a question to answer about its intended function f; for an ifthenelse whose predicate is p, whose then-part is g and whose else-part is h, the question is: "When p is true does g do f, and when p is false does h do f?" Because every program has only finitely many control structures, the theorem "reduces verification to a finite number of checks".

ExamReg's late_band, from Chapter Seventy, on cyclomatic complexity, is two ifthenelse structures in a row; each if that returns acts as an ifthenelse whose else-part is the rest of the function. Its intended function comes from the fee rules: for a whole number of days from 0 to 15, give on time for 0, 1 to 7 days for 1 to 7, and 8 to 15 days for 8 to 15. The precondition, that the number is a whole number from 0 to 15, is established by total_fee, which refuses everything else before calling it.

StructureConditionAnswer, and why
First ifthenelse, days_late == 0When true, does returning on time do f?Yes: f gives on time for 0
When false, does the rest do f?It must do f for 1 to 15; this is the second structure's intended function
Second ifthenelse, days_late <= 7When true, does returning 1 to 7 days do f for 1 to 15?Yes: the value is not 0 and, by the precondition, not negative, so it is 1 to 7
When false, does returning 8 to 15 days do f for 1 to 15?Yes: the value is over 7 and, by the precondition, at most 15

Both conditions of both structures hold, so late_band is correct with respect to its specification. In cleanroom this reasoning is a verbal proof agreed by the team in a review. Because the domain here is tiny, sixteen whole numbers, the proof can also be confirmed by exhaustion: checking every input is itself a proof when the inputs are finite, the idea that model checking scales up.

munotes.in513

Approaches to SQA: Formal Methods, Cleanroom and Process Models

def late_band(days_late):                    # ExamReg's function, exactly as in Chapter 70
    if days_late == 0:
        return "on time"
    if days_late <= 7:
        return "1 to 7 days"
    return "8 to 15 days"

# the specification: the late fee bands of the fee rules, for 0 to 15 days (total_fee refuses the rest)
BANDS = [(0, 0, "on time"), (1, 7, "1 to 7 days"), (8, 15, "8 to 15 days")]
domain = [(d, band) for low, high, band in BANDS for d in range(low, high + 1)]
wrong = [d for d, band in domain if late_band(d) != band]
print(f"every valid input checked: {len(domain)}; disagreements with the specification: {wrong or 'none'}")
print(f"outside the precondition, late_band(-1) gives {late_band(-1)!r}")
every valid input checked: 16; disagreements with the specification: none
outside the precondition, late_band(-1) gives '1 to 7 days'

All sixteen valid inputs agree with the specification. The last line shows the second fact about proofs: given -1, outside the precondition, the function answers 1 to 7 days, and a student early by a day would be charged a late fee if anything ever called it that way. The proof never claimed otherwise; it rested on total_fee refusing negative days, and the second condition's answer used that precondition explicitly. This is exactly the defect class that the exam cell's ruling on negative days (Chapter Eight, on test design techniques) and the input validation defects of Chapter Eighty, on using defect data, were about: the proof makes the dependency visible instead of leaving it to chance.

Cleanroom software engineering

The SEI's reference model explains the name: "The Cleanroom name is borrowed from hardware Cleanrooms, with their emphasis on rigorous engineering discipline and focus on defect prevention rather than defect removal." Its objective is ambitious: "A principal objective of the Cleanroom process is development of software that exhibits zero failures in use."

Four practices make up the process.

  • Incremental development. The software is specified, built and certified in a pipeline of increments that accumulate into the final system, so quality is measured early and continually.
  • Box structures. Specification and design proceed from black boxes (behaviour as seen from outside), to state boxes (with the data kept between uses), to clear boxes (procedures), each checked against the one before.
  • Correctness verification before execution. "Correctness verification by development teams is used to identify and eliminate defects prior to any execution of the software." The developers do not test their own code; they verify it, structure by structure, in team reviews.
  • Statistical testing by an independent team. "Software execution is controlled by an independent certification team that uses statistical testing methods to evaluate software quality."
munotes.in514

Approaches to SQA: Formal Methods, Cleanroom and Process Models

The results reported are strong. The SEI's report cites quality "Improvements of 10 to 20X and substantially more over baseline performance", and an IBM device controller that "exhibited no failures in three years use at over 300 customer locations". Independent of the SEI, Boehm and Basili (2001) report that "Data from the use of Cleanroom at NASA have shown 25 to 75 percent reductions in failure rates during testing."

Worked example 2: statistical usage testing

Cleanroom's view of testing starts from a fact: "The set of possible executions of a software system is an infinite population. All testing is really sampling from that infinite population." If the sample is random and drawn according to how the software will be used, the results support statistical statements about quality in use. The certification team therefore builds a usage model, usually a Markov chain whose states are states of use and whose arcs carry the probability of each next step. Then "Test cases are randomly generated from the usage models, so that every test case represents a possible use of the software as defined by the models."

ExamReg's usage model below is the book's own. A student opens the portal and logs in; most fill the exam form, some only download a hall ticket, a few log out at once; a payment sometimes fails and is retried. The program does two things the reference model describes: a standard calculation on the model, the expected number of visits to each state in one scenario (and so the expected test case length), and the generation of 1,000 test cases from it.

import random

# ExamReg's usage model: from each state of use, the next state and its probability (the book's own)
MODEL = {"invoke":           [("log in", 1.0)],
         "log in":           [("fill form", 0.55), ("hall ticket", 0.30), ("log out", 0.15)],
         "fill form":        [("pay fee", 0.80), ("log out", 0.20)],
         "pay fee":          [("hall ticket", 0.70), ("payment failed", 0.10), ("log out", 0.20)],
         "payment failed":   [("pay fee", 0.60), ("log out", 0.40)],
         "hall ticket":      [("log out", 1.0)],
         "log out":          [("exit", 1.0)]}
STATES = list(MODEL)

# a standard calculation: expected visits to each state in one scenario, summed step by step
expected, now = dict.fromkeys(STATES, 0.0), {"invoke": 1.0}
for _ in range(200):
    step = {}
    for state, p in now.items():
        expected[state] += p
        for nxt, q in MODEL[state]:
            if nxt != "exit":
                step[nxt] = step.get(nxt, 0.0) + p * q
    now = step

# statistical usage testing: test cases generated at random from the model
rng = random.Random(88)
def test_case():
    path, state = [], "invoke"
    while state != "exit":
        path.append(state)
        nexts, weights = zip(*MODEL[state])
        state = rng.choices(nexts, weights)[0]
    return path

cases = [test_case() for _ in range(1000)]
print(f"{'state of use':<16}{'expected visits':>16}{'in 1,000 cases':>16}")
for state in STATES:
    seen = sum(case.count(state) for case in cases) / len(cases)
    print(f"{state:<16}{expected[state]:>16.3f}{seen:>16.3f}")
print(f"test case length: expected {sum(expected.values()):.3f} states,"
      f" generated {sum(map(len, cases)) / len(cases):.3f}")
print("cases that include a failed payment:", sum("payment failed" in c for c in cases))
munotes.in515

Approaches to SQA: Formal Methods, Cleanroom and Process Models

state of use     expected visits  in 1,000 cases
invoke                     1.000           1.000
log in                     1.000           1.000
fill form                  0.550           0.545
pay fee                    0.468           0.473
payment failed             0.047           0.044
hall ticket                0.628           0.634
log out                    1.000           1.000
test case length: expected 4.693 states, generated 4.696
cases that include a failed payment: 42

The calculation and the sample agree. The model says a scenario visits the payment page 0.468 times on average (the 0.55 who fill the form times the 0.80 who go on to pay, raised a little by retries after failed payments), and the 1,000 generated test cases visited it 0.473 times each. The expected test case length, 4.693 states, is matched to within 0.003. That agreement is what makes the results of such testing statistically meaningful: the test cases are a sample of use in the proportions of use.

What usage testing emphasises. Test effort follows use: the generated cases visit the payment page 0.473 times each on average, and 634 of the 1,000 reach the hall ticket, which a scenario can visit only once. The reference model notes the benefit: "Because statistical usage testing tends to detect errors with high failure rates, it is an efficient approach to improving software reliability." It also shows the cost. Only 42 of the 1,000 cases include a failed payment, the state where a defect could charge a student twice. For such states the reference model allows separate models "to provide independent certification of stress situations or infrequently used functions with high consequences of failure", and ExamReg's certification team would build one for payment failures.

Process models: CMMI and ISO 9001

The process approach assures quality by assuring the process. CMMI's introduction states the premise it inherited from the quality movement: "the quality of a system or product is highly influenced by the quality of the process used to develop and maintain it." The ISTQB syllabus states the same assumption for quality assurance: "if a good process is followed correctly, then it will generate a good product."

Two families of process model dominate. CMMI describes the practices of capable organisations in process areas, such as the process and product quality assurance of Chapter Eighty-Six, on SQA activities, grouped into maturity levels (Chapter One Hundred, on Lean and CMMI). ISO 9001 states requirements for a quality management system, against which an organisation can be certified (Chapters Ninety-Four and Ninety-Five, on the ISO 9000 family and ISO 9001). Neither says how to write a particular program. Both ask whether the organisation plans, controls, measures and improves the way it writes programs.

munotes.in516

Approaches to SQA: Formal Methods, Cleanroom and Process Models

The approach has a limit that the audits of Chapter Eighty-Six, on SQA activities, met. A process can be followed faithfully and still be a poor process; compliance measures adherence, not results. Process models therefore insist on measuring outcomes and improving the process when they fall short.

The product approach

The product approach examines what was made: reviews and inspections of work products (Chapters Twenty-Eight to Thirty, on reviews), testing at every level, and measurement of the product against a quality model such as ISO/IEC 25010:2023 (Chapter Twenty, on the quality model today) with the metrics of Chapters Sixty-Five to Seventy-Two, from what a metric is to complexity and basis path testing. Its measure of success is the one IEEE 730-2014 gives for software quality: the "degree to which a software product meets established requirements".

Its strength is directness: it looks at the thing the users will get. Its limit is Dijkstra's: examination finds the defects it samples, and cannot show that none remain. That is why no serious SQA programme uses the product approach alone.

The approaches compared

ApproachWhat it assuresEvidenceMain limit
Formal methodsThe program meets its specification for every input satisfying the preconditionsProofs; model checkingOnly as good as the specification; costly, so kept for critical parts
CleanroomVery few defects enter testing; reliability in use is certifiedVerification reviews; statistical usage test resultsNeeds discipline and training; certification is only as good as the usage model
Process modelsThe organisation's process is defined, followed and improvedAudits; assessments; process measuresA followed process can still be a poor one
Product approachThe product, as built, meets its requirements where it was examinedReview and test results; product metricsShows the presence of defects, never their absence

The approaches are not rivals. ExamReg's project uses all four in proportion: a verbal proof for a function as small and critical as late_band, usage-based test selection, a process audited against its SQA plan, and reviews and tests of the product.

What it does not mean

A proof does not make a program correct in every sense. It shows that the program meets its specification under its preconditions; a wrong specification, or an input outside the preconditions, is outside the proof.

Cleanroom does not mean no testing. It means developers verify instead of testing their own code, and an independent team tests statistically; the certification team's aim is "to provide scientific certification of software fitness for use".

munotes.in517

Approaches to SQA: Formal Methods, Cleanroom and Process Models

Usage-based testing is not testing only the common paths. Common paths get most tests by design; rare paths with serious consequences get models and tests of their own.

A certified process is not a certified product. ISO 9001 certification says the quality management system meets the standard's requirements, not that any particular release is free of defects.

Quick revision

  • Four approaches: formal methods; cleanroom; process models (CMMI, ISO 9001); the product approach (reviews, testing, product metrics).
  • Dijkstra (1972): testing shows the presence of bugs, not their absence; let "correctness proof and program grow hand in hand".
  • Proof of correctness (ISO/IEC/IEEE 24765): proves mathematically that a program satisfies its specified requirements; relative to the specification and its preconditions.
  • Cleanroom (SEI 1996): "zero failures in use"; defect prevention rather than removal; incremental development; box structures; correctness verification before execution by the Correctness Conditions (ifthenelse: when p is true does g do f, and when p is false does h do f?); statistical usage testing by an independent certification team. Reported gains of 10 to 20 times; NASA 25 to 75 per cent fewer failures in testing (Boehm and Basili).
  • Usage model: a Markov chain of states of use; test cases generated at random from it; standard calculations such as expected visits and expected test case length.
  • Worked examples: late_band verified by the Correctness Conditions and confirmed on all 16 valid inputs (and -1, outside the precondition, gives 1 to 7 days); 1,000 generated cases match the model (0.473 against 0.468 visits to payment); 42 include a failed payment.

Test yourself

1. Describe four approaches to software quality assurance. Formal methods prove mathematically that the program meets its specification. Cleanroom prevents defects by correctness verification during development and certifies reliability by statistical usage testing. Process models such as CMMI and ISO 9001 assure the process that produces the software. The product approach examines the software itself by reviews, testing and measurement against a quality model.

2. What is a proof of correctness, and what are its limits? A formal technique that proves mathematically that a program satisfies its specified requirements, for every input satisfying its preconditions. It is only as good as the specification, which it cannot check; it says nothing about inputs outside its preconditions; and it costs effort, so it is used for critical software and critical parts.

3. What is cleanroom software engineering? State its main practices. A theory-based, team-oriented process for developing and certifying high-reliability software under statistical quality control, aiming at zero failures in use and preventing defects rather than removing them. Its practices are incremental development, box-structure specification and design, correctness verification by the development team before any execution, and statistical usage testing by an independent certification team.

munotes.in518

Approaches to SQA: Formal Methods, Cleanroom and Process Models

4. State the correctness condition for an if-then-else structure and apply it to a small function. For an ifthenelse with predicate p, then-part g, else-part h and intended function f: when p is true, does g do f, and when p is false, does h do f? For ExamReg's late_band on 0 to 15 days: when the days are 0, returning on time is right; otherwise, when the days are at most 7 they are 1 to 7 by the precondition, and 1 to 7 days is right; when over 7 they are 8 to 15, and 8 to 15 days is right.

5. What is statistical usage testing, and why is it efficient? Testing with test cases generated at random from a usage model, such as a Markov chain of states of use with the probabilities of each transition, so that the tests are a sample of real use and support statistical estimates of reliability. It is efficient because it tends to find first the failures users would meet most often; rare, high-consequence functions are given separate models.

6. Why is the process approach not enough on its own? Because following a process faithfully does not guarantee a good product if the process itself is weak; compliance measures adherence, not results. Process approaches must therefore be combined with measurement of outcomes and with examination of the product.

Contents This chapter on its own page

munotes.in519

Chapter Eighty-Nine

Statistical Software Quality Assurance and Six Sigma

Syllabus topic Module 2, "Software Quality Assurance: ... Statistical Quality Assurance and Software Reliability"

In one line

Statistical software quality assurance makes quality decisions from defect data instead of impressions: collect and categorise the defects, trace each to its underlying cause, rank the causes to find the vital few, remove those causes, and measure again; Six Sigma turns the same idea into a disciplined project method, DMAIC, with a demanding target of 3.4 defects per million opportunities.

In the wording a student can write in an examination: statistical quality control is "The application of statistical techniques to control quality" (ASQ), and statistical process control the "statistically based analysis of a process and measures of process performance, which identify common and special causes of variation in process performance and maintain process performance within limits" (ISO/IEC/IEEE 24765). Applied to software, statistical SQA collects and categorises defect data, traces each defect to its cause, uses the Pareto principle ("most effects come from relatively few causes") to isolate the vital few causes, corrects them, and measures the result. Six Sigma "is a method that provides organizations tools to improve the capability of their business processes"; its projects follow DMAIC (Define, Measure, Analyze, Improve, Control), and "Six Sigma quality performance means 3.4 defects per million opportunities (accounting for a 1.5-sigma shift in the mean)" (ASQ).

Statistics instead of impressions

Every project has opinions about why its software fails: the testers blame rushed code, the developers blame changing requirements, the managers blame the testers. Statistical quality assurance replaces the opinions with counts. ISO's sixth quality management principle states the reason: "Decisions based on the analysis and evaluation of data and information are more likely to produce desired results."

For software the method has five steps, each already met in Module 2.

  1. Collect and categorise. Record every defect with its attributes: type, origin, severity, where it was found (Chapter Seventy-Three, on what a defect is).
  2. Trace each defect to its underlying cause. Not the symptom, not the person: the cause the process can change (Chapter Eighty, on using defect data to improve the process).
  3. Rank the causes. ASQ's glossary gives the principle: the Pareto chart, "named after 19th century economist Vilfredo Pareto", suggests "that most effects come from relatively few causes".
  4. Correct the vital few. Improvement effort goes to the few causes that account for most of the defects, Juran's "few vital breakthroughs" (Chapter Eighty-Two, on Shewhart, Deming and Juran).
  5. Measure again. Compare the next release with this one, against the causes that were not targeted, as Chapter Eighty did.

Worked example 1: release 2.0's defects by cause

When release 2.0's causal analysis traced each of its 200 defects to an underlying cause, nine kinds of cause emerged (the last row gathers eight rarer ones). The program ranks them.

# release 2.0's 200 defects by the underlying cause that causal analysis recorded (FINDINGS 5.2.7)
causes = {"requirement silent or ambiguous": 52, "rules misunderstood by developer": 30,
          "no shared validation": 28, "interface changed without notice": 24, "coding slip": 22,
          "screens never reviewed with users": 16, "test data unlike real records": 10,
          "documents not updated": 10, "eight other causes": 8}
total = sum(causes.values())

print(f"{'cause':<36}{'defects':>8}{'share':>7}{'cumulative':>12}")
running, vital = 0, []
for cause, n in sorted(causes.items(), key=lambda item: -item[1]):
    running += n
    if running - n < total / 2:                  # causes needed to account for half the defects
        vital.append(cause)
    print(f"{cause:<36}{n:>8}{n / total:>7.0%}{running / total:>12.0%}")
share = sum(causes[c] for c in vital) / total
print(f"the vital few: {len(vital)} of {len(causes)} causes, {share:.0%} of the defects")
munotes.in520

Statistical Software Quality Assurance and Six Sigma

cause                                defects  share  cumulative
requirement silent or ambiguous           52    26%         26%
rules misunderstood by developer          30    15%         41%
no shared validation                      28    14%         55%
interface changed without notice          24    12%         67%
coding slip                               22    11%         78%
screens never reviewed with users         16     8%         86%
test data unlike real records             10     5%         91%
documents not updated                     10     5%         96%
eight other causes                         8     4%        100%
the vital few: 3 of 9 causes, 55% of the defects

The vital few. Three causes of the nine account for 55 per cent of the defects, and all three concern the fee and eligibility rules: requirements silent or ambiguous about a rule, developers misunderstanding a rule, and no shared validation of the rules' inputs. The distribution is concentrated, though less than the 80/20 rule suggests: five causes are needed to pass three quarters. As Chapter Eighty-Two, on Shewhart, Deming and Juran, found, the rule describes a tendency, and the data decides.

What the ranking changes. Without it, a natural reaction to 200 defects is to test harder. The ranking points elsewhere: the largest cause lies in the requirements, before any code is written, and the top three share a single remedy, stating the rules once, precisely, and reviewing them with the exam cell. The actions of Chapter Eighty, on using defect data to improve the process (a shared validation module, a requirements review checklist item on invalid input), attacked part of this, and the release 2.1 figures showed the targeted defects falling. The ranking counts defects equally; Chapter One Hundred Four, on Pareto diagrams, weights them by cost and draws the chart.

Six Sigma

Six Sigma began at Motorola: in ASQ's history, "a methodology developed by Motorola to improve its business processes by minimizing defects", which "evolved into an organizational approach that achieved breakthroughs and significant bottom-line results." ASQ lists the threads its definitions share: teams assigned well-defined projects with a direct effect on the organisation; training in statistical thinking at every level, with specialists, called Black Belts, trained in advanced statistics and project management; the DMAIC approach to solving problems; and management support for the whole as a business strategy.

munotes.in521

Statistical Software Quality Assurance and Six Sigma

The name. In ASQ's glossary a sigma is "One standard deviation in a normally distributed process", and Six Sigma quality is "A term generally used to indicate process capability in terms of process spread measured by standard deviations in a normally distributed process." ASQ's Six Sigma page makes the picture concrete: a process at Six Sigma quality keeps its own variation within its control limits, three standard deviations from the centre line, while the requirement's tolerance limits lie six standard deviations away. Its output almost never falls outside the tolerance.

The number. Six Sigma's measure is defects per million opportunities (DPMO): the defects counted, divided by the opportunities for a defect, times a million, the same normalisation as ASQ's parts per million, "the number of defects normalized to a population of one million for ease of comparison." The target is ASQ's: "Six Sigma quality performance means 3.4 defects per million opportunities (accounting for a 1.5-sigma shift in the mean)". The shift is a convention of the calculation: a six sigma process is taken to operate with its mean 1.5 standard deviations off centre, leaving 4.5 standard deviations to the nearest limit.

DMAIC

ASQ describes DMAIC as "a data-driven quality strategy used to improve processes." Its five phases, with the tools ASQ lists for each:

PhaseWhat it does (ASQ)Typical tools (ASQ)
DefineThe problem, the improvement activity, the project goals, and the customer's requirementsProject charter; voice of the customer; value stream map
Measure"Measure process performance."Process map; capability analysis; Pareto chart
Analyze"Analyze the process to determine root causes of variation, poor performance (defects)."Root cause analysis; failure mode and effects analysis; multi-vari chart
Improve"Improve process performance by addressing and eliminating the root causes."Design of experiments; kaizen event
Control"Control the improved process and future process performance."Control plan; statistical process control; 5S; mistake proofing (poka-yoke)

For a new product or process, rather than an existing one, Six Sigma uses DMADV: Define, Measure, Analyze, Design and Verify. And ASQ notes a lineage: "The DMAIC process easily lends itself to the project approach to quality improvement encouraged and promoted by Juran."

Chapter Eighty's causal analysis of input validation defects was, in effect, a DMAIC project, and reading it phase by phase shows the method on software.

PhaseExamReg's input validation project (Chapter Eighty)
DefineThe problem: input validation, the largest defect type, 58 of release 2.0's 200 defects; the customer's requirement: every invalid input refused
Measure2.90 input validation defects per KLOC in release 2.0
AnalyzeThe five whys: no shared validation; requirements silent on invalid input
ImproveA shared validation module; a requirements review checklist item; unit test templates from partitions and boundaries
ControlThe checklist item and templates made standard; the rate watched release by release (1.12 per KLOC in release 2.1)
munotes.in522

Statistical Software Quality Assurance and Six Sigma

Worked example 2: sigma levels, and ExamReg's fees

The program first computes the defects per million opportunities at each sigma level, with the 1.5-sigma shift, from the normal distribution itself; the last row should reproduce ASQ's 3.4. It then measures release 2.0's first month of fees: the exam cell reconciled 18,000 fees, each made of three parts (the form fee, the backlog fee and the late fee), and found 9 parts wrong, on 9 different forms.

from statistics import NormalDist

def dpmo(sigma_level, shift=1.5):
    """Defects per million opportunities at a sigma level, with the conventional 1.5-sigma shift."""
    return 1e6 * (1 - NormalDist().cdf(sigma_level - shift))

def sigma_level(defects_per_million, shift=1.5):
    return NormalDist().inv_cdf(1 - defects_per_million / 1e6) + shift

for k in range(1, 7):
    print(f"{k} sigma: {dpmo(k):>11,.1f} defects per million opportunities")

forms, parts, wrong = 18_000, 3, 9        # release 2.0's first month of fees (FINDINGS 5.2.7)
for name, opportunities in [("every fee part", forms * parts), ("every form", forms)]:
    d = 1e6 * wrong / opportunities
    print(f"opportunity = {name}: {opportunities:,} opportunities, {d:.1f} DPMO,"
          f" sigma level {sigma_level(d):.2f}")
1 sigma:   691,462.5 defects per million opportunities
2 sigma:   308,537.5 defects per million opportunities
3 sigma:    66,807.2 defects per million opportunities
4 sigma:     6,209.7 defects per million opportunities
5 sigma:       232.6 defects per million opportunities
6 sigma:         3.4 defects per million opportunities
opportunity = every fee part: 54,000 opportunities, 166.7 DPMO, sigma level 5.09
opportunity = every form: 18,000 opportunities, 500.0 DPMO, sigma level 4.79

The scale. The sixth row reproduces ASQ's 3.4 exactly, which confirms how the number is built: the chance of a value beyond 4.5 standard deviations of a normal distribution, six sigma less the 1.5-sigma shift. The scale is steep. Moving from four to five sigma cuts the rate from 6,209.7 to 232.6 per million, and from five to six, to 3.4.

The month of fees. Counting every fee part as an opportunity, 9 wrong parts in 54,000 make 166.7 DPMO, a sigma level of 5.09. Counting every form as one opportunity, the same 9 errors make 500.0 DPMO and 4.79 sigma. Nothing about the portal changed between the two lines; only the definition of an opportunity did. A sigma level is therefore meaningless until the opportunity is defined and held fixed, the same lesson as the four defect counts of Chapter Seventy-Three and the line counts of Chapter Sixty-Seven, on size metrics.

Six Sigma and software. The fee example also shows where the metric fits software best: repeated outputs such as fees computed, forms processed or transactions completed, where each output is a natural opportunity. For defects in code there is no natural unit of opportunity (a line? a function? a requirement?), and a sigma level computed per line of code says more about the choice of unit than about quality. DMAIC, by contrast, is a problem-solving method that ASQ says "can be implemented as a standalone quality improvement procedure", and it carries over to software unchanged, as the input validation project shows.

munotes.in523

Statistical Software Quality Assurance and Six Sigma

Six Sigma and lean

ASQ distinguishes the two approaches often combined as lean Six Sigma: "Lean focuses on waste reduction, whereas Six Sigma emphasizes variation reduction." Lean, and CMMI, are the subject of Chapter One Hundred, on Lean, CMMI and choosing a methodology.

What it does not mean

Statistical SQA is not counting for its own sake. Its purpose is the ranking of causes and the action on the vital few; a count that changes no decision is not worth collecting.

The vital few are not always 20 per cent. The Pareto principle describes a concentration whose size the data decides; in release 2.0, 3 causes of 9 held 55 per cent.

A sigma level is not comparable without its definition. The same month of fees is 5.09 or 4.79 sigma depending on what counts as an opportunity.

Six Sigma is not the number six. It is a project method with trained people, management support and a measure; the 3.4 per million is its target, not its content.

Quick revision

  • Statistical quality control (ASQ): "The application of statistical techniques to control quality"; it includes acceptance sampling, which statistical process control does not.
  • Statistical SQA: collect and categorise defects; trace each to its underlying cause; rank the causes (Pareto: "most effects come from relatively few causes"); correct the vital few; measure again.
  • Six Sigma (Motorola; ASQ): teams on well-defined projects, Black Belts, DMAIC, management support; a sigma is one standard deviation; the target is 3.4 defects per million opportunities with the 1.5-sigma shift.
  • DPMO = defects divided by opportunities, times a million.
  • DMAIC: Define, Measure, Analyze, Improve, Control; DMADV (Define, Measure, Analyze, Design, Verify) for new products.
  • Lean against Six Sigma (ASQ): waste reduction against variation reduction.
  • Worked examples: 3 of 9 causes gave 55 per cent of release 2.0's defects, all about the fee rules; sigma table 6,209.7 (4), 232.6 (5), 3.4 (6); the month of fees 166.7 DPMO and 5.09 sigma per fee part, or 500.0 and 4.79 per form.

Test yourself

1. What is statistical software quality assurance? List its steps. The use of statistics on defect data to decide where to improve. Information about defects is collected and categorised; each defect is traced to its underlying cause; the causes are ranked, using the Pareto principle, to isolate the vital few; the vital few are corrected; and the effect is measured in the next release.

munotes.in524

Statistical Software Quality Assurance and Six Sigma

2. In the worked example, what did the ranking of causes show, and what action follows? That three of nine causes, all concerning the fee and eligibility rules (silent or ambiguous requirements, misunderstood rules, and no shared validation), accounted for 55 per cent of the defects. The action is to state the rules once, precisely, review them with the exam cell and validate them in one shared place, rather than simply to test harder.

3. What is Six Sigma, and what does "3.4 defects per million opportunities" mean? A method, begun at Motorola, that improves process capability through well-defined projects led by statistically trained staff, following DMAIC, with management support. The figure is its target: the rate of defects when the tolerance limits are six standard deviations from the process centre and, by convention, the mean is taken as shifted by 1.5 standard deviations, leaving 4.5 standard deviations to the nearest limit.

4. Explain the DMAIC phases, with a software example. Define the problem, goals and customer requirements; Measure process performance; Analyze to find the root causes of defects; Improve by eliminating the root causes; Control the improved process. For ExamReg: define the input validation problem (58 of 200 defects), measure 2.90 per KLOC, analyze with the five whys, improve with a shared validation module and a review checklist item, and control by making them standard and watching the rate (1.12 per KLOC in the next release).

5. How is DPMO computed, and why must the opportunity be defined? Defects divided by the number of opportunities for a defect, multiplied by a million. The opportunity must be defined and held fixed because the same data gives different results under different definitions: ExamReg's 9 wrong fee parts in 18,000 fees are 166.7 DPMO (5.09 sigma) per fee part but 500.0 DPMO (4.79 sigma) per form.

6. Distinguish Six Sigma from lean. Lean focuses on reducing waste, work that adds no value, and on standard work and flow; Six Sigma focuses on reducing variation, using statistical analysis. They are often combined as lean Six Sigma.

Contents This chapter on its own page

munotes.in525

Chapter Ninety

Software Reliability

Syllabus topic Module 2, "Software Quality Assurance: ... Statistical Quality Assurance and Software Reliability"

In one line

Software reliability is how long, and how probably, software runs without failing in the conditions it is actually used in; it is measured by failures over time, summarised as mean time to failure, mean time to repair and availability, and it behaves unlike hardware reliability because software does not wear out but changes.

In the wording a student can write in an examination: software reliability is the "extent to which a system has operated without countable failures in a specified environment for a specified time span" (IEEE 982:2024), classically "the probability of failure-free software operation for a specified period of time in a specified environment" (the ANSI definition, as Lyu gives it). A failure is a "departure of system behavior from system requirements" (IEEE 982:2024). The mean time to failure (MTTF) is the expected time until the next failure; the mean time to repair (MTTR) the "expected or observed duration required to return a malfunctioning system or component to normal operations" (ISO/IEC/IEEE 24765); the mean time between failures (MTBF) the "expected or observed time between consecutive failures in a system or component" (the same standard). Availability is the "ratio of uptime divided by the sum of uptime plus downtime" (IEEE 982:2024), usually computed as MTTF / (MTTF + MTTR).

What reliability is, and is not

Reliability is one of the nine quality characteristics of ISO/IEC 25010:2023, which defines it as the "capability of a product to perform specified functions under specified conditions for a specified period of time without interruptions and failures" (Chapter Twenty, on the quality model today). Three parts of every definition matter.

  • Failures, not faults. Reliability counts what users experience. A fault that is never executed causes no failure and costs no reliability; one fault on the most used path can fail thousands of times.
  • A specified environment. Reliability depends on how the software is used. Lyu defines the operational profile as "the set of operations that the software can execute along with the probability with which they will occur", the same idea as the usage model of Chapter Eighty-Eight, on approaches to SQA. The same program can be highly reliable for one group of users and unreliable for another.
  • A time span. Reliability is always over a period: an hour, a session, a registration window. A number without its period means nothing.

Reliability is not correctness. Correctness is the "degree to which a system or component is free from faults in its specification, design, and implementation" (ISO/IEC/IEEE 24765). A program can be incorrect and still very reliable, if its faults lie where users rarely go: ExamReg's version that charged the wrong late fee at exactly 7 days (Chapter Fifty-Five, on boundary value analysis) failed only for students exactly a week late. And a program proved correct against a wrong specification can fail every day. Correctness asks about the program against its specification; reliability asks about the program in use.

munotes.in526

Software Reliability

Hardware fails by wearing out; software does not

The NIST/SEMATECH e-Handbook describes the failure rate of most manufactured products over their lives as the bathtub curve. It begins with an early failure period, "a high but rapidly decreasing failure rate"; then "the failure rate levels off and remains roughly constant" through the intrinsic failure period, where "most systems spend most of their lifetimes"; and finally the wearout failure period, when "the failure rate begins to increase as materials wear out and degradation failures occur at an ever increasing rate."

Software has no wearout period. In Lyu's words, software "does not wear out, burn out, or deteriorate, i.e., its reliability does not decrease with time." Instead, "software generally enjoys reliability growth during testing and operation since software faults can be detected and removed when software failures occur." And there is a way down as well: "software may experience reliability decrease due to abrupt changes of its operational usage or incorrect modifications to the software." Chapter Twenty-Three, on quality in software development, tabulated these differences; the figure shows them as curves.

Two failure-rate curves over time: hardware's bathtub curve, with early failures, a long flat period and rising wear-out; software's curve, falling as faults are removed and jumping at each change, with no wear-out

Figure 90.1 Hardware wears out; software's failure rate falls as faults are removed and jumps when it is changed

The software curve explains a practical rule. Every change to software, a fix, a new feature, a new kind of user, can move its failure rate up, so reliability figures belong to one version in one environment, and a new release is measured again.

The measures

Lyu defines the MTTF as the expected time until the next failure, adding that it "is also known as MTBF", the name ISO/IEC/IEEE 24765 gives to the "time between consecutive failures". The MTTR is the expected time to repair the system after a failure. From the two, Lyu gives availability, "the probability that a system is available when needed", as

Availability = MTTF / (MTTF + MTTR)

which is the same as IEEE 982:2024's uptime divided by uptime plus downtime, since over a period the failures divide the uptime into MTTF-sized pieces and the downtime into MTTR-sized ones.

Reliability over a time span needs a model of how failures arrive. The simplest is the exponential model, in which the failure rate is constant: the NIST/SEMATECH e-Handbook notes that "The exponential distribution is the only distribution to have a constant failure rate", that its reliability is R(t) = e^(-t/MTTF), and that "another name for the exponential mean is the Mean Time To Fail". It fits the flat part of the bathtub curve, and for software a version whose failure rate is not changing: no fixes, no new usage.

munotes.in527

Software Reliability

When a task needs several components working at once, the handbook's rule for independent components applies: "to calculate the reliability of a system of independent components, multiply the reliability functions of all the components together."

Worked example: release 2.0's first 60 days in use

During release 2.0's first 60 days in use, 1,440 hours, the ExamReg portal stopped serving students six times; the payment gateway it depends on, a third-party service, failed three times in the same days. These are failures, not defects: the log records when the service stopped and for how long, not which fault caused it. The program computes the measures and the reliability over two time spans, assuming a constant failure rate.

from math import exp

HOURS = 60 * 24                            # release 2.0's first 60 days in use (FINDINGS 5.5)
portal = [0.5, 1.0, 0.25, 2.0, 0.75, 1.5]  # hours each outage of the portal lasted
gateway = [0.5, 0.5, 1.0]                  # the payment gateway's outages in the same 60 days

def measures(outages):
    up = HOURS - sum(outages)
    return up / len(outages), sum(outages) / len(outages), up / HOURS   # MTTF, MTTR, availability

for name, outages in [("portal", portal), ("payment gateway", gateway)]:
    mttf, mttr, available = measures(outages)
    print(f"{name}: {len(outages)} failures; MTTF {mttf:.1f} h, MTTR {mttr:.2f} h;"
          f" availability {available:.2%}, and MTTF / (MTTF + MTTR) = {mttf / (mttf + mttr):.2%}")

mttf_portal, mttf_gateway = measures(portal)[0], measures(gateway)[0]
for span, hours in [("a 1-hour session", 1), ("the 15-day registration window", 15 * 24)]:
    r_portal = exp(-hours / mttf_portal)                 # exponential model: constant failure rate
    r_payment = r_portal * exp(-hours / mttf_gateway)    # paying needs both: reliabilities multiply
    print(f"{span}: portal without a failure {r_portal:.4f}; portal and gateway {r_payment:.4f}")
portal: 6 failures; MTTF 239.0 h, MTTR 1.00 h; availability 99.58%, and MTTF / (MTTF + MTTR) = 99.58%
payment gateway: 3 failures; MTTF 479.3 h, MTTR 0.67 h; availability 99.86%, and MTTF / (MTTF + MTTR) = 99.86%
a 1-hour session: portal without a failure 0.9958; portal and gateway 0.9937
the 15-day registration window: portal without a failure 0.2217; portal and gateway 0.1046

The measures. The portal ran 1,434 of the 1,440 hours: an MTTF of 239.0 hours, an MTTR of 1.00 hour and an availability of 99.58 per cent, which the formula MTTF / (MTTF + MTTR) reproduces exactly. The gateway failed half as often and recovered faster, at 99.86 per cent.

The time span decides. For one student in a 1-hour session, the portal gets through without a failure with probability 0.9958, and a payment, which needs the gateway too, with probability 0.9937. Over the whole 15-day registration window, the chance of no portal failure at all is only 0.2217, and of neither failing 0.1046. Both statements describe the same portal. A 99.58 per cent available system will very probably fail at least once during a fortnight, and a reliability requirement must therefore say which span it means: no failure during a student's session and no failure during the registration window are very different promises.

munotes.in528

Software Reliability

The model's limit. The exponential model assumes a constant failure rate, which holds for release 2.0 only as long as nobody changes it. The software curve in the figure is the warning: each fix or new release restarts the measurement. Chapter Ninety-Two, on measuring software reliability, fits models in which the failure rate falls as faults are removed.

What it does not mean

Reliability is not the absence of faults. It is the absence of failures in use; faults in unused code cost no reliability, and one fault on a common path costs a great deal.

Availability is not reliability. A system that fails often but recovers in seconds can be highly available and still unreliable; ExamReg was 99.58 per cent available and yet likely to fail during any fortnight.

Software does not wear out. Its reliability falls when it is changed or used differently, not with age.

MTTF is not a guarantee. It is an average; with a constant failure rate, a system with an MTTF of 239 hours fails within its first 239 hours more often than not.

Quick revision

  • Software reliability (IEEE 982:2024): operation "without countable failures in a specified environment for a specified time span"; (ANSI, via Lyu) "the probability of failure-free software operation for a specified period of time in a specified environment".
  • Failure (IEEE 982:2024): "departure of system behavior from system requirements". Operational profile (Lyu): the operations and their probabilities.
  • Reliability against correctness: correctness is freedom from faults against the specification; reliability is freedom from failures in use.
  • Hardware: the bathtub curve (early failure, intrinsic failure, wearout). Software: no wear-out; reliability grows as faults are removed; it drops with changes in usage or faulty modifications.
  • Measures: MTTF, MTTR, MTBF; availability = MTTF / (MTTF + MTTR) = uptime / (uptime + downtime); exponential model R(t) = e^(-t/MTTF) for a constant failure rate; independent components multiply.
  • Worked example: 6 outages in 1,440 hours: MTTF 239.0 h, MTTR 1.00 h, availability 99.58 per cent; a 1-hour session 0.9958 (0.9937 with the gateway); the 15-day window 0.2217 (0.1046).

Test yourself

1. Define software reliability, and explain why its definition names an environment and a time span. The probability of failure-free operation of software for a specified period of time in a specified environment. The environment matters because reliability depends on how the software is used, its operational profile; the time span matters because the probability of getting through without a failure falls as the period grows.

munotes.in529

Software Reliability

2. Distinguish reliability from correctness. Correctness is the degree to which software is free from faults in its specification, design and implementation; reliability is the degree to which it operates without failures in use. Software with faults in rarely used paths can be very reliable; software correct against a wrong specification can fail constantly.

3. Compare the failure curves of hardware and software. Hardware follows the bathtub curve: a falling early failure rate, a long flat intrinsic period, and a rising wear-out period. Software does not wear out: its failure rate falls as failures reveal faults that are removed, and rises when the software is changed or used in new ways, so its curve falls with jumps at changes.

4. Define MTTF, MTTR and availability, and give the formula connecting them. MTTF is the expected time until the next failure; MTTR the expected time to repair the system after a failure; availability the proportion of time the system is available when needed. Availability = MTTF / (MTTF + MTTR), equivalently uptime / (uptime + downtime).

5. A portal failed 6 times in 1,440 hours, with 6 hours of outages in all. Compute its MTTF, MTTR and availability. Uptime 1,434 hours; MTTF 1,434 / 6 = 239.0 hours; MTTR 6 / 6 = 1.00 hour; availability 1,434 / 1,440, or 239 / (239 + 1), which is 99.58 per cent.

6. Why is a system with 99.58 per cent availability still likely to fail during a two-week period? Because availability measures the share of time it is up, not the chance of an unbroken run. With an MTTF of 239 hours and a constant failure rate, the probability of no failure in 360 hours is e^(-360/239), about 0.22, so a failure during the period is more likely than not.

Contents This chapter on its own page

munotes.in530

Chapter Ninety-One

Statistical Process Control: Control Charts for Software

Syllabus topic Module 2, "Software Quality Assurance: ... Statistical process control techniques"

In one line

Statistical process control watches a process through its own data, plotted in time order on a control chart whose centre line and limits are computed from that data; points outside the limits, or patterns inside them, signal special causes to be found and removed, while everything else is the common-cause variation that only a change to the process can reduce.

In the wording a student can write in an examination: statistical process control (SPC) is the "statistically based analysis of a process and measures of process performance, which identify common and special causes of variation in process performance and maintain process performance within limits" (ISO/IEC/IEEE 24765). A control chart is "a graph used to study how a process changes over time", with "a central line for the average, an upper line for the upper control limit, and a lower line for the lower control limit", the lines "determined from historical data" (ASQ). For single measurements the individuals chart uses the moving range, the difference between consecutive values: its limits are the mean plus or minus 3 times the average moving range divided by 1.128 (NIST/SEMATECH). A point beyond the limits, or one of the Western Electric rules (2 of 3 points beyond 2 sigma, 4 of 5 beyond 1 sigma, 8 in a row on one side, 6 in a row rising or falling), signals a special cause.

What a control chart is for

Chapter Eighty-One, on quality concepts, separated common causes, "built into the process", from special causes, "non-routine events", and left a question open: were all thirty-eight of the fee page's served times common-cause variation? A histogram could not say, because it discards the order of the data. The control chart keeps the order, and answers.

ASQ lists what the chart is used for: "When controlling ongoing processes by finding and correcting problems as they occur", "When predicting the expected range of outcomes from a process", "When determining whether a process is stable (in statistical control)", and "When determining whether your quality improvement project should aim to prevent specific problems or to make fundamental changes to the process." The last is the decision of Chapter Eighty-One, on quality concepts: a special cause is removed where it happened; common-cause variation needs a change to the process.

Control limits are not specification limits. The NIST/SEMATECH e-Handbook separates them plainly: "Control Limits are used to determine if the process is in a state of statistical control", while "Specification Limits are used to determine if the product will function in the intended fashion." A control limit comes from the process's own data; a requirement such as the fee page within 2 seconds comes from the users. A process can be in control and still fail its specification, which is the stable but slow process of Chapter Eighty-One, on quality concepts.

munotes.in531

Statistical Process Control: Control Charts for Software

The individuals chart

Software data often arrive one value at a time: one response time per request, one repair time per defect, one defect count per release. For such data the handbook's individuals chart estimates the process's variation from the moving range, "the absolute value of the first difference (e.g., the difference between two consecutive data points) of the data." With x-bar the mean of the values and MR-bar the mean of the moving ranges:

  • centre line = x-bar
  • upper control limit (UCL) = x-bar + 3 × MR-bar / 1.128
  • lower control limit (LCL) = x-bar - 3 × MR-bar / 1.128

The constant 1.128 is, as the handbook notes, the value of d2 for n = 2: the factor that turns the average range of pairs into an estimate of the standard deviation. The moving ranges can be charted too; their upper limit is 3.267 × MR-bar, the handbook's factor D4 for ranges of two values, and their lower limit is 0.

Reading the chart. The first signal is a point beyond the limits. The handbook adds the Western Electric rules, patterns inside the limits that are about as unlikely as a point outside them in a stable process: 2 of the last 3 points beyond 2 sigma on the same side; 4 of the last 5 beyond 1 sigma on the same side; 8 consecutive points on one side of the centre line; 6 in a row rising or falling. ASQ's page lists the same kinds of signal. The rules have a price: with the single 3-sigma rule a stable process gives a false alarm about "every 371 points on the average", and adding the rules raises that to "about once in every 91.75 points".

Limits need a first phase. Limits computed from data that contain a special cause are distorted by it. ASQ describes the practice: limits from the first points are conditional, and the chart is recomputed once points from a period in control are available. When a point is outside, its cause is investigated; if a cause is found and removed, the point is dropped and the limits recomputed.

Worked example 1: the fee page's forty requests

The program computes the individuals chart for the forty fee page requests of Chapter Forty-Five, on load testing, in the order they were sent; drops the points outside the limits whose causes have been found; recomputes; and applies the Western Electric rules to the result.

from statistics import mean

# Chapter 45's forty requests for ExamReg's fee page, in order: (elapsed ms, HTTP response code)
runs = [(380, 200), (410, 200), (417, 200), (463, 200), (454, 200), (516, 200), (491, 200),
        (569, 200), (528, 200), (622, 200), (565, 200), (675, 200), (602, 200), (728, 200),
        (639, 200), (431, 200), (676, 200), (484, 200), (413, 200), (537, 200), (450, 200),
        (590, 200), (487, 200), (643, 200), (524, 200), (696, 200), (561, 200), (749, 200),
        (598, 200), (452, 200), (635, 200), (505, 200), (672, 200), (558, 200), (409, 200),
        (120, 503), (446, 200), (664, 200), (95, 503), (717, 200)]

def xmr_limits(x):
    """Individuals chart: x-bar +/- 3 MR-bar / 1.128; moving range upper limit 3.267 MR-bar."""
    mr = [abs(b - a) for a, b in zip(x, x[1:])]
    centre, mr_bar = mean(x), mean(mr)
    sigma = mr_bar / 1.128
    return centre, centre - 3 * sigma, centre + 3 * sigma, sigma, 3.267 * mr_bar

times = [ms for ms, _ in runs]
centre, lcl, ucl, sigma, mr_ucl = xmr_limits(times)
print(f"all 40: centre {centre:.1f}, limits {lcl:.1f} to {ucl:.1f} ms")
outside = [i for i, ms in enumerate(times, 1) if not lcl <= ms <= ucl]
print("   outside the limits:", ", ".join(f"request {i} ({times[i - 1]} ms)" for i in outside))

served = [ms for i, ms in enumerate(times, 1) if i not in outside]   # causes found: 503 refusals
centre, lcl, ucl, sigma, mr_ucl = xmr_limits(served)
print(f"without them: centre {centre:.1f}, limits {lcl:.1f} to {ucl:.1f} ms,"
      f" moving range limit {mr_ucl:.1f} ms")

def signals(x, centre, sigma):
    """The Western Electric rules in the NIST/SEMATECH e-Handbook: the point each first fires at."""
    z = [(v - centre) / sigma for v in x]

    def same_side(window, beyond, needed):
        return any(sum(s * v > beyond for v in window) >= needed for s in (1, -1))
    found = {}
    for i in range(len(z)):
        tests = {"a point beyond 3 sigma": abs(z[i]) > 3,
                 "2 of 3 beyond 2 sigma, same side": i >= 2 and same_side(z[i - 2:i + 1], 2, 2),
                 "4 of 5 beyond 1 sigma, same side": i >= 4 and same_side(z[i - 4:i + 1], 1, 4),
                 "8 in a row on one side": i >= 7 and same_side(z[i - 7:i + 1], 0, 8),
                 "6 in a row rising or falling": i >= 5 and same_side(
                     [b - a for a, b in zip(x[i - 5:i], x[i - 4:i + 1])], 0, 5)}
        for rule, fired in tests.items():
            if fired and rule not in found:
                found[rule] = i + 1
    return found

fired = signals(served, centre, sigma)
print("rules that fire on the 38 served pages:", fired or "none")
moving = [abs(b - a) for a, b in zip(served, served[1:])]
print("moving ranges over their limit:", sum(m > mr_ucl for m in moving))
munotes.in532

Statistical Process Control: Control Charts for Software

all 40: centre 529.3, limits 130.3 to 928.3 ms
   outside the limits: request 36 (120 ms), request 39 (95 ms)
without them: centre 551.5, limits 254.2 to 848.7 ms, moving range limit 365.1 ms
rules that fire on the 38 served pages: none
moving ranges over their limit: 0
munotes.in533

Statistical Process Control: Control Charts for Software

Phase one: two special causes. With all forty requests, the limits run from 130.3 to 928.3 ms, and two points fall below the lower limit: requests 36 and 39, the two 503 responses. The investigation that Chapter Eighty-One, on quality concepts, called for found their cause: a database backup job ran during the test and twice made the database refuse connections for a moment. The job was moved to a quiet hour, and the two points are removed.

Phase two: a stable process. Without them, the chart centres on 551.5 ms with limits from 254.2 to 848.7 ms. Every served page lies inside, no moving range exceeds its limit of 365.1 ms, and none of the Western Electric rules fires. By the control chart, the thirty-eight served times are common-cause variation: the answer to Chapter Eighty-One's open question. Chapter One Hundred Seven, on run charts, puts the same data to a different, more sensitive test.

An individuals control chart of the forty fee page requests: the served pages vary between the limits of 254.2 and 848.7 ms around a centre of 551.5; the two 503 responses fall far below the lower limit

Figure 91.1 The fee page's individuals chart: the served pages are in control; the two 503 responses are special causes

What the chart promises. A stable process is predictable: as long as nothing changes, the next page will very probably be served in 254 to 849 ms. Whether that is good enough is a different question, answered by the specification; the chart only says that the variation is the process's own. To make every page faster, the process itself has to change.

Worked example 2: repair times, a process measure

Control charts serve software processes as well as products. Chapter Sixty-Eight, on quality, process and test metrics, recorded how long each of release 2.0's twelve after-release defects took from report to a fix in production. The program charts them. A time cannot be negative, so a lower limit below zero is shown as zero.

from statistics import mean

# hours from report to a fix in production for release 2.0's 12 after-release defects (FINDINGS 5.2.2)
hours = [4, 6, 3, 30, 8, 5, 72, 10, 6, 4, 12, 20]

def limits(x):
    sigma = mean(abs(b - a) for a, b in zip(x, x[1:])) / 1.128
    centre = mean(x)
    return centre, max(0.0, centre - 3 * sigma), centre + 3 * sigma    # no time is below zero

for label, x in [("all 12 fixes", hours), ("without the 7th", hours[:6] + hours[7:])]:
    centre, lcl, ucl = limits(x)
    out = [h for h in x if not lcl <= h <= ucl]
    print(f"{label}: centre {centre:.1f} h, limits {lcl:.1f} to {ucl:.1f} h, outside: {out or 'none'}")
munotes.in534

Statistical Process Control: Control Charts for Software

all 12 fixes: centre 15.0 h, limits 0.0 to 65.3 h, outside: [72]
without the 7th: centre 9.8 h, limits 0.0 to 32.2 h, outside: none

The seventh fix, 72 hours, is above the upper limit of 65.3 hours. Its cause was outside the team's process: the fix waited for the payment gateway's provider to change its side. With it removed, the remaining eleven fixes centre on 9.8 hours with an upper limit of 32.2, and all lie inside; the 30-hour fix, though long, is within the process's own variation. Twelve points make only conditional limits, as ASQ warns, but the chart already separates the one repair that needs a different response (an agreement with the provider on response times) from the ones that need none.

Using control charts in a software project

What to chartWhy
Response times, error rates, build timesProduct and operations behaviour, request by request
Repair times, review rates, defects found per reviewProcess behaviour, event by event
Defects per release, per KLOCProcess results, release by release (few points; limits stay conditional for long)

Three cautions apply to software. The data must come from one process: charting a fee page and a hall ticket page together mixes two processes and hides both. Many software measures change every release, and a chart spanning a change should be restarted after it, which is exactly what makes a shift visible (Chapter Eighty, on using defect data, measured one). And a signal is a reason to investigate, not a verdict: a rule fires by chance about once in 92 points when all the rules are used.

What it does not mean

In control does not mean good. It means predictable; the fee page is in control between 254 and 849 ms, and whether that meets its requirement is a separate question.

A point outside the limits is not automatically an error. It is a signal to look for a special cause; only when one is found is the point removed.

Control limits are not targets. They describe what the process does, computed from its data; nobody sets them.

More rules are not always better. Each rule adds sensitivity and false alarms; the handbook reports false alarms rising from about 1 in 371 points to about 1 in 92.

Quick revision

  • SPC (ISO/IEC/IEEE 24765): statistical analysis of a process to "identify common and special causes of variation" and "maintain process performance within limits".
  • Control chart (ASQ): time-ordered data; centre line at the average; upper and lower control limits from historical data.
  • Individuals chart (NIST/SEMATECH): moving range = difference between consecutive values; limits x-bar plus or minus 3 × MR-bar / 1.128 (d2 for n = 2); moving range limit 3.267 × MR-bar (D4).
  • Signals: a point beyond 3 sigma; 2 of 3 beyond 2 sigma; 4 of 5 beyond 1 sigma; 8 in a row on one side; 6 in a row rising or falling. False alarms: about 1 in 371 points with the first rule alone, about 1 in 92 with all.
  • Control against specification limits: the first come from the process, the second from the requirements.
  • Worked examples: fee page phase one limits 130.3 to 928.3 ms, requests 36 and 39 (the 503s) below; phase two 254.2 to 848.7 ms around 551.5, no rule fires; repair times: the 72-hour fix beyond 65.3 h, the rest within 32.2 h.
munotes.in535

Statistical Process Control: Control Charts for Software

Test yourself

1. What is statistical process control, and what does a control chart show? The use of statistics to analyse a process and its measures of performance, to identify common and special causes of variation and keep performance within limits. A control chart plots the process's data in time order with a centre line and upper and lower control limits computed from the data, showing whether the variation is stable or affected by special causes.

2. How are the limits of an individuals chart computed? From the values' mean, x-bar, and the mean of the moving ranges, MR-bar, the differences between consecutive values: the centre line is x-bar and the limits are x-bar plus or minus 3 × MR-bar / 1.128, where 1.128 is d2 for ranges of two values. The moving range chart's upper limit is 3.267 × MR-bar.

3. State the rules that signal a special cause. A point beyond the 3-sigma limits; 2 of the last 3 points beyond 2 sigma on the same side; 4 of the last 5 beyond 1 sigma on the same side; 8 consecutive points on one side of the centre line; 6 points in a row steadily rising or falling.

4. Distinguish control limits from specification limits. Control limits are computed from the process's own data and show whether it is in statistical control; specification limits come from the requirements and show whether the product will do its job. A process can be in control and still fail its specification.

5. Describe what happened in the fee page example. With all forty requests, the limits were 130.3 to 928.3 ms and the two 503 responses fell below the lower limit; their cause, a database backup job running during the test, was found and removed, so they were dropped. Recomputed on the thirty-eight served pages, the limits were 254.2 to 848.7 ms around 551.5, every point was inside and no rule fired: the served times are common-cause variation.

munotes.in536

Statistical Process Control: Control Charts for Software

6. Why can a control chart be useful for a process measure such as repair time? It separates repairs that took long for a special reason from the process's ordinary variation. Release 2.0's 72-hour fix was beyond the upper limit of 65.3 hours and had an outside cause, a wait for the payment gateway's provider, which calls for a specific response; the other repairs, including one of 30 hours, were within the process's own variation.

Contents This chapter on its own page

munotes.in537

Chapter Ninety-Two

Measuring Software Reliability

Syllabus topic Module 2, "Software Quality Assurance: ... Statistical process control techniques: Software reliability measurement and improvement"

In one line

Software reliability is measured by recording failures against time, during system test and in use, and fitting a reliability model to the record; the model estimates the reliability reached so far and predicts how much more testing is needed to reach an objective, which turns is it ready? into a question with a number for an answer.

In the wording a student can write in an examination: failure intensity is the "quantitative characterization of in-service reliability as a ratio of the number of observed failures to the duration of an observed span" (IEEE 982:2024); the rate of occurrence of failures (ROCOF) is Lyu's other name for the failure rate function. A failure intensity objective (FIO) is the "failure intensity level to be achieved during pre-release testing", as part of a system's release criteria (IEEE 982:2024). Reliability data come as failure-count data (failures per period) or time-between-failures data. Estimation determines the reliability achieved so far; prediction forecasts it. Musa's basic execution time model assumes the expected number of failures by execution time t is β0(1 - e^(-β1 t)), so the failure intensity β0β1e^(-β1 t) falls exponentially as faults are removed.

What is measured

Chapter Ninety, on software reliability, measured a running system by its outages. Measuring reliability during development needs the same raw material, failures and times, collected while the software is tested. Lyu names the two forms such data take.

  • Failure-count data track "the number of failures detected per unit of time": for example, 4 failures in the first 8 hours of testing, 4 in the next 8, 3 in the next.
  • Time-between-failures data track "the intervals between consecutive failures": 0.5 hours to the first failure, then 1.2 hours to the second, and so on.

Either can be converted into the other. The time can be calendar time or execution time. Musa's model uses execution time, because, as Lyu reports, "Musa feels that execution time is more reflective of the actual stress induced on the software system than the amount of calendar time that has elapsed."

Failure intensity is the measure that summarises the data: failures per unit of time. Lyu's closely related failure rate function, "also called the rate of occurrence of failures", is the probability of a failure per unit time just after time t, given none before it. For software under test, failure intensity should fall as failures reveal faults that are removed; that falling curve is the reliability growth of Chapter Ninety's figure.

Estimation and prediction

Lyu separates two activities.

  • Estimation "determines current software reliability by applying statistical inference techniques to failure data obtained during system test or during system operation": the reliability achieved up to now.
  • Prediction "determines future software reliability based upon available software metrics and measures": from failure data once testing has begun, or from product and process measures before it (early prediction).
munotes.in538

Measuring Software Reliability

A software reliability model makes prediction possible. It "specifies the general form of the dependence of the failure process on the principal factors that affect it: fault introduction, fault removal, and the operational environment." Its purpose, in Lyu's words, "is twofold: (1) to predict the extra time needed to test the software to achieve a specified objective; (2) to predict the expected reliability of the software when the testing is finished."

Musa's basic execution time model

Lyu calls Musa's basic model the one that "has had the widest distribution among the software reliability models", developed by John Musa of AT&T Bell Laboratories. Its main assumptions, as Lyu lists them:

  1. The cumulative number of failures by execution time t follows a Poisson process with mean value function μ(t) = β0(1 - e^(-β1 t)): the expected number of failures in a period is proportional to the expected number of faults still undetected. As t grows, μ(t) approaches β0, "the total number of faults that would be detected in the limit".
  2. The execution times between failures are exponentially distributed piece by piece: "the hazard rate for a single fault is constant."

From these follow the failure intensity λ(t) = β0β1e^(-β1 t), which "decreases exponentially to 0", and the reliability over a further Δt hours, R(Δt | t) = exp(-β0e^(-β1 t)(1 - e^(-β1 Δt))). The two parameters are estimated from the data by maximum likelihood. With n failures at times t1 to tn, and T the total time observed (the last failure time plus any failure-free time after it), the estimates satisfy

  • β0 = n / (1 - e^(-β1 T)), and
  • n / β1 - nT / (e^(β1 T) - 1) - (t1 + t2 + ... + tn) = 0,

the second equation being solved for β1 first.

Worked example: checking the model, then using it

A program that fits a model should be checked on data with a published answer before it is trusted with new data. Lyu's Example 3.3 gives ten failure times in CPU hours, 15 failure-free hours after the last, and its answer: β0 = 13.6 and β1 = 0.006. The program solves the same equations, checks that it gets Lyu's answer, and then fits ExamReg release 2.1's system test, in which 16 failures were logged over 260 hours of test execution. The release criterion is a failure intensity objective of 0.01 failures per test hour, one failure per 100 hours.

from math import exp, log

def fit_basic(times, total):
    """Musa's basic execution time model: the maximum likelihood estimates of beta0 and beta1
    (Lyu, section 3.3.4.4) from the failure times and the total time observed."""
    n, s = len(times), sum(times)

    def g(b1):                                   # the equation for beta1; it falls as beta1 grows
        return n / b1 - n * total / (exp(b1 * total) - 1) - s
    lo, hi = 1e-9, 1.0
    for _ in range(200):                         # bisection
        mid = (lo + hi) / 2
        lo, hi = (mid, hi) if g(mid) > 0 else (lo, mid)
    b1 = (lo + hi) / 2
    return n / (1 - exp(-b1 * total)), b1

# Lyu's Example 3.3: ten failures, at these CPU hours, then 15 more hours without a failure
b0, b1 = fit_basic([10, 18, 32, 49, 64, 86, 105, 132, 167, 207], 222)
print(f"Lyu's example 3.3: beta0 {b0:.1f}, beta1 {b1:.3f} (the handbook gives 13.6 and 0.006)")

# ExamReg release 2.1's system test (FINDINGS 5.7): failures at these test hours, 260 hours in all
times = [3, 8, 14, 20, 28, 37, 47, 58, 71, 86, 103, 122, 145, 171, 200, 232]
total, objective = 260, 0.01                     # objective: 0.01 failures per test hour
b0, b1 = fit_basic(times, total)
now = b0 * b1 * exp(-b1 * total)                 # failure intensity b0 b1 exp(-b1 t), at t = 260
print(f"ExamReg 2.1: beta0 {b0:.2f} failures in all, beta1 {b1:.5f};"
      f" {b0 - len(times):.2f} failures expected still to come")
print(f"failure intensity now {now:.4f} per test hour (MTTF {1 / now:.1f} h)")
r24 = exp(-b0 * exp(-b1 * total) * (1 - exp(-b1 * 24)))
print(f"probability of no failure in the next 24 test hours: {r24:.3f}")
print(f"test hours still needed to reach {objective}: {log(b0 * b1 / objective) / b1 - total:.1f}")
munotes.in539

Measuring Software Reliability

Lyu's example 3.3: beta0 13.6, beta1 0.006 (the handbook gives 13.6 and 0.006)
ExamReg 2.1: beta0 17.78 failures in all, beta1 0.00885; 1.78 failures expected still to come
failure intensity now 0.0158 per test hour (MTTF 63.4 h)
probability of no failure in the next 24 test hours: 0.711
test hours still needed to reach 0.01: 51.5

The check. The program reproduces Lyu's answer, β0 = 13.6 and β1 = 0.006, so its equations and its solver can be trusted with new data.

Estimation. For release 2.1 the model estimates 17.78 failures in all, of which 16 have been seen, so about 1.78 remain to be found by testing of this kind. The failure intensity now is 0.0158 failures per test hour, a mean time to failure of 63.4 hours: the gaps between failures grew from 5 hours at the start to over 30 at the end, and the model has turned that growth into a current rate.

Prediction. The objective is 0.01 failures per test hour, and the current 0.0158 has not reached it. Setting β0β1e^(-β1 t) equal to 0.01 and solving for t gives the answer to Lyu's first purpose: about 51.5 more hours of system testing. For the second, the chance of getting through the next 24 test hours without a failure is 0.711. Instead of the testers feel it is nearly ready, the release decision can say on this model, 52 more test hours reach the objective the college agreed to.

munotes.in540

Measuring Software Reliability

The limits. The numbers are only as good as the model's assumptions. Execution time, a constant hazard per fault and testing that resembles real use (the operational profile of Chapter Ninety, on software reliability) must all roughly hold; the likelihood equation for β1 has a solution only when failures are thinning out; and a model should be checked against how well it predicted before its forecasts are relied on. Lyu notes that Musa recommends this model particularly when predicting early reliability, when the program is changing substantially during the data collection, and when studying the effect of a new engineering technique.

Measuring reliability in a project

StepWhat is doneFor ExamReg release 2.1
Define failureWhat counts as a failure, by severityAny departure from the requirements a user would notice
Set an objectiveA failure intensity objective for release0.01 failures per test hour
Collect dataFailure times (or counts) in execution time, during testing that follows the operational profile16 failure times over 260 test hours
Fit a modelEstimate its parameters; check it on known data firstMusa's basic model; checked on Lyu's Example 3.3
Estimate and predictCurrent intensity; further testing needed; reliability ahead0.0158 per hour now; about 51.5 more hours
DecideRelease when the objective is met, or continue testingContinue testing

What it does not mean

A reliability model does not find faults. It summarises failures already found and forecasts the next ones; testing finds faults.

A predicted number is not a promise. It holds only while the model's assumptions hold; a change to the software or to how it is tested restarts the measurement.

Fewer failures in a week is not proof of reliability growth. It may mean less testing was done; that is why time is measured in execution time, not calendar time.

Failures are not defects. One defect can cause many failures, and reliability counts what users would experience, not what the code contains.

Quick revision

  • Failure intensity (IEEE 982:2024): observed failures divided by the duration observed. ROCOF: rate of occurrence of failures. Failure intensity objective: the level to be reached in pre-release testing as a release criterion.
  • Data (Lyu): failure-count data (failures per period) and time-between-failures data; execution time preferred to calendar time.
  • Estimation (reliability achieved so far) and prediction (future reliability; the extra testing needed to reach an objective).
  • Musa's basic execution time model (Lyu 3.3.4): μ(t) = β0(1 - e^(-β1 t)); λ(t) = β0β1e^(-β1 t); β0 = total failures in the limit; R(Δt | t) = exp(-β0e^(-β1 t)(1 - e^(-β1 Δt))); maximum likelihood estimates from failure times and total time.
  • Worked example: Lyu's Example 3.3 reproduced (13.6, 0.006); release 2.1: 17.78 failures expected in all (1.78 to come), 0.0158 failures per hour now (MTTF 63.4 h), 0.711 probability of 24 failure-free hours, about 51.5 more hours to reach 0.01.
munotes.in541

Measuring Software Reliability

Test yourself

1. What data are collected to measure software reliability? Failures and the times at which they occur, either as failure-count data, the number of failures in each period, or as time-between-failures data, the intervals between consecutive failures. Time is preferably execution time, and failures are counted according to a stated definition and severity.

2. Define failure intensity and the failure intensity objective. Failure intensity is the number of observed failures divided by the duration observed, the rate at which failures occur. The failure intensity objective is the level of failure intensity to be reached during pre-release testing as part of the release criteria.

3. Distinguish reliability estimation from reliability prediction. Estimation applies statistical inference to failure data from testing or operation to determine the reliability achieved so far. Prediction determines future reliability, from failure data through a reliability model, or, before testing, from product and process measures (early prediction).

4. State Musa's basic execution time model. The cumulative number of failures by execution time t is a Poisson process with mean value function β0(1 - e^(-β1 t)), where β0 is the total number of failures that would be seen in the limit; the failure intensity β0β1e^(-β1 t) falls exponentially as faults are removed, and the time between failures is exponential with a constant hazard for each fault.

5. How does a reliability model support the release decision? By estimating the current failure intensity from the failure data and predicting the further testing needed to reach the failure intensity objective. For ExamReg release 2.1 the intensity was 0.0158 failures per test hour against an objective of 0.01, and the model predicted about 51.5 more hours of testing.

6. Why was the program run on Lyu's example first? To check that its equations and its numerical solution are right: it reproduced the published answer, β0 = 13.6 and β1 = 0.006, before being trusted with release 2.1's data, whose answer nobody knows in advance.

Contents This chapter on its own page

munotes.in542

Chapter Ninety-Three

Improving Software Reliability

Syllabus topic Module 2, "Software Quality Assurance: ... Software reliability measurement and improvement"

In one line

Software becomes more reliable in four ways, which Lyu names fault prevention, fault removal, fault tolerance and fault/failure forecasting; testing to remove faults works but with diminishing returns, the operational profile tells a team which faults users would actually meet, and where failures can cause harm, reliability is not enough and safety must be engineered in its own right.

In the wording a student can write in an examination: reliability growth is the "improvement in reliability that results from correction of faults" (ISO/IEC/IEEE 24765). Lyu's four technical areas are: fault prevention, "To avoid, by construction, fault occurrences"; fault removal, "To detect, by verification and validation, the existence of faults and eliminate them"; fault tolerance, "To provide, by redundancy, service complying with the specification in spite of faults having occurred or occurring"; and fault/failure forecasting, "To estimate, by evaluation, the presence of faults and the occurrence and consequences of failures." The operational profile, "the set of operations that the software can execute along with the probability with which they will occur" (Lyu), directs testing and fixing to what users do most. Safety, the "expectation that a system does not, under defined conditions, lead to a state in which human life, health, property, or the environment is endangered" (ISO/IEC/IEEE 12207:2026), is a different property from reliability.

Four ways to more reliable software

Area (Lyu)What it doesTechniques Lyu namesOn ExamReg
Fault preventionStops faults being madeRefined requirements, good design methods, structured programming, clear code, formal methods, reuseThe shared validation module and requirements checklist item (Chapter Eighty, on using defect data)
Fault removalFinds and removes faults that were madeTesting, formal inspectionReviews, and every test level of Module 1
Fault toleranceKeeps the service right when a fault is triggeredMonitoring, atomicity of actions, decision verification, exception handling; recovery blocks, N-version programmingA payment that either completes or is rolled back whole
Fault/failure forecastingEstimates faults remaining and failures to comeReliability models, failure data, toolsMusa's model on release 2.1 (Chapter Ninety-Two, on measuring reliability)

NIST Special Publication 500-235 describes the effect of multiple versions: "any disagreement between versions can be reported during testing but a majority voting mechanism helps reduce the likelihood of incorrect output after delivery." Lyu's summary of prevention lists "The interactive refinement of the user's system requirement, the engineering of the software specification process, the use of good software design methods, the enforcement of a structured programming discipline, and the encouragement of writing clear code". For removal, besides testing, he names formal inspection, "a rigorous process focused on finding faults, correcting faults, and verifying the corrections" (Chapter Twenty-Nine, on inspection). His conclusion is that no single area is enough: "software reliability engineers must apply a combination of the above methods for the delivery of reliable software systems."

munotes.in543

Improving Software Reliability

Worked example 1: what more testing buys

Reliability growth testing is fault removal guided by forecasting: test, record failures, remove their faults, and watch the failure intensity fall. Chapter Ninety-Two fitted Musa's basic model to release 2.1's system test. Under that model the failure intensity falls exponentially with test time, so it halves every ln 2 / β1 hours. The program refits the model and follows four halvings.

from math import exp, log

def fit_basic(times, total):                     # Musa's basic model, as in Chapter 92
    n, s = len(times), sum(times)
    lo, hi = 1e-9, 1.0
    for _ in range(200):
        mid = (lo + hi) / 2
        g = n / mid - n * total / (exp(mid * total) - 1) - s
        lo, hi = (mid, hi) if g > 0 else (lo, mid)
    b1 = (lo + hi) / 2
    return n / (1 - exp(-b1 * total)), b1

# release 2.1's system test (FINDINGS 5.7): failure times in test hours, 260 hours in all
times, total = [3, 8, 14, 20, 28, 37, 47, 58, 71, 86, 103, 122, 145, 171, 200, 232], 260
b0, b1 = fit_basic(times, total)
halving = log(2) / b1                            # hours for the failure intensity to halve
print(f"each halving of the failure intensity takes {halving:.1f} more test hours")
t = total
for step in range(4):
    found = b0 * (exp(-b1 * t) - exp(-b1 * (t + halving)))     # expected failures in the step
    print(f"   {t:>5.0f} to {t + halving:>5.0f} h: intensity {b0 * b1 * exp(-b1 * t):.4f}"
          f" -> {b0 * b1 * exp(-b1 * (t + halving)):.4f}; failures expected {found:.2f}")
    t += halving
each halving of the failure intensity takes 78.3 more test hours
     260 to   338 h: intensity 0.0158 -> 0.0079; failures expected 0.89
     338 to   417 h: intensity 0.0079 -> 0.0039; failures expected 0.45
     417 to   495 h: intensity 0.0039 -> 0.0020; failures expected 0.22
     495 to   573 h: intensity 0.0020 -> 0.0010; failures expected 0.11

Each halving of the failure intensity costs the same 78.3 hours of testing, but finds half as many failures as the one before: 0.89 expected failures in the first 78 hours, then 0.45, 0.22 and 0.11. Testing buys reliability at a steadily rising price per failure found. That is the quantitative reason for Lyu's "combination": after a point, the next improvement is cheaper through prevention, which stops faults being made, or tolerance, which stops them becoming failures, than through yet more testing.

Worked example 2: fixing what students would meet

Not every fault costs the same reliability. A fault in an operation students use constantly causes more failures than one in an operation they rarely reach, and the operational profile says which is which. Release 2.1's system test ran 2,000 sessions generated from Chapter Eighty-Eight's usage model and logged the operation in which each of its 16 failures occurred. The program combines the failures per use of each operation with how often a student uses it in a session.

munotes.in544

Improving Software Reliability

# release 2.1's system test (FINDINGS 5.7): times each operation ran, and the failures in it
ran = {"log in": 2000, "fill form": 1100, "pay fee": 936, "payment failed": 94, "hall ticket": 1256}
failed = {"log in": 1, "fill form": 6, "pay fee": 5, "payment failed": 1, "hall ticket": 3}
# the operational profile: expected uses of each operation in one student's session (Chapter 88)
uses = {"log in": 1.0, "fill form": 0.55, "pay fee": 0.468, "payment failed": 0.047,
        "hall ticket": 0.628}

per_use = {op: failed[op] / ran[op] for op in ran}
per_session = {op: uses[op] * per_use[op] for op in ran}
total = sum(per_session.values())
print(f"{'operation':<16}{'failures per use':>18}{'share of the failures a student meets':>39}")
for op in sorted(ran, key=lambda o: -per_session[o]):
    print(f"{op:<16}{per_use[op]:>18.4f}{per_session[op] / total:>39.0%}")
print(f"probability that a session meets a failure: {total:.4f}")
operation         failures per use  share of the failures a student meets
fill form                   0.0055                                    38%
pay fee                     0.0053                                    31%
hall ticket                 0.0024                                    19%
log in                      0.0005                                     6%
payment failed              0.0106                                     6%
probability that a session meets a failure: 0.0080

The order of work. Filling the form and paying the fee account for 38 and 31 per cent of the failures a student would meet; improving those two operations improves reliability as students experience it most. The retry after a failed payment has the highest failure rate per use, 0.0106, yet only 6 per cent of the failures students meet, because few sessions reach it. Ranked by failures per use alone, it would come first; ranked by the operational profile, it comes last.

And yet. A failure in the retry could charge a student twice. The operational profile ranks faults by how often they would be met, not by what they would do, and a rare failure with serious consequences needs attention of its own. That is where reliability ends and safety begins.

Reliability is not safety

The standards define safety as the "expectation that a system does not, under defined conditions, lead to a state in which human life, health, property, or the environment is endangered" (ISO/IEC/IEEE 12207:2026), and a hazard as an "intrinsic property or condition that has the potential to cause harm or damage" (IEEE 1012-2024). Reliability counts failures; safety asks what a failure would do.

Leveson and Turner's account of the Therac-25, the radiation therapy machine of Chapter Three, on why software must be tested, names the confusion among its lessons: "This software was highly reliable. It worked tens of thousands of times before overdosing anyone, and occurrences of erroneous behavior were few and far between. AECL assumed that their software was safe because it was reliable, and this led to complacency." A second lesson is about tolerance: "The software did not contain self-checks or other error-detection and error-handling features" that would have caught the inconsistencies and coding errors.

munotes.in545

Improving Software Reliability

Safety is therefore engineered and assured separately. Hazards are identified and analysed; the software that could lead to them is found; defensive design (self-checks, interlocks, fault tolerance) is added where they could occur; and assurance confirms it. NASA's software assurance standard lists among its purposes "Ensuring that the software systems are safe and that the software safety-critical requirements are followed." ExamReg endangers no one's life, but it handles students' money, and the same reasoning applies at its scale: the payment retry is ExamReg's hazard, and it gets a model, tests and an atomic design of its own.

What it does not mean

Reliability is not improved only by testing. Testing removes faults with diminishing returns; prevention and tolerance supply what testing cannot afford.

The highest failure rate is not the first fix. Faults are ranked by the failures users would meet, which combines the rate with how often the operation is used, and then by their consequences.

Reliable is not safe. The Therac-25 worked tens of thousands of times; the rare failure killed people.

Fault tolerance is not an excuse for faults. It keeps the service right when a fault is triggered; the fault is still found and removed.

Quick revision

  • Reliability growth (ISO/IEC/IEEE 24765): improvement in reliability from correcting faults.
  • Four areas (Lyu): fault prevention (avoid by construction), fault removal (detect by V&V and eliminate), fault tolerance (service despite faults, by redundancy), fault/failure forecasting (estimate by evaluation); a combination is needed.
  • Fault tolerance techniques (Lyu): monitoring, atomicity of actions, decision verification, exception handling; recovery blocks and N-version programming by design diversity.
  • Operational profile (Lyu): operations and their probabilities; rank faults by the failures users would meet.
  • Safety (ISO/IEC/IEEE 12207:2026) is not reliability (Therac-25: "safe because it was reliable"); hazards need analysis and defensive design.
  • Worked examples: each halving of release 2.1's failure intensity takes 78.3 test hours and finds 0.89, 0.45, 0.22, 0.11 failures; form and payment give 38 and 31 per cent of the failures a student meets; the payment retry has the highest rate per use (0.0106) but a 6 per cent share; a session meets a failure with probability 0.0080.

Test yourself

1. Name and explain the four technical areas for achieving reliable software. Fault prevention avoids faults by construction, through good requirements, design methods and coding discipline; fault removal detects faults by verification and validation, such as testing and inspection, and eliminates them; fault tolerance uses redundancy and defensive techniques to keep the service correct when a fault is triggered; fault/failure forecasting estimates the faults remaining and the failures to come with reliability models.

munotes.in546

Improving Software Reliability

2. What is reliability growth testing, and why does it have diminishing returns? Testing in which failures are recorded and their faults removed, so that reliability grows and is tracked with a model. Under an exponential model each halving of the failure intensity takes the same test time but finds half as many failures as the previous halving; for release 2.1, 78.3 hours per halving, finding 0.89, then 0.45, then 0.22 failures.

3. How does the operational profile help improve reliability? It gives the probability with which each operation is used, so failures per use can be weighted by use to show which faults users would actually meet. For ExamReg, filling the form and paying caused 38 and 31 per cent of the failures students would meet, while the payment retry, with the highest failure rate per use, caused only 6 per cent.

4. Describe two fault tolerance techniques. In a single version of the software, exception handling and atomic actions (an operation that completes whole or not at all) partially tolerate faults, keeping a triggered fault from becoming a wrong result. With design diversity, functionally equivalent versions are developed independently; with several versions, a majority vote among their outputs reduces the chance of an incorrect output, the idea behind N-version programming (Lyu also names recovery blocks).

5. Distinguish reliability from safety, with an example. Reliability is how rarely a system fails; safety is whether its failures can lead to harm to life, health, property or the environment. The Therac-25's software worked tens of thousands of times, so it was highly reliable, but its rare failures overdosed patients; its maker assumed it was safe because it was reliable.

6. Why does the payment retry on ExamReg need attention although it contributes few failures? Because its consequence, charging a student twice, is serious even if rare. The operational profile ranks by how often failures would be met, not by their consequences, so hazardous rare functions are given their own analysis, tests and defensive design.

Contents This chapter on its own page

munotes.in547

Chapter Ninety-Four

The ISO 9000 Family and the Seven Principles

Syllabus topic Module 2, "Software Quality Assurance: ... ISO 9000 Quality Standards"

In one line

The ISO 9000 family is the set of international standards for quality management: ISO 9000 gives its fundamentals and vocabulary, ISO 9001 its requirements (the only standard of the family an organisation can be certified to), ISO 9004 guidance for sustained success, ISO 19011 guidance for auditing, and sector standards such as ISO/IEC/IEEE 90003 adapt it to software; all of them rest on seven quality management principles.

In the wording a student can write in an examination: the ISO 9000 family is ISO's family of quality management standards, now in 2026 editions for its two core members. ISO 9000:2026 "provides the fundamentals and vocabulary for quality management systems (QMS)"; ISO 9001:2026's "requirements define how to establish, implement, maintain, and continually improve a quality management system (QMS)", and it "is the only standard that can be certified to" (ISO). The seven quality management principles are customer focus, leadership, engagement of people, process approach, improvement, evidence-based decision making and relationship management; "ISO 9000, ISO 9001 and related ISO quality management standards are based on these seven QMPs."

The family

StandardEdition now currentWhat it is
ISO 9000ISO 9000:2026 (May 2026)Quality management: fundamentals and vocabulary; the principles and terms, including the definition of quality
ISO 9001ISO 9001:2026 (16 September 2026)Quality management systems: requirements; what a QMS must do; the only member that can be certified to
ISO 9004ISO 9004:2018 (confirmed 2023)Guidance "for enhancing an organization's ability to achieve sustained success", with a self-assessment tool
ISO 19011ISO 19011:2026 (May 2026)"Guidelines for auditing management systems", including quality management systems
Sector standardsSeveralISO 9001 adapted to a sector; for software, ISO/IEC/IEEE 90003

The table gives ISO's titles with a colon in place of their dash. ISO's page for ISO 9001:2026 lists the sector adaptations: ISO 13485 on medical devices, ISO 22163 on railways, ISO 29001 on petroleum and gas, ISO 18091 on local government, ISO/TS 54001 on elections, and "ISO/IEC/IEEE 90003 on computer software". Chapter Ninety-Five, on ISO 9001, certification and software, takes up ISO 9001's requirements and ISO/IEC/IEEE 90003.

A currency warning. The textbooks MU lists were written against ISO 9001:2008 or ISO 9001:2015. Both are withdrawn. ISO's page records that the 2026 edition "focuses on improving clarity", "emphasizes the importance of quality culture and leadership and separates risk and opportunities", and that certified organisations "will have to transition to the new version within the timeframe set by their certification cycle." ISO 19011 also has a 2026 edition, which replaced the 2018 one this year. An answer that names the 2015 edition as current is out of date.

How the family grew

ASQ's history of quality dates the first edition: "The ISO 9000 series of quality-management standards, for example, were published in 1987", the same year the Baldrige award was established. It records two revisions of emphasis: "In 2000, the ISO 9000 series of quality management standards was revised to increase emphasis on customer satisfaction", and "in 2015, the ISO 9001 standard was revised to increase emphasis on risk management." ISO's own page records the 2026 revision, which followed "global consultation in 2023" and "a consensus confirmed that revising the standard would enhance its value". Chapter Eighty-Three, on Feigenbaum, Ishikawa, Crosby and TQM, placed the family in the quality movement: ASQ notes that TQM's principles now live on in quality management systems and standards such as the ISO 9000 series.

munotes.in548

The ISO 9000 Family and the Seven Principles

The seven quality management principles

ISO's brochure explains what a principle is here: "a basic belief, theory or rule that has a major influence on the way in which something is done", and it stresses that "These principles are not listed in priority order." For each it gives a statement.

PrincipleISO's statement
1. Customer focus"The primary focus of quality management is to meet customer requirements and to strive to exceed customer expectations."
2. Leadership"Leaders at all levels establish unity of purpose and direction and create conditions in which people are engaged in achieving the organization's quality objectives."
3. Engagement of people"Competent, empowered and engaged people at all levels throughout the organization are essential to enhance its capability to create and deliver value."
4. Process approach"Consistent and predictable results are achieved more effectively and efficiently when activities are understood and managed as interrelated processes that function as a coherent system."
5. Improvement"Successful organizations have an ongoing focus on improvement."
6. Evidence-based decision making"Decisions based on the analysis and evaluation of data and information are more likely to produce desired results."
7. Relationship management"For sustained success, an organization manages its relationships with interested parties, such as suppliers."

The principles gather ideas this module has met separately. Customer focus is Juran's fitness for use; leadership and engagement are Deming's points about management and fear (Chapter Eighty-Two, on Shewhart, Deming and Juran); the process approach is the premise of all quality assurance; improvement is the PDSA cycle and causal analysis; evidence-based decision making is statistical quality assurance; relationship management extends quality to suppliers and partners, the concern of Deming's fourth point.

Worked example: ExamReg's project against the seven principles

ISO 9004 offers organisations a self-assessment of how far they have adopted its concepts. A simple version asks, for each principle, what evidence the organisation can show. For ExamReg's software house, the evidence is what this book's earlier chapters recorded.

PrincipleEvidence from ExamReg's projectWhat is missing
Customer focusThe exam cell's ruling on invalid input; its part in reviews and acceptance testing (Chapter Forty-Two, on acceptance testing)A measure of student satisfaction after each registration season
LeadershipThe head of delivery acted on the escalated noncompliance NC-10 (Chapter Eighty-Six, on SQA activities)Quality objectives set by leadership, not only by the project
Engagement of peopleA requirements checklist and test templates built by the team after causal analysis (Chapter Eighty, on using defect data)Training records; a regular retrospective
Process approachAn SQA plan whose audits cover every standard the project follows (Chapter Eighty-Seven, on the quality assurance plan)Process measures watched with control charts across releases
ImprovementCausal analysis of input validation; release 2.1's rate fell from 2.90 to 1.12 per KLOCThe same loop for the next largest cause
Evidence-based decision makingThe Pareto of defect causes (Chapter Eighty-Nine, on statistical SQA); the release decision from a reliability model (Chapter Ninety-Two, on measuring reliability)Data kept consistently from release to release
Relationship managementThe payment gateway's provider, whose wait caused the 72-hour repair (Chapter Ninety-One, on statistical process control)An agreement with the provider on response times
munotes.in549

The ISO 9000 Family and the Seven Principles

The table shows two things the principles are for. First, they are practical: every principle already has evidence in a small college portal's project. Second, they find gaps. The weakest row is relationship management: the one repair that fell outside the process's limits depended on a supplier with whom no agreement existed.

What it does not mean

ISO 9000 is not the standard organisations are certified to. ISO 9000 is the vocabulary and fundamentals; certification is to ISO 9001, and only to ISO 9001.

The principles are not a ranking. ISO says they are not in priority order, and their relative importance varies between organisations and over time.

The 2015 edition is not current. ISO 9001:2026 replaced it on 16 September 2026, and ISO 9000:2026 replaced ISO 9000:2015 in May.

A family standard is not a software method. ISO 9001 says what a quality management system must achieve, not how to write or test software; ISO/IEC/IEEE 90003 gives the software guidance.

Quick revision

  • The family: ISO 9000:2026 (fundamentals and vocabulary); ISO 9001:2026 (requirements; the only certifiable one); ISO 9004:2018 (sustained success; self-assessment); ISO 19011:2026 (auditing); sector standards, including ISO/IEC/IEEE 90003:2018 for software.
  • History (ASQ; ISO): first published 1987; revised 2000 (customer satisfaction) and 2015 (risk); 2026 editions (clarity, quality culture and leadership, risk separated from opportunities).
  • Seven principles: customer focus; leadership; engagement of people; process approach; improvement; evidence-based decision making; relationship management. Not in priority order.
  • Worked example: ExamReg's project shows evidence for every principle, and its weakest is relationship management (the 72-hour repair that waited for the payment gateway's provider).
munotes.in550

The ISO 9000 Family and the Seven Principles

Test yourself

1. What is the ISO 9000 family? Name its main members. ISO's family of international standards for quality management. ISO 9000 gives the fundamentals and vocabulary; ISO 9001 the requirements for a quality management system, and it is the only one that can be certified to; ISO 9004 guidance for sustained success; ISO 19011 guidelines for auditing management systems; and sector standards adapt ISO 9001, such as ISO/IEC/IEEE 90003 for computer software.

2. Which editions are current, and why does it matter? ISO 9000:2026 (May 2026) and ISO 9001:2026 (16 September 2026) replaced the 2015 editions; ISO 19011:2026 replaced the 2018 edition; ISO 9004:2018 remains current. It matters because certified organisations must transition to the new edition, and an answer that treats the 2015 edition as current is out of date.

3. State the seven quality management principles. Customer focus; leadership; engagement of people; process approach; improvement; evidence-based decision making; relationship management.

4. Explain the process approach, with a software example. Consistent and predictable results come more effectively when activities are understood and managed as interrelated processes working as a coherent system. In software, requirements, design, coding, review and testing are managed as linked processes with defined inputs and outputs; for example, ExamReg's SQA plan audits each process against its standard.

5. Explain evidence-based decision making, with a software example. Decisions based on the analysis and evaluation of data are more likely to produce the desired results. For ExamReg, the improvement effort was directed by a Pareto analysis of defect causes, and the release decision by a reliability model's prediction of the testing still needed.

6. How did ExamReg's project show a gap in relationship management? One repair took 72 hours, far beyond the process's usual variation, because it waited for the payment gateway's provider; the principle calls for managing such supplier relationships, for example with an agreement on response times.

Contents This chapter on its own page

munotes.in551

Chapter Ninety-Five

ISO 9001: The Requirements, Certification and Software

Syllabus topic Module 2, "Software Quality Assurance: ... ISO 9000 Quality Standards"

In one line

ISO 9001 states what an organisation's quality management system must do, in seven topics shared with ISO's other management system standards, from understanding the organisation's context to improving the system; an organisation may have an accredited, independent body certify that its system meets the requirements; and ISO/IEC/IEEE 90003 explains how software organisations apply them, without adding any.

In the wording a student can write in an examination: ISO 9001:2026, whose title is Quality management systems: requirements, specifies requirements that "define how to establish, implement, maintain, and continually improve a quality management system (QMS)" (ISO). Its topics are context of the organization, leadership, planning, support, operation, performance evaluation and improvement. Certification is "third-party attestation related to an object of conformity assessment, with the exception of accreditation" (ISO/IEC 29110-1-2:2024); certification bodies are themselves accredited, when "an accreditation body has provided independent confirmation of the certification body's competence" (ISO). ISO/IEC/IEEE 90003:2018 "provides guidance for organizations in the application of ISO 9001:2015 to the acquisition, supply, development, operation and maintenance of computer software and related support services."

The requirements in seven topics

ISO's page summarises what the standard covers. The requirements themselves are in the standard, which is sold by ISO; the table gives each topic in ISO's public words and what it means for a software house such as the one that builds ExamReg.

TopicISO's summaryIn a software house
Context of the organizationDetermine "the external and internal factors that affect their ability to achieve the intended results of their quality management system"Who the customers and interested parties are (the college, its students, the payment gateway's provider) and what they need
Leadership"The standard emphasizes the importance of leadership in implementing and maintaining a quality management system."Management owns the quality policy and objectives, and acts on escalations
Planning"The quality management system must include measures designed to achieve an organization's quality objectives and continuously improve the system's effectiveness."Quality objectives with measures; risks and opportunities addressed
Support"ISO 9001 addresses issues such as resources, competence, awareness, communication and documented information."Trained people, tools, and controlled documents and records
Operation"The processes necessary to meet customer requirements and increase customer satisfaction must be planned, implemented and controlled."The life cycle: requirements, design, coding, reviews, testing, release, support
Performance evaluationOrganisations must "monitor, measure, analyze and evaluate the performance and effectiveness of their quality management system."Defect, process and reliability measures; internal audits; management review
Improvement"continuously increasing the effectiveness of the quality management system based on the results of performance evaluation and other data sources"Corrective action on nonconformities; causal analysis; improvement projects

ISO's page adds that "ISO standards that look at different types of management systems, such as ISO 9001 for quality and ISO 14001 for environmental management, are all structured in the same way." That is why ISO describes the 2026 edition as easier to integrate into an organisation's existing management systems, such as one for the environment, and why the definitions in this chapter can be taken from another management system standard on the same structure.

munotes.in552

ISO 9001: The Requirements, Certification and Software

A few of those shared terms matter for what follows. An audit is a "systematic, independent and documented process for obtaining audit evidence and evaluating it objectively to determine the extent to which audit criteria are fulfilled"; a nonconformity is the "non-fulfillment of a requirement"; corrective action is "action to eliminate the cause of a nonconformity and to prevent recurrence"; and documented information is "information required to be controlled and maintained by an organization and the medium on which it is contained" (all ISO/IEC 19770-1:2017).

What changed in 2026

ISO's page describes the sixth edition, published on 16 September 2026: it "focuses on improving clarity to help organizations of all sizes, maturity and purpose to better understand the requirements", "emphasizes the importance of quality culture and leadership and separates risk and opportunities to ensure organizations proactively take actions to pursue beneficial results", and improves "alignment to other ISO management system standards". It replaced the fifth edition, "published in 2015", and its 2024 amendment, both now withdrawn.

For an organisation already certified, ISO's answer is: "Certified organizations will have to transition to the new version within the timeframe set by their certification cycle. For further information, they should contact their certification body."

Certification and accreditation

Certification is voluntary. "As with other ISO management system standards, companies implementing ISO 9001 can choose whether they want to go through a certification process or not." It is, in ISO's words, "one way to demonstrate to stakeholders and customers that you are committed and able to consistently deliver high quality products or services." ISO reports "more than one million certificates issued to organizations in 189 countries".

Three parties are involved.

  1. The organisation implements a quality management system that meets ISO 9001.
  2. A certification body, independent of it (a third party), audits the system against the standard and, if it conforms, issues a certificate.
  3. An accreditation body confirms that the certification body itself is competent: "Holding a certificate issued by an accredited conformity assessment body may bring an additional layer of confidence, as an accreditation body has provided independent confirmation of the certification body's competence."

The audits follow the guidance of ISO 19011, "Guidelines for auditing management systems", whose current edition is ISO 19011:2026 (Chapter Ninety-Four, on the ISO 9000 family). An auditor looks for objective evidence, "data supporting the existence or verity of something" (ISO/IEC 33001:2015): records, measures and documents that show the system does what it claims. A nonconformity found in the audit calls for corrective action: action to remove its cause, so that it does not recur.

munotes.in553

ISO 9001: The Requirements, Certification and Software

ISO 9001 for software: ISO/IEC/IEEE 90003

ISO 9001 applies to any organisation, so it says nothing specific about software. ISO/IEC/IEEE 90003:2018 fills the gap as guidance: it "provides guidance for organizations in the application of ISO 9001:2015 to the acquisition, supply, development, operation and maintenance of computer software and related support services", and "It does not add to or otherwise change the requirements of ISO 9001:2015." It is not a certification standard: its guidelines "are not intended to be used as assessment criteria in quality management system registration/certification", although an organisation may use it alongside ISO 9001 to judge a software quality management system.

Two currency facts belong in any answer about it. It was confirmed in 2025 and "remains current", so it is the edition to cite; and it is still written for ISO 9001:2015, which ISO 9001:2026 has replaced. ISO's page lists no edition for the 2026 standard yet, so a software organisation today applies ISO/IEC/IEEE 90003's guidance to requirements that have since been revised, and checks each point against the new edition.

Worked example: a gap analysis for ExamReg's software house

Before inviting a certification body, an organisation checks itself against each topic. For ExamReg's software house, the evidence is what this book's chapters recorded; the gaps are what a certification audit would be likely to find.

TopicEvidence already recordedGap to close
ContextThe exam cell, students and the payment gateway's provider as interested partiesNo documented review of external factors, such as the university's rule changes
LeadershipHead of delivery acted on an escalated noncompliance (Chapter Eighty-Six, on SQA activities)No written quality policy or organisation-wide quality objectives
PlanningRelease criteria such as the failure intensity objective (Chapter Ninety-Two, on measuring reliability)Risks and opportunities not recorded separately, as the 2026 edition expects
SupportSQA plan and audit records kept in version control (Chapter Eighty-Seven, on the quality assurance plan)Competence and training records
OperationA defined life cycle with reviews, test levels and acceptance testing (Module 1)Supplier control for the payment gateway
Performance evaluationDefect metrics (Chapter Seventy-Nine, on defect metrics), control charts (Chapter Ninety-One, on statistical process control) and SQA audits (Chapter Eighty-Six, on SQA activities)No management review of the whole system
ImprovementCausal analysis and release 2.1's measured result (Chapter Eighty, on using defect data)Corrective action records linked to each nonconformity

The analysis shows what certification checks: not whether ExamReg is free of defects, but whether the organisation has a system that plans, controls, measures and improves its work, with objective evidence for each part. Most of the evidence already exists because the project followed its SQA plan; the gaps are in the organisation-wide layer (policy, objectives, management review, supplier control) that a single project does not provide.

munotes.in554

ISO 9001: The Requirements, Certification and Software

What it does not mean

ISO 9001 certification does not certify a product. It attests that the quality management system conforms to the standard; a certified organisation can still ship a defective release.

ISO 9001 does not prescribe how to build software. It states what the management system must achieve; how, for software, is the organisation's choice, guided by ISO/IEC/IEEE 90003.

ISO/IEC/IEEE 90003 is not a certification standard. Certification is to ISO 9001; 90003 is guidance.

Certification is not compulsory. ISO leaves it to each organisation; customers or contracts may require it.

Quick revision

  • ISO 9001:2026: requirements to "establish, implement, maintain, and continually improve" a QMS; the only certifiable member of the family; published 16 September 2026, replacing ISO 9001:2015.
  • Seven topics: context of the organization, leadership, planning, support, operation, performance evaluation, improvement; the same structure as ISO's other management system standards.
  • 2026 changes (ISO): clarity; quality culture and leadership; risks separated from opportunities; alignment with other management system standards; certified organisations transition within their certification cycle.
  • Certification: third-party attestation (ISO/IEC 29110-1-2:2024); voluntary; over one million certificates in 189 countries. Accreditation: an accreditation body confirms the certification body's competence. Audits follow ISO 19011:2026 and look for objective evidence.
  • ISO/IEC/IEEE 90003:2018: guidance for applying ISO 9001 to software; adds no requirements; not for certification; still written for ISO 9001:2015.
  • Worked example: ExamReg's project has evidence for every topic; its gaps are organisation-wide (policy, objectives, management review, supplier control, training and corrective action records).

Test yourself

1. What does ISO 9001 specify? Name the topics its requirements cover. The requirements for a quality management system: how to establish, implement, maintain and continually improve it. The topics are context of the organization, leadership, planning, support, operation, performance evaluation and improvement.

2. What is the difference between certification and accreditation? Certification is third-party attestation that something, here an organisation's quality management system, conforms to specified requirements, ISO 9001. Accreditation is independent confirmation, by an accreditation body, that a certification body is competent to certify. A certificate from an accredited body therefore carries more confidence.

3. What changed with ISO 9001:2026, and what must certified organisations do? ISO describes clearer requirements, more emphasis on quality culture and leadership, risks and opportunities treated separately, and closer alignment with other management system standards. Certified organisations must transition to the new edition within the timeframe set by their certification cycle, in consultation with their certification body.

munotes.in555

ISO 9001: The Requirements, Certification and Software

4. What is ISO/IEC/IEEE 90003, and how does it relate to ISO 9001? Guidance for applying ISO 9001 to the acquisition, supply, development, operation and maintenance of computer software and related services. It adds no requirements and is not used as certification criteria; certification remains to ISO 9001. Its current edition, 2018, is still written for ISO 9001:2015.

5. Does ISO 9001 certification guarantee good software? Explain. No. It attests that the organisation's quality management system conforms to the standard: that it plans, controls, measures and improves its work with objective evidence. A certified organisation can still release defective software; certification gives confidence in the system, not a guarantee about a product.

6. What would a certification audit be likely to find in ExamReg's software house? Project-level evidence is strong: an SQA plan and audit records, defect and reliability measures, causal analysis. The likely nonconformities are at organisation level: no written quality policy or objectives, no management review of the whole system, no supplier control for the payment gateway, and incomplete training and corrective action records.

Contents This chapter on its own page

munotes.in556

Chapter Ninety-Six

Why Reviews Pay: The Cost of a Late Defect

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: Formal Technical Reviews and their benefits"

In one line

A defect costs more to fix the later it is found, often a hundred times more after delivery than during requirements and design, and a defect that escapes early leads to more defects in the work built on it; reviews find defects early, so they save far more than they cost, and ExamReg's own data show both effects.

In the wording a student can write in an examination: Boehm and Basili (2001) report that "Finding and fixing a software problem after delivery is often 100 times more expensive than finding and fixing it during the requirements and design phase", while for small, noncritical systems the ratio is "more like 5:1 than 100:1". Defect amplification is the name this book gives to the effect the ISTQB syllabus describes: "Defects in work products produced earlier in the SDLC, if undetected, often lead to defective work products later in the lifecycle." Reviews attack both: they find defects early, when fixes are cheap, and they stop defects propagating; Boehm and Basili report that peer reviews catch "from 31 to 93 percent of the defects, with a median of around 60 percent."

The cost of a late defect

The earliest chapters of this book cited the finding; this chapter uses it. Boehm and Basili put it first in their list of what the evidence shows about reducing defects: "Finding and fixing a software problem after delivery is often 100 times more expensive than finding and fixing it during the requirements and design phase." They qualify it carefully. The word "often" was added deliberately, because "the cost-escalation factor for small, noncritical software systems" is "more like 5:1 than 100:1", and because "good architectural practices can significantly reduce the cost-escalation factor even for large critical systems." Fagan, reporting IBM's inspection results (Chapter Twenty-Nine, on inspection), gave a similar range: rework at the early levels "is 10 to 100 times less expensive than if it is done in the last half of the process."

Why should a late fix cost so much more? A defect found in a requirements review is corrected in one document. The same defect found after release has been designed around, coded, tested and shipped: the fix must change the requirement, the design, the code and the tests, retest everything they touch, redeploy, and deal with the users it affected.

The finding also has a cost in effort. Boehm and Basili report that "Current software projects spend about 40 to 50 percent of their effort on avoidable rework", effort "spent fixing software difficulties that could have been discovered earlier and fixed less expensively or avoided altogether", and that "About 80 percent of avoidable rework comes from 20 percent of the defects."

munotes.in557

Why Reviews Pay: The Cost of a Late Defect

Defect amplification

A defect in an early work product does not stay one defect. The ISTQB syllabus states it in two places: "Defects in work products produced earlier in the SDLC, if undetected, often lead to defective work products later in the lifecycle", and, as its third principle of testing, "Defects that are removed early in the process will not cause subsequent defects in derived work products." A wrong fee rule in the requirements becomes a wrong design, wrong code, wrong tests that expect the wrong fee, and a wrong user guide. This book calls the effect defect amplification: each work product built on a defective one can multiply the defect.

Worked example: what release 2.0's reviews saved

Release 2.0's 200 defects were recorded with the work product each was in and the activity that found it; its three reviews found 74 of them. The program asks what fixing the defects cost, in units of one requirements-review fix, and what it would have cost if the reviews had not been held. Two assumptions are needed, and both are stated. Without reviews, each work product's review finds are taken to have been found later, spread in proportion to where that product's other defects were found. And the cost of a fix at each activity rises from 1 at the requirements review to Boehm and Basili's end ratio after release, evenly on a logarithmic scale between; they give only the end ratio, so the program shows both of theirs, 100:1 and 5:1. Finally, it counts the requirements and design defects that reached the code, the ones that could amplify.

# release 2.0's 200 defects: where each was made, and the activity that found it (FINDINGS 5.2.6)
STAGES = ["req review", "design review", "code review", "unit", "integration", "system",
          "acceptance", "after release"]
found = {"requirements": [16, 3, 1, 1, 1, 2, 3, 1],
         "design":       [0, 21, 5, 3, 4, 4, 1, 2],
         "code":         [0, 0, 28, 40, 26, 20, 3, 3],
         "documents":    [0, 0, 0, 0, 1, 2, 3, 6]}
REVIEWS = 3                                     # the first three activities are reviews

def without_reviews(row):
    """Move a row's review finds to the later activities, in proportion to its own later finds."""
    moved, later = sum(row[:REVIEWS]), row[REVIEWS:]
    return [0] * REVIEWS + [n + moved * n / sum(later) for n in later]

def cost(rows, factor):
    return sum(n * f for row in rows for n, f in zip(row, factor))

for ratio in (100, 5):                          # the cost of a fix after release : in requirements
    factor = [ratio ** (k / (len(STAGES) - 1)) for k in range(len(STAGES))]   # even on a log scale
    with_r = cost(found.values(), factor)
    without_r = cost((without_reviews(row) for row in found.values()), factor)
    print(f"cost ratio {ratio}:1 -> fixes cost {with_r:,.0f} units with reviews,"
          f" {without_r:,.0f} without ({without_r / with_r:.2f} times)")

# amplification: requirements and design defects that reach the code, with and without reviews
reach_with = (sum(found["requirements"]) - sum(found["requirements"][:2])
              + sum(found["design"]) - found["design"][1])
reach_without = sum(found["requirements"]) + sum(found["design"])
print(f"requirements and design defects reaching the code: {reach_with} with reviews,"
      f" {reach_without} without")
munotes.in558

Why Reviews Pay: The Cost of a Late Defect

cost ratio 100:1 -> fixes cost 3,419 units with reviews, 5,365 without (1.57 times)
cost ratio 5:1 -> fixes cost 456 units with reviews, 576 without (1.26 times)
requirements and design defects reaching the code: 28 with reviews, 68 without

The cost. At Boehm and Basili's 100:1, fixing release 2.0's defects cost 3,419 units with its reviews and would have cost 5,365 without them, 1.57 times as much. Even at the small-system ratio of 5:1, the reviews save a fifth of the fixing cost: 456 units against 576. The saving is large because the reviews found 74 defects at the cheapest end of the scale.

The amplification. With the reviews, 28 requirements and design defects reached the code; without them, all 68 would have. Each of the 40 extra ones would have been built into code, tests and documents, which is the ISTQB statement in numbers: the program's cost figures count each defect once, so they understate what skipping the reviews would have cost.

What the comparison leaves out. The reviews themselves cost effort: 62 person-hours in release 2.0, against 296 for testing, and Chapter Sixty-Eight, on quality, process and test metrics, found them finding 1.19 defects per hour against testing's 0.39. Chapter One Hundred Two, on using quality costs for decision making, puts hours and money on both sides of the decision.

Review metrics

A project that holds reviews should measure them, for the same reason it measures testing. Three measures recur in this book.

MeasureHow it is computedRelease 2.0
Review effectivenessDefects the review found divided by the defects present when it ran (Fagan's error detection efficiency)Requirements review 57.1, design review 46.2, code review 21.2 per cent (Chapter Seventy-Nine, on defect metrics)
Review yield per hourDefects found divided by the person-hours spent1.19 defects per hour across the reviews, against 0.39 for testing (Chapter Sixty-Eight)
Where defects escapeDefects of a work product found after its own review28 requirements and design defects reached the code

Boehm and Basili's range gives a benchmark: "Numerous studies confirm that peer review provides an effective technique that catches from 31 to 93 percent of the defects, with a median of around 60 percent." ExamReg's requirements review, at 57.1 per cent, is near that median; its code review, at 21.2 per cent, is below the whole range. That is the measure doing its job: it points to the code review as the review to strengthen, with the formal procedure of Chapter Ninety-Seven, on formal technical reviews, and the reading techniques Boehm and Basili also report, whose perspective-based form catches "35 percent more defects than nondirected reviews."

munotes.in559

Why Reviews Pay: The Cost of a Late Defect

What it does not mean

The 100:1 ratio is not a constant. Boehm and Basili say "often", report about 5:1 for small, noncritical systems, and note that good architecture reduces the ratio; the principle holds at either ratio.

Reviews do not replace testing. They find different defects: Boehm and Basili report "that peer reviews, analysis tools, and testing catch different classes of defects at different points in the development cycle."

Not all rework is avoidable. Boehm and Basili distinguish avoidable rework from changes that emerge from prototyping and learning, which "should not be discouraged by classifying them as avoidable defects."

A cheap review is not a free one. Reviews cost effort; the case for them is that they cost less than the late fixes they prevent.

Quick revision

  • Late defects (Boehm and Basili 2001): after delivery "often 100 times more expensive" than in requirements and design; "more like 5:1" for small, noncritical systems; Fagan: early rework "10 to 100 times less expensive".
  • Avoidable rework: about 40 to 50 per cent of effort; about 80 per cent of it from 20 per cent of the defects.
  • Defect amplification (this book's name for the ISTQB statement): undetected defects in earlier work products "often lead to defective work products later"; early removal prevents "subsequent defects in derived work products".
  • Peer reviews catch 31 to 93 per cent of defects, median about 60; perspective-based reviews 35 per cent more than nondirected ones.
  • Review metrics: effectiveness (found over present), yield per hour, escapes.
  • Worked example (cost factors spread evenly on a log scale, an assumption): fixing cost 3,419 units with reviews against 5,365 without at 100:1 (1.57 times), 456 against 576 at 5:1; 28 requirements and design defects reached the code with reviews, 68 without.

Test yourself

1. Why does a defect cost more to fix the later it is found? Because by then other work has been built on it: a requirements defect found after release must be corrected in the requirement, the design, the code, the tests and the documents, everything affected retested and redeployed, and the users it affected dealt with. Boehm and Basili report that fixing after delivery is often 100 times as expensive as in the requirements and design phase, about 5 times for small, noncritical systems.

2. What is defect amplification? The effect that a defect left undetected in an early work product leads to defects in the work products derived from it, so that one requirements or design defect becomes several in design, code, tests and documents. The ISTQB syllabus states it, and its principle that early testing saves time and money rests on it.

munotes.in560

Why Reviews Pay: The Cost of a Late Defect

3. How do reviews reduce the cost of defects? They find defects in requirements, designs and code before execution, at the cheapest point to fix them, and they stop those defects propagating into later work products. Peer reviews catch from 31 to 93 per cent of the defects present, with a median of around 60 per cent.

4. In the worked example, what did release 2.0's reviews save? At a 100:1 cost ratio, fixing its defects cost 3,419 units with the reviews against an estimated 5,365 without, a factor of 1.57; at 5:1, 456 against 576. And 28 requirements and design defects reached the code instead of 68, so 40 fewer defects could amplify.

5. Name three review metrics and what each shows. Review effectiveness, defects found over defects present, shows how thoroughly a review works; yield per hour, defects found per person-hour, shows its efficiency; escapes, a work product's defects found after its review, show what it missed.

6. What did the review metrics show about ExamReg's code review, and what follows? Its effectiveness, 21.2 per cent, was below the 31 to 93 per cent range Boehm and Basili report for peer reviews, while the requirements review, at 57.1 per cent, was near the median. The code review should be strengthened, for example with a formal technical review procedure and directed reading techniques.

Contents This chapter on its own page

munotes.in561

Chapter Ninety-Seven

Formal Technical Reviews

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: Formal Technical Reviews and their benefits"

In one line

A formal technical review is a planned, structured examination of a work product by technically qualified peers, run by a defined process with defined roles and written records: it finds defects and other issues, records them in an issues list and a review report, decides whether the product can go on, and keeps data that let the organisation improve the reviews themselves.

In the wording a student can write in an examination: a technical review is a "formal peer review of a work product by a team of technically qualified personnel that examines the suitability of the work product for its intended use and identifies discrepancies from specifications and standards" (ISO/IEC 20246:2017), and a formal review one that "follows a defined process with formal documented output" (the same standard). Its objective, in CMMI's words, is "to identify defects for removal and to recommend other changes that are needed". Its outputs are an issues list (an issue is an "observation that deviates from expectations") and a review report with the decision and the data. CMMI's guidelines: "there should be sufficient preparation, the conduct should be managed and controlled, consistent and sufficient data should be recorded (an example is conducting a formal inspection), and action items should be recorded"; and "The focus of the peer review should be on the work product in review, not on the person who produced it."

Reviews seen from quality assurance

Chapters Twenty-Eight to Thirty, on reviews, inspection and walkthrough, taught reviews as a testing technique: the process, the roles and the four review types of the ISTQB syllabus, and Fagan's inspection. This chapter looks at the same activity from the side of quality assurance. There the question is not only what did this review find? but are the project's reviews planned, held as the plan says, recorded, and improving? CMMI places peer reviews in its verification process area as "an important and effective verification method implemented via inspections, structured walkthroughs, or a number of other collegial review methods." The SQA plan schedules them (Chapter Eighty-Seven, on the quality assurance plan), and the SQA group audits that they happen as planned (Chapter Eighty-Six, on SQA activities).

The "formal" in formal technical review is exactly ISO/IEC 20246's meaning: a defined process with documented output. The "technical" is the reviewers: peers "qualified to do the same work", which is the standard's definition of a peer review, examining a product for its intended use and against its specifications and standards. CMMI adds a distinction that matters: "These reviews are structured and are not management reviews." A formal technical review examines a product, not a project's progress or a person's performance.

Preparing the review

CMMI's practice "Prepare for peer reviews of selected work products" lists what is decided before anyone meets. In order:

munotes.in562

Formal Technical Reviews

  1. The type of review: an inspection, a structured walkthrough or another method, chosen for the product and the risk.
  2. The data to be collected during the review, decided in advance so that every review records the same things.
  3. Entry and exit criteria: when the product is ready to be reviewed, and when the review is finished.
  4. Criteria for requiring another review, such as the amount of rework.
  5. Checklists "to ensure that work products are reviewed consistently", covering items such as "Rules of construction", "Design guidelines", "Completeness", "Correctness", "Maintainability" and "Common defect types".
  6. A schedule, including when materials will be available.
  7. Checking the entry criteria before the product is distributed.
  8. Distributing the product early, "early enough to enable them to adequately prepare for the peer review."
  9. Assigning roles. CMMI's examples are "Leader", "Reader", "Recorder" and "Author".
  10. Individual preparation: each reviewer reviews the product before the meeting.

The roles have the duties Chapter Twenty-Nine, on inspection, gave Fagan's moderator, reader, recorder and author: the leader plans and runs the review, the reader leads the team through the product, the recorder writes down every issue, and the author answers questions and later fixes the product.

The review meeting

CMMI's practice "Conduct peer reviews of selected work products and identify issues resulting from these reviews" is the meeting. Its subpractices are: perform the assigned roles; "Identify and document defects and other issues in the work product"; "Record results of the peer review, including action items"; collect the review data; communicate issues to the people concerned; hold an additional review if needed; and "Ensure that the exit criteria for the peer review are satisfied."

Three rules keep the meeting productive. It finds and records issues; it does not solve them, which is the author's work afterwards (Fagan, in Chapter Twenty-Nine, on inspection). It is kept short, since Fagan found detection falls off after two hours. And its subject is the product: "The focus of the peer review should be on the work product in review, not on the person who produced it." When issues arise, "they should be communicated to the primary developer of the work product for correction."

The issues list and the review report

A formal review is defined by its documented output. Two documents come out of it.

  • The issues list records each issue: an identifier, where in the product it is, a description, its severity (for example major, minor, or a question to be answered), and, after the meeting, its resolution. ISO/IEC 20246's definition is deliberately broad: an issue is any "observation that deviates from expectations", so a question the author must answer is an issue as much as a defect is.
  • The review report records the review as a whole: the product and its size, the team and their roles, the preparation and meeting times, the number and kinds of issues, the action items, and the decision against the exit criteria (accept, accept once the issues are fixed and checked, or review again).
munotes.in563

Formal Technical Reviews

The report's data are CMMI's third practice: "Analyze data about the preparation, conduct, and results of the peer reviews." CMMI lists typical data: "product name, product size, composition of the peer review team, type of peer review, preparation time per reviewer, length of the review meeting, number of defects found, type and origin of defect". And it warns how they must not be used: "Examples of the inappropriate use of peer review data include using data to evaluate the performance of people and using data for attribution."

Worked example: analysing a review of the hall ticket module

In release 2.1 the hall ticket module was restructured, as the waiver on NC-04 required (Chapter Eighty-Six, on SQA activities), and given a formal technical review. Four people took part: a leader, a reader and a recorder, all reviewers, and the author. The program analyses the review's data the way CMMI asks: preparation against the expected rate, the meeting against its limit, the issues by severity, and the rework against the criterion for reviewing again. The expected rate and the limits are Fagan's, as Chapter Twenty-Nine, on inspection, gave them; treating a rate more than twice the expected one as not prepared is this book's rule.

# release 2.1's formal technical review of the hall ticket module (FINDINGS 5.8)
lines = 420
preparation = {"leader": 3.2, "reader": 3.5, "recorder": 1.0}          # hours, each alone
meeting_hours = 2.5
issues = [("I-1", "major"), ("I-2", "major"), ("I-3", "major"), ("I-4", "minor"), ("I-5", "minor"),
          ("I-6", "minor"), ("I-7", "minor"), ("I-8", "minor"), ("I-9", "question")]
reworked = 38

# the guidelines used to analyse the review (Fagan, as Chapter 29 quoted him)
EXPECTED_PREP = 125          # lines per hour for code preparation
MAX_SESSION = 2              # hours: no session longer than two hours
REINSPECT_OVER = 0.05        # reinspect if more than 5 per cent of the material was reworked

for role, hours in preparation.items():
    rate = lines / hours
    note = "   <- over twice the expected rate" if rate > 2 * EXPECTED_PREP else ""
    print(f"{role:<9} prepared {hours:.1f} h: {rate:>4.0f} lines per hour{note}")
verdict = "over the limit" if meeting_hours > MAX_SESSION else "within it"
print(f"meeting {meeting_hours} h: {verdict}")
for severity in ("major", "minor", "question"):
    print(f"{severity:<9}{sum(s == severity for _, s in issues):>2}")
share = reworked / lines
print(f"reworked {reworked} of {lines} lines, {share:.1%}:",
      "another review required" if share > REINSPECT_OVER else "moderator follow-up only")
munotes.in564

Formal Technical Reviews

leader    prepared 3.2 h:  131 lines per hour
reader    prepared 3.5 h:  120 lines per hour
recorder  prepared 1.0 h:  420 lines per hour   <- over twice the expected rate
meeting 2.5 h: over the limit
major     3
minor     5
question  1
reworked 38 of 420 lines, 9.0%: another review required

What the review found. Nine issues: three major (the hall ticket shown before the fee is confirmed, a status code the fee module never sends, an expired session treated as a paid one), five minor, and one question about the exam cell's rule for detained students, which goes to the exam cell as an action item.

What the data say about the review itself. The leader and reader prepared at 131 and 120 lines an hour, close to the expected 125. The recorder spent one hour on 420 lines, 420 an hour: at that speed the product was skimmed, not studied, and the recorder's contribution to finding issues was probably small. The meeting ran 2.5 hours, beyond the two hours after which Fagan found detection falls off. Neither observation is about blaming the recorder or the leader; both are process data, and the SQA group raises them as improvements to how reviews are scheduled (preparation time booked in advance, the module split into two sessions).

The decision. The rework changed 38 of the 420 lines, 9.0 per cent, above the 5 per cent at which Fagan's rule calls for the material to be reviewed again. The report's decision is therefore a second review after rework, not a follow-up check by the leader alone.

Guidelines for formal technical reviews

GuidelineSource
Review the product, not the producerCMMI: "The focus of the peer review should be on the work product in review"
Prepare sufficiently, with the product distributed in timeCMMI's first guideline; SP 2.1
Manage and control the conduct: roles, an agenda, a time limit of about two hoursCMMI; Fagan
Find and record issues; do not solve them in the meetingFagan (Chapter Twenty-Nine, on inspection)
Use checklists, and keep them up to date from defect dataCMMI SP 2.1
Record consistent data and every action itemCMMI's third and fourth guidelines
Set entry and exit criteria, and a criterion for reviewing againCMMI SP 2.1; Fagan's 5 per cent rule
Never use review data to judge peopleCMMI SP 2.3

What it does not mean

A formal technical review is not a management review. CMMI says so directly: peer reviews "are structured and are not management reviews". Managers learn the outcome from the report.

Formal does not mean long. It means a defined process with documented output; a well-run technical review of a small product can take an hour.

munotes.in565

Formal Technical Reviews

The review does not fix the product. It finds and records issues; the author fixes them afterwards, and the leader or a second review checks the fixes.

Review data are not performance data. Using them to rate the author or the reviewers is, in CMMI's words, an inappropriate use, and it would stop people reporting honestly.

Quick revision

  • Technical review (ISO/IEC 20246:2017): a formal peer review by technically qualified people of a product's suitability for its intended use and its discrepancies from specifications and standards. Formal review: a defined process with formal documented output.
  • Objective (CMMI): identify defects for removal and recommend other changes needed.
  • Prepare (CMMI SP 2.1): type; data to collect; entry and exit criteria; criteria for another review; checklists; schedule; entry check; early distribution; roles (leader, reader, recorder, author); individual preparation.
  • Conduct (SP 2.2): roles performed; issues identified and documented; results and action items recorded; data collected; issues communicated; another review if needed; exit criteria met.
  • Outputs: the issues list and the review report. Analyse (SP 2.3): preparation, conduct and results data; never to evaluate people.
  • Guidelines (CMMI): sufficient preparation; managed and controlled conduct; consistent and sufficient data; action items recorded; focus on the product, not the person.
  • Worked example: 9 issues (3 major); the recorder prepared at 420 lines an hour against 125 expected; a 2.5-hour meeting; 9.0 per cent reworked, so a second review.

Test yourself

1. What is a formal technical review, and what are its objectives? A formal peer review of a work product by technically qualified people, following a defined process with documented output, that examines the product's suitability for its intended use and identifies discrepancies from its specifications and standards. Its objectives are to identify defects for removal and to recommend other changes that are needed.

2. Describe how a formal technical review is prepared. The type of review is chosen; the data to be collected, entry and exit criteria and criteria for another review are set; checklists are prepared; the review is scheduled; the product is checked against the entry criteria and distributed early; roles (leader, reader, recorder, author) are assigned; and each reviewer studies the product before the meeting.

3. What happens in the review meeting, and what does it produce? The participants perform their roles, the reader leads through the product, and the issues found are identified and recorded, not solved. The review produces an issues list, with each issue's location, description and severity, and a review report recording the product, the team, the times, the issues, the action items and the decision against the exit criteria.

4. State the guidelines for conducting formal technical reviews. Review the product, not the producer; prepare sufficiently, with the product distributed in time; manage and control the conduct, with roles and a time limit; find issues rather than solving them; use checklists; record consistent data and every action item; set entry, exit and re-review criteria; and never use review data to judge people.

munotes.in566

Formal Technical Reviews

5. What data should a review record, and how may they not be used? The product and its size, the team, the type of review, each reviewer's preparation time, the length of the meeting, and the number, types and origins of the defects found. They must not be used to evaluate people's performance or to attribute blame.

6. In the worked example, why was a second review required, and what else did the data show? Because the rework changed 9.0 per cent of the module, above the 5 per cent at which the material is reviewed again. The data also showed that one reviewer prepared at 420 lines an hour, more than three times the expected rate, and that the meeting ran beyond two hours: both are process improvements for how reviews are scheduled, not faults of the people.

Contents This chapter on its own page

munotes.in567

Chapter Ninety-Eight

The Benefits of Formal Technical Reviews

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: Formal Technical Reviews and their benefits"

In one line

Formal technical reviews pay in three currencies: they find defects early, when fixes are cheapest; they find defects that testing cannot, in documents and in code that never runs; and they leave people who understand the product, a team that learns from its own errors, and data that improve the process.

In the wording a student can write in an examination: the benefits of formal technical reviews are, first, early detection: static testing "can detect defects in the earliest phases of the SDLC" (ISTQB), where Fagan found rework "10 to 100 times less expensive"; second, defects dynamic testing cannot detect, "e.g., unreachable code, design patterns not implemented as desired, defects in non-executable work products"; third, lower overall cost: "the overall project costs are usually much lower than when no reviews are performed because less time and effort needs to be spent on fixing defects later in the project"; and fourth, benefits beyond defects: confidence in work products, "a shared understanding" among stakeholders, better communication, and feedback that helps each author make fewer errors, which Fagan called "One of the most significant benefits of inspections".

The evidence

Fagan's results. Chapter Twenty-Nine, on inspection, set out the measurements Fagan reported. In IBM systems programming, inspections produced "a 23 percent increase in the productivity of the coding operation alone", and the inspected code had "38 percent less errors" in testing than comparable code checked by walk-throughs. At Aetna Life and Casualty, inspections were the only change to two programmers' process, and "The resulting saving in programmer resources was 25 percent." Fagan himself warned that the results "cannot be considered representative of every situation".

Across many studies. Boehm and Basili (2001) summarise the wider evidence: peer review "catches from 31 to 93 percent of the defects, with a median of around 60 percent" (Chapter Ninety-Six, on why reviews pay, set ExamReg's reviews against that range).

What the ISTQB syllabus concludes. "Even though reviews can be costly to implement, the overall project costs are usually much lower than when no reviews are performed because less time and effort needs to be spent on fixing defects later in the project."

Worked example: what release 2.0's reviews saved in finding effort

Chapter Ninety-Six measured the saving in fixing: defects found early are cheaper to fix. There is a second saving, in finding. The program compares the reviews' and testing's rates of finding defects in release 2.0, then asks how long testing would have needed to find the reviews' 74 defects at its own average rate. That assumption is generous to testing: some of those defects, in requirements and designs, are the kind testing finds late or not at all.

# release 2.0: person-hours and defects found, by activity (FINDINGS 5.2 and 5.2.2)
reviews = {"requirements review": (12, 16), "design review": (20, 24), "code review": (30, 34)}
testing = {"unit": (80, 44), "integration": (64, 32), "system": (112, 28), "acceptance": (40, 10)}

def totals(activities):
    return sum(h for h, _ in activities.values()), sum(d for _, d in activities.values())

review_hours, review_found = totals(reviews)
test_hours, test_found = totals(testing)
review_rate, test_rate = review_found / review_hours, test_found / test_hours
print(f"reviews: {review_found} defects in {review_hours} hours, {review_rate:.2f} per hour")
print(f"testing: {test_found} defects in {test_hours} hours, {test_rate:.2f} per hour")

instead = review_found / test_rate                   # if testing had to find them, at its own rate
print(f"finding the reviews' {review_found} defects by testing instead: about {instead:.0f} hours,"
      f" {instead - review_hours:.0f} more than the reviews took")
print(f"share of all 200 defects found by reviews: {review_found / 200:.0%}")
munotes.in568

The Benefits of Formal Technical Reviews

reviews: 74 defects in 62 hours, 1.19 per hour
testing: 114 defects in 296 hours, 0.39 per hour
finding the reviews' 74 defects by testing instead: about 192 hours, 130 more than the reviews took
share of all 200 defects found by reviews: 37%

The reviews found 74 defects, 37 per cent of all release 2.0's defects, in 62 person-hours: 1.19 defects an hour, three times testing's 0.39. Left to testing, the same defects would have taken about 192 hours to find, 130 hours more than the reviews cost, and that is before counting the cheaper fixes of Chapter Ninety-Six, which at Boehm and Basili's 100:1 made the fixing of release 2.0's defects 1.57 times dearer without the reviews. On this project, the reviews repaid their hours several times over.

Defects that testing cannot find

Some defects are invisible to execution. The ISTQB syllabus names examples: "unreachable code, design patterns not implemented as desired, defects in non-executable work products". A requirement that never says what happens to an invalid input cannot fail a test until someone writes the test, and nobody writes it, because the requirement does not ask for it; a review with a checklist item does every input say what happens to an invalid value? finds the gap directly (Chapter Eighty, on using defect data). Boehm and Basili make the general point: "peer reviews, analysis tools, and testing catch different classes of defects at different points in the development cycle." Reviews and testing are complements, not substitutes.

Benefits that are not defects

Shared understanding and communication. "Since static testing can be performed early in the SDLC, a shared understanding can be created among the involved stakeholders. Communication will also be improved between the involved stakeholders." ExamReg's review of the hall ticket module produced a question, I-9, about the exam cell's rule for detained students (Chapter Ninety-Seven, on formal technical reviews): not a defect, but a gap in the team's understanding of a rule, closed before it could become one.

munotes.in569

The Benefits of Formal Technical Reviews

Confidence in work products. Static testing "provides the ability to evaluate the quality of, and to build confidence in work products", and "By verifying the documented requirements, the stakeholders can also make sure that these requirements describe their actual needs."

Learning. Fagan counted feedback among "the most significant benefits of inspections": "The programmer finds out what error types he is most prone to make and their quantity and how to find them. This feedback takes place within a few days of writing the program." His comparison with walk-throughs lists, for inspections but not walk-throughs, fewer future errors from detailed feedback to each programmer.

Process improvement. The same comparison credits inspections with improving "inspection efficiency from analysis of results" and with the analysis of data to find process problems and improvements. CMMI builds the analysis into the practice (Chapter Ninety-Seven, on formal technical reviews), and its data feed the causal analysis of Chapter Eighty.

Less rework downstream. Every requirements or design defect removed in review is one that cannot be built into code, tests and documents: in release 2.0, 28 such defects reached the code with reviews, against 68 without (Chapter Ninety-Six).

The costs, honestly

Reviews are not free, and their benefits depend on doing them well. They take people's time: release 2.0's cost 62 person-hours. They depend on preparation: a reviewer who skims, like the recorder at 420 lines an hour in the example of Chapter Ninety-Seven, on formal technical reviews, adds little. And their data are useful only if they are never turned against the people reviewed, which Fagan and CMMI both insist on. The benefits in this chapter are the result of reviews held under the guidelines of Chapter Ninety-Seven; reviews held carelessly give their costs without their benefits.

What it does not mean

Reviews do not replace testing. They find different classes of defects; release 2.0 needed both.

Fagan's percentages are not guarantees. He reported them for particular projects and warned against treating them as representative; each organisation should measure its own.

The benefit is not only the defect count. Understanding, communication, learning and process data are benefits too, even when a review finds few defects.

A review that finds nothing did not necessarily waste its time. It may confirm a good product, but its preparation rate and checklist should be checked before concluding that.

Quick revision

  • Early detection (ISTQB): defects found in the earliest phases; fixes "10 to 100 times less expensive" early (Fagan).
  • Defects testing cannot find (ISTQB): unreachable code, design patterns not implemented as desired, defects in non-executable work products.
  • Lower overall cost (ISTQB): project costs "usually much lower than when no reviews are performed".
  • Evidence: Fagan (23 per cent coding productivity, 38 per cent fewer errors than walk-throughs, Aetna 25 per cent saving); Boehm and Basili (31 to 93 per cent of defects, median about 60).
  • Other benefits: confidence in work products; shared understanding and communication; feedback and learning; process data for improvement; less downstream rework.
  • Worked example: reviews 74 defects in 62 hours (1.19 per hour) against testing 0.39 per hour; testing would need about 192 hours to find them, 130 more; reviews found 37 per cent of all defects.
munotes.in570

The Benefits of Formal Technical Reviews

Test yourself

1. List the benefits of formal technical reviews. Early detection of defects, when fixes are cheapest; detection of defects testing cannot find, in non-executable work products and code that never runs; lower overall project cost; confidence in work products; shared understanding and better communication among stakeholders; feedback that helps authors make fewer errors; data for improving the process; and fewer defects propagating into later work products.

2. What evidence did Fagan report for inspections? At IBM, a 23 per cent increase in coding productivity and 38 per cent fewer errors in testing than comparable code checked by walk-throughs; at Aetna, a 25 per cent saving in programmer resources. He warned that the results are not representative of every situation.

3. Give examples of defects a review can find that testing cannot. Unreachable code, which no test executes; design patterns not implemented as intended; and defects in non-executable work products such as requirements and designs, for example a requirement that never says what happens to an invalid input.

4. In the worked example, what did release 2.0's reviews save in finding effort? They found 74 defects in 62 person-hours, 1.19 an hour against testing's 0.39. At testing's rate, finding the same defects would have taken about 192 hours, 130 more than the reviews took, before counting the lower cost of fixing defects found early.

5. Describe two benefits of reviews that are not defects found. Shared understanding: a review brings stakeholders to the same understanding of a product early, as ExamReg's review did when a question about detained students exposed a gap in the team's knowledge of a rule. Learning: feedback on the kinds of error each author makes helps them make fewer, which Fagan counted among the most significant benefits.

6. Why do reviews not always deliver these benefits? Because the benefits depend on how reviews are held: reviewers must prepare at a sensible rate, meetings must be managed, data must be recorded, and results must never be used to judge people. Reviews held carelessly incur their costs without their benefits.

Contents This chapter on its own page

munotes.in571

Chapter Ninety-Nine

Quality Improvement Methodologies: PDCA and Kaizen

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Introduction to quality, improvement methodologies"

In one line

PDCA, plan-do-check-act, is a four-step cycle for improving a process through small, tested changes; Deming preferred to call it plan-do-study-act, to stress comparing the results with what was predicted; and kaizen is the Japanese practice of gradual, unending improvement by everyone, one small change after another, which the cycle turns into a method.

In the wording a student can write in an examination: the plan-do-check-act (PDCA) cycle is "A four-step process for quality improvement. In the first step (plan), a way to effect improvement is developed. In the second step (do), the plan is carried out. In the third step (check), a study takes place between what was predicted and what was observed in the previous step. In the last step (act), action should be taken to correct or improve the process" (ASQ). The Deming cycle is "Another term for the plan-do-study-act cycle. Walter Shewhart created it (calling it the plan-do-check-act cycle), but W. Edwards Deming popularized it, calling it plan-do-study-act" (ASQ). Kaizen is "A Japanese term that means gradual, unending improvement by doing little things better and setting and achieving increasingly higher standards" (ASQ).

Continuous improvement

ASQ defines continuous improvement as "the ongoing improvement of products, services or processes through incremental and breakthrough improvements." Some practitioners distinguish two terms. Continual improvement is "A broader term preferred by W. Edwards Deming to refer to general processes of improvement and encompassing 'discontinuous' improvements", and continuous improvement "A subset of continual improvement, with a more specific focus on linear, incremental improvement within an existing process." Every method in this chapter and the next is a way of organising improvement so that it happens by design rather than by accident. ISO's fifth quality management principle, improvement, states the goal: "Successful organizations have an ongoing focus on improvement" (Chapter Ninety-Four, on the ISO 9000 family).

The PDCA cycle

ASQ calls PDCA "Among the most widely used tools for the continuous improvement model", and describes its steps in practical terms:

  1. Plan: "Identify an opportunity and plan for change."
  2. Do: "Implement the change on a small scale."
  3. Check: "Use data to analyze the results of the change and determine whether it made a difference."
  4. Act: "If the change was successful, implement it on a wider scale and continuously assess your results. If the change did not work, begin the cycle again."

The cycle repeats: each Act is the starting point of the next Plan. Two features make it work. The change is tried small first, so a bad idea costs little. And the result is judged by data, so the organisation learns something whether the change works or not.

PDSA: study, not check

The Deming Institute describes the cycle Deming taught as PDSA, "a systematic process for gaining valuable learning and knowledge for the continual improvement of a product, process, or service" (Chapter Eighty-Two, on Shewhart, Deming and Juran, introduced it). "Dr. Deming emphasized the PDSA Cycle, not the PDCA Cycle, with a third step emphasis on Study (S), not Check (C)." The difference is not a word. In the Institute's account, Check is about whether a change succeeded or failed; Deming's focus "was on predicting the results of an improvement effort, studying the actual results, and comparing them to possibly revise the theory." A cycle run as PDSA therefore writes down its predictions in the Plan step, so that the Study step has something to compare the results with.

munotes.in572

Quality Improvement Methodologies: PDCA and Kaizen

ASQ's glossary gives the history in one line: Shewhart created the cycle, calling it plan-do-check-act, and Deming popularised it as plan-do-study-act. The two names are used interchangeably in practice; the examination answer should know both and the reason for Deming's word.

Kaizen

ASQ's glossary defines kaizen as gradual, unending improvement "by doing little things better and setting and achieving increasingly higher standards", and notes that "Masaaki Imai made the term famous in his book, Kaizen: The Key to Japan's Competitive Success." A process kaizen is one made "at an individual process or in a specific area", also called a point kaizen; a kaizen event, in ASQ's list of improvement tools, introduces "rapid change by focusing on a narrow project and using the ideas and motivation of the people who do the work."

Toyota, whose production system made kaizen widely known, describes it as the work of people, not machines: "Only humans can implement kaizen for the sake of evolution", and "all of Toyota is implementing kaizen to TPS day and night to ensure its continued evolution." Its account of jidoka, "automation with a human touch", shows kaizen and quality working together: work is first done well by hand while people "implement kaizen" and eliminate waste, inconsistency and unreasonable requirements (muda, mura and muri), and machines then detect abnormalities and stop, which "eliminates the outflow of defective products while also making it possible to build quality into processes".

Kaizen and PDCA together. Kaizen is the attitude: small improvements, every day, by the people doing the work. PDCA is the procedure each improvement follows: plan it, try it small, check it against the prediction, then adopt it or rethink it.

Kaizen in a software team. In this book's reading, the same ideas appear in modern software practice: the retrospective, whose purpose in the Scrum Guide is "to plan ways to increase quality and effectiveness" (Chapter Eighty-Three, on Feigenbaum, Ishikawa, Crosby and TQM), is a kaizen meeting held every sprint; and a build that stops when a test fails is jidoka for code, detecting the abnormality and stopping the line before the defect flows on (Chapter Thirty-Nine, on regression testing and continuous integration).

munotes.in573

Quality Improvement Methodologies: PDCA and Kaizen

Worked example: a PDSA cycle on ExamReg's code reviews

Chapter Ninety-Seven, on formal technical reviews, left two process problems: a reviewer who prepared at 420 lines an hour, and a meeting of 2.5 hours. The SQA group ran a PDSA cycle on them.

  • Plan. The change: preparation time booked in the schedule, one hour per 125 lines for each reviewer, and every session capped at two hours, longer reviews split. The predictions, written down in advance: every reviewer prepares within twice the expected rate; no session runs over two hours; and the reviews find at least as many major issues per 100 lines as the hall ticket review did.
  • Do. The change was tried on a small scale: the next three code reviews, of parts of the fee module.
  • Study. The program compares each prediction with the result.
# PLAN (FINDINGS 5.9): preparation booked in the schedule, sessions capped at two hours; predictions
EXPECTED_PREP, MAX_SESSION = 125, 2.0           # lines per hour; hours per session
before = {"lines": 420, "majors": 3}            # the hall ticket review of Chapter 97

# DO: the next three code reviews, (lines, preparation hours per reviewer, sessions in hours, majors)
reviews = {"A": (300, [2.5, 2.4, 2.2], [1.8], 3),
           "B": (360, [3.0, 2.8, 1.2], [1.5, 1.5], 3),
           "C": (280, [2.3, 2.2, 2.4], [1.6], 2)}

# STUDY: compare each prediction with what happened
rates = [lines / h for lines, prep, _, _ in reviews.values() for h in prep]
slow_enough = sum(r <= 2 * EXPECTED_PREP for r in rates)
print(f"prediction 1, every reviewer within twice the expected rate:"
      f" {slow_enough} of {len(rates)} -> {'met' if slow_enough == len(rates) else 'not met'}")
longest = max(s for _, _, sessions, _ in reviews.values() for s in sessions)
print(f"prediction 2, no session over {MAX_SESSION:.0f} hours: longest {longest} h ->"
      f" {'met' if longest <= MAX_SESSION else 'not met'}")
old = 100 * before["majors"] / before["lines"]
new = 100 * sum(r[3] for r in reviews.values()) / sum(r[0] for r in reviews.values())
print(f"prediction 3, major issues per 100 lines at least {old:.2f}: {new:.2f} ->"
      f" {'met' if new >= old else 'not met'}")
prediction 1, every reviewer within twice the expected rate: 8 of 9 -> not met
prediction 2, no session over 2 hours: longest 1.8 h -> met
prediction 3, major issues per 100 lines at least 0.71: 0.85 -> met

Study. Two predictions were met: no session ran over two hours, the longest being 1.8, and the reviews found 0.85 major issues per 100 lines against 0.71 before. The first was not: 8 of the 9 reviewers prepared within twice the expected rate, but one, in review B, prepared 360 lines in 1.2 hours. Studying why, rather than only checking pass or fail, found the cause: that reviewer had been pulled onto an urgent fix the day before, so the booked preparation time existed on the schedule but not in practice. The theory needs revising: booking time is necessary but not sufficient.

munotes.in574

Quality Improvement Methodologies: PDCA and Kaizen

Act. The session cap is adopted as standard and written into the review procedure the SQA plan audits (Chapter Eighty-Seven, on the quality assurance plan). The preparation rule goes round the cycle again with a revised plan: a new entry criterion that a review is postponed if any reviewer has not prepared, with the prediction that no review is held with an unprepared reviewer. That is kaizen in practice: two small improvements, each tested, one kept and one refined.

What it does not mean

PDCA is not a one-time project. The cycle repeats; each Act feeds the next Plan.

Check is not a formality. Without predictions written down in the Plan, there is nothing to compare, which is why Deming insisted on Study.

Kaizen is not only small changes. It is continual small improvement by everyone; breakthrough improvements (Juran's, Chapter Eighty-Two) and kaizen events for narrow projects sit beside it.

A failed prediction is not a failed cycle. Review B's unprepared reviewer taught the team more than a success would have: the plan's theory was incomplete.

Quick revision

  • PDCA (ASQ): plan a way to improve; do it (on a small scale); check, comparing what was predicted with what was observed; act to correct or improve the process; repeat.
  • PDSA (Deming Institute): Study rather than Check, with the emphasis on predicting results, studying the actual ones and revising the theory; Shewhart created the cycle, Deming popularised it (ASQ).
  • Continuous improvement (ASQ): ongoing improvement through incremental and breakthrough changes; continual improvement is Deming's broader term.
  • Kaizen (ASQ): gradual, unending improvement by doing little things better; made famous by Masaaki Imai; process (point) kaizen; kaizen event. At Toyota, "Only humans can implement kaizen"; jidoka builds quality into processes.
  • Worked example: session cap met (longest 1.8 h); major issues 0.85 per 100 lines against 0.71, met; preparation 8 of 9 reviewers, not met (a reviewer pulled onto an urgent fix); act: adopt the cap, re-plan preparation with an entry criterion.

Test yourself

1. Explain the PDCA cycle. A four-step cycle for improving a process. Plan: identify an opportunity and develop a way to improve, with predictions. Do: carry out the change, on a small scale first. Check: study what happened, comparing the observed results with the predicted ones. Act: if the change worked, adopt it more widely and keep assessing; if not, correct the plan and begin the cycle again.

munotes.in575

Quality Improvement Methodologies: PDCA and Kaizen

2. What is the difference between PDCA and PDSA? PDSA is Deming's name for the cycle, with Study in place of Check. Check tends to ask only whether a change succeeded or failed; Study compares the results with the predictions made in the Plan step and uses the comparison to revise the theory behind the change, so that each cycle adds knowledge.

3. What is kaizen? A Japanese term for gradual, unending improvement by doing little things better and setting and achieving increasingly higher standards, carried out continually by the people who do the work. It was made widely known by Masaaki Imai and by Toyota's production system, where it underlies jidoka, building quality into processes.

4. How are kaizen and PDCA related? Kaizen is the practice of continual small improvement; PDCA is the procedure each improvement follows, so that every small change is planned, tried on a small scale, checked against a prediction, and adopted or refined.

5. Give examples of kaizen in a software team. The sprint retrospective, whose purpose is to plan ways to increase quality and effectiveness; small changes to the review procedure or the coding standard, tried and measured; and a build that stops when a test fails, which, like jidoka, stops a defect from flowing on.

6. In the worked example, which prediction failed, and what did the team do? The prediction that every reviewer would prepare within twice the expected rate: one reviewer in review B prepared 360 lines in 1.2 hours after being pulled onto an urgent fix. The team kept the change that worked, the two-hour session cap, and ran the cycle again for preparation with a revised plan: a review is postponed if a reviewer has not prepared.

Contents This chapter on its own page

munotes.in576

Chapter One Hundred

Lean, CMMI and Choosing a Methodology

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Introduction to quality, improvement methodologies"

In one line

Lean improves a process by removing waste, work that adds no value for the customer; CMMI improves an organisation's processes level by level, from ad hoc to optimising, against a model of what capable organisations practise; and choosing among PDCA, Six Sigma, Lean and CMMI starts from the problem: a change to try, variation to reduce, waste to remove, or an organisation's whole way of working to build.

In the wording a student can write in an examination: lean is "a set of management practices to improve efficiency and effectiveness by eliminating waste" (ASQ), and waste (muda) "the performance of unnecessary work as a result of errors, poor organization, or communication"; ASQ's eight wastes spell DOWNTIME: defects, overproduction, waiting, non-utilized talent, transportation, inventory, motion and extra-processing. CMMI (Capability Maturity Model Integration) rates an organisation's processes on five maturity levels: 1 Initial, 2 Managed, 3 Defined, 4 Quantitatively Managed, 5 Optimizing (CMMI for Development v1.3); "Maturity levels are used to characterize organizational improvement relative to a set of process areas, and capability levels characterize organizational improvement relative to an individual process area." Improvement with CMMI can be planned with the IDEAL model: Initiating, Diagnosing, Establishing, Acting and Learning.

Lean

ASQ defines lean as a set of management practices whose "core principle" is "to reduce and eliminate non-value adding activities and waste." It grew out of manufacturing: Toyota describes the objective of its production system as "to thoroughly eliminate waste and shorten lead times to deliver vehicles to customers quickly, at a low cost, and with high quality", built on two pillars, jidoka and Just-in-Time (Chapter Ninety-Nine, on PDCA and kaizen, described jidoka). Lean enterprise extends the idea "through the entire value stream or supply chain", because, in ASQ's words, "The leanest factory cannot achieve its full potential if it has to work with non-lean suppliers and subcontractors."

The eight wastes. ASQ lists the wastes whose initials spell DOWNTIME. The right-hand column is this book's reading of each for a software project, with examples from ExamReg.

Waste (ASQ)In softwareOn ExamReg
DefectsRework to find and fix defects200 defects in release 2.0; fixes after release costing the most (Chapter Ninety-Six, on why reviews pay)
OverproductionWork produced before or beyond needA coverage report for every module on every build that nobody read (NC-08, Chapter Eighty-Six, on SQA activities)
WaitingWork stopped for someone or somethingThe 72-hour repair that waited for the payment gateway's provider (Chapter Ninety-One, on statistical process control)
Non-utilized talentPeople's knowledge not usedTesters not invited to requirements reviews
TransportationHanding work between teams and toolsA defect passed between teams before anyone owns it
InventoryWork started but not finishedUntested builds; defects open for weeks (Chapter Seventy-Seven, on tracking defects to closure)
MotionEffort spent searching and switchingA reviewer pulled onto an urgent fix mid-review (Chapter Ninety-Nine, on PDCA and kaizen)
Extra-processingWork beyond what is neededA duplicated check found in the hall ticket review (Chapter Ninety-Seven, on formal technical reviews)
munotes.in577

Lean, CMMI and Choosing a Methodology

Lean and Six Sigma. ASQ's comparison is short: "Lean focuses on waste reduction, whereas Six Sigma emphasizes variation reduction" (Chapter Eighty-Nine, on statistical SQA and Six Sigma). The two are often combined as lean Six Sigma.

CMMI

CMMI is a model of the practices of organisations that develop products well; this book has used its process areas throughout Module 2, from process and product quality assurance (Chapter Eighty-Six, on SQA activities) to causal analysis and resolution (Chapter Eighty, on using defect data). Its owners describe it as "Originally created for the U.S. Department of Defense to assess the quality and capability of their software contractors". The version quoted here, CMMI for Development v1.3 (SEI, 2010), has 22 process areas; later versions exist, and the CMMI Institute, now part of ISACA, says the model "is continuously updated", so the levels and process areas below are v1.3's.

Two representations. In the staged representation an organisation is rated on maturity levels, each defined by a set of process areas; in the continuous representation each process area is rated separately on capability levels. In v1.3's words: "Maturity levels are used to characterize organizational improvement relative to a set of process areas, and capability levels characterize organizational improvement relative to an individual process area."

The five maturity levels.

LevelNameCMMI v1.3's description
1Initial"processes are usually ad hoc and chaotic"
2Managed"the projects have ensured that processes are planned and executed in accordance with policy"
3Defined"processes are well characterized and understood, and are described in standards, procedures, tools, and methods"
4Quantitatively Managed"the organization and projects establish quantitative objectives for quality and process performance and use them as criteria in managing projects"
5Optimizing"an organization continually improves its processes based on a quantitative understanding of its business objectives and performance needs"

Level 1 organisations, CMMI notes, often produce products that work, but "frequently exceed the budget and schedule documented in their plans", and success "depends on the competence and heroics of the people in the organization and not on the use of proven processes." CMMI's advice on moving up is firm: "each maturity level forms a necessary foundation for the next level, trying to skip maturity levels is usually counterproductive."

IDEAL. CMMI suggests that planning an improvement programme can begin "with an improvement approach such as the" IDEAL model, whose name stands for Initiating, Diagnosing, Establishing, Acting and Learning. The five phases are a PDCA cycle at the scale of a whole organisation: start the programme, find where the organisation stands, plan the changes, make them, and learn from the result.

munotes.in578

Lean, CMMI and Choosing a Methodology

Worked example: ExamReg's software house on the staged scale

The earlier chapters of this book contain evidence of which CMMI v1.3 process areas ExamReg's software house practises: an SQA group that audits processes, a causal analysis loop, reviews, testing at every level, measurement. They also contain its gaps: no agreement with its payment gateway's provider, no organisation-wide standard processes or training records (Chapter Ninety-Five, on ISO 9001). The program records this book's reading of the evidence for each of the 22 process areas and computes the maturity level the staged representation would give. It is not an appraisal, which is done by trained appraisers with the SEI's appraisal method; it shows how the staged rule works.

# CMMI for Development v1.3's 22 process areas by maturity level, and this book's reading of the
# evidence for ExamReg's software house from earlier chapters (True: practised; False: not yet)
LEVELS = {
    2: {"Requirements Management": True, "Project Planning": True,
        "Project Monitoring and Control": True, "Supplier Agreement Management": False,
        "Measurement and Analysis": True, "Process and Product Quality Assurance": True,
        "Configuration Management": True},
    3: {"Requirements Development": True, "Technical Solution": True, "Product Integration": True,
        "Verification": True, "Validation": True, "Organizational Process Focus": False,
        "Organizational Process Definition": False, "Organizational Training": False,
        "Integrated Project Management": False, "Risk Management": False,
        "Decision Analysis and Resolution": False},
    4: {"Organizational Process Performance": False, "Quantitative Project Management": False},
    5: {"Causal Analysis and Resolution": True, "Organizational Performance Management": False},
}

def staged_level(levels):
    """The highest maturity level whose process areas, and all those below it, are practised."""
    reached = 1
    for level in sorted(levels):
        if not all(levels[level].values()):
            break
        reached = level
    return reached

for level, areas in LEVELS.items():
    missing = [name for name, met in areas.items() if not met]
    print(f"level {level}: {len(areas) - len(missing)} of {len(areas)} practised;"
          f" missing: {', '.join(missing) or 'none'}")
print("maturity level reached (staged):", staged_level(LEVELS))
LEVELS[2]["Supplier Agreement Management"] = True        # an agreement with the gateway's provider
print("with a supplier agreement for the payment gateway:", staged_level(LEVELS))
level 2: 6 of 7 practised; missing: Supplier Agreement Management
level 3: 5 of 11 practised; missing: Organizational Process Focus, Organizational Process Definition, Organizational Training, Integrated Project Management, Risk Management, Decision Analysis and Resolution
level 4: 0 of 2 practised; missing: Organizational Process Performance, Quantitative Project Management
level 5: 1 of 2 practised; missing: Organizational Performance Management
maturity level reached (staged): 1
with a supplier agreement for the payment gateway: 2

The staged answer. The software house practises 6 of the 7 level 2 process areas, 5 of the 11 at level 3, and even causal analysis and resolution at level 5. Yet under the staged representation it is at maturity level 1, because one level 2 process area, supplier agreement management, is missing. That process area applies here: CMMI's scope covers acquiring "products, services, and product and service components that can be delivered to the project's customer or included in a product or service system", and the payment gateway is part of ExamReg's service. One agreement with the gateway's provider would move the house to level 2.

munotes.in579

Lean, CMMI and Choosing a Methodology

The continuous view. The same evidence, read process area by process area, shows a more capable organisation than its level suggests: strong verification, validation and quality assurance, a working improvement loop, and weaknesses in organisation-wide process definition, training and risk management. That is why CMMI offers both representations. The staged rating says what the house cannot yet claim; the process area profile says where to work.

The next step. CMMI's advice not to skip levels points the same way as Chapter Ninety-Four's reading of ISO's principles: the first fix is relationship management with the supplier, then the organisation-level processes of level 3.

Choosing a methodology

The methods of this module answer different questions, and the table is this book's summary of what the sources say each is for.

MethodThe problem it fitsWhat it needs
PDCA or PDSA (Chapter Ninety-Nine)A specific change to try and learn fromA prediction, a small trial, data on the result
Kaizen (Chapter Ninety-Nine)Many small everyday improvementsPeople empowered to change their own work
Six Sigma and DMAIC (Chapter Eighty-Nine)A process with too much variation or too many defectsData, statistical skills, a defined project, management support
LeanA process with too much waste: waiting, rework, handoffsA view of the value stream, from request to delivery
CMMIAn organisation whose processes are ad hoc or inconsistent across projectsA long-term programme, a model to measure against, appraisal
ISO 9001 (Chapter Ninety-Five)A need to demonstrate a working quality management system, often to customersDocumented processes, internal audits, certification

Three points guide the choice. Match the method to the problem: a defect rate that varies widely is a Six Sigma problem; a slow flow of work full of waiting is a lean one; an organisation where every project works differently is a CMMI one. Combine them: ASQ notes that the distinction between lean and Six Sigma "has blurred", and every method uses a PDCA-like cycle inside. And start small: the IDEAL model and PDCA both begin by diagnosing and trying before committing the whole organisation.

What it does not mean

Lean is not cutting people. It removes work that adds no value; ASQ even counts non-utilized talent among the wastes.

munotes.in580

Lean, CMMI and Choosing a Methodology

A maturity level is not a measure of product quality. Level 1 organisations "often produce products and services that work"; the level describes how predictably they do.

The staged level does not tell the whole story. ExamReg's house is level 1 on the staged scale with practices up to level 5; the continuous representation shows the profile.

No method is best for every problem. The choice depends on whether the problem is a change to test, variation, waste, or an organisation's whole way of working.

Quick revision

  • Lean (ASQ): management practices to improve efficiency and effectiveness by eliminating waste; lean enterprise extends it along the value stream; Toyota's objective: eliminate waste and shorten lead times.
  • Eight wastes (DOWNTIME): defects, overproduction, waiting, non-utilized talent, transportation, inventory, motion, extra-processing.
  • Lean against Six Sigma (ASQ): waste reduction against variation reduction.
  • CMMI (v1.3): 22 process areas; staged (maturity levels) and continuous (capability levels) representations; levels 1 Initial, 2 Managed, 3 Defined, 4 Quantitatively Managed, 5 Optimizing; do not skip levels; IDEAL: Initiating, Diagnosing, Establishing, Acting, Learning. Later versions (V2.0 and after) are kept by the CMMI Institute (ISACA).
  • Choosing: fit the method to the problem; combine methods; start small.
  • Worked example: ExamReg's house practises 6 of 7 level 2 areas, 5 of 11 at level 3 and 1 of 2 at level 5, but is level 1 staged, for lack of supplier agreement management; with an agreement for the payment gateway, level 2.

Test yourself

1. What is lean, and what are its eight wastes? A set of management practices to improve efficiency and effectiveness by eliminating waste, work that adds no value for the customer. The eight wastes, remembered as DOWNTIME, are defects, overproduction, waiting, non-utilized talent, transportation, inventory, motion and extra-processing.

2. Give software examples of three of the wastes. Defects: the rework of finding and fixing bugs, especially after release. Waiting: work stopped for a supplier, a review or an approval, like ExamReg's 72-hour wait for the payment gateway's provider. Overproduction: reports produced that nobody reads, like a coverage report for every module on every build.

3. Describe CMMI's five maturity levels. Level 1, Initial: processes are ad hoc and chaotic, and success depends on individual heroics. Level 2, Managed: projects plan and execute processes according to policy, and monitor and control them. Level 3, Defined: processes are described in organisation-wide standards and tailored for each project. Level 4, Quantitatively Managed: quantitative objectives for quality and process performance are used to manage projects. Level 5, Optimizing: processes are continually improved from a quantitative understanding of the organisation's objectives and performance.

4. Distinguish CMMI's staged and continuous representations. The staged representation rates an organisation on maturity levels, each requiring a set of process areas at that level and below. The continuous representation rates each process area separately on capability levels, giving a profile rather than a single number.

munotes.in581

Lean, CMMI and Choosing a Methodology

5. Why was ExamReg's software house at maturity level 1 despite practising causal analysis? Because the staged representation requires every process area of level 2 before level 2 is reached, and one, supplier agreement management, was missing: there was no agreement with the payment gateway's provider. Practices at higher levels do not count until the levels below are complete.

6. How should an organisation choose among PDCA, Six Sigma, lean and CMMI? By the problem: PDCA for a specific change to try and learn from; Six Sigma for too much variation or too many defects in a process; lean for too much waste, such as waiting and rework; CMMI for an organisation whose processes are ad hoc or inconsistent. The methods can be combined, and each change is best started small.

Contents This chapter on its own page

munotes.in582

Chapter One Hundred One

The Cost of Quality

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Utilizing quality costs for decision making"

In one line

The cost of quality adds up everything an organisation spends because quality matters: preventing defects, appraising work to find them, and paying for the failures, those caught before the customer and those the customer meets; set out as a statement, it shows where the money goes and whether it is being spent at the right end.

In the wording a student can write in an examination: the cost of quality (COQ) is the "life-cycle costs associated with assuring that a product or service conforms to requirements, plus failure costs from non-conformance to requirements" (ISO/IEC/IEEE 24774:2021). ASQ describes it as a methodology to determine "the extent to which its resources are used for activities that prevent poor quality, that appraise the quality of the organization's products or services, and that result from internal and external failures". It has four categories: prevention costs, "incurred to prevent or avoid quality problems"; appraisal costs, "associated with measuring and monitoring activities related to quality"; internal failure costs, "incurred to remedy defects discovered before the product or service is delivered to the customer"; and external failure costs, "incurred to remedy defects discovered by customers" (ASQ).

Why count the cost of quality

Quality work is often seen as a cost with no visible return: testers, reviews and audits appear in the budget, while the failures they prevent do not. The cost of quality puts both sides on the same page. ASQ explains the purpose: the information "allows an organization to determine the potential savings to be gained by implementing process improvements", and an analysis of quality costs "provides a method of assessing the effectiveness of the management of quality and a means of determining problem areas, opportunities, savings, and action priorities."

Two people shaped the idea. Armand Feigenbaum's book "was the first text to characterize quality costs as the costs of prevention, appraisal, and internal and external failure" (Chapter Eighty-Three, on Feigenbaum, Ishikawa, Crosby and TQM). Philip Crosby made it a communication tool: ASQ records that "He referred to the measure as the 'price of nonconformance' and argued that organizations choose to pay for poor quality."

The four categories

CategoryASQ's descriptionASQ's examplesIn a software project
PreventionIncurred to prevent or avoid quality problems; planned and incurred before actual operationProduct requirements, quality planning, quality assurance, trainingQuality plans, standards and checklists, training, process improvement
AppraisalMeasuring and monitoring activities related to qualityVerification, quality audits, supplier ratingReviews, testing, SQA audits
Internal failureRemedying defects found before deliveryWaste, scrap, rework, failure analysisFixing and retesting defects found in reviews and testing; debugging
External failureRemedying defects found by customersRepairs and servicing, warranty claims, complaints, returnsFixing defects in production; help desk work; correcting wrong results for users
munotes.in583

The Cost of Quality

The first two are the costs of doing quality work: in the standard's words, the costs "associated with assuring that a product or service conforms to requirements". The last two are the "failure costs from non-conformance to requirements", Crosby's price of nonconformance. ASQ also names the cost of poor quality (COPQ), "the costs associated with providing poor quality products or services", which it divides into appraisal, internal failure and external failure costs; the two groupings answer different questions, and an answer should say which one it uses.

Worked example: release 2.0's cost of quality statement

The software house records effort in person-hours and charges Rs 800 an hour. Release 2.0's quality-related work comes from the book's data: the reviews, tests and fix counts of earlier chapters, and the average effort to fix a defect where it was found (1.0 hour in a review, 4.0 in testing, 8.0 after release). The program sets out the cost of quality statement.

# release 2.0's cost of quality, in person-hours (FINDINGS 5.10)
RATE, PROJECT_HOURS = 800, 4800                # Rs per person-hour; the whole project's effort
statement = {
    "prevention": {"SQA plan and procedures": 24, "training in reviews and test techniques": 40,
                   "standards and checklists kept up to date": 16},
    "appraisal": {"reviews": 62, "testing": 296, "SQA audits": 48},
    "internal failure": {"fixing 74 review finds": 74 * 1.0, "fixing 114 test finds": 114 * 4.0,
                         "redoing 13 reopened fixes": 30},
    "external failure": {"fixing 12 after-release defects": 12 * 8.0, "handling complaints": 40,
                         "correcting 9 wrong fees by hand": 12},
}
total = sum(sum(items.values()) for items in statement.values())
for category, items in statement.items():
    hours = sum(items.values())
    print(f"{category:<17}{hours:>6.0f} h  Rs {hours * RATE:>9,.0f}  {hours / total:>4.0%}")
    for item, h in items.items():
        print(f"   {item:<40}{h:>6.0f} h")
conformance = sum(statement["prevention"].values()) + sum(statement["appraisal"].values())
print(f"cost of quality {total:,.0f} h, Rs {total * RATE:,.0f}:"
      f" {total / PROJECT_HOURS:.1%} of the project's {PROJECT_HOURS:,} hours")
print(f"conformance (prevention + appraisal) {conformance / total:.0%},"
      f" nonconformance (failures) {1 - conformance / total:.0%}")
prevention           80 h  Rs    64,000    7%
   SQA plan and procedures                     24 h
   training in reviews and test techniques     40 h
   standards and checklists kept up to date    16 h
appraisal           406 h  Rs   324,800   34%
   reviews                                     62 h
   testing                                    296 h
   SQA audits                                  48 h
internal failure    560 h  Rs   448,000   47%
   fixing 74 review finds                      74 h
   fixing 114 test finds                      456 h
   redoing 13 reopened fixes                   30 h
external failure    148 h  Rs   118,400   12%
   fixing 12 after-release defects             96 h
   handling complaints                         40 h
   correcting 9 wrong fees by hand             12 h
cost of quality 1,194 h, Rs 955,200: 24.9% of the project's 4,800 hours
conformance (prevention + appraisal) 41%, nonconformance (failures) 59%

Where the money went. Release 2.0's quality costs came to 1,194 person-hours, about Rs 9.55 lakh, a quarter (24.9 per cent) of the whole project's effort. Almost half, 47 per cent, was internal failure: fixing and refixing the defects found before release. External failure was 12 per cent, appraisal 34 per cent, and prevention only 7 per cent.

munotes.in584

The Cost of Quality

What the shape says. Failures took 59 per cent of the cost of quality, the work of doing quality only 41 per cent, and prevention, the category whose purpose is to stop the others arising, was the smallest by far. That is the pattern Boehm and Basili describe across software, where projects "spend about 40 to 50 percent of their effort on avoidable rework". Much of release 2.0's 560 hours of internal failure was avoidable: 58 of its defects were input validation defects that one shared module and one checklist item later cut sharply (Chapter Eighty, on using defect data). The statement therefore points where Chapter One Hundred Two, on using quality costs for decision making, looks next: moving hours from failure to prevention and early appraisal.

What the figures leave out. The statement counts the software house's hours. External failure also costs the customer: the students who were charged wrong fees, the exam cell's lost time, and the trust in the portal. ASQ's list of external failure costs includes "Complaints: All work and costs associated with handling and servicing customers' complaints", but the harm to users rarely appears on anyone's timesheet, which is one reason external failure is worse than its hours suggest.

Reading ASQ's rule of thumb carefully

ASQ writes that "Many organizations will have true quality-related costs as high as 15-20% of sales revenue, some going as high as 40% of total operations", and that "A general rule of thumb is that costs of poor quality in a thriving company will be about 10-15% of operations." Those figures are percentages of sales revenue or of operations, and they are not directly comparable with ExamReg's 24.9 per cent of one project's effort. A cost of quality figure means something only against its own base and its own history: the useful comparison for ExamReg is release against release, which is the trend Chapter One Hundred Two computes.

What it does not mean

The cost of quality is not the cost of the quality department. It includes every hour spent because quality matters, most of it by developers fixing defects.

A low cost of quality is not automatically good. Cutting appraisal and prevention can lower it for one release while defects move to the customer, where failure costs are highest.

Prevention and appraisal are not waste. They are the costs of conformance; the goal is to spend them where they save more in failure costs.

munotes.in585

The Cost of Quality

The four categories are not fixed labels for activities. A review is appraisal; improving the review checklist so that defects are not made is prevention; the classification follows the purpose of the work.

Quick revision

  • Cost of quality (ISO/IEC/IEEE 24774:2021): the life-cycle costs of assuring conformance to requirements, plus the failure costs of non-conformance.
  • Four categories (ASQ): prevention (plans, training, standards, QA); appraisal (verification, reviews, testing, audits); internal failure (rework before delivery); external failure (repairs, complaints, returns after delivery).
  • Conformance and nonconformance: prevention and appraisal against internal and external failure; Crosby's "price of nonconformance".
  • History: Feigenbaum first classified quality costs into the four categories.
  • Worked example: release 2.0: 1,194 hours (Rs 9,55,200), 24.9 per cent of the project; prevention 7, appraisal 34, internal failure 47, external failure 12 per cent; failures 59 per cent of the cost of quality.

Test yourself

1. Define the cost of quality and name its four categories. The total cost of everything done because quality matters: the costs of assuring that the product conforms to its requirements plus the costs of failing to conform. Its categories are prevention, appraisal, internal failure and external failure.

2. Give two examples of each category in a software project. Prevention: writing the quality plan and standards; training developers and testers. Appraisal: reviews and inspections; testing and SQA audits. Internal failure: fixing defects found in testing; retesting reopened fixes. External failure: fixing defects users report; handling complaints and correcting wrong results for users.

3. What is the difference between the cost of conformance and the cost of nonconformance? The cost of conformance is what is spent to make and check that the product conforms: prevention and appraisal. The cost of nonconformance is what is spent because it did not: internal and external failure, which Crosby called the price of nonconformance.

4. Draw up a cost of quality statement from given figures. List the items under the four categories with their hours or amounts, total each category, express each as a share of the total cost of quality, and compare the total with the project's effort. For ExamReg release 2.0: prevention 80 hours, appraisal 406, internal failure 560, external failure 148, a total of 1,194 hours, 24.9 per cent of the project.

5. What did release 2.0's statement show, and what does it suggest? That failures took 59 per cent of the cost of quality and prevention only 7 per cent, with internal failure the largest category at 47 per cent. It suggests shifting effort towards prevention and early appraisal, since many of the failures, such as the input validation defects, were avoidable.

6. Why is external failure worse than its hours suggest? Because its costs fall partly on the customer and are not in the supplier's accounts: students charged wrong fees, the exam cell's time, and lost trust in the product. A defect that reaches the user also costs the most to fix.

Contents This chapter on its own page

munotes.in586

Chapter One Hundred Two

Using Quality Costs for Decision Making

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Utilizing quality costs for decision making"

In one line

Quality costs become useful when they drive decisions: tracked release after release they show whether quality spending is moving from failure to prevention, and set against an improvement's cost they show whether it pays, which depends on the organisation's own cost of a late defect, not on anyone else's.

In the wording a student can write in an examination: an analysis of quality costs, ASQ says, "provides a method of assessing the effectiveness of the management of quality and a means of determining problem areas, opportunities, savings, and action priorities." It supports three kinds of decision. Trend analysis compares the cost of quality over time, normalised for size, and its mix between conformance (prevention, appraisal) and failure. Investment decisions compare an improvement's cost with the failure costs it is expected to save, with a break-even point. Life-cycle decisions weigh extra quality cost now against lower costs later; Boehm and Basili (2001) found that high-dependability software costs about 50 per cent more to develop but can cost about the same over its life, because it is cheaper to maintain.

From a statement to a decision

Chapter One Hundred One, on the cost of quality, drew up release 2.0's statement: 1,194 person-hours, 59 per cent of it failure. A statement on its own records what happened. ASQ asks for more: the costs "must be a true measure of the quality effort", and "The quality cost system, once established, should become dynamic and have a positive impact on the achievement of the organization's mission, goals, and objectives." In practice that means three uses: following the trend, deciding on improvements, and understanding the cost of quality over a product's whole life.

Worked example: a trend, a decision, and the life cycle

The program does all three. First, it compares releases 2.0 and 2.1, per KLOC so that a smaller release does not look better merely for being smaller. Second, it evaluates a proposal for release 2.2: a Fagan-style inspection of the fee module's 4,000 lines of code by four people, at the preparation and meeting rates of Chapter Twenty-Nine, on inspection, replacing the light code review now used. The inputs are ExamReg's own: release 2.1's defect density, the share of defects still present at code review in release 2.0, the code review's effectiveness of 21.2 per cent (Chapter Seventy-Nine, on defect metrics) against the median of about 60 per cent that Boehm and Basili report for peer reviews, and the cost of fixing a defect where it is found (1.0 hour in a review, 4.0 in testing, 8.0 after release, plus the complaints and corrections that come with each defect students meet). Third, it redoes Boehm and Basili's life-cycle arithmetic.

# 1. The trend: the cost of quality of releases 2.0 and 2.1, in person-hours (FINDINGS 5.10)
releases = {"2.0": (20, {"prevention": 80, "appraisal": 406, "internal failure": 560,
                         "external failure": 148}),
            "2.1": (8, {"prevention": 70, "appraisal": 180, "internal failure": 150,
                        "external failure": 30})}
for name, (kloc, coq) in releases.items():
    total = sum(coq.values())
    failure = coq["internal failure"] + coq["external failure"]
    print(f"release {name}: {total} h, {total / kloc:.1f} h per KLOC; prevention"
          f" {coq['prevention'] / total:.0%}, failures {failure / total:.0%}")

# 2. The decision: a Fagan-style code inspection of the fee module in release 2.2
lines, people = 4000, 4
cost = people * (lines / 125 + lines / 150)     # preparation and meeting at Fagan's rates
cost -= 30 * lines / 20000                      # less the light code review it replaces
present = lines / 1000 * (69 / 8) * (160 / 200)  # 2.1's density; the share present at code review
extra = present * (0.60 - 0.212)                # effectiveness 21.2% now; 60% median for reviews
# a defect found by the inspection costs 1.0 h to fix; missed, 90.5% are found in testing
# (4.0 h) and 9.5% after release (8.0 h to fix + 52 / 12 h of complaints and corrections)
saving = extra * (0.905 * (4.0 - 1.0) + 0.095 * (8.0 + 52 / 12 - 1.0))
print(f"inspection: {cost:.1f} extra hours to find {extra:.1f} more defects early;"
      f" saving {saving:.1f} h; net {saving - cost:+.1f} h")
print(f"break-even: each defect found early would have to save {cost / extra:.1f} h")

# 3. Boehm and Basili: is quality free over the life cycle? (per instruction; 30% build, 70% keep)
low_build, high_build = 1.0, 1.5                # high dependability costs 50% more to build
low_keep, high_keep = 1.5 * low_build, 0.85 * high_build
print(f"life cycle per instruction: low dependability {0.3 * low_build + 0.7 * low_keep:.3f},"
      f" high {0.3 * high_build + 0.7 * high_keep:.3f}")
munotes.in587

Using Quality Costs for Decision Making

release 2.0: 1194 h, 59.7 h per KLOC; prevention 7%, failures 59%
release 2.1: 430 h, 53.8 h per KLOC; prevention 16%, failures 42%
inspection: 228.7 extra hours to find 10.7 more defects early; saving 40.6 h; net -188.1 h
break-even: each defect found early would have to save 21.4 h
life cycle per instruction: low dependability 1.350, high 1.342

The trend. Release 2.1 cost 53.8 hours of quality work per KLOC against 59.7 for release 2.0, and its mix moved the right way: prevention rose from 7 to 16 per cent of the cost of quality, and failures fell from 59 to 42 per cent. That is the evidence that release 2.0's causal analysis, which added prevention (Chapter Eighty, on using defect data), paid in the next release's costs, and it is the kind of evidence ASQ means by the cost system becoming "dynamic".

munotes.in588

Using Quality Costs for Decision Making

The decision, and why the answer is no. The inspection would cost 228.7 more hours and find about 10.7 more defects early. On ExamReg's own costs, moving a defect from testing to inspection saves 3 hours of fixing, and from after release about 11, so the expected saving is only 40.6 hours: the proposal loses 188.1 hours. The break-even line explains why: the inspection pays only if each defect found early saves 21.4 hours. In a large, critical system, where Boehm and Basili's "often 100 times" applies, it easily would; in ExamReg, a small, noncritical portal whose defects are cheap to fix, it does not. The decision follows from the organisation's own data, which is exactly why the data are collected. The same analysis also points at better options: a lighter, checklist-based review of the fee module's riskiest parts costs far fewer hours, and prevention, which removed defects outright in release 2.1, needs no finding at all.

The life cycle. Boehm and Basili asked whether Crosby's "quality is free" is right, since "it costs 50 percent more per source instruction to develop high-dependability software products than to develop low-dependability software products." Their answer rests on maintenance: low-dependability software "costs about 50 percent per instruction more to maintain than to develop, whereas high-dependability software costs about 15 percent less to maintain than to develop", and with 30 per cent of life-cycle cost in development and 70 in maintenance, "low-dependability software becomes about the same in cost per instruction as high-dependability software". Weighting each per-instruction cost 30 to 70, which is one reading of their arithmetic, reproduces the conclusion: 1.350 against 1.342. Their overall verdict on Crosby: "Maybe for some low-criticality, short-lifetime software, but not for the most important cases."

Using quality costs well

DecisionWhat the cost of quality contributesExamReg
Where to actThe largest failure costs, like a Pareto of costs (Chapter One Hundred Four, on Pareto diagrams)Internal failure, 47 per cent of release 2.0's cost of quality
Whether an action workedThe trend, normalised for size, and the shift in mix59.7 to 53.8 hours per KLOC; failures from 59 to 42 per cent
Whether to investCost against expected savings, with a break-even pointThe fee module inspection: rejected, break-even 21.4 hours a defect
How much quality to build inLife-cycle cost, not development cost aloneBoehm and Basili: about equal per instruction for high and low dependability

What it does not mean

Quality costs do not make the decision alone. They put numbers on it; safety, reputation and harm to users (the external failures students suffer) may outweigh the hours.

An improvement that does not pay here may pay elsewhere. The inspection fails on ExamReg's costs and would succeed where late defects cost far more; ratios from other organisations must not be imported unexamined.

munotes.in589

Using Quality Costs for Decision Making

Quality is not free in every case. Boehm and Basili show higher development cost can be repaid over the life cycle, but not for all software, as they say themselves.

A falling cost of quality is not proof of success. It must be normalised for size and read with its mix; cutting appraisal lowers the cost now and raises failure costs later.

Quick revision

  • Uses of quality costs (ASQ): assess the management of quality; find problem areas, opportunities, savings and action priorities; the cost system should become "dynamic".
  • Trend analysis: cost of quality per unit of size, release by release, and its mix of conformance against failure.
  • Investment decisions: an improvement's cost against the failure costs it saves; the break-even saving per defect.
  • Life-cycle cost (Boehm and Basili 2001): high dependability costs 50 per cent more to develop, maintenance 15 per cent below development cost against 50 per cent above for low dependability; over a 30/70 life cycle, about the same; "quality is free" fails only for "some low-criticality, short-lifetime software".
  • Worked example: 59.7 to 53.8 hours per KLOC, prevention 7 to 16 per cent, failures 59 to 42 per cent; the fee module inspection costs 228.7 hours, saves 40.6, break-even 21.4 hours a defect: rejected on ExamReg's costs; life cycle 1.350 against 1.342.

Test yourself

1. How are quality costs used for decision making? To find where quality money is lost (the largest failure costs), to judge whether improvements worked (trends normalised for size and the shift from failure to prevention), to decide whether a proposed improvement pays (its cost against the failure costs it saves, with a break-even point), and to weigh extra quality cost now against lower costs over the product's life.

2. What did the trend between releases 2.0 and 2.1 show? The cost of quality per KLOC fell from 59.7 to 53.8 hours, prevention rose from 7 to 16 per cent of it, and failures fell from 59 to 42 per cent: the added prevention of release 2.1 was paying back in lower failure costs.

3. Why was the proposed fee module inspection rejected? On ExamReg's own costs it would cost 228.7 more hours and save only about 40.6 hours of fixing, because the defects it would find early are cheap to fix later in this small, noncritical system. It would pay only if each defect found early saved at least 21.4 hours.

4. What is a break-even point in a quality cost decision, and why is it useful? The value at which an improvement's savings equal its cost, here the hours each early-found defect must save. It shows what would have to be true for the improvement to pay, so the decision can be judged against the organisation's own data or revisited when those data change.

munotes.in590

Using Quality Costs for Decision Making

5. Is quality free? Give Boehm and Basili's answer. Not always. High-dependability software costs about 50 per cent more per instruction to develop, but is cheaper to maintain, so over a life cycle of 30 per cent development and 70 per cent maintenance it costs about the same as low-dependability software; in their words, "the investment is more than worth it if the project involves significant operations and maintenance costs." Crosby's claim may fail only for some low-criticality, short-lifetime software.

6. Why should an organisation not use another organisation's cost ratios for its decisions? Because the value of finding defects early depends on how much a late defect costs, which varies from about 5 to 1 in small, noncritical systems to often 100 to 1 in large, critical ones. ExamReg's inspection decision is negative on its own ratios and would be positive on a large system's.

Contents This chapter on its own page

munotes.in591

Chapter One Hundred Three

The Seven Basic Quality Tools

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Introduction to quality improvement tools"

In one line

The seven basic quality tools are simple, mostly graphical methods that let anyone collect data about a problem and see what it says: a check sheet to count, a histogram to see a distribution, a Pareto chart to find the few causes that matter, a cause-and-effect diagram to list possible causes, a scatter diagram to test a relationship, a control chart to judge stability, and stratification (or, in some lists, a flowchart or run chart) to separate the data into groups.

In the wording a student can write in an examination: ASQ lists the seven basic quality tools as the cause-and-effect diagram ("It illustrates the main causes and subcauses leading to an effect"), the check sheet ("a simple data recording tool"), the control chart, the histogram ("used for displaying the distribution of data graphically"), the Pareto chart, the scatter diagram ("used to analyze and identify potential relationships between two variables") and stratification ("breaking down data into categories to more easily make sense of them"); "more current views typically include flowchart and leave off stratification", and some include the run chart. They "were first highlighted in Kaoru Ishikawa's classic book Guide to Quality Control."

Tools for everyone

Chapter Eighty-Three, on Feigenbaum, Ishikawa, Crosby and TQM, told how Ishikawa's company-wide quality control spread quality work from specialists to everyone, and how his cause-and-effect diagram "has provided a powerful tool that can easily be used by non-specialists to analyze and solve problems." The seven basic tools are that idea as a kit. None needs more than counting and drawing; each answers one question; and ASQ notes that "Using these tools in combination further multiplies their value."

The seven tools and the question each answers

ToolThe question it answersIn this book
Check sheetHow often does each kind of problem occur, and when or where?This chapter
HistogramHow is a measurement spread out?The fee page's response times (Chapter Eighty-One, on quality concepts)
Pareto chartWhich few problems account for most of the effect?Chapter One Hundred Four, on Pareto diagrams
Cause-and-effect diagramWhat could be causing this problem?Chapter One Hundred Five, on cause-effect diagrams
Scatter diagramAre two variables related?Chapter One Hundred Six, on scatter diagrams
Control chartIs the process stable, or are there special causes?Chapter Ninety-One, on statistical process control
StratificationDoes the picture change when the data are split into groups?This chapter
(Flowchart)What are the steps of the process?The life cycle and test process diagrams of Module 1
(Run chart)How does a measure change over time?Chapter One Hundred Seven, on run charts

The check sheet

ASQ describes the check sheet as "a structured, prepared form for collecting and analyzing data", and warns that it is "Not to be confused with checklists": a checklist, like those of Chapter Sixty-Four, on checklist-based testing, lists things to check; a check sheet records how often something happened. It is used "When collecting data on the frequency or patterns of events, problems, defects, defect location, defect causes, or similar issues", especially "When data can be observed and collected repeatedly by the same person or at the same location".

munotes.in592

The Seven Basic Quality Tools

ASQ's steps for making one:

  1. "Decide what event or problem will be observed. Develop operational definitions."
  2. "Decide when data will be collected and for how long."
  3. Design the form so that data "can be recorded simply by making check marks or X's or similar symbols and so that data do not have to be recopied for analysis."
  4. "Label all spaces on the form."
  5. "Test the check sheet for a short trial period to be sure it collects the appropriate data and is easy to use."
  6. "Each time the targeted event or problem occurs, record data on the check sheet."

Worked example: the help desk's check sheet

During release 2.0's first registration week, ExamReg's help desk kept a check sheet: one tally mark for each complaint, under its type and its day. The six days end on the last day before the late fee begins. The program prints the sheet as the help desk kept it, then stratifies the same data by day.

# the help desk's check sheet for release 2.0's first registration week (FINDINGS 5.11):
# complaints by type, one count per day; day 6 is the last day before the late fee starts
sheet = {"payment taken, form not submitted": [3, 4, 2, 5, 6, 14],
         "session expired while filling form": [2, 1, 3, 2, 3, 8],
         "account locked after wrong passwords": [1, 2, 1, 1, 2, 5],
         "wrong fee shown": [1, 0, 1, 1, 0, 2],
         "other": [1, 1, 0, 2, 1, 3]}

def tally(n):
    """Tally marks in groups of five, as a check sheet records them."""
    return " ".join(["|||||"] * (n // 5) + (["|" * (n % 5)] if n % 5 else []))

print("the check sheet, as kept for the week")
for complaint, counts in sheet.items():
    print(f"   {complaint:<38}{tally(sum(counts)):<42}{sum(counts):>3}")

print("the same data stratified by day")
print(f"   {'':<38}" + "".join(f"{d:>4}" for d in range(1, 7)))
for complaint, counts in sheet.items():
    print(f"   {complaint:<38}" + "".join(f"{n:>4}" for n in counts))
days = [sum(day) for day in zip(*sheet.values())]
print(f"   {'all complaints':<38}" + "".join(f"{n:>4}" for n in days))
print(f"day 6 alone: {days[-1]} of {sum(days)} complaints ({days[-1] / sum(days):.0%})")
the check sheet, as kept for the week
   payment taken, form not submitted     ||||| ||||| ||||| ||||| ||||| ||||| ||||   34
   session expired while filling form    ||||| ||||| ||||| ||||                     19
   account locked after wrong passwords  ||||| ||||| ||                             12
   wrong fee shown                       |||||                                       5
   other                                 ||||| |||                                   8
the same data stratified by day
                                            1   2   3   4   5   6
   payment taken, form not submitted        3   4   2   5   6  14
   session expired while filling form       2   1   3   2   3   8
   account locked after wrong passwords     1   2   1   1   2   5
   wrong fee shown                          1   0   1   1   0   2
   other                                    1   1   0   2   1   3
   all complaints                           8   8   7  11  12  32
day 6 alone: 32 of 78 complaints (41%)
munotes.in593

The Seven Basic Quality Tools

What the check sheet shows. In a week, 78 complaints. One kind dominates: payment taken, form not submitted, 34 of them, more than the next two kinds together. That is a Pareto question, and Chapter One Hundred Four, on Pareto diagrams, asks it properly. It is also the problem most worth a cause-and-effect diagram, which Chapter One Hundred Five, on cause-effect diagrams, draws for failed fee payments.

What stratification adds. Split by day, the complaints are steady at 7 to 12 a day for five days and then jump to 32 on day 6, 41 per cent of the week's total, with every kind of complaint rising. Two readings are possible, and the data alone do not decide between them: the portal may fail more under the last day's load, or more students may simply use it that day. The next tools separate them: a scatter diagram of complaints against daily logins (Chapter One Hundred Six, on scatter diagrams) and a run chart of complaints over the season (Chapter One Hundred Seven, on run charts).

Why the check sheet came first. Without it, the help desk's impression might have been students keep getting logged out, the most irritating complaint to handle. The sheet's counts put the payment problem first and showed the deadline effect, which is what ASQ's first step, deciding exactly what to observe and defining it, makes possible.

The tools together

The tools form a natural sequence for a quality problem, and this module's last chapters follow it.

  1. Count with a check sheet: what happens, how often, when.
  2. Rank with a Pareto chart: which problems matter most.
  3. Explore causes with a cause-and-effect diagram: what could produce the top problem.
  4. Test a suspected cause with a scatter diagram or stratification: is it related, and in which groups?
  5. Watch the process with a run chart or control chart: did the fix work, and does it stay fixed?

What it does not mean

A check sheet is not a checklist. A checklist says what to check; a check sheet counts what happened.

Basic does not mean weak. The tools are simple so that everyone can use them; most quality problems need nothing more.

munotes.in594

The Seven Basic Quality Tools

The list is not fixed. ASQ notes that newer lists include the flowchart instead of stratification, and some include the run chart.

A tool does not supply the cause. The check sheet showed the deadline effect; deciding what causes it needs further data and judgement.

Quick revision

  • The seven basic quality tools (ASQ; first highlighted in Ishikawa's Guide to Quality Control): cause-and-effect diagram, check sheet, control chart, histogram, Pareto chart, scatter diagram, stratification (or flowchart; some lists add the run chart).
  • Each answers one question: how often (check sheet), how spread (histogram), which few (Pareto), what causes (cause-and-effect), related (scatter), stable (control chart), different by group (stratification).
  • Check sheet (ASQ): a prepared form for collecting data by tally marks; not a checklist; steps: define what to observe, decide when and how long, design and label the form, trial it, record each event.
  • Together: count, rank, explore causes, test a cause, watch the process.
  • Worked example: 78 complaints in the first week; payment taken but form not submitted 34; day 6 alone 32 (41 per cent), every kind of complaint rising.

Test yourself

1. Name the seven basic quality tools and say what each is for. The check sheet records how often problems occur; the histogram shows how a measurement is distributed; the Pareto chart ranks problems to find the few that matter most; the cause-and-effect diagram lists possible causes of a problem; the scatter diagram shows whether two variables are related; the control chart shows whether a process is stable; stratification splits data into groups to see whether the picture differs. Some lists use a flowchart or a run chart instead of stratification.

2. What is a check sheet, and how is one made? A structured, prepared form for collecting data by marking tallies as events occur. To make one: decide exactly what will be observed and define it; decide when and for how long to collect; design and label the form so that marks need no recopying; test it for a short period; then record every occurrence.

3. How does a check sheet differ from a checklist? A checklist lists items to check or steps to follow, such as the questions in a review; a check sheet records how many times each kind of event happened, producing data for analysis.

4. What is stratification, and what did it show in the worked example? Breaking data down into categories to make sense of them. Split by day, the help desk's complaints were steady at 7 to 12 a day for five days and jumped to 32 on the last day before the late fee, 41 per cent of the week, a deadline effect the weekly totals hid.

munotes.in595

The Seven Basic Quality Tools

5. In what order are the basic tools typically used on a problem? Count occurrences with a check sheet; rank them with a Pareto chart; explore possible causes of the top problem with a cause-and-effect diagram; test a suspected cause with a scatter diagram or stratification; and watch the process after a change with a run chart or control chart.

6. Why are the seven tools called basic, and why did Ishikawa promote them? Because they need only counting and drawing, so anyone can use them without statistical training. Ishikawa promoted them so that quality work could be done by everyone in the organisation, not only by specialists, which is the idea of company-wide quality control.

Contents This chapter on its own page

munotes.in596

Chapter One Hundred Four

Pareto Diagrams

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Pareto Diagrams"

In one line

A Pareto chart ranks a set of problems from largest to largest-so-far, so that a plain bar chart with a rising cumulative line shows at a glance which few categories account for most of the total, which is where limited effort should go first.

In the wording a student can write in an examination: ASQ's Pareto chart (also called a Pareto diagram or Pareto analysis) "is a bar graph. The lengths of the bars represent frequency or cost (time or money), and are arranged with longest bars on the left and the shortest to the right. In this way the chart visually depicts which situations are more significant." It is named for the Pareto principle, which Joseph Juran first stated in 1950 and which ASQ gives as the rule that "80% of the effects come from 20% of the possible causes", sometimes called the "80-20" rule. Juran called the small share of causes that account for most of the effect the vital few, and the rest the useful many, a name he chose over his first term, "trivial many", once he judged that no cause in quality work is truly trivial.

Why rank before you act

Chapter One Hundred Three, on the seven basic quality tools, ended its check sheet with 78 complaints in five categories and one clear leader, but a check sheet only counts; it does not say how much of the total the leader is, or whether the next category is worth a second effort. ASQ gives the Pareto chart's purpose directly: use it "When there are many problems or causes and you want to focus on the most significant", and "When analyzing broad causes by looking at their specific components". Chapter Eighty, on using defect data, already used this reasoning once without the chart: it picked input validation for release 2.0's first improvement because it was the largest defect type, 58 of the 200, so one successful action on it would remove more defects than an action on any other type, and promised that this chapter would set out the same reasoning for every type at once. A Pareto chart is that reasoning, made visible and applied to every category together rather than to one category chosen by inspection.

Building one

ASQ's procedure, condensed to its steps:

  1. "Decide what categories you will use to group items."
  2. "Decide what measurement is appropriate. Common measurements are frequency, quantity, cost and time."
  3. "Decide what period of time the Pareto chart will cover".
  4. "Collect the data, recording the category each time, or assemble data that already exist."
  5. "Subtotal the measurements for each category."
  6. "Determine the appropriate scale for the measurements you have collected. The maximum value will be the largest subtotal", unless the optional cumulative line is drawn, in which case "the maximum value will be the sum of all subtotals". Mark this scale on the left.
  7. "Construct and label bars for each category. Place the tallest at the far left, then the next tallest to its right, and so on." Small categories "can be grouped as 'other'."
  8. (Optional) "Calculate the percentage for each category" and draw a right-hand scale in percentages, lined up so that, in ASQ's example, "the left measurement that corresponds to one-half should be exactly opposite 50% on the right scale."
  9. (Optional) "Calculate and draw cumulative sums": place a dot above the second bar at the sum of the first two categories, a dot above the third bar at that sum plus the third category, and so on, then "Connect the dots, starting at the top of the first bar. The last dot should reach 100% on the right scale."
munotes.in597

Pareto Diagrams

Worked example: release 2.0's defects, ranked two ways

The chart needs only a count (or a cost) by category. The program uses release 2.0's 200 defects by type (FINDINGS 5.2), which ASQ's steps 1 to 5 already fixed: the category is the defect type, the measurement is frequency, and the period is the whole release. It ranks them twice: once by how many defects of each type occurred, and once by how many hours they took to fix, using the average fix effort of FINDINGS 5.12. The second ranking is ASQ's first listed variation, the weighted Pareto chart, used when a rarer category may still matter more because each occurrence costs more.

# release 2.0's 200 defects by type (FINDINGS 5.2) and the average hours to fix one (FINDINGS 5.12)
defects = {"input validation": (58, 1.2), "logic and computation": (44, 4.6),
           "interface": (30, 3.6), "user interface": (24, 1.0), "data and database": (18, 5.5),
           "documentation": (12, 0.5), "performance": (8, 7.5), "security": (6, 9.5)}

def pareto(values, title):
    """Sort the categories largest first; print each with its share and the cumulative share."""
    total, running = sum(values.values()), 0
    print(f"{title}: {total:g} in all")
    for name, value in sorted(values.items(), key=lambda item: -item[1]):
        running += value
        print(f"   {name:<23}{value:>7g}{value / total:>6.0%}{running / total:>7.0%}")

pareto({name: n for name, (n, _) in defects.items()}, "defects by type")
pareto({name: n * hours for name, (n, hours) in defects.items()}, "hours to fix, by type")
defects by type: 200 in all
   input validation            58   29%    29%
   logic and computation       44   22%    51%
   interface                   30   15%    66%
   user interface              24   12%    78%
   data and database           18    9%    87%
   documentation               12    6%    93%
   performance                  8    4%    97%
   security                     6    3%   100%
hours to fix, by type: 626 in all
   logic and computation    202.4   32%    32%
   interface                  108   17%    50%
   data and database           99   16%    65%
   input validation          69.6   11%    77%
   performance                 60   10%    86%
   security                    57    9%    95%
   user interface              24    4%    99%
   documentation                6    1%   100%
munotes.in598

Pareto Diagrams

Pareto chart of release 2.0's 200 defects by type: bars falling from input validation (58) to security (6), with a cumulative line reaching 100 per cent; the four largest types, shaded, make up 78 per cent of the defects

Figure 104.1 Release 2.0's defects by type, ranked, with the cumulative line ASQ's construction adds

By count, the vital few are four types. Input validation, logic and computation, interface and user interface together are 156 of 200 defects, 78 per cent, the four bars the figure shades. Data and database brings the running total to 87 per cent with a fifth category. The three smallest types, documentation, performance and security, are 13 per cent of the count between them: Juran's useful many, not worth ignoring, but not where a first action should go.

By hours, the ranking changes. Weighted by the effort to fix one, logic and computation moves to the top on its own, 202.4 of 626 hours, 32 per cent, because each one takes nearly four times as long to fix as an input validation defect (4.6 hours against 1.2). Input validation, the largest category by count, drops to fourth by hours; security, the smallest category by count at only 6 defects, is the fifth-largest by hours because each one averages 9.5 hours. The two charts agree that logic, interface and data defects all belong near the top, but the count chart puts input validation first and the hours chart puts it fourth.

Which ranking to use. ASQ's first measurement choice is "frequency, quantity, cost and time"; both are legitimate Pareto charts on the same data, and the right one depends on what is scarce. Chapter Eighty acted on the count, correctly: the goal there was fewer defects reaching the customer, for which the largest count is the largest target. Chapter One Hundred One, on the cost of quality, and Chapter One Hundred Two, on using quality costs for decision making, act on cost and hours, for which the weighted chart is the one that answers the question being asked. A Pareto chart does not choose the measurement; the decision it is meant to support does.

What the two charts do not say

Neither chart says input validation defects are unimportant, or that logic defects are the only ones worth fixing. A Pareto chart ranks a fixed period's data; it does not predict next release's ranking, and Chapter Eighty showed that a targeted action can move a category from first to fourth in one release (2.90 to 1.12 defects per KLOC) while leaving every other type's rate unchanged or slightly higher. Ranking is a starting point for where to look, repeated each time there is new data, not a verdict fixed for all time.

munotes.in599

Pareto Diagrams

What it does not mean

The vital few are not fixed for every measurement. The same data, counted differently, can rank differently: release 2.0's chart by count and by hours agree on three types out of five in the top half, not all.

80-20 is a tendency, not an exact split. Chapter Eighty-Two, on the Quality Movement's Shewhart, Deming and Juran, already recorded Boehm and Basili's range of 60 to 90 per cent of software defects from 20 per cent of modules across studies; ExamReg's own split by type, 78 per cent in four of eight types (50 per cent of the types), is one instance of the same tendency, not a proof of the exact 80/20 figure.

A Pareto chart ranks; it does not explain. It shows that input validation defects are the largest category, not why they occur. Chapter One Hundred Five, on cause-effect diagrams, asks why for the largest category of a different kind of event, the help desk's complaints about failed fee payments.

"Other" is a convenience, not a category to act on. ASQ allows grouping small categories as "other" to keep the chart readable; an action cannot be aimed at "other" because it is not one cause.

Quick revision

  • Pareto chart (ASQ): a bar graph of frequency or cost, longest bar to shortest, showing which few categories matter most; variations include the weighted Pareto chart (by cost or time, not just count) and comparative Pareto charts.
  • Pareto principle (Juran, 1950; ASQ's 80-20 rule): about 80 per cent of effects from about 20 per cent of causes; the vital few causes against the useful many (Juran's later term for what he first called the "trivial many").
  • Construction (ASQ): fix the categories, the measurement and the period; subtotal each category; scale the left axis to the largest subtotal (or, with a cumulative line, to the total); draw bars largest to smallest; optionally add a right-hand percentage scale and a cumulative line from dots at each bar's running total, reaching 100 per cent.
  • Worked example: release 2.0's 200 defects by type; by count, four of eight types make 78 per cent; by hours to fix, the ranking changes and logic and computation alone is 32 per cent.
  • A Pareto chart ranks a measurement for a fixed period; it does not explain a cause or predict the next period's ranking.

Test yourself

1. What is a Pareto chart, and when is it used? A bar graph that ranks categories from the largest measurement to the smallest, usually with a cumulative percentage line, used when there are many possible problems or causes and the most significant few need to be found.

2. State the Pareto principle and Juran's two names for its two groups. About 80 per cent of effects come from about 20 per cent of possible causes. Juran called the 20 per cent of causes the "vital few" and the rest the "useful many", a term he adopted in place of "trivial many" once he judged no quality problem to be trivial.

munotes.in600

Pareto Diagrams

3. List the main steps in constructing a Pareto chart. Decide the categories, the measurement (frequency, quantity, cost or time) and the period to cover; collect or assemble the data; subtotal each category; scale the chart (to the largest subtotal, or to the total if a cumulative line is added); draw the bars from largest to smallest; optionally add a percentage scale and a cumulative line of dots from each bar's running total to 100 per cent.

4. What is a weighted Pareto chart, and when would you use one? A Pareto chart ranked by cost or time rather than by count, used when a category with few occurrences still matters more because each occurrence is expensive. Release 2.0's defects ranked by hours to fix put logic and computation defects first, even though input validation had more defects, because each logic defect took nearly four times as long to fix.

5. Using release 2.0's defect data, which types are the vital few by count, and how does the ranking change by cost? By count, input validation, logic and computation, interface, and user interface are 78 per cent of the 200 defects. Weighted by average hours to fix, the ranking changes: logic and computation alone is 32 per cent of the 626 hours, and input validation drops to fourth.

6. Why does a Pareto chart not explain why a problem occurs, and what tool does? A Pareto chart only ranks categories by a measurement; it says which problem is largest, not what causes it. Chapter One Hundred Five, on cause-effect diagrams, explores the possible causes of a problem once the Pareto chart has identified which one to examine.

Contents This chapter on its own page

munotes.in601

Chapter One Hundred Five

Cause-Effect Diagrams

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Cause-effect Diagrams"

In one line

A cause-and-effect diagram takes one clearly stated problem and, by sorting ideas into a handful of broad categories, structures a team's brainstorm so that no obvious family of causes is skipped and each idea finds its place; the five whys then push one branch of that brainstorm down to a cause specific enough to act on.

In the wording a student can write in an examination: the fishbone diagram, also called a cause-and-effect diagram or an Ishikawa diagram after its creator Kaoru Ishikawa, "can help users identify the many possible causes for a problem by sorting ideas into useful categories and is especially useful in structuring brainstorming sessions" (ASQ). ASQ's glossary adds that "The diagram illustrates the main causes and subcauses leading to an effect (symptom)." It is drawn with the problem statement in a box at the head, a horizontal spine leading to it, and the main categories of cause as branches off the spine, "like a fish's ribs", which is why "the chart is shaped like a fish skeleton." The five whys is a related technique: asking "why" repeatedly, usually five times, to follow one cause down to the one worth fixing. Sakichi Toyoda, who originated it, said "by repeating why five times, the nature of the problem as well as its solution becomes clear" (ASQ).

Why brainstorm before concluding

Chapter One Hundred Three, on the seven basic quality tools, and Chapter One Hundred Four, on Pareto diagrams, both stopped at ranking: the help desk's check sheet found "payment taken, form not submitted" the largest complaint of release 2.0's first registration week, 34 of 78, and named it "the problem most worth a cause-and-effect diagram." Neither tool says why it happens. A team that skips straight to a guess risks fixing the first plausible story instead of the real one; ASQ's categories exist to stop that, by forcing a sweep across every broad kind of cause before anyone commits to one.

Building one

ASQ's procedure:

  1. "Agree on the problem statement or effect being analyzed."
  2. Draw the spine: "a horizontal, right-facing arrow", with the problem statement in a box at its head.
  3. "Identify the main categories of the problem's causes and draw these as branches emanating from the central arrow, like a fish's ribs." Ishikawa's generic labels are the 6 M's: Materials, Machinery (which "may need to be a separate category" for software), Methods, Measurement, Manpower, and Mother Nature (environment and externalities), with Money sometimes added as a seventh. ASQ adds that Ishikawa "encouraged creativity in naming these categories to communicate more clearly to those who would be using the diagram."
  4. Brainstorm against each category by asking "Why does this happen?", writing each idea as a branch; "Causes can be written in several places if they relate to several categories."
  5. "Continue asking 'Why?' to generate deeper levels of causes, writing subcauses as branches off the causes."
  6. "When the group runs out of ideas, focus on the places where there are fewer ideas."
  7. Analyse the causes "to determine those that should be addressed further", remembering "that the purpose is to cure the problem, not the symptoms."
munotes.in602

Cause-Effect Diagrams

The five whys

The five whys drills one branch further than a fishbone session usually goes by hand. ASQ: "Use the five whys technique when you want to push a team investigating a problem to delve into more details of the root causes. The five whys can be used with brainstorming or the cause-and-effect diagram." The method is plain: write the problem at the top of a page, then "Ask 'why' or 'how' five times and write the answers on the lines drawn from number one to five", noting that "It may take less or more than five times to reach the root cause or solution."

Worked example: why fee payment is taken but the form is not submitted

The team runs both tools on release 2.0's largest complaint. The fishbone sorts what the team already suspects into four categories (Materials and Mother Nature turned up nothing specific to a software portal, so the session did not force entries there); the five whys then follows the Machinery branch's first cause, the one whose mechanism, a race between the gateway's callback and the form's session timer, is precise enough to test.

# release 2.0's first registration week: brainstormed causes of "payment taken,
# form not submitted" (34 of the help desk's 78 complaints, FINDINGS 5.11)
causes = {
    "Machinery": ["the gateway's success callback can arrive after the form session has timed out",
                  "a late callback retry is not matched back to the order it belongs to",
                  "the confirmation page re-reads the wrong order id after back and forward"],
    "Methods": ["the form never re-checks payment status before asking the student to pay again",
                "nothing routinely matches the gateway's ledger against ExamReg's session log"],
    "Manpower": ["the help desk cannot see the gateway's side of a transaction",
                 "a student who reloads mid-payment is not a case the session design expected"],
    "Measurement": ["no alert fires when a callback arrives after its session has timed out"],
}

def plural(n):
    return "cause" if n == 1 else "causes"

print('effect: "payment taken, form not submitted" (34 of 78 complaints, FINDINGS 5.11)')
for category, items in causes.items():
    print(f"\n{category} ({len(items)} {plural(len(items))} so far)")
    for item in items:
        print(f"   - {item}")
fewest = min(causes, key=lambda k: len(causes[k]))
print(f"\nfewest ideas so far: {fewest} ({len(causes[fewest])}); ASQ: look there next")

whys = [
    ("Why does payment sometimes get taken but the form stay unsubmitted?",
     "the gateway's success callback sometimes arrives after the form's session has timed out"),
    ("Why does the session time out before the callback arrives?",
     "its timer is set to ordinary form-fill time, not to how long the gateway can take to confirm"),
    ("Why can the gateway take longer than the session allows?",
     "confirmation depends on the payment gateway's own provider, whose timing ExamReg does not control"),
    ("Why did this not surface before the first registration week?",
     "release 2.0's tests checked response time, never a callback arriving after a session had ended"),
    ("Why did the requirement not cover a late callback?",
     "the payment confirmation requirement never stated what to do if it arrived after the session ended"),
]
print("\nfive whys, on the Machinery branch's first cause:")
for i, (why, because) in enumerate(whys, 1):
    print(f"{i}. {why}\n   {because}")
munotes.in603

Cause-Effect Diagrams

effect: "payment taken, form not submitted" (34 of 78 complaints, FINDINGS 5.11)

Machinery (3 causes so far)
   - the gateway's success callback can arrive after the form session has timed out
   - a late callback retry is not matched back to the order it belongs to
   - the confirmation page re-reads the wrong order id after back and forward

Methods (2 causes so far)
   - the form never re-checks payment status before asking the student to pay again
   - nothing routinely matches the gateway's ledger against ExamReg's session log

Manpower (2 causes so far)
   - the help desk cannot see the gateway's side of a transaction
   - a student who reloads mid-payment is not a case the session design expected

Measurement (1 cause so far)
   - no alert fires when a callback arrives after its session has timed out

fewest ideas so far: Measurement (1); ASQ: look there next

five whys, on the Machinery branch's first cause:
1. Why does payment sometimes get taken but the form stay unsubmitted?
   the gateway's success callback sometimes arrives after the form's session has timed out
2. Why does the session time out before the callback arrives?
   its timer is set to ordinary form-fill time, not to how long the gateway can take to confirm
3. Why can the gateway take longer than the session allows?
   confirmation depends on the payment gateway's own provider, whose timing ExamReg does not control
4. Why did this not surface before the first registration week?
   release 2.0's tests checked response time, never a callback arriving after a session had ended
5. Why did the requirement not cover a late callback?
   the payment confirmation requirement never stated what to do if it arrived after the session ended
A fishbone diagram: a spine leading to "payment taken, form not submitted" (34 of 78 complaints), with four rib categories, Machinery, Methods, Manpower and Measurement, each carrying short brainstormed causes; the Machinery rib is drawn heavier as the branch the five whys follow

Figure 105.1 Release 2.0's cause-and-effect diagram for its largest first-week complaint

What the fishbone found. Eight causes across four categories, none of them a guess about which student or which day: a family of Machinery causes about the gateway's callback and session timing, a family of Methods causes about what the system fails to re-check, a family of Manpower causes about what the help desk cannot see, and one Measurement cause about what nobody is alerted to. Measurement has the fewest entries, one, which by ASQ's rule is where the team should brainstorm further next, a separate question from which branch to chase first.

munotes.in604

Cause-Effect Diagrams

What the five whys found. Starting from Machinery's first cause, the chain of five answers does not stop at the gateway, a third party ExamReg does not control; it stops one step further back, at the session timer's length and, beneath that, at a requirement that never said what should happen to a late confirmation. That is a cause the team can act on: change the requirement and the timer, not merely apologise for the gateway's timing.

Where this has been seen before. FINDINGS 5.2.2 records that one of release 2.0's twelve after-release defects took 72 hours to fix because the fix had to wait on the payment gateway's own provider to change its side of the callback, and Chapter Ninety-One, on statistical process control, found that same repair time sitting outside the control limits of the other eleven: a special cause, needing an agreement with the provider on response times rather than a change to ExamReg's own process. That record and this chapter's five whys were built from different data, a repair-time log against a structured brainstorm, and they arrive at the same mechanism. Neither proves the two are the identical defect; together they are two independent readings that agree on where release 2.0's gateway integration is weak, which is the kind of agreement a real investigation would treat as confirmation, not coincidence.

What it does not mean

A fishbone diagram does not prove a cause; it organises candidates. Every branch in the worked example is a brainstormed hypothesis until it is checked; only the five whys branch was followed to something specific enough to test, and even that needs confirming before a fix is built on it.

The 6 M's are a checklist for completeness, not a form to fill in mechanically. ASQ's own advice is to rename or drop categories that do not fit; this chapter used four of six because release 2.0 is software with no physical materials or weather to blame.

Five whys is not exactly five. ASQ says explicitly that it may take more or fewer questions; five is a name and a habit, not a rule that stops a team one question early or pushes it one question past a good answer.

munotes.in605

Cause-Effect Diagrams

A cause found by one branch does not rule out the others. The Methods, Manpower and Measurement branches were not chased in this chapter; they may hold real, separate causes of the same complaint, still open.

Quick revision

  • Cause-and-effect diagram (fishbone, Ishikawa diagram; ASQ): sorts a brainstorm's causes for one stated effect into categories, drawn as ribs off a spine leading to the effect, shaped like a fish skeleton.
  • Ishikawa's 6 M's: Materials, Machinery, Methods, Measurement, Manpower, Mother Nature (sometimes Money); rename or drop categories to fit the problem.
  • Procedure (ASQ): state the problem; draw the spine and categories; brainstorm causes per category, asking "why" for subcauses; focus further brainstorming where ideas are fewest; then decide which causes to pursue.
  • Five whys (Toyoda, via ASQ): ask "why" repeatedly, typically five times, to drill from a symptom to a specific, actionable cause; pairs naturally with a cause-and-effect diagram.
  • Worked example: 34 of 78 complaints were "payment taken, form not submitted"; the fishbone found eight candidate causes in four categories; five whys on the Machinery branch reached a requirement gap about late payment confirmations, the same mechanism as the 72-hour outlier fix of FINDINGS 5.2.2 and Chapter Ninety-One's statistical process control.

Test yourself

1. What is a cause-and-effect diagram, and why is it also called a fishbone diagram? A tool that sorts a team's brainstormed causes of one stated problem into a small number of categories, drawn as branches off a spine leading to the problem; it is called a fishbone diagram because the finished drawing looks like a fish skeleton.

2. Name Ishikawa's 6 M's and say what each covers. Materials (parts, supplies); Machinery (equipment, including software); Methods (procedures and processes); Measurement (indicators and data capture); Manpower (people and their training); Mother Nature (environment and externalities); a less common seventh, Money, is sometimes added.

3. Outline the procedure for building a fishbone diagram. Agree the problem statement; draw the spine with the problem in a box at its head; add the main categories as ribs; brainstorm causes under each category, going deeper by asking why; when ideas run out, concentrate on the categories with fewest of them; then decide which causes to act on.

4. What is the five whys technique, and when is it used? Repeatedly asking "why", usually about five times, to follow a symptom down to a specific root cause; used to push a team past a first answer, often together with a cause-and-effect diagram, when a deeper cause is wanted than a one-line brainstorm entry gives.

5. In the worked example, which branch did the five whys follow, and what did it find? The Machinery branch's first cause, that the payment gateway's callback can arrive after the form's session has timed out. Five whys traced it past the gateway itself to the session timer's length and, beneath that, to a requirement that never specified what should happen if confirmation arrived late.

munotes.in606

Cause-Effect Diagrams

6. How does the five whys' finding connect to release 2.0's repair-time data? FINDINGS 5.2.2 records an after-release fix that took 72 hours because it waited on the payment gateway's provider to change its side of the callback, which Chapter Ninety-One's control chart flagged as a special cause. The five whys, built independently from a structured brainstorm, arrived at the same mechanism, two different methods agreeing on the same weak point.

Contents This chapter on its own page

munotes.in607

Chapter One Hundred Six

Scatter Diagrams

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Scatter Diagrams"

In one line

A scatter diagram plots one variable against another to show whether they move together, letting a team see a relationship, and often its shape and its exceptions, before running any statistic; the diagram can only show that two things are associated, never that one causes the other.

In the wording a student can write in an examination: a scatter diagram (also called a scatter plot or X-Y graph) "graphs pairs of numerical data, with one variable on each axis, to look for a relationship between them. If the variables are correlated, the points will fall along a line or curve. The better the correlation, the tighter the points will hug the line" (ASQ). The NIST/SEMATECH handbook states its purpose as checking "for Relationship": a scatter plot "reveals relationships or association between two variables", and can show whether two variables are related, whether that relation is linear or not, whether the spread of one changes with the other, and whether there are outliers. Both sources give the same warning in different words: NIST states plainly that "There is no statistical procedure, the scatter plot included, that proves cause-and-effect", and ASQ warns "do not assume that one variable caused the other. Both may be influenced by a third variable."

Building one, and testing what it shows

ASQ's procedure: collect paired data where a relationship is suspected; draw the independent variable on the horizontal axis and the dependent variable on the vertical; plot a point, or a touching pair of points, for every pair of values; then look at the pattern. If a line or curve is clear, ASQ says a team "may stop because variables are correlated" and "may wish to use regression or correlation analysis now." If it is not clear, ASQ gives a manual significance test: split the points into four quadrants at their medians, add the diagonally opposite quadrants' counts to get two sums, take the smaller (Q) and the total (N), and compare Q against a trend-test table for that N; a small enough Q says the pattern is unlikely to be chance. This chapter takes the first path ASQ names, computing a correlation coefficient by program, because release 2.0's data is exactly the "clear pattern, check it numerically" case.

Worked example: does the day 6 spike track logins, or something else?

Chapter One Hundred Three, on the seven basic quality tools, left a question open: the help desk's complaints jumped on day 6 of the first registration week, and it could not say from the check sheet alone whether the portal simply failed more under that day's load, or whether more students were merely using it that day. A scatter diagram of that day's logins against that day's complaints, across all six days, answers the first half; fitting the first five days alone and checking what they would have predicted for day 6 answers the second.

munotes.in608

Scatter Diagrams

from statistics import correlation

# release 2.0's first registration week: daily logins (FINDINGS 5.14) and help desk
# complaints (FINDINGS 5.11), by day; day 6 is the last day before the late fee starts
days = [1, 2, 3, 4, 5, 6]
logins =     [430, 390, 470, 400, 460, 900]
complaints = [8,   8,   7,   11,  12,  32]

print(f"{'day':>3}{'logins':>8}{'complaints':>12}{'per 100 logins':>16}")
for d, lg, cp in zip(days, logins, complaints):
    print(f"{d:>3}{lg:>8}{cp:>12}{100 * cp / lg:>16.2f}")

r_all = correlation(logins, complaints)
r_five = correlation(logins[:5], complaints[:5])
print(f"\nPearson's r, all six days: {r_all:.3f}")
print(f"Pearson's r, days 1 to 5 only: {r_five:.3f}")

# the least-squares line through days 1 to 5, extended to day 6's logins
n = 5
mean_x = sum(logins[:5]) / n
mean_y = sum(complaints[:5]) / n
slope = (sum((lg - mean_x) * (cp - mean_y) for lg, cp in zip(logins[:5], complaints[:5]))
         / sum((lg - mean_x) ** 2 for lg in logins[:5]))
intercept = mean_y - slope * mean_x
predicted_day6 = intercept + slope * logins[5]
print(f"\ndays 1 to 5 alone predict {predicted_day6:.1f} complaints at day 6's {logins[5]} logins;"
      f" day 6 actually had {complaints[5]}")
day  logins  complaints  per 100 logins
  1     430           8            1.86
  2     390           8            2.05
  3     470           7            1.49
  4     400          11            2.75
  5     460          12            2.61
  6     900          32            3.56

Pearson's r, all six days: 0.965
Pearson's r, days 1 to 5 only: -0.033

days 1 to 5 alone predict 8.3 complaints at day 6's 900 logins; day 6 actually had 32
A scatter diagram of daily complaints against daily logins: five ordinary days clustered near 400 logins, a dashed near-flat line fitted through them, and day 6 far above and to the right, well clear of the point that line predicts at its login count

Figure 106.1 Release 2.0's first registration week: complaints against logins, and what days 1 to 5 alone would have predicted for day 6

All six days look strongly related. Pearson's r across all six is 0.965, close to the ASQ description of points that "hug the line" tightly. A team that stopped there could read it as confirmation that complaints simply track traffic.

Without day 6, there is nothing to see. The same statistic on days 1 to 5 alone is minus 0.033, no relationship at all. Those five days' logins move within a narrow band, 390 to 470, which is exactly the case ASQ's considerations warn about: "consider whether the independent (x-axis) variable has been varied widely. Sometimes a relationship is not apparent because the data do not cover a wide enough range." The high overall r is not five ordinary days confirming a pattern; it is one distant point, high on both axes, dominating a statistic computed from only six.

What days 1 to 5 would have predicted, against what happened. Fitting days 1 to 5 alone and reading off their line at day 6's login count gives 8.3 complaints. Day 6 had 32, nearly four times that. If day 6 were only a bigger version of an ordinary day, this is roughly what it would have produced instead.

munotes.in609

Scatter Diagrams

Association, and the explanation a scientist supplies. NIST is explicit that a scatter plot can never prove cause and effect, and that "it is ultimately only the researcher (relying on the underlying science/engineering) who can conclude that causality actually exists." The statistic here only shows that day 6 breaks the pattern the other five days set; it does not say why. Chapter One Hundred Five, on cause-effect diagrams, already supplied a mechanism that fits: the gateway's confirmation callback racing the form's session timer, a race more likely to be lost as concurrent load rises near the deadline. The complaint rate of 3.56 per 100 logins on day 6, against 1.49 to 2.75 across the other five, is consistent with a rate that worsens under load, not merely a count that grows with it; the scatter diagram supplies the association, that chapter's fishbone and five whys supply the engineering reason a reader can trust it.

What it does not mean

A strong overall correlation is not evidence by itself. With only six points, one extreme pair can produce a high r even when the rest show nothing; always check what the statistic looks like with that point removed, as Chapter Ninety-One's control chart did with its own 72-hour outlier.

No correlation in a narrow range does not mean no relationship exists. Days 1 to 5's flat line reflects a login count that barely changed across those days; it says nothing about what would happen at day 6's volume, which is exactly why day 6 is worth plotting rather than assumed away.

Correlation is not causation, however strong it looks. Both logins and complaints could rise together on day 6 because a third factor, the approaching deadline, drives both independently; the scatter diagram cannot rule that out by itself.

A relationship found here does not transfer to every release. Release 2.1's fix for the gateway callback (were one made) would need its own week of data before the same chart could be trusted again.

Quick revision

  • Scatter diagram (ASQ; also scatter plot, X-Y graph): plots paired data, one variable per axis, to look for a relationship; a tighter hug to a line or curve means a stronger correlation.
  • Procedure (ASQ): plot the pairs; if a pattern is clear, use regression or correlation analysis; otherwise split into quadrants at the medians and compare Q, the smaller diagonal sum, against N on a trend-test table.
  • What it answers (NIST/SEMATECH): whether two variables are related, linearly or not, whether the spread of one depends on the other, and whether there are outliers.
  • Not proof of cause: NIST states directly that no statistical procedure, the scatter plot included, proves cause and effect; ASQ warns that a third variable may drive both.
  • Worked example: all six days, r = 0.965; days 1 to 5 alone, r = minus 0.033; their line predicts 8.3 complaints at day 6's login count, against 32 actual, consistent with the gateway-callback mechanism Chapter One Hundred Five's cause-effect diagram found worsening under load.
munotes.in610

Scatter Diagrams

Test yourself

1. What is a scatter diagram, and what does a tighter clustering of points along a line mean? A plot of paired data, one variable on each axis, used to look for a relationship between them. Points that hug a line or curve tightly indicate a stronger correlation; a wide scatter indicates a weak one.

2. Outline ASQ's procedure for testing whether a scatter diagram's pattern is real. Plot the pairs. If a line or curve is obvious, correlation or regression analysis may be used directly. Otherwise, divide the points into four quadrants at their medians, sum the counts in diagonally opposite quadrant pairs, take the smaller sum (Q) and the total (N), and compare Q against a trend-test table for that N to judge whether the pattern could be chance.

3. Why can a scatter diagram never prove that one variable causes another? Because association only shows that two variables move together; a scatter diagram cannot rule out a third variable driving both, or the causation running the other way. NIST states this directly: no statistical procedure, the scatter plot included, proves cause and effect.

4. In the worked example, why did the correlation change so much between all six days and the first five alone? Across all six days, one point, day 6, is far higher on both axes than the rest and dominates the statistic, giving r = 0.965. Among the first five days alone, logins vary little and show no relationship to complaints, r = minus 0.033; the apparent strong correlation came from a single point, not from five days confirming a pattern.

5. What did fitting days 1 to 5 alone predict for day 6, and what actually happened? Their least-squares line predicted about 8.3 complaints at day 6's login count. Day 6 actually had 32 complaints, nearly four times the prediction, showing day 6 was not simply a larger version of an ordinary day.

6. How does Chapter One Hundred Five's cause-and-effect work relate to this chapter's finding? The scatter diagram shows only that day 6 breaks the pattern the other days set, an association; it cannot say why. The five whys of Chapter One Hundred Five identified a specific mechanism, the gateway's callback racing the session timer under load, that plausibly explains why the complaint rate, not just the count, rises on the highest-traffic day.

Contents This chapter on its own page

munotes.in611

Chapter One Hundred Seven

Run Charts

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Run charts"

In one line

A run chart plots data in the order it happened against its own median, and four simple rules, about shifts, trends, the number of runs and standout points, turn that picture into an objective check for patterns too small to trust by eye but too real to be chance.

In the wording a student can write in an examination: a run chart is "a graphical display of data plotted in some type of order", usually time, against a median centreline rather than calculated limits (Perla, Provost and Murray). Its advantage, in the authors' words, is that "it preserves the time order of the data", unlike a significance test that only compares separate, already-aggregated summaries: the same mean and standard deviation can come from a process that improved and held, one that had already improved before a change, or one that improved and slid back, and only the time-ordered picture tells them apart. Four rules find non-random patterns against the median: Rule 1, shift, "Six or more consecutive points either all above or all below the median"; Rule 2, trend, "Five or more consecutive points all going up or all going down"; Rule 3, runs, too few or too many crossings of the median line, judged against tabled critical values; and Rule 4, astronomical point, a point "obviously, even blatantly, different from the rest", which is "subjective" where the first three are "probability based".

Why a median, and why these four rules

The median, not a calculated limit, is the run chart's centreline, for two reasons Perla, Provost and Murray give: "it provides the point at which half the observations are expected to be above and below the centreline" and "the median is not influenced by extreme values in the data." That second reason matters directly for ExamReg: a control chart's limits move if the data used to set them contains an outlier, which is why Chapter Ninety-One had to remove the two 503 responses before trusting its limits at all. A median barely moves.

The three probability-based rules (shift, trend, runs) are built for exactly this kind of small-sample question: is a pattern real, or could ordinary chance have produced it, at about a 5 per cent risk of a false alarm. Rule 4 is different by design, a place for judgement where the other three are silent; the authors note it should not be confused with a chart's highest or lowest point, which every run chart has whether or not anything is wrong.

Worked example: the fee page's 38 served times, run-ordered

Chapter Ninety-One's control chart found the fee page's 38 served times in control: every point inside 254.2 to 848.7 ms, no moving range too large, no Western Electric rule fired, and it promised that this chapter would put the same data to a different, more sensitive test. The program runs that test: the same 38 times Chapter Forty-Five's load testing recorded, in the order they were sent, against their own median.

munotes.in612

Run Charts

from statistics import median

# Chapter Forty-Five's fee page load test, all forty requests (elapsed ms, response code);
# the two 503s are the special causes Chapter Ninety-One's control chart already removed
FEE_PAGE = [(380, 200), (410, 200), (417, 200), (463, 200), (454, 200), (516, 200), (491, 200),
            (569, 200), (528, 200), (622, 200), (565, 200), (675, 200), (602, 200), (728, 200),
            (639, 200), (431, 200), (676, 200), (484, 200), (413, 200), (537, 200), (450, 200),
            (590, 200), (487, 200), (643, 200), (524, 200), (696, 200), (561, 200), (749, 200),
            (598, 200), (452, 200), (635, 200), (505, 200), (672, 200), (558, 200), (409, 200),
            (120, 503), (446, 200), (664, 200), (95, 503), (717, 200)]
times = [ms for ms, code in FEE_PAGE if code == 200]

m = median(times)
print(f"n = {len(times)}, median = {m}")

# Rule 1, shift: 6 or more consecutive points all above, or all below, the median
def longest_shift(values, m):
    best, side, run = 0, None, 0
    for v in values:
        if v == m:
            continue                          # on the median: skip, neither breaks nor extends
        s = v > m
        if s == side:
            run += 1
        else:
            side, run = s, 1
        best = max(best, run)
    return best

shift = longest_shift(times, m)
print(f"rule 1, shift: longest run on one side of the median = {shift} "
      f"({'SIGNAL' if shift >= 6 else 'no signal'})")

# Rule 2, trend: 5 or more consecutive points all rising, or all falling (repeats do not count)
def longest_trend(values):
    best, direction, run = 0, None, 0
    prev = None
    for v in values:
        if prev is not None and v != prev:
            d = v > prev
            if d == direction:
                run += 1
            else:
                direction, run = d, 2         # the pair that started the new direction
            best = max(best, run)
        prev = v
    return best

trend = longest_trend(times)
print(f"rule 2, trend: longest run rising or falling = {trend} "
      f"({'SIGNAL' if trend >= 5 else 'no signal'})")

# Rule 3, runs: count crossings of the median, plus one; compare with table 1 (n = 38: 14 to 26)
def count_runs(values, m):
    sides = [v > m for v in values if v != m]
    return 1 + sum(a != b for a, b in zip(sides, sides[1:]))

runs = count_runs(times, m)
lower, upper = 14, 26
verdict = "too few" if runs < lower else "too many" if runs > upper else "within range"
print(f"rule 3, runs: {runs} runs (table 1 for n=38: {lower} to {upper}, {verdict})")

# Rule 4, astronomical point: subjective; none of the 38 stands out as obviously different
print("rule 4, astronomical point: none of the 38 is obviously unlike the rest, by eye")

first_seven = times[:7]
print(f"\nfirst seven times: {first_seven}")
print(f"all seven below the median: {all(t < m for t in first_seven)}")
munotes.in613

Run Charts

n = 38, median = 547.5
rule 1, shift: longest run on one side of the median = 7 (SIGNAL)
rule 2, trend: longest run rising or falling = 4 (no signal)
rule 3, runs: 18 runs (table 1 for n=38: 14 to 26, within range)
rule 4, astronomical point: none of the 38 is obviously unlike the rest, by eye

first seven times: [380, 410, 417, 463, 454, 516, 491]
all seven below the median: True
A run chart of the fee page's 38 served times against their median of 547.5 ms: the first seven, shaded, all fall below the median, a shift by Rule 1; the rest cross back and forth normally

Figure 107.1 The fee page's 38 served times, run-ordered against their median: the first seven are a shift the control chart did not flag

Rule 1 fires. The first seven served times are 380, 410, 417, 463, 454, 516 and 491 ms, all seven below the median of 547.5. Seven consecutive points on one side clears Rule 1's threshold of six, a signal at about the 5 per cent risk level the rule is built for.

Rules 2 and 3 do not. The longest rising or falling run is four points, one short of Rule 2's five. The 38 times cross the median 18 times, inside table 1's range of 14 to 26 for a sample this size, so Rule 3 finds nothing unusual about how often the line changes sides. Rule 4 finds no single point that stands out from the rest by eye; the two extreme low values are the 503s, already removed as special causes, not something the run chart itself had to catch.

A signal the control chart missed. Chapter Ninety-One's individuals chart judged the same 38 points in control: every one inside its limits, no rule of its own fired. A control chart's rules are built around distance from the centre and the size of successive differences; a run that stays on one side of the median without ever straying far from it can pass every one of those tests and still not be random. That is what "a different, more sensitive test" meant: not a better chart, a chart built to catch a different shape of pattern.

A plausible reading, not a proven one. A run of low times at the very start of a load test is a familiar shape in practice: the first few requests can benefit from a warm cache, a freshly opened connection, or a server not yet carrying its full test load, before the run settles into whatever the rest of the test looks like. Rule 1 only says the first seven are unlikely to be chance; it does not say why, and this book's data gives no independent check, the way Chapter One Hundred Six's scatter diagram checked the seven basic quality tools' open question of Chapter One Hundred Three with a second dataset. A real investigation would rerun the test and watch whether the first few requests are low again.

munotes.in614

Run Charts

The run chart against the control chart

Control chart (individuals)Run chart
CentrelineMean of the dataMedian of the data
What it needsEnough data for a stable mean and moving range; sensitive to outliers while computing limitsUsable from a handful of points; the median resists outliers
What it catchesPoints too far from the centre, or patterns in how far successive points differLong runs on one side of the median, long rises or falls, too few or too many crossings, one obviously odd point
What it does not catch wellA run that stays close to the centre but consistently to one sideThe size of a shift, or how far outside a normal range a point falls
This chapter's data38 points in control: every one inside 254.2 to 848.7 msThe same 38 points: a shift, the first seven all below the median

Neither chart is the better one in general; they are built to catch different shapes of non-random pattern, and Perla, Provost and Murray note that a run chart is often the first, simplest display drawn before a more demanding chart like this book's individuals chart is built at all.

What it does not mean

A shift is not a magnitude. Rule 1 says seven points are unlikely to be on one side of the median by chance; it says nothing about how far below the median they are, which is a question for the control chart's limits, not the run chart's rules.

Passing three rules is not proof of nothing wrong. Rule 4 is deliberately subjective; a chart can pass every counted rule and still show something a reader's judgement catches that no rule was built to count.

The median rule needs enough points. Perla, Provost and Murray state plainly that "The shift and run rules require more than 10 points before they are applicable"; a run chart with five or six points can still be drawn and read by eye, but Rules 1 and 3 should not be applied to it.

A run chart signal is a reason to look, not a finished explanation. This chapter's shift is real by the rule's own test; the warm-start explanation offered for it is plausible, not established, exactly the caution Chapter One Hundred Six gave its own scatter diagram's association.

munotes.in615

Run Charts

Quick revision

  • Run chart (Perla, Provost and Murray, 2011): data plotted in time order against the median; preserves order that a summary statistic destroys.
  • Median as centreline: half the points expected above and below it; unlike a mean, not pulled by extreme values.
  • Rule 1, shift: 6 or more consecutive points all on one side of the median.
  • Rule 2, trend: 5 or more consecutive points all rising or all falling; repeats do not count.
  • Rule 3, runs: too few or too many crossings of the median, against table 1's critical values for the sample size.
  • Rule 4, astronomical point: a point obviously unlike the rest; subjective, unlike the first three.
  • Worked example: the fee page's 38 served times, in control by Chapter Ninety-One's control chart, fail Rule 1: the first seven are all below the median of 547.5, a shift the control chart's own rules did not raise.

Test yourself

1. What is a run chart, and what is its main advantage over a summary statistic? Data plotted in the order it occurred, usually against a median centreline. Its advantage is that it preserves time order: the same mean and standard deviation can arise from a sustained improvement, an improvement that had already started, or one that did not hold, and only the time-ordered chart tells these apart.

2. Why does a run chart use the median rather than the mean as its centreline? Because the median is the point half the observations are expected to fall above and below, and, unlike the mean, it is not pulled by extreme values in the data, so a single outlier does not shift where "normal" is drawn.

3. State Rules 1 and 2 and what each detects. Rule 1, shift: six or more consecutive points all above or all below the median signals a shift in the process. Rule 2, trend: five or more consecutive points all rising or all falling signals a trend; repeated equal values count once and do not break the run.

4. In the worked example, which rule fired, and on what evidence? Rule 1, shift. The first seven of the fee page's 38 served times, 380 to 491 ms, all fall below the median of 547.5 ms, clearing Rule 1's threshold of six consecutive points on one side.

5. Why could a control chart judge the same data in control while the run chart found a signal? A control chart's rules are built around distance from the centreline and the size of successive differences; a run that stays close to the centre but consistently on one side of it can satisfy those rules while still failing a rule built specifically to detect runs relative to the median.

munotes.in616

Run Charts

6. Why is the warm-start explanation for the shift described as plausible rather than proven? Because Rule 1 only shows that seven points in a row below the median is unlikely to be chance; it gives no reason on its own. The warm-cache or fresh-connection explanation fits common experience of load tests, but nothing in this book's data independently confirms it, so it remains a reading of the signal, not a demonstrated cause.

Contents This chapter on its own page

munotes.in617

Chapter One Hundred Eight

What the Examination Asks, and How to Answer It

Syllabus topic Not a module topic: MU's evaluation scheme for this course, pages 104 and 105 of the syllabus circular, read against this book's own structure.

In one line

This course is examined in a one-hour paper worth 30 marks, three questions of four parts each, two parts answered per question at five marks apiece, plus 20 marks of internal assessment; every chapter in this book was written to stand alone as one of those five-mark answers, so revising the book chapter by chapter is revising for the paper directly.

In the wording a student can write about the scheme itself: the course carries 50 marks, 20 internal and 30 external. Internal assessment is two class tests, one on each module, 10 marks each, averaged to 10, plus an assignment on each module, 5 marks each, totalling 10: 20 marks. The external paper is a Semester End Examination of 1 hour for 30 marks: Q.1 on Module 1, Q.2 on Module 2, and Q.3 on both modules together, each question offering four parts headed "Answer any 2 of the following", worth 10 marks. Twelve parts are printed; six are answered; each is worth five marks.

The scheme, exactly as printed

ComponentMarks
Internal (20)Class Test 1, Module 110, averaged with Class Test 2 to 10
Class Test 2, Module 210, averaged with Class Test 1 to 10
Assignment, Module 15
Assignment, Module 25
External (30), 1 hourQ.1, Module 1, any 2 of 410
Q.2, Module 2, any 2 of 410
Q.3, Modules 1 and 2, any 2 of 410

The paired practical, Software Testing and Quality Assurance Practical, is a separate paper with its own 50 marks and its own two-hour examination; this book, and this scheme, cover the theory paper only.

What a five-mark answer contains

A paper of twelve parts with six answered in one hour gives about ten minutes to a part: not long enough for everything a topic could say, long enough for a complete answer if it has a shape. Every chapter in this book was written to that shape, and it is the shape a five-mark answer wants:

  1. The definition, in examination wording. Each chapter's second paragraph, opening with the words in the wording a student can write in an examination, is written to be reproduced close to as printed: a clear definition, then its key terms, each in the source's own words.
  2. The substance. Two or three points that say what the definition means and why it matters, not a restatement of the definition in different words.
  3. One worked thing. A short example, a small computation, a table, or a named illustration, the difference between an answer that defines a term and one that shows it understood.
  4. A closing line. One sentence tying the answer back to the question asked, which is what a chapter's "Quick revision" box already compresses each topic into.
munotes.in618

What the Examination Asks, and How to Answer It

Worked example: turning a chapter into a five-mark answer

Chapter One Hundred Four, on Pareto diagrams, is one topic among many; the program below only checks the arithmetic of the scheme itself, the same arithmetic for every topic.

# MU's evaluation scheme for this 2-credit theory course (authorities/mu-software-testing-syllabus.txt)
internal = {"class tests, averaged": 10, "two assignments": 10}
external = {"Q.1, Module 1": 10, "Q.2, Module 2": 10, "Q.3, Modules 1 and 2": 10}

print("internal assessment")
for item, marks in internal.items():
    print(f"   {item:<24}{marks:>3} marks")
print(f"   {'total':<24}{sum(internal.values()):>3} marks")

print("external examination, 1 hour")
for item, marks in external.items():
    print(f"   {item:<24}{marks:>3} marks")
print(f"   {'total':<24}{sum(external.values()):>3} marks")

total = sum(internal.values()) + sum(external.values())
print(f"\ncourse total: {total} marks")

# each external question offers 4 parts; the student answers any 2, each worth 5 marks
parts_per_question, answered_per_question = 4, 2
questions = len(external)
printed = questions * parts_per_question
answered = questions * answered_per_question
mark_each = external["Q.1, Module 1"] // answered_per_question
print(f"\n{questions} questions x {parts_per_question} parts printed = {printed} parts on the paper")
print(f"{questions} questions x {answered_per_question} answered = {answered} answered, "
      f"{mark_each} marks each = {answered * mark_each} marks")

# the book's own shape: which module carries which question
modules = {"I": (51, "Q.1"), "II": (57, "Q.2, and with Module I, Q.3")}
for name, (chapters, carries) in modules.items():
    print(f"Module {name}: {chapters} chapters, carries {carries}")
print(f"book total: {sum(c for c, _ in modules.values())} chapters")
internal assessment
   class tests, averaged    10 marks
   two assignments          10 marks
   total                    20 marks
external examination, 1 hour
   Q.1, Module 1            10 marks
   Q.2, Module 2            10 marks
   Q.3, Modules 1 and 2     10 marks
   total                    30 marks

course total: 50 marks

3 questions x 4 parts printed = 12 parts on the paper
3 questions x 2 answered = 6 answered, 5 marks each = 30 marks
Module I: 51 chapters, carries Q.1
Module II: 57 chapters, carries Q.2, and with Module I, Q.3
book total: 108 chapters

Applying the shape to one topic. Take Chapter One Hundred Four's own material: definition, a Pareto chart ranks categories by frequency or cost, largest to smallest, with a cumulative line; substance, it rests on the 80-20 rule, Juran's vital few against the useful many, and can be weighted by cost rather than count; worked thing, release 2.0's 200 defects, four types making 78 per cent by count, the ranking changing when weighted by hours to fix; closing line, a Pareto chart ranks, it does not explain, which is why a cause-and-effect diagram comes next. Read aloud at an unhurried pace, that is a five-mark, ten-minute answer, and it is already sitting in the chapter's own "In one line", "Quick revision" and worked example, not invented for the exam.

munotes.in619

What the Examination Asks, and How to Answer It

Q.3: where the modules cross

Q.3 draws from both modules in one answer, and this syllabus was built to allow it: three places where a Module 1 topic and a Module 2 topic are, in substance, the same idea taught twice, once in general and once inside software quality assurance specifically.

Module 1Module 2The crossing
Chapters 28 to 30: reviews, inspection, walkthroughChapters 97 to 98: formal technical reviewsA review is the general mechanism; a formal technical review is that mechanism run to a defined procedure, with defect logging and a rework decision.
Chapters 24 to 25: quality control, quality assurance, quality managementChapters 84 to 89: software quality assurance's background, challenges, activities, plan, approaches and statisticsModule 1 defines QA in general; Module 2 is QA applied to a software project specifically: its own plan, its own audits, its own statistical tools.
Chapters 33 to 36: unit testing and its techniquesChapters 59 to 60: statement and branch testingUnit testing is the test level; statement and branch coverage are structural techniques for deciding what a unit test must exercise, most often used at that same level.

A Q.3 answer on any of these three pairs draws one part from each module's chapter and states the relationship in a sentence, which is the "crossing" itself, before giving each side's own five-mark content.

Internal assessment: class tests and assignments

Two class tests, each confined to one module, are the same five-mark shape at twice the length: a definition, substance and one worked thing for each of two topics, in the time a class test allows. An assignment rewards depth on fewer topics than a class test or the paper can reach: a full worked example carried further than this book's own, with its own numbers, checked the way every number in this book was checked, by running it.

A revision order

The scheme itself sets the order, not a guess at what will be asked. Module 1's fifty-one chapters answer Q.1; revise them first, using each chapter's "Quick revision" box as the pass. Module 2's fifty-six syllabus chapters answer Q.2 (this closing chapter is about the exam, not Module 2's content, and is one more); revise them the same way, second. Then revisit the three crossings above for Q.3, since they are the one place a single answer must hold two chapters' material at once. Internal assessment needs nothing beyond the module it tests, done in the same pass as that module.

What it does not mean

Nothing in this book, or in this chapter, predicts a question. The scheme says how many marks, how many parts, and which modules a question draws on; it does not say which topic within a module will be asked, and no pattern across past papers changes what this book teaches or how much space a topic gets. A claim that a topic is "frequently asked" is not made here, because it would not be true in any sense a program could check.

munotes.in620

What the Examination Asks, and How to Answer It

A revision order is not a priority order. Revising Module 1 first is not a claim that Module 1 matters more; the order only follows the paper's own Q.1, Q.2, Q.3 structure.

The five-mark shape is a starting discipline, not a ceiling. A student with time to write more may; the shape describes what a complete answer needs in ten minutes, not the most a strong answer could say.

The practical paper has its own scheme. Two hours, 50 marks, a different structure entirely; this chapter, like this book, covers the theory paper only.

Quick revision

  • Scheme: 50 marks total, 20 internal (class tests averaged to 10, two assignments totalling 10), 30 external.
  • External paper: 1 hour; Q.1 (Module 1), Q.2 (Module 2), Q.3 (Modules 1 and 2), each "any 2 of 4" parts, 10 marks; 12 parts printed, 6 answered, 5 marks each.
  • A five-mark answer: definition in exam wording, two or three points of substance, one worked thing, a closing line, the shape every chapter in this book already has.
  • Q.3's three crossings: reviews/inspection/walkthrough with formal technical reviews; quality and QA in general with software quality assurance specifically; unit testing with statement and branch testing.
  • Revision order: Module 1 for Q.1, Module 2 for Q.2, then the three crossings for Q.3; nothing here predicts a question.

Test yourself

1. What are this course's total marks, and how do they split between internal and external assessment? 50 marks: 20 internal (two class tests averaged to 10, two assignments totalling 10) and 30 external (a one-hour Semester End Examination).

2. Describe the structure of the external paper. Three questions, Q.1 on Module 1, Q.2 on Module 2, Q.3 on both modules, each offering four parts headed "Answer any 2 of the following" for 10 marks; twelve parts printed in all, six answered, five marks each.

3. What four things does this book say a five-mark answer should contain? A definition in examination wording, two or three points of substance explaining what the definition means, one worked example or small computation, and a closing line tying the answer to the question.

4. Name the three places this book's Module 1 and Module 2 content cross, as Q.3 might draw on them. Reviews, inspection and walkthrough (Module 1) with formal technical reviews (Module 2); quality control, assurance and management in general (Module 1) with software quality assurance specifically (Module 2); unit testing and its techniques (Module 1) with statement and branch testing (Module 2).

munotes.in621

What the Examination Asks, and How to Answer It

5. In what order does this chapter suggest revising the book, and why that order? Module 1 first, for Q.1; Module 2 second, for Q.2; then the three Q.3 crossings. The order follows the external paper's own Q.1, Q.2, Q.3 structure, not a judgement about which module matters more.

6. Why does this book refuse to say a topic is "frequently asked"? Because the evaluation scheme fixes how many marks and how many parts a question carries, and which modules it draws on, but not which topic within a module is chosen; a frequency claim about topics would be a prediction this book cannot check by running a program, so it is not made.

Contents This chapter on its own page

munotes.in622

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!