munotes®

Why Reviews Pay: The Cost of a Late Defect

Get access to whole semester resourcesSemester Pass

Chapter Ninety-Six

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: Formal Technical Reviews and their benefits"

Pages 557 to 561 of 622

In one line

A defect costs more to fix the later it is found, often a hundred times more after delivery than during requirements and design, and a defect that escapes early leads to more defects in the work built on it; reviews find defects early, so they save far more than they cost, and ExamReg's own data show both effects.

In the wording a student can write in an examination: Boehm and Basili (2001) report that "Finding and fixing a software problem after delivery is often 100 times more expensive than finding and fixing it during the requirements and design phase", while for small, noncritical systems the ratio is "more like 5:1 than 100:1". Defect amplification is the name this book gives to the effect the ISTQB syllabus describes: "Defects in work products produced earlier in the SDLC, if undetected, often lead to defective work products later in the lifecycle." Reviews attack both: they find defects early, when fixes are cheap, and they stop defects propagating; Boehm and Basili report that peer reviews catch "from 31 to 93 percent of the defects, with a median of around 60 percent."

The cost of a late defect

The earliest chapters of this book cited the finding; this chapter uses it. Boehm and Basili put it first in their list of what the evidence shows about reducing defects: "Finding and fixing a software problem after delivery is often 100 times more expensive than finding and fixing it during the requirements and design phase." They qualify it carefully. The word "often" was added deliberately, because "the cost-escalation factor for small, noncritical software systems" is "more like 5:1 than 100:1", and because "good architectural practices can significantly reduce the cost-escalation factor even for large critical systems." Fagan, reporting IBM's inspection results (Chapter Twenty-Nine, on inspection), gave a similar range: rework at the early levels "is 10 to 100 times less expensive than if it is done in the last half of the process."

Why should a late fix cost so much more? A defect found in a requirements review is corrected in one document. The same defect found after release has been designed around, coded, tested and shipped: the fix must change the requirement, the design, the code and the tests, retest everything they touch, redeploy, and deal with the users it affected.

The finding also has a cost in effort. Boehm and Basili report that "Current software projects spend about 40 to 50 percent of their effort on avoidable rework", effort "spent fixing software difficulties that could have been discovered earlier and fixed less expensively or avoided altogether", and that "About 80 percent of avoidable rework comes from 20 percent of the defects."

munotes.in557

Why Reviews Pay: The Cost of a Late Defect

Defect amplification

A defect in an early work product does not stay one defect. The ISTQB syllabus states it in two places: "Defects in work products produced earlier in the SDLC, if undetected, often lead to defective work products later in the lifecycle", and, as its third principle of testing, "Defects that are removed early in the process will not cause subsequent defects in derived work products." A wrong fee rule in the requirements becomes a wrong design, wrong code, wrong tests that expect the wrong fee, and a wrong user guide. This book calls the effect defect amplification: each work product built on a defective one can multiply the defect.

Worked example: what release 2.0's reviews saved

Release 2.0's 200 defects were recorded with the work product each was in and the activity that found it; its three reviews found 74 of them. The program asks what fixing the defects cost, in units of one requirements-review fix, and what it would have cost if the reviews had not been held. Two assumptions are needed, and both are stated. Without reviews, each work product's review finds are taken to have been found later, spread in proportion to where that product's other defects were found. And the cost of a fix at each activity rises from 1 at the requirements review to Boehm and Basili's end ratio after release, evenly on a logarithmic scale between; they give only the end ratio, so the program shows both of theirs, 100:1 and 5:1. Finally, it counts the requirements and design defects that reached the code, the ones that could amplify.

# release 2.0's 200 defects: where each was made, and the activity that found it (FINDINGS 5.2.6)
STAGES = ["req review", "design review", "code review", "unit", "integration", "system",
          "acceptance", "after release"]
found = {"requirements": [16, 3, 1, 1, 1, 2, 3, 1],
         "design":       [0, 21, 5, 3, 4, 4, 1, 2],
         "code":         [0, 0, 28, 40, 26, 20, 3, 3],
         "documents":    [0, 0, 0, 0, 1, 2, 3, 6]}
REVIEWS = 3                                     # the first three activities are reviews

def without_reviews(row):
    """Move a row's review finds to the later activities, in proportion to its own later finds."""
    moved, later = sum(row[:REVIEWS]), row[REVIEWS:]
    return [0] * REVIEWS + [n + moved * n / sum(later) for n in later]

def cost(rows, factor):
    return sum(n * f for row in rows for n, f in zip(row, factor))

for ratio in (100, 5):                          # the cost of a fix after release : in requirements
    factor = [ratio ** (k / (len(STAGES) - 1)) for k in range(len(STAGES))]   # even on a log scale
    with_r = cost(found.values(), factor)
    without_r = cost((without_reviews(row) for row in found.values()), factor)
    print(f"cost ratio {ratio}:1 -> fixes cost {with_r:,.0f} units with reviews,"
          f" {without_r:,.0f} without ({without_r / with_r:.2f} times)")

# amplification: requirements and design defects that reach the code, with and without reviews
reach_with = (sum(found["requirements"]) - sum(found["requirements"][:2])
              + sum(found["design"]) - found["design"][1])
reach_without = sum(found["requirements"]) + sum(found["design"])
print(f"requirements and design defects reaching the code: {reach_with} with reviews,"
      f" {reach_without} without")
munotes.in558

Why Reviews Pay: The Cost of a Late Defect

cost ratio 100:1 -> fixes cost 3,419 units with reviews, 5,365 without (1.57 times)
cost ratio 5:1 -> fixes cost 456 units with reviews, 576 without (1.26 times)
requirements and design defects reaching the code: 28 with reviews, 68 without

The cost. At Boehm and Basili's 100:1, fixing release 2.0's defects cost 3,419 units with its reviews and would have cost 5,365 without them, 1.57 times as much. Even at the small-system ratio of 5:1, the reviews save a fifth of the fixing cost: 456 units against 576. The saving is large because the reviews found 74 defects at the cheapest end of the scale.

The amplification. With the reviews, 28 requirements and design defects reached the code; without them, all 68 would have. Each of the 40 extra ones would have been built into code, tests and documents, which is the ISTQB statement in numbers: the program's cost figures count each defect once, so they understate what skipping the reviews would have cost.

What the comparison leaves out. The reviews themselves cost effort: 62 person-hours in release 2.0, against 296 for testing, and Chapter Sixty-Eight, on quality, process and test metrics, found them finding 1.19 defects per hour against testing's 0.39. Chapter One Hundred Two, on using quality costs for decision making, puts hours and money on both sides of the decision.

Review metrics

A project that holds reviews should measure them, for the same reason it measures testing. Three measures recur in this book.

MeasureHow it is computedRelease 2.0
Review effectivenessDefects the review found divided by the defects present when it ran (Fagan's error detection efficiency)Requirements review 57.1, design review 46.2, code review 21.2 per cent (Chapter Seventy-Nine, on defect metrics)
Review yield per hourDefects found divided by the person-hours spent1.19 defects per hour across the reviews, against 0.39 for testing (Chapter Sixty-Eight)
Where defects escapeDefects of a work product found after its own review28 requirements and design defects reached the code

Boehm and Basili's range gives a benchmark: "Numerous studies confirm that peer review provides an effective technique that catches from 31 to 93 percent of the defects, with a median of around 60 percent." ExamReg's requirements review, at 57.1 per cent, is near that median; its code review, at 21.2 per cent, is below the whole range. That is the measure doing its job: it points to the code review as the review to strengthen, with the formal procedure of Chapter Ninety-Seven, on formal technical reviews, and the reading techniques Boehm and Basili also report, whose perspective-based form catches "35 percent more defects than nondirected reviews."

munotes.in559

Why Reviews Pay: The Cost of a Late Defect

What it does not mean

The 100:1 ratio is not a constant. Boehm and Basili say "often", report about 5:1 for small, noncritical systems, and note that good architecture reduces the ratio; the principle holds at either ratio.

Reviews do not replace testing. They find different defects: Boehm and Basili report "that peer reviews, analysis tools, and testing catch different classes of defects at different points in the development cycle."

Not all rework is avoidable. Boehm and Basili distinguish avoidable rework from changes that emerge from prototyping and learning, which "should not be discouraged by classifying them as avoidable defects."

A cheap review is not a free one. Reviews cost effort; the case for them is that they cost less than the late fixes they prevent.

Quick revision

  • Late defects (Boehm and Basili 2001): after delivery "often 100 times more expensive" than in requirements and design; "more like 5:1" for small, noncritical systems; Fagan: early rework "10 to 100 times less expensive".
  • Avoidable rework: about 40 to 50 per cent of effort; about 80 per cent of it from 20 per cent of the defects.
  • Defect amplification (this book's name for the ISTQB statement): undetected defects in earlier work products "often lead to defective work products later"; early removal prevents "subsequent defects in derived work products".
  • Peer reviews catch 31 to 93 per cent of defects, median about 60; perspective-based reviews 35 per cent more than nondirected ones.
  • Review metrics: effectiveness (found over present), yield per hour, escapes.
  • Worked example (cost factors spread evenly on a log scale, an assumption): fixing cost 3,419 units with reviews against 5,365 without at 100:1 (1.57 times), 456 against 576 at 5:1; 28 requirements and design defects reached the code with reviews, 68 without.

Test yourself

1. Why does a defect cost more to fix the later it is found? Because by then other work has been built on it: a requirements defect found after release must be corrected in the requirement, the design, the code, the tests and the documents, everything affected retested and redeployed, and the users it affected dealt with. Boehm and Basili report that fixing after delivery is often 100 times as expensive as in the requirements and design phase, about 5 times for small, noncritical systems.

2. What is defect amplification? The effect that a defect left undetected in an early work product leads to defects in the work products derived from it, so that one requirements or design defect becomes several in design, code, tests and documents. The ISTQB syllabus states it, and its principle that early testing saves time and money rests on it.

munotes.in560

Why Reviews Pay: The Cost of a Late Defect

3. How do reviews reduce the cost of defects? They find defects in requirements, designs and code before execution, at the cheapest point to fix them, and they stop those defects propagating into later work products. Peer reviews catch from 31 to 93 per cent of the defects present, with a median of around 60 per cent.

4. In the worked example, what did release 2.0's reviews save? At a 100:1 cost ratio, fixing its defects cost 3,419 units with the reviews against an estimated 5,365 without, a factor of 1.57; at 5:1, 456 against 576. And 28 requirements and design defects reached the code instead of 68, so 40 fewer defects could amplify.

5. Name three review metrics and what each shows. Review effectiveness, defects found over defects present, shows how thoroughly a review works; yield per hour, defects found per person-hour, shows its efficiency; escapes, a work product's defects found after its review, show what it missed.

6. What did the review metrics show about ExamReg's code review, and what follows? Its effectiveness, 21.2 per cent, was below the 31 to 93 per cent range Boehm and Basili report for peer reviews, while the requirements review, at 57.1 per cent, was near the median. The code review should be strengthened, for example with a formal technical review procedure and directed reading techniques.

munotes.in561

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!