munotes®

Defect Metrics

Get access to whole semester resourcesSemester Pass

Chapter Seventy-Nine

Syllabus topic Module 2, "Defect Management: ... Metrics related to defects"

Pages 452 to 457 of 622

In one line

Defect metrics turn a release's defect records into answers: how dense the defects are and where, how many were caught before users met them and by which activity, how many slipped past the review meant to catch them, how long they survived, how many fixes failed, and how much triage effort went on reports that were not defects.

In the wording a student can write in an examination: the main metrics related to defects are defect density, the "number of defects per unit of product size" (ISO/IEC/IEEE 24765); defect removal efficiency (DRE), the defects found before release as a percentage of all defects found, before and after, which the ISTQB syllabus lists among defect metrics as the "defect detection percentage"; the effectiveness of each stage, which Fagan defined for inspections as "Errors found by an inspection" divided by "Total errors in the product before inspection"; leakage, the defects that escape the activity meant to catch them; defect age, how long or how many stages a defect survives; the reopen rate, the fixes that fail their confirmation test; the share of reports that are not defects; and a severity index, a weighted count whose weights are a policy, not a measurement.

One dataset, many questions

Chapter Sixty-Eight, on quality, process and test metrics, computed release 2.0's overall defect density and its defects found per hour of each activity. This chapter goes further into the defects themselves, and it needs one more table: for each of the 200 defects, the work product it was in and the stage that found it.

OriginReq reviewDesign reviewCode reviewUnitIntegrationSystemAcceptanceAfter releaseTotal
Requirements16311123128
Design02153441240
Code002840262033120
Documents0000123612
Total1624344432281012200

The row totals are the problem types of Chapter Seventy-Three, on what a defect is; the column totals are the activities that found them, used since Chapter Twenty-Four, on quality control and quality assurance. The table adds only the link between them.

Removal efficiency, overall and stage by stage

Defect removal efficiency asks the question every release is judged by: of all the defects it contained, what share did the team catch before users did? It can only be computed after release, once users have had time to find what testing missed, and it is always provisional: a defect found next year lowers it.

Stage effectiveness asks the same question of each activity. Fagan defined it for inspections: "Error detection efficiency" is "Errors found by an inspection" divided by the "Total errors in the product before inspection". The denominator is the hard part: the errors present at a stage are those already in the work products it can see, minus those earlier stages removed, and it is known only in hindsight, when later stages have found the rest.

munotes.in452

Defect Metrics

Worked example 1: efficiency, leakage and age

STAGES = ["req review", "design review", "code review", "unit", "integration",
          "system", "acceptance", "after release"]
# where release 2.0's defects came from (rows) and which stage found them (columns): FINDINGS 5.2.6
FOUND = {"requirements": [16, 3, 1, 1, 1, 2, 3, 1],
         "design":       [0, 21, 5, 3, 4, 4, 1, 2],
         "code":         [0, 0, 28, 40, 26, 20, 3, 3],
         "documents":    [0, 0, 0, 0, 1, 2, 3, 6]}
BORN = {"requirements": 0, "design": 1, "code": 2, "documents": 2}   # the first stage that can see them

total = sum(sum(row) for row in FOUND.values())
delivered = sum(row[-1] for row in FOUND.values())
print(f"defect removal efficiency: {total - delivered} of {total} found before release"
      f" = {100 * (total - delivered) / total:.1f}%")

print(f"{'stage':<14}{'present':>8}{'found':>6}{'effectiveness':>15}")
found_so_far = 0
for i, stage in enumerate(STAGES[:-1]):
    present = sum(sum(row) for o, row in FOUND.items() if BORN[o] <= i) - found_so_far
    found = sum(row[i] for row in FOUND.values())
    print(f"{stage:<14}{present:>8}{found:>6}{100 * found / present:>14.1f}%")
    found_so_far += found

print("leakage: defects that escaped the review of the work product they were in")
for origin, stage in (("requirements", 0), ("design", 1), ("code", 2)):
    row = FOUND[origin]
    print(f"   {origin:<13} {sum(row) - row[stage]:>3} of {sum(row):>3} escaped {STAGES[stage]}:"
          f" {100 * (sum(row) - row[stage]) / sum(row):.1f}%")

print("age: how many stages each kind of defect survived before it was found, on average")
for origin, row in FOUND.items():
    stages_survived = sum(n * (i - BORN[origin]) for i, n in enumerate(row))
    print(f"   {origin:<13} {stages_survived / sum(row):.2f}")
defect removal efficiency: 188 of 200 found before release = 94.0%
stage          present found  effectiveness
req review          28    16          57.1%
design review       52    24          46.2%
code review        160    34          21.2%
unit               126    44          34.9%
integration         82    32          39.0%
system              50    28          56.0%
acceptance          22    10          45.5%
leakage: defects that escaped the review of the work product they were in
   requirements   12 of  28 escaped req review: 42.9%
   design         19 of  40 escaped design review: 47.5%
   code           92 of 120 escaped code review: 76.7%
age: how many stages each kind of defect survived before it was found, on average
   requirements  1.68
   design        1.40
   code          1.49
   documents     4.17

Overall. 188 of 200 defects were found before release: a removal efficiency of 94.0 per cent, the figure Chapter Sixty-Six's goal, question, metric model judged against its target of 90.

Stage by stage. The requirements review found 16 of the 28 defects the requirements held, 57.1 per cent. The design review saw the 12 requirements defects that escaped plus the 40 design defects, 52 in all, and caught 24. The code review is the weakest link at 21.2 per cent: 160 defects were present when it ran, most of them in 120 freshly written code units, and it found 34. The four test levels (34.9, 39.0, 56.0 and 45.5 per cent) are exactly Chapter Thirty-One's figures, on a strategic approach to software testing, computed there from the same data; system testing was the most effective single stage of the whole release.

munotes.in453

Defect Metrics

Leakage. Leakage looks at the same data from each work product's side. Of the 28 requirements defects, 12 escaped the requirements review, 42.9 per cent; of the design defects, 47.5 per cent escaped the design review; of the code defects, 76.7 per cent escaped the code review. Every defect that leaks is found later, if at all, by an activity that costs more per defect (Chapter Ninety-Six, on why reviews pay, has the published evidence).

Age. Defects in the documentation survived more than four stages on average, far longer than any other kind: no review looked at the user documents, so they waited for testers and users to trip over them. That single number is a finding with an obvious remedy.

Worked example 2: density by module, reopening, rejection and severity

# release 2.0 (FINDINGS 5.2, 5.2.3, 5.2.4, 5.2.6)
by_module = {"Exam form": (74, 5.5), "Fee payment": (58, 4.0), "Admin reports": (20, 3.5),
             "Hall ticket": (18, 2.5), "Login": (12, 1.5), "Profile": (10, 1.5),
             "Notifications": (8, 1.5)}                         # (defects, KLOC)
print("defect density by module, highest first:")
for module, (d, kloc) in sorted(by_module.items(), key=lambda m: -m[1][0] / m[1][1]):
    print(f"   {module:<14} {d:>3} defects in {kloc:>3} KLOC: {d / kloc:>5.1f} per KLOC")

reports, product, duplicates, rejected = 250, 200, 18, 10
fixes, reopened = 198, 14
print(f"reopen rate: {reopened} of {fixes} fixes = {100 * reopened / fixes:.1f}%")
print(f"reports that were not product defects: {reports - product} of {reports}"
      f" = {100 * (reports - product) / reports:.1f}%"
      f" (duplicates {duplicates}, rejected as mistakes {rejected})")

severity = {"critical": 8, "major": 46, "minor": 98, "cosmetic": 48}
weights = {"critical": 10, "major": 5, "minor": 2, "cosmetic": 1}        # the team's policy
index = sum(severity[s] * weights[s] for s in severity) / sum(severity.values())
print(f"severity index with weights 10/5/2/1: {index:.2f} per defect")
print("   the counts it summarises:", ", ".join(f"{s} {n}" for s, n in severity.items()))
defect density by module, highest first:
   Fee payment     58 defects in 4.0 KLOC:  14.5 per KLOC
   Exam form       74 defects in 5.5 KLOC:  13.5 per KLOC
   Login           12 defects in 1.5 KLOC:   8.0 per KLOC
   Hall ticket     18 defects in 2.5 KLOC:   7.2 per KLOC
   Profile         10 defects in 1.5 KLOC:   6.7 per KLOC
   Admin reports   20 defects in 3.5 KLOC:   5.7 per KLOC
   Notifications    8 defects in 1.5 KLOC:   5.3 per KLOC
reopen rate: 14 of 198 fixes = 7.1%
reports that were not product defects: 50 of 250 = 20.0% (duplicates 18, rejected as mistakes 10)
severity index with weights 10/5/2/1: 2.77 per defect
   the counts it summarises: critical 8, major 46, minor 98, cosmetic 48
munotes.in454

Defect Metrics

Density by module. The exam form had the most defects, 74, but it is also the largest module; per KLOC, the fee payment module is worse, 14.5 against 13.5. Counts say where most of the defects are; densities say where the code is weakest. Both point at the same two modules, which between them held two-thirds of the release's defects.

Reopen rate. 14 of 198 fixes failed their confirmation test, 7.1 per cent. Each was a defect the developer believed fixed; the rate measures how well fixes are tested before they are handed back, and a rising rate is an early warning.

Reports that were not defects. One report in five was a duplicate, a user's or operator's mistake, a test error, an enhancement request or unrepeatable (Chapter Seventy-Five, on the defect management process). The figure measures the load on triage, and a high duplicate share says testers cannot easily search the existing reports.

Severity index. Weighted 10, 5, 2 and 1, release 2.0's defects score 2.77 per defect. The number is only as meaningful as the weights: Chapter Sixty-Five, on what a software metric is, showed that averaging ranked values can reverse a comparison when the codes change, and a severity index is such an average. It is useful as long as its weights are a stated policy (here, the team's rough view of relative cost) and used unchanged from release to release; it is never a measurement of severity, and the counts it summarises are always reported beside it.

What each metric is for

MetricThe question it answersWatch out for
Defect densityHow defective is the product, or each part of it?The size measure and the counting rule for defects
Removal efficiencyWhat share of defects did we catch before users?It needs time after release, and it only ever falls
Stage effectivenessWhich activities catch the defects present when they run?The denominator is known only in hindsight
LeakageWhich reviews let their own kind of defect through?Small counts per work product
AgeHow long do defects survive?Measured in stages or in days: say which
Reopen rateHow often does a fix fail?Count reopenings of the same defect once or each time
Reports not defectsHow much triage effort goes on non-defects?Duplicates, mistakes and requests mean different things
Severity indexA single weighted figure for severityIt depends on the weights; report the counts too
munotes.in455

Defect Metrics

What it does not mean

A high removal efficiency does not mean few defects. It means few escaped; a release full of defects can have a high efficiency if testing was thorough.

A module with many defects is not necessarily the worst. Density, not count, compares modules of different sizes.

Stage effectiveness is not a score for the people. It measures an activity on one release, with a denominator that later stages supplied.

A severity index is not an average severity. It is a weighted count under a policy, and the counts behind it must be shown.

Quick revision

  • Defect density: defects per unit of size (ISO/IEC/IEEE 24765); by module, it shows the weakest code.
  • Defect removal efficiency: defects found before release over all defects; release 2.0: 188 of 200, 94.0 per cent.
  • Stage effectiveness (Fagan's error detection efficiency): errors found by a stage over errors present before it; release 2.0: requirements review 57.1, design review 46.2, code review 21.2, unit 34.9, integration 39.0, system 56.0, acceptance 45.5 per cent.
  • Leakage: requirements 42.9, design 47.5, code 76.7 per cent escaped their own review.
  • Age in stages: documents 4.17, far above the rest.
  • Density: fee payment 14.5 and exam form 13.5 defects per KLOC; reopen rate 7.1 per cent; not defects 20 per cent of reports; severity index 2.77 with weights 10, 5, 2, 1.

Test yourself

1. Define defect removal efficiency and compute it for release 2.0. The defects found before release as a percentage of all defects found, before and after release. Release 2.0: 188 found before release of 200 in all, 94.0 per cent.

2. What is the effectiveness of a stage, and why is its denominator hard to know? The defects a stage finds divided by the defects present in the product when it runs, as Fagan defined error detection efficiency for inspections. The defects present include those no stage has yet found, which are known only later, when subsequent stages and users find them.

3. A module of 2 KLOC has 30 defects and one of 6 KLOC has 48. Which is more defective, and why? The first: 30 defects in 2 KLOC is 15 per KLOC, against 48 in 6 KLOC, 8 per KLOC. The second has more defects only because it is larger.

4. What is defect leakage? Give release 2.0's figures. The share of a work product's defects that escape the activity meant to catch them. In release 2.0, 12 of 28 requirements defects escaped the requirements review (42.9 per cent), 19 of 40 design defects escaped the design review (47.5 per cent), and 92 of 120 code defects escaped the code review (76.7 per cent).

munotes.in456

Defect Metrics

5. What does the reopen rate measure, and what was release 2.0's? The share of fixes that fail their confirmation test and are reopened, a measure of the quality of fixes and of the testing done before they are handed back. Release 2.0: 14 of 198, 7.1 per cent.

6. Why must a severity index be reported with the counts it summarises? Because it is a weighted sum of ranked categories, and its value depends on the weights chosen, which are a policy, not a measurement; the counts at each severity level are the data, and show what the index hides.

munotes.in457

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!