munotes®

Quality, Process and Test Metrics

Get access to whole semester resourcesSemester Pass

Chapter Sixty-Eight

Syllabus topic Module 2, "Software Metrics: ... utilizing different types of metrics"

Pages 391 to 395 of 622

In one line

Metrics about software fall into three groups: product metrics say how good the thing is, process metrics say how well it was made and tested, and test metrics say how far the testing has got and what it has found; each is useful only when it answers a question and is read alongside the others.

In the wording a student can write in an examination: in the ISTQB syllabus's words, "Test metrics are gathered to show progress against the planned test schedule and budget, the current quality of the test object, and the effectiveness of the test activities with respect to the test objectives or an iteration goal." The common kinds are project progress metrics, test progress metrics, product quality metrics, defect metrics, risk metrics, coverage metrics and cost metrics. Among product quality metrics, defect density is the "number of defects per unit of product size", and the mean time to repair (sometimes called the mean time to change) is "the mean time the maintenance team requires to implement a change and restore the system to working order" (ISO/IEC/IEEE 24765). Among process metrics, the defects each activity finds and the effort it takes show which activities work best.

Three questions, three groups

Chapter Sixty-Five, on what a software metric is, divided the objects of measurement into products, processes and resources. For a tester the same division becomes three questions:

  • How good is the product? Product quality metrics: defects in it, how it performs, how reliable it is, how quickly it can be changed.
  • How well was it made and tested? Process metrics: where defects were found, at what cost in effort, how many escaped.
  • Where has the testing got to? Test metrics: tests run and passed, coverage reached, defects found and fixed, risks remaining.

The groups overlap on purpose. Defects found after release describe the product (they are in it) and the process (it let them through).

The ISTQB list of test metrics

The syllabus gives seven kinds, each with examples:

KindThe syllabus's examplesWhere this book meets it
Project progress"task completion, resource usage, test effort"Effort by activity, below
Test progresstest cases implemented, environment readiness, "number of test cases run/not run, passed/failed, test execution time"Chapter Ten's report of 160 of 240 run, on test reporting
Product quality"availability, response time, mean time to failure"Chapter Forty-Five's response times under load testing; Chapter Ninety's reliability
Defect"number and priorities of defects found/fixed, defect density, defect detection percentage"Below, and Chapter Seventy-Nine, on defect metrics
Risk"residual risk level"The risks still open at release
Coverage"requirements coverage, code coverage"Chapter Forty-One's traceability, on validation testing; Chapters Fifty-Eight to Sixty on code coverage
Cost"cost of testing, organizational cost of quality"Chapter One Hundred One, on the cost of quality
munotes.in391

Quality, Process and Test Metrics

CMMI's list of commonly used derived measures overlaps it: "Defect density", "Peer review coverage", "Test or verification coverage" and "Reliability measures (e.g., mean time to failure)".

Product quality metrics

Defect density divides the defects by the size of the product, so that a large system and a small one can be compared. It depends entirely on both counts being defined: which defects (all found, or only those after release; which severities) and which size (lines by what rule, or function points). Delivered defect density counts only the defects that reached users, and is the one that describes what users got.

Mean time to repair, in its maintenance sense, measures how quickly the product can be changed: from a problem being reported to a working fix in production. It belongs to maintainability, a product characteristic, although the team's process shows in it too.

The other product quality metrics the syllabus names, availability, response time and mean time to failure, are measured by running the system: Chapter Forty-Five measured ExamReg's response times under load, and Chapter Ninety defines reliability and computes them.

Process metrics

A process metric says how an activity performed. For defect detection the basic ones are:

  • Defects found by each activity and the share of all defects that each found, which Chapter Sixty-Six used in its goal, question, metric model.
  • Effort spent in each activity, in person-hours.
  • Defects found per person-hour, the activity's efficiency at finding defects.
  • Defect removal efficiency, the share of all defects found before release, which Chapter Seventy-Nine defines fully.

Worked example: release 2.0's product and process metrics

The program computes release 2.0's product metrics, by both measures of size, and its process metrics by activity. The effort figures and change times come from the project's time records.

from statistics import mean, median

KLOC, FUNCTION_POINTS = 20, 72.8          # release 2.0's size: FINDINGS 5.2, and Chapter 67's count
found_by = {"requirements review": 16, "design review": 24, "code review": 34,
            "unit testing": 44, "integration testing": 32, "system testing": 28,
            "acceptance testing": 10, "after release": 12}            # FINDINGS 5.2
hours = {"requirements review": 12, "design review": 20, "code review": 30, "unit testing": 80,
         "integration testing": 64, "system testing": 112, "acceptance testing": 40}  # FINDINGS 5.2.2
change_hours = [4, 6, 3, 30, 8, 5, 72, 10, 6, 4, 12, 20]  # the 12 fixes after release (FINDINGS 5.2.2)

total, delivered = sum(found_by.values()), found_by["after release"]
print("PRODUCT")
print(f"   defect density            {total / KLOC:.1f} per KLOC, {total / FUNCTION_POINTS:.2f} per function point")
print(f"   delivered defect density  {delivered / KLOC:.2f} per KLOC, {delivered / FUNCTION_POINTS:.2f} per function point")
print(f"   mean time to change       {mean(change_hours):.1f} hours (median {median(change_hours):.1f},"
      f" longest {max(change_hours)})")

print("PROCESS")
for activity, h in hours.items():
    print(f"   {activity:<21} {found_by[activity]:>3} defects in {h:>3} hours:"
          f" {found_by[activity] / h:.2f} per hour")
for kind in ("review", "testing"):
    d = sum(n for a, n in found_by.items() if a.endswith(kind))
    h = sum(n for a, n in hours.items() if a.endswith(kind))
    print(f"   all {kind + 's' if kind == 'review' else kind:<17} {d:>3} defects in {h:>3} hours: {d / h:.2f} per hour")
munotes.in392

Quality, Process and Test Metrics

PRODUCT
   defect density            10.0 per KLOC, 2.75 per function point
   delivered defect density  0.60 per KLOC, 0.16 per function point
   mean time to change       15.0 hours (median 7.0, longest 72)
PROCESS
   requirements review    16 defects in  12 hours: 1.33 per hour
   design review          24 defects in  20 hours: 1.20 per hour
   code review            34 defects in  30 hours: 1.13 per hour
   unit testing           44 defects in  80 hours: 0.55 per hour
   integration testing    32 defects in  64 hours: 0.50 per hour
   system testing         28 defects in 112 hours: 0.25 per hour
   acceptance testing     10 defects in  40 hours: 0.25 per hour
   all reviews            74 defects in  62 hours: 1.19 per hour
   all testing           114 defects in 296 hours: 0.39 per hour

The product. Release 2.0 had 10 defects per KLOC over its whole life, and 0.6 per KLOC reached students; by function points, 2.75 and 0.16. The two sizes give different numbers for the same defects, which is why a density is never reported without its size measure. The mean time to change is 15 hours, but the median is 7: one change took 72 hours and pulls the mean up. Reporting both, and the longest, tells the reader that most changes are quick and one was not, which a mean alone hides.

The process. Each review found more than one defect per person-hour; unit and integration testing found about half a defect per hour; system and acceptance testing a quarter. Taken together, the reviews found 1.19 defects per hour against 0.39 for testing, about three times as many. That is a real finding about this project, and a reason to look at where review effort could be increased, which Chapters Ninety-Six to Ninety-Eight, on reviews, examine with published evidence.

It is also easy to misread. Later activities find fewer defects per hour partly because earlier ones have already removed the easy ones, and partly because they look for different kinds of defect: a system test finds failures of the whole portal that no review of one document could have seen. Defects per hour compares how the activities performed on this project; it does not say that any activity could be dropped.

Reading metrics well

A few habits keep metrics honest.

  • Pair a rate with its volume. 1.33 defects per hour sounds best of all, but the requirements review ran for 12 hours; the largest number of defects came from unit testing.
  • Report the distribution, not only the mean, when a few values are extreme, as the change times were.
  • Watch trends, not single values. A density this release means most beside the same density last release, measured the same way.
  • Compare against decision criteria agreed in advance, the thresholds and targets of ISO/IEC/IEEE 15939 (Chapter Sixty-Six's goal, question, metric example set 90 per cent of defects found before release).
munotes.in393

Quality, Process and Test Metrics

What it does not mean

A test metric is not a quality metric. The number of tests passed says how far testing has got; the product's quality is in its defects, performance and reliability.

A low defect density is not proof of quality. It may mean the product is good, or that testing found little; it needs the process metrics beside it.

Defects per hour does not rank activities for removal. Each activity finds kinds of defect the others miss.

A mean is not always the typical value. With a skewed distribution, the median says more.

Quick revision

  • Test metrics show "progress against the planned test schedule and budget, the current quality of the test object, and the effectiveness of the test activities" (ISTQB).
  • ISTQB's seven kinds: project progress, test progress, product quality, defect, risk, coverage, cost.
  • Defect density: "number of defects per unit of product size"; mean time to repair in the maintenance sense: the mean time to implement a change and restore working order (ISO/IEC/IEEE 24765).
  • Process metrics: defects by activity, effort by activity, defects per person-hour, defect removal efficiency.
  • Worked example: 10.0 defects per KLOC (2.75 per function point), 0.60 delivered per KLOC (0.16 per function point); mean time to change 15.0 hours, median 7.0; reviews 1.19 defects per hour against testing 0.39.
  • Read metrics with their volumes, distributions, trends and decision criteria.

Test yourself

1. Why are test metrics gathered? Name the kinds ISTQB lists. To show progress against the planned test schedule and budget, the current quality of the test object, and the effectiveness of the test activities against their objectives. The kinds are project progress, test progress, product quality, defect, risk, coverage and cost metrics.

2. Distinguish product, process and test metrics, with an example of each. Product metrics describe the software itself, such as defect density or response time. Process metrics describe how it was developed and tested, such as defects found per person-hour of review. Test metrics describe the state of the testing, such as test cases run against planned, or requirements coverage.

3. A system of 50 KLOC had 150 defects, 9 of them found after release. Compute its defect density and delivered defect density. Defect density: 150 defects in 50 KLOC, 3.0 per KLOC. Delivered defect density: 9 in 50 KLOC, 0.18 per KLOC.

munotes.in394

Quality, Process and Test Metrics

4. What is mean time to change, and why report the median with it? The mean time needed to implement a change and restore the system to working order, a measure of maintainability. The median is reported because a few long changes can pull the mean far above the typical case, as a single 72-hour change did in the worked example (mean 15.0 hours, median 7.0).

5. In the worked example, reviews found 1.19 defects per hour and testing 0.39. What can and cannot be concluded? It can be concluded that on this project reviewing found defects more efficiently per hour of effort, a reason to consider more review. It cannot be concluded that testing could be reduced or dropped, because later activities find kinds of defect reviews cannot see, and they find fewer per hour partly because earlier activities removed the easier ones.

6. Why should a defect density always be reported with its size measure? Because the same defects give different densities by different sizes (10.0 per KLOC and 2.75 per function point for the same release), and the density is comparable only with others measured the same way.

munotes.in395

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!