Using Quality Costs for Decision Making
Chapter One Hundred Two
Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Utilizing quality costs for decision making"
Pages 587 to 591 of 622
In one line
Quality costs become useful when they drive decisions: tracked release after release they show whether quality spending is moving from failure to prevention, and set against an improvement's cost they show whether it pays, which depends on the organisation's own cost of a late defect, not on anyone else's.
In the wording a student can write in an examination: an analysis of quality costs, ASQ says, "provides a method of assessing the effectiveness of the management of quality and a means of determining problem areas, opportunities, savings, and action priorities." It supports three kinds of decision. Trend analysis compares the cost of quality over time, normalised for size, and its mix between conformance (prevention, appraisal) and failure. Investment decisions compare an improvement's cost with the failure costs it is expected to save, with a break-even point. Life-cycle decisions weigh extra quality cost now against lower costs later; Boehm and Basili (2001) found that high-dependability software costs about 50 per cent more to develop but can cost about the same over its life, because it is cheaper to maintain.
From a statement to a decision
Chapter One Hundred One, on the cost of quality, drew up release 2.0's statement: 1,194 person-hours, 59 per cent of it failure. A statement on its own records what happened. ASQ asks for more: the costs "must be a true measure of the quality effort", and "The quality cost system, once established, should become dynamic and have a positive impact on the achievement of the organization's mission, goals, and objectives." In practice that means three uses: following the trend, deciding on improvements, and understanding the cost of quality over a product's whole life.
Worked example: a trend, a decision, and the life cycle
The program does all three. First, it compares releases 2.0 and 2.1, per KLOC so that a smaller release does not look better merely for being smaller. Second, it evaluates a proposal for release 2.2: a Fagan-style inspection of the fee module's 4,000 lines of code by four people, at the preparation and meeting rates of Chapter Twenty-Nine, on inspection, replacing the light code review now used. The inputs are ExamReg's own: release 2.1's defect density, the share of defects still present at code review in release 2.0, the code review's effectiveness of 21.2 per cent (Chapter Seventy-Nine, on defect metrics) against the median of about 60 per cent that Boehm and Basili report for peer reviews, and the cost of fixing a defect where it is found (1.0 hour in a review, 4.0 in testing, 8.0 after release, plus the complaints and corrections that come with each defect students meet). Third, it redoes Boehm and Basili's life-cycle arithmetic.
# 1. The trend: the cost of quality of releases 2.0 and 2.1, in person-hours (FINDINGS 5.10)
releases = {"2.0": (20, {"prevention": 80, "appraisal": 406, "internal failure": 560,
"external failure": 148}),
"2.1": (8, {"prevention": 70, "appraisal": 180, "internal failure": 150,
"external failure": 30})}
for name, (kloc, coq) in releases.items():
total = sum(coq.values())
failure = coq["internal failure"] + coq["external failure"]
print(f"release {name}: {total} h, {total / kloc:.1f} h per KLOC; prevention"
f" {coq['prevention'] / total:.0%}, failures {failure / total:.0%}")
# 2. The decision: a Fagan-style code inspection of the fee module in release 2.2
lines, people = 4000, 4
cost = people * (lines / 125 + lines / 150) # preparation and meeting at Fagan's rates
cost -= 30 * lines / 20000 # less the light code review it replaces
present = lines / 1000 * (69 / 8) * (160 / 200) # 2.1's density; the share present at code review
extra = present * (0.60 - 0.212) # effectiveness 21.2% now; 60% median for reviews
# a defect found by the inspection costs 1.0 h to fix; missed, 90.5% are found in testing
# (4.0 h) and 9.5% after release (8.0 h to fix + 52 / 12 h of complaints and corrections)
saving = extra * (0.905 * (4.0 - 1.0) + 0.095 * (8.0 + 52 / 12 - 1.0))
print(f"inspection: {cost:.1f} extra hours to find {extra:.1f} more defects early;"
f" saving {saving:.1f} h; net {saving - cost:+.1f} h")
print(f"break-even: each defect found early would have to save {cost / extra:.1f} h")
# 3. Boehm and Basili: is quality free over the life cycle? (per instruction; 30% build, 70% keep)
low_build, high_build = 1.0, 1.5 # high dependability costs 50% more to build
low_keep, high_keep = 1.5 * low_build, 0.85 * high_build
print(f"life cycle per instruction: low dependability {0.3 * low_build + 0.7 * low_keep:.3f},"
f" high {0.3 * high_build + 0.7 * high_keep:.3f}")Using Quality Costs for Decision Making
release 2.0: 1194 h, 59.7 h per KLOC; prevention 7%, failures 59%
release 2.1: 430 h, 53.8 h per KLOC; prevention 16%, failures 42%
inspection: 228.7 extra hours to find 10.7 more defects early; saving 40.6 h; net -188.1 h
break-even: each defect found early would have to save 21.4 h
life cycle per instruction: low dependability 1.350, high 1.342The trend. Release 2.1 cost 53.8 hours of quality work per KLOC against 59.7 for release 2.0, and its mix moved the right way: prevention rose from 7 to 16 per cent of the cost of quality, and failures fell from 59 to 42 per cent. That is the evidence that release 2.0's causal analysis, which added prevention (Chapter Eighty, on using defect data), paid in the next release's costs, and it is the kind of evidence ASQ means by the cost system becoming "dynamic".
Using Quality Costs for Decision Making
The decision, and why the answer is no. The inspection would cost 228.7 more hours and find about 10.7 more defects early. On ExamReg's own costs, moving a defect from testing to inspection saves 3 hours of fixing, and from after release about 11, so the expected saving is only 40.6 hours: the proposal loses 188.1 hours. The break-even line explains why: the inspection pays only if each defect found early saves 21.4 hours. In a large, critical system, where Boehm and Basili's "often 100 times" applies, it easily would; in ExamReg, a small, noncritical portal whose defects are cheap to fix, it does not. The decision follows from the organisation's own data, which is exactly why the data are collected. The same analysis also points at better options: a lighter, checklist-based review of the fee module's riskiest parts costs far fewer hours, and prevention, which removed defects outright in release 2.1, needs no finding at all.
The life cycle. Boehm and Basili asked whether Crosby's "quality is free" is right, since "it costs 50 percent more per source instruction to develop high-dependability software products than to develop low-dependability software products." Their answer rests on maintenance: low-dependability software "costs about 50 percent per instruction more to maintain than to develop, whereas high-dependability software costs about 15 percent less to maintain than to develop", and with 30 per cent of life-cycle cost in development and 70 in maintenance, "low-dependability software becomes about the same in cost per instruction as high-dependability software". Weighting each per-instruction cost 30 to 70, which is one reading of their arithmetic, reproduces the conclusion: 1.350 against 1.342. Their overall verdict on Crosby: "Maybe for some low-criticality, short-lifetime software, but not for the most important cases."
Using quality costs well
| Decision | What the cost of quality contributes | ExamReg |
|---|---|---|
| Where to act | The largest failure costs, like a Pareto of costs (Chapter One Hundred Four, on Pareto diagrams) | Internal failure, 47 per cent of release 2.0's cost of quality |
| Whether an action worked | The trend, normalised for size, and the shift in mix | 59.7 to 53.8 hours per KLOC; failures from 59 to 42 per cent |
| Whether to invest | Cost against expected savings, with a break-even point | The fee module inspection: rejected, break-even 21.4 hours a defect |
| How much quality to build in | Life-cycle cost, not development cost alone | Boehm and Basili: about equal per instruction for high and low dependability |
What it does not mean
Quality costs do not make the decision alone. They put numbers on it; safety, reputation and harm to users (the external failures students suffer) may outweigh the hours.
An improvement that does not pay here may pay elsewhere. The inspection fails on ExamReg's costs and would succeed where late defects cost far more; ratios from other organisations must not be imported unexamined.
Using Quality Costs for Decision Making
Quality is not free in every case. Boehm and Basili show higher development cost can be repaid over the life cycle, but not for all software, as they say themselves.
A falling cost of quality is not proof of success. It must be normalised for size and read with its mix; cutting appraisal lowers the cost now and raises failure costs later.
Quick revision
- Uses of quality costs (ASQ): assess the management of quality; find problem areas, opportunities, savings and action priorities; the cost system should become "dynamic".
- Trend analysis: cost of quality per unit of size, release by release, and its mix of conformance against failure.
- Investment decisions: an improvement's cost against the failure costs it saves; the break-even saving per defect.
- Life-cycle cost (Boehm and Basili 2001): high dependability costs 50 per cent more to develop, maintenance 15 per cent below development cost against 50 per cent above for low dependability; over a 30/70 life cycle, about the same; "quality is free" fails only for "some low-criticality, short-lifetime software".
- Worked example: 59.7 to 53.8 hours per KLOC, prevention 7 to 16 per cent, failures 59 to 42 per cent; the fee module inspection costs 228.7 hours, saves 40.6, break-even 21.4 hours a defect: rejected on ExamReg's costs; life cycle 1.350 against 1.342.
Test yourself
1. How are quality costs used for decision making? To find where quality money is lost (the largest failure costs), to judge whether improvements worked (trends normalised for size and the shift from failure to prevention), to decide whether a proposed improvement pays (its cost against the failure costs it saves, with a break-even point), and to weigh extra quality cost now against lower costs over the product's life.
2. What did the trend between releases 2.0 and 2.1 show? The cost of quality per KLOC fell from 59.7 to 53.8 hours, prevention rose from 7 to 16 per cent of it, and failures fell from 59 to 42 per cent: the added prevention of release 2.1 was paying back in lower failure costs.
3. Why was the proposed fee module inspection rejected? On ExamReg's own costs it would cost 228.7 more hours and save only about 40.6 hours of fixing, because the defects it would find early are cheap to fix later in this small, noncritical system. It would pay only if each defect found early saved at least 21.4 hours.
4. What is a break-even point in a quality cost decision, and why is it useful? The value at which an improvement's savings equal its cost, here the hours each early-found defect must save. It shows what would have to be true for the improvement to pay, so the decision can be judged against the organisation's own data or revisited when those data change.
Using Quality Costs for Decision Making
5. Is quality free? Give Boehm and Basili's answer. Not always. High-dependability software costs about 50 per cent more per instruction to develop, but is cheaper to maintain, so over a life cycle of 30 per cent development and 70 per cent maintenance it costs about the same as low-dependability software; in their words, "the investment is more than worth it if the project involves significant operations and maintenance costs." Crosby's claim may fail only for some low-criticality, short-lifetime software.
6. Why should an organisation not use another organisation's cost ratios for its decisions? Because the value of finding defects early depends on how much a late defect costs, which varies from about 5 to 1 in small, noncritical systems to often 100 to 1 in large, critical ones. ExamReg's inspection decision is negative on its own ratios and would be positive on a large system's.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.