Quality in Software Development
Chapter Twenty-Three
Syllabus topic Module 1, "Definition of Quality and Quality Assurance: Understanding quality, in software development"
Pages 120 to 128 of 622
In one line
Software quality is quality as the last chapter defined it, applied to a product that is designed but never manufactured, that does not wear out but changes constantly, and that can be judged from three places: its code, its behaviour under test, and its use by real people.
In the wording a student can write in an examination: software quality is the "capability of a software product to satisfy stated and implied needs when used under specified conditions" (ISO/IEC 25000:2014), or, more narrowly, the "degree to which a software product meets established requirements" (IEEE 730-2014, since replaced by a 2026 edition). Understanding quality in software development means understanding two things. First, why software is different: its faults are design faults, not physical ones; it does not wear out, but it declines when its use or its code changes; and, in Frederick Brooks's analysis, it is complex, must conform to interfaces other people designed, is constantly changed, and is invisible. Second, that its quality has three views: internal quality (the code and documents), external quality (the software running) and quality in use (the software in its users' hands).
Why software strains the ordinary idea of quality
Much of quality management grew up in manufacturing. ASQ's history of quality records that factory quality was long kept by inspecting products, and that during the Second World War the United States armed forces moved from inspecting every unit to sampling inspection. Both ideas assume many copies of one design, any of which can come out slightly wrong.
Software has no such copies. Every copy of ExamReg is the same program; if the fee rule is wrong in one, it is wrong in all of them, and inspecting a sample of copies would find nothing that inspecting one would not. The quality of software is decided almost entirely before the first copy exists, in its requirements, design and code. Michael Lyu's Handbook of Software Reliability Engineering (1996) puts the difference in one sentence: "Unlike hardware faults which are mostly physical faults, software faults are design faults", which are "harder to visualize, classify, detect, and correct."
The consequence runs through this whole course. Quality in software development cannot be inspected into the product at the end; it has to be built in while the product is designed and written. That is why so much of the syllabus is about reviews, process and prevention, the subject of quality assurance in Chapter Twenty-Four, and not only about testing the finished program.
Software does not wear out, but it does decline
The second difference is time. A pump, a tyre or a hard disk wears out: its parts age, and it fails more often until it is replaced. Lyu states the contrast plainly: software reliability differs from hardware reliability "in the sense that software does not wear out, burn out, or deteriorate, i.e., its reliability does not decrease with time." An unchanged program, given the same input in the same conditions, does on its thousandth run what it did on its first.
Quality in Software Development
That does not mean software quality stays fixed. Lyu names the two ways it falls: "software may experience reliability decrease due to abrupt changes of its operational usage or incorrect modifications to the software." Both happen to ExamReg every session.
- A change in use. The code that handled a few dozen forms a day in the first week meets 300 on the last date. Nothing in the program has changed, but it is being used as it never was before, and a fault that no ordinary day reached is reached.
- A change in the code. The exam cell revises the fee table, a developer edits the fee function, and the edit breaks a case that used to work. Every correction is itself a change to a design, and can bring a new fault with it. Guarding against that is the job of regression testing (Chapter Thirty-Nine).
The same passage names the good news. Software "generally enjoys reliability growth during testing and operation", because each fault found and removed is gone for good from every copy.
| Hardware | Software | |
|---|---|---|
| Where its faults come from | Mostly physical: wear, material, manufacture | Design: requirements, design and code |
| Copies | Each unit can come out differently | Every copy is identical |
| Left unchanged over time | Wears out, and fails more often | Does not wear out |
| What lowers its reliability | Age and wear | A change in how it is used, or a faulty change to it |
| Finding and fixing a fault | Restores the unit to its design | Changes the design: reliability can grow, or a new fault can enter |
Chapter Ninety, on software reliability, returns to this contrast with the failure curves of hardware and software.
Brooks: four properties that make software quality hard
In "No Silver Bullet", a paper first given at the IFIP World Computing Conference in 1986 and later reprinted in his book The Mythical Man-Month, Frederick Brooks separated the difficulties of software into accidental ones, which better tools can remove, and essential ones, which belong to software itself. He named "the inherent properties of this irreducible essence of modern software systems: complexity, conformity, changeability, and invisibility." Each is a reason quality is hard to achieve in software development, and each makes a demand on testing.
Complexity. "Software entities are more complex for their size than perhaps any other human construct, because no two parts are alike (at least above the statement level)." Brooks traced unreliability straight to it: "From the complexity comes the difficulty of enumerating, much less understanding, all the possible states of the program, and from that comes the unreliability." For a tester this is why exhaustive testing is impossible, the second of the seven principles in Chapter Four, and why test design techniques exist: to choose, from far more states than anyone could try, the few tests most likely to find faults.
Quality in Software Development
Conformity. A physicist can hope that nature obeys a few simple laws. Software must instead fit interfaces that other people designed for their own reasons, and Brooks observed that "much complexity comes from conformation to other interfaces; this cannot be simplified out by any redesign of the software alone." ExamReg must fit the paper codes it is given, the exam cell's fee rules, the payment gateway's protocol and the browsers students use. Its developers can simplify none of them, and each is a place where the program can be right by its own logic and wrong against the thing it must fit. Integration testing and compatibility testing exist for exactly these seams.
Changeability. "All successful software gets changed." Software, Brooks wrote, "is pure thought-stuff, infinitely malleable", and "The pressures for extended function come chiefly from users who like the basic function and invent new uses for it." A program that will certainly change has to be built to be changed, which is why maintainability and flexibility are quality characteristics in their own right in the quality model today (Chapter Twenty), and why every change needs its regression tests.
Invisibility. "Software is invisible and unvisualizable." A building has a floor plan, on which "Contradictions become obvious, omissions can be caught"; a program's structure, when anyone tries to draw it, turns out to be "not one, but several, general directed graphs, superimposed one upon another." Nobody can look at a program and see whether it is good, the way an inspector can look at a weld. Its quality has to be made visible by other means: documents that can be reviewed, models, the measures of Module 2, and tests that turn behaviour into evidence.
| Brooks's property | What it means | What it asks of quality work |
|---|---|---|
| Complexity | More states than anyone can list; no two parts alike | Test design techniques, risk-based choice of tests, reviews |
| Conformity | Must fit interfaces designed by others | Clear interface requirements; integration and compatibility testing |
| Changeability | All successful software gets changed | Maintainable design, regression testing, control of changes |
| Invisibility | No single drawing shows the whole of it | Reviews of documents and models, measures, tests as evidence |
Three views of quality: internal, external and in use
The last difference is where quality is seen from. The standards distinguish three views, and each is judged by different people with different evidence.
Quality in Software Development
Internal quality is the "totality of attributes of a product that determine its ability to satisfy stated and implied needs when used under specified conditions" (ISO/IEC/IEEE 24765). It is the quality of the code, design and documents themselves, judged without running the program: by reviews, by static analysis, and by internal measures, each a "measure of the product itself, either direct or indirect", such as the size of a module or its complexity.
External quality is the "extent to which a product satisfies stated and implied needs when used under specified conditions". It is judged by running the software, as testers do. An external measure is an "indirect measure of a product derived from measures of the behavior of the system of which it is a part", such as the number of failures in a test cycle or the response time under load.
Quality in use is, in ISO/IEC 25019:2023, the "extent to which the system or product, when it is used in a specified context of use, satisfies or exceeds" what its stakeholders need "to achieve specified beneficial goals or outcomes". The context of use is the "combination of users, goals and tasks, resources, and environment" (ISO TR 25060:2023). Quality in use is judged by the people who use the software, for their own purposes, in their own conditions. ISO's own page describes the 2023 standard as "a quality-in-use model composed of three characteristics". They are beneficialness, the "extent of benefit resulting from the use of a product, system, or service"; freedom from risk, the "extent to which a product or system mitigates the potential risk to economic status, human life, health, society, financial values, enterprise activities, or the environment"; and acceptability.
Figure 23.1 Three views of software quality: internal, external and in use
The older product quality standard, ISO/IEC 9126, linked the three views in the chain the figure shows. Internal quality influences external quality: well-structured, well-reviewed code is more likely to behave well when it runs. External quality influences quality in use: software that behaves well under test is more likely to serve its users. Read the other way, each depends on the one before it: quality in use cannot be had without external quality, nor external quality without internal quality. The influence is real, but it is not a guarantee, and the gaps between the views are where some of the most instructive failures live. The worked example below is one of them.
Who judges software quality
| Who | What they see | View of quality | Their evidence |
|---|---|---|---|
| Developer and reviewer | Code, design and documents | Internal | Review findings, static analysis, internal measures |
| Tester | The running system, in a test environment | External | Test results, failures found, external measures |
| User (a student filling in the form) | The system in real use | In use | Whether the task got done, and without harm |
| Customer (the college that paid for it) | The system's cost and benefit | In use, and value | Benefit for the money, complaints, risks avoided |
| Operator (the IT cell) | The system in production | External and in use | Availability, incidents, how hard it is to run |
Quality in Software Development
Each of these people can be satisfied while another is not. The developers can be proud of clean code that the students find confusing; the students can be happy with a form that the IT cell must restart every night. This is Garvin's point from Chapter Twenty-Two, on what quality means, seen from inside a software project: a quality plan has to ask every one of them.
Worked example: one test result, two very different qualities in use
Two builds of ExamReg's fee function each contain one defect. Build A has lost the on-time case, Chapter One's defect from the chapter on what software testing is: it charges a form submitted on or before the last date as if it were late. Build B refuses a form submitted on the fifteenth day late, which the rule still accepts. The portal records any on-time form as 0 days late, and the test team runs one test for each value from 0 to 20 days late; both builds fail exactly one test. By that external measure, they are equally good.
The exam cell also has last session's record of how many forms arrived on each day. The counts below are this book's illustration, not a real college's data. The program weighs each build's failures by how many students would actually have met them.
def rule(days_late): # the exam cell's rule; 0 means on or before the last date
if days_late < 0 or days_late > 15:
raise ValueError("form not accepted")
if days_late == 0:
return 0
if days_late <= 7:
return 100
return 500
def build_a(days_late): # defect A: Chapter One's, the on-time case lost
if days_late < 0 or days_late > 15:
raise ValueError("form not accepted")
if days_late <= 7:
return 100
return 500
def build_b(days_late): # defect B: the fifteenth day refused
if days_late < 0 or days_late >= 15:
raise ValueError("form not accepted")
if days_late == 0:
return 0
if days_late <= 7:
return 100
return 500
def outcome(fee, days_late):
try:
return fee(days_late)
except ValueError:
return "refused"
days = list(range(0, 21)) # one test for each value of days late, 0 to 20
forms = [720, # last session's forms on or before the last date
50, 30, 25, 20, 15, 12, 28, # 1 to 7 days late
18, 10, 8, 7, 6, 5, 9, 14, # 8 to 15 days late
8, 6, 4, 3, 2] # 16 to 20 days late (refused, as the rule says)
total = sum(forms)
print("tests:", len(days), " forms in the session:", total)
for name, build in [("build A", build_a), ("build B", build_b)]:
wrong = [d for d in days if outcome(build, d) != outcome(rule, d)]
hit = sum(n for d, n in zip(days, forms) if d in wrong)
print(f"{name}: wrong on day {', '.join(map(str, wrong))};"
f" tests failed {len(wrong)} of {len(days)} ({len(wrong) / len(days):.1%});"
f" students hit {hit} of {total} ({hit / total:.1%})")Quality in Software Development
tests: 21 forms in the session: 1000
build A: wrong on day 0; tests failed 1 of 21 (4.8%); students hit 720 of 1000 (72.0%)
build B: wrong on day 15; tests failed 1 of 21 (4.8%); students hit 14 of 1000 (1.4%)By the test results the two builds are identical: each fails 1 of its 21 tests, about 4.8 per cent. In use they are nothing alike. Build A's defect meets every student who submitted on time, which is most of them: it would have charged 720 of the 1,000 students a late fee they did not owe, 72 per cent of all users. Build B's defect would have met 14 students, 1.4 per cent.
But a count of students is not the whole of quality in use either. Build A's harm is Rs 100 taken wrongly, which can be refunded. Build B's harm is a refused form: under the rule, those 14 students cannot sit the examination unless somebody notices in time. Quality in use includes freedom from risk, and a defect that meets few users can still do the most damage. A tester who reports only 1 of 21 tests failed for each build has told the exam cell almost nothing it needs to know. The defect report has to say who is hit and how badly, which is what severity and priority in a defect report (Chapter Seventy-Six) exist to record.
The general lesson is about test design. A test count treats every input as equally important; users do not. The on-time case deserves more tests, and more careful ones, than the eleventh day late, because that is where the students are.
When is software good enough?
No release is free of defects; testing can show that defects are present but never that they are absent, the first of the seven principles in Chapter Four. Every release is therefore a decision that the software is good enough. The NIST study of 2002 called this the central difficulty: "The major problem for the software industry is deciding when a firm should stop testing". It reported that commercial developers decide with a combination of rules of thumb: a sufficient percentage of test cases passing, statistics from a code coverage tool, counts and trends of defects by severity, beta testing by real users, and the number of new problem reports falling below a threshold. NIST described these as nonanalytical: none of them calculates the risk that remains.
Quality in Software Development
The three views make the decision more honest. Internal quality asks whether any serious review finding is still open. External quality asks whether the exit criteria are met: the planned tests run, the pass rate reached, no critical defect open. Quality in use asks whether the software has been tried in its real context of use, with its real users and its busiest day. The worked example shows why the third question cannot be skipped. A release rule of at least 95 per cent of tests passed would let build A through, with 20 of 21 tests passed, about 95.2 per cent, and it would still overcharge 720 students.
The cost of getting it wrong
Chapter Three, on why software must be tested, showed what poor quality has cost when software fails in the field, from a lost rocket to a national estimate of $59.5 billion a year. ExamReg's scale is smaller but the pattern is the same. Released, build A's defect would have taken 720 × 100 = 72,000 rupees from students who owed nothing, and the college would then pay again: in refunds, corrected receipts, apologies and an emergency release made under pressure. Found in a review of the code, the same defect would have cost one changed line and a retest. Chapter One Hundred One, on the cost of quality, turns this into the four kinds of quality cost and shows how to count them.
What it does not mean
Software not wearing out does not mean its quality is permanent. Its reliability falls when its use changes or when a change to the code brings in a new fault.
Internal quality is not a private concern of developers. External quality and quality in use are built on it, and it decides how costly every later change will be.
Passing tests is not the same as quality in use. Tests measure external quality over the inputs someone chose; users meet the inputs their context produces.
Good enough is not an excuse for poor quality. It is a decision, made with evidence, that the remaining risk is acceptable to the people who will carry it.
Quick revision
- Software quality: "capability of a software product to satisfy stated and implied needs when used under specified conditions" (ISO/IEC 25000:2014); "degree to which a software product meets established requirements" (IEEE 730-2014).
- Software faults are design faults, and every copy is identical, so quality must be built in during development, not inspected in at the end.
- Software does not wear out (Lyu 1996), but its reliability falls with a change in use or a faulty change to the code, and grows as faults are found and removed.
- Brooks (1986), the essential properties of software: complexity, conformity, changeability, invisibility.
- Internal quality: the product itself, judged without running it. External quality: the software running, judged by tests. Quality in use: in a context of use (users, goals and tasks, resources, environment); ISO/IEC 25019:2023 has three characteristics: beneficialness, freedom from risk, acceptability.
- Internal quality influences external quality, which influences quality in use (the chain of ISO/IEC 9126).
- Worked example: two builds each failed 1 of 21 tests; in use one would have hit 720 of 1,000 students, the other 14 students with a worse harm.
- Deciding when software is good enough is, in NIST's words of 2002, the major problem; decide with all three views.
Quality in Software Development
Test yourself
1. What is software quality? Give two standard definitions. ISO/IEC 25000:2014 defines it as the capability of a software product to satisfy stated and implied needs when used under specified conditions. IEEE 730-2014 defines it more narrowly as the degree to which a software product meets established requirements.
2. Why is the quality of software different from the quality of a manufactured product? Software's faults are design faults, not physical ones, and every copy is identical, so quality cannot be controlled by inspecting copies; it must be built in during development. Software does not wear out, but its reliability falls when its use changes or when a change to the code introduces a fault.
3. Explain Brooks's four essential properties of software and their effect on testing. Complexity: software has more states than can be listed, so exhaustive testing is impossible and tests must be designed. Conformity: it must fit interfaces designed by others, which calls for integration and compatibility testing. Changeability: all successful software gets changed, which calls for maintainable design and regression testing. Invisibility: its structure cannot be seen, so quality must be made visible through reviews, measures and tests.
4. Distinguish internal quality, external quality and quality in use, with an example of each for an online registration portal. Internal quality is the quality of the product itself, judged without running it, for example the complexity of the fee code found in a review. External quality is the quality of the running software, judged by testing, for example the number of failures in a test cycle. Quality in use is the quality experienced by real users in their context of use, for example whether students register correctly and without harm on the last date.
Quality in Software Development
5. Two builds each fail one of 21 tests. Why might one be far worse than the other? Because the test count weighs every input equally, while real users do not arrive equally. A defect in the commonest case, such as an on-time form, can hit hundreds of users while a defect on a rare day hits a few, and a defect that hits few can still do more harm, such as refusing a registration. Quality in use depends on the context of use and on the severity of the harm.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.