Why Software Must Be Tested
Chapter Three
Syllabus topic Module 1, "Software Testing Fundamentals: Nature of errors and the need for testing"
Pages 15 to 22 of 622
In one line
Software must be tested because software that fails costs money, time and trust, and sometimes lives, and testing is the cheapest way to find a failure before it reaches the people who depend on the software.
In the wording a student can write in an examination: testing is necessary because defects are inevitable in human work, because a defect found late costs far more to fix than one found early, because failures in operation cause financial, legal, reputational and human harm, and because testing provides the evidence on which a release decision, a contract or a regulator relies. Testing reduces the risk of failure in operation; it cannot remove it.
Why the question is worth asking
Every project is under pressure to ship. Testing takes time, needs people, and produces reports that nobody enjoys reading. It is tempting to treat it as the part that can be squeezed. The ISTQB syllabus puts the other side of the argument plainly: software that does not work correctly "can lead to many problems, including loss of money, time or business reputation, and, in extreme cases, even injury or death."
The honest way to answer "why test?" is not with a slogan but with what actually happened when testing was missing. The four failures below are among the best documented in the history of software, and each is told here from the official investigation, not from retellings, because retellings drift.
Ariane 5, 4 June 1996: code reused without the test that mattered
The European Space Agency's inquiry board opens its report with the event: on 4 June 1996 the maiden flight of the Ariane 5 launcher ended in failure, and "Only about 40 seconds after initiation of the flight sequence, at an altitude of about 3700 m, the launcher veered off its flight path, broke up and exploded."
The cause, as the board found it, was in the inertial reference system (the SRI), whose software was reused from Ariane 4. An internal exception "was caused during execution of a data conversion from 64-bit floating point to 16-bit signed integer value." The number being converted, a horizontal bias value called BH, was larger than a 16-bit signed integer can hold. It was larger because "the early part of the trajectory of Ariane 5 differs from that of Ariane 4 and results in considerably higher horizontal velocity values." And the function doing the conversion was not even needed: it performed alignment, and "As soon as the launcher lifts off, this function serves no purpose."
The conversion is easy to reproduce. A 16-bit signed integer holds values from -32768 to 32767, and Python's struct module refuses to pack anything larger into that format, which is exactly the kind of refusal the SRI's software raised.
Why Software Must Be Tested
import struct
def to_16_bit(value):
"""Pack a number into a 16-bit signed integer, as the SRI's conversion did."""
return struct.pack(">h", int(value))
for horizontal_bias in [1250.0, 32767.0, 40000.0]:
try:
to_16_bit(horizontal_bias)
print(f"{horizontal_bias:>9}: fits in 16 bits")
except struct.error as problem:
print(f"{horizontal_bias:>9}: conversion refused ({problem})") 1250.0: fits in 16 bits
32767.0: fits in 16 bits
40000.0: conversion refused ('h' format requires -32768 <= number <= 32767)The three values are this book's, chosen to straddle the limit; the report does not publish the flight's value of BH. What the board does say about testing is the lesson. Equipment testing of the SRI was rigorous, "However, no test was performed to verify that the SRI would behave correctly when being subjected to the count-down and flight time sequence and the trajectory of Ariane 5." And then: "Had such a test been performed by the supplier or as part of the acceptance test, the failure mechanism would have been exposed." The test was never written because the SRI's specification did not contain the Ariane 5 trajectory data as a requirement: a defect in a requirement, exactly the kind Chapter Two, on errors, faults and failures, warned propagates.
The Patriot at Dhahran, 25 February 1991: an error too small to see, run for too long
The United States General Accounting Office (GAO) report IMTEC-92-26 begins: "On February 25, 1991, a Patriot missile defense system operating at Dhahran, Saudi Arabia, during Operation Desert Storm failed to track and intercept an incoming Scud. This Scud subsequently hit an Army barracks, killing 28 Americans."
The Patriot's clock counted time "in tenths of seconds", and to predict where a Scud would appear next, the count had to be converted to a real number. Its registers were 24 bits long, so the conversion "cannot be any more precise than 24 bits." One tenth has no exact binary form, so every tick of the clock was stored very slightly short, and the shortfall accumulated for as long as the system ran. At Dhahran the battery "had been operating continuously for over 100 hours."
How short, exactly? GAO says only that the registers were 24 bits long; it does not print the format in which the tenth was held. So the program below tries two readings and lets GAO's own published figures decide. In the first, the tenth is kept to 24 binary places after the point and the rest chopped off. In the second, it is kept to 23, which is what a 24-bit register leaves for the fraction if one bit holds the sign. The figures GAO printed in its Appendix II are typed in beside the program, from the page image of the report.
Why Software Must Be Tested
from fractions import Fraction
tenth = Fraction(1, 10)
def drift(places, hours):
"""Seconds lost after `hours` if each tenth is chopped to `places` binary places."""
stored = Fraction(int(tenth * 2**places), 2**places)
return hours * 3600 * 10 * (tenth - stored)
for places in (24, 23):
print(f"kept to {places} binary places: drift after 100 h = {float(drift(places, 100)):.4f} s")
print()
# GAO/IMTEC-92-26, Appendix II: hours, calculated time, inaccuracy (seconds)
gao = [(1, "3599.9966", "0.0034"), (8, "28799.9725", "0.0275"),
(20, "71999.9313", "0.0687"), (48, "172799.8352", "0.1648"),
(72, "259199.7528", "0.2472"), (100, "359999.6667", "0.3433")]
for hours, gao_time, gao_drift in gao:
ours = f"{float(drift(23, hours)):.4f}"
check = hours * 3600 - Fraction(gao_drift)
note = "" if f"{float(check):.4f}" == gao_time else f" <- GAO's time should be {float(check):.4f}"
print(f"{hours:>3} h: drift {ours} s, GAO printed {gao_drift}{note}")kept to 24 binary places: drift after 100 h = 0.1287 s
kept to 23 binary places: drift after 100 h = 0.3433 s
1 h: drift 0.0034 s, GAO printed 0.0034
8 h: drift 0.0275 s, GAO printed 0.0275
20 h: drift 0.0687 s, GAO printed 0.0687
48 h: drift 0.1648 s, GAO printed 0.1648
72 h: drift 0.2472 s, GAO printed 0.2472
100 h: drift 0.3433 s, GAO printed 0.3433 <- GAO's time should be 359999.6567The first reading gives a drift of about an eighth of a second at 100 hours, which is not what GAO measured, so it is rejected. The second reproduces every figure in GAO's inaccuracy column, to four decimal places: after 100 hours the clock was a third of a second out. In a third of a second a missile flying at "approximately MACH 5" covers hundreds of metres, and GAO's table gives the shift of the radar's range gate as 687 metres, enough for the system to "look in the wrong place for the incoming Scud."
Two small lessons sit inside this one. A model is tested against data, and the first one here failed its test: that is how the 23 was found, not assumed. And the program found something in the report itself. In GAO's last row the calculated time (359999.6667) and the inaccuracy (.3433) cannot both be right, because 360000 minus .3433 is 359999.6567. The inaccuracy column is the one the drift model confirms. It is a small slip in a careful report, and it is why a table of figures is worth recomputing rather than trusting.
The report's account of the timing is the sharpest lesson. Israeli data had shown a loss of accuracy after 8 consecutive hours, and "Army officials modified the software to improve the system's accuracy. However, the modified software did not reach Dhahran until February 26, 1991", the day after the Scud struck. No test had run the system for as long as it was actually used. GAO notes that the Patriot had never before been used against Scud missiles "nor was it expected to operate continuously for long periods of time."
Why Software Must Be Tested
Therac-25, 1985 to 1987: software trusted to be safe
Nancy Leveson's account "Medical Devices: The Therac-25", drawn from her investigation with Clark S. Turner (IEEE Computer, July 1993), begins: "Between June 1985 and January 1987, a computer-controlled radiation therapy machine, called the Therac-25, massively overdosed six people." Eleven machines had been installed, five in the United States and six in Canada.
The design had shifted safety from hardware to software. Earlier machines in the family had "independent protective circuits" and "mechanical interlocks for policing the machine", but "In the Therac-25, software checks were substituted for many of the traditional hardware interlocks." Code was reused from the earlier machines too, and one of the bugs was later found in the older Therac-20's software, where it had hurt nobody because that machine "includes hardware safety interlocks."
Two software defects stand out, and each was invisible to ordinary testing. At Tyler, Texas, the overdose happened only when an experienced operator edited the prescription quickly: "If the prescription data was edited at a fast pace (as is natural for someone who has repeated the procedure a large number of times), the overdose occurred." The hospital physicist had to practise before he could reproduce it, and the manufacturer's engineer at first "could not reproduce the error" at all. At Yakima, Washington, a one-byte variable called Class3 was "incremented by one in each pass" through a setup routine, and a zero value meant a safety check was skipped: "on every 256th pass through the Set Up Test code, the variable will overflow and have a zero value."
The second defect takes four lines to reproduce.
def skipped_checks(passes, fixed=False):
"""Passes on which the upper collimator check is skipped."""
class3, skipped = 0, []
for pass_number in range(1, passes + 1):
class3 = 1 if fixed else (class3 + 1) % 256 # one byte: 255 + 1 wraps to 0
if class3 == 0: # zero is read as "nothing to check"
skipped.append(pass_number)
return skipped
print("as shipped:", skipped_checks(1000))
print("as fixed :", skipped_checks(1000, fixed=True))as shipped: [256, 512, 768]
as fixed : []The check is skipped on three passes in a thousand, and only when the operator happens to press the button on exactly one of them. The fix, as the account reports it, was "simple: the program is changed so that the Class3 variable is set to some fixed nonzero value each time through Set Up Test instead of being incremented." A defect with a one-line fix killed and injured people because no test drove the setup routine through hundreds of passes, and none drove the keyboard as fast as a practised operator did.
Why Software Must Be Tested
Knight Capital, 1 August 2012: one server missed, and dead code woken
The United States Securities and Exchange Commission's order (Release 70694, 2013) records that on August 1, 2012 Knight Capital's automated order router, SMARS, "routed millions of orders into the market over a 45-minute period, and obtained over 4 million executions in 154 stocks for more than 397 million shares." It had been processing just 212 small retail orders. "Knight lost over $460 million from these unwanted positions."
The chain of events is a lesson in two kinds of testing at once. New code was deployed to replace an old, unused function called Power Peg, and the new code "repurposed a flag that was formerly used to activate the Power Peg code." During deployment "one of Knight's technicians did not copy the new code to one of the eight SMARS computer servers," and "Knight did not have a second technician review this deployment". On the eighth server the repurposed flag woke the old code. And that old code had itself been changed years earlier without a retest: in 2005 a function it relied on was moved, and "Knight did not retest the Power Peg code after moving the cumulative quantity function to determine whether Power Peg would still function correctly if called."
Two things were missing: a regression test after a change to old code, and a review (a static check) of the deployment. Chapter Thirty-Nine, on regression testing, smoke testing and continuous integration, and Chapter Twenty-Eight, on software reviews, are about exactly these.
What failures cost a whole economy
Famous disasters are rare; ordinary failures are constant, and their cost adds up. In 2002 the United States National Institute of Standards and Technology (NIST) surveyed software developers and users and estimated that "the national annual cost estimates of an inadequate infrastructure for software testing are estimated to be $59.5 billion." It estimated that "The potential cost reduction from feasible infrastructure improvements is $22.2 billion." Software users bore about 60 per cent of the cost and developers about 40 per cent.
The figures are more than twenty years old and were for one country, so the exact amounts are history. The shape of the finding is what lasts: most of the cost of poor testing is paid by users, not by the people who wrote the software, and a large part of it could be avoided by testing better and earlier.
Why late is expensive
The failures above were found in operation, the most expensive place to find anything. Barry Boehm and Victor Basili, in their 2001 "Software Defect Reduction Top 10 List", put the first item on the list this way: "Finding and fixing a software problem after delivery is often 100 times more expensive than finding and fixing it during the requirements and design phase."
Why Software Must Be Tested
The reason is easy to see on ExamReg. A missing case in the late-fee rule, caught in a review of the requirement, costs a sentence. Caught in testing, it costs a code change and a retest. Caught after release, it costs refunds to students who were overcharged, a correction to every receipt, a public apology and a patch deployed under pressure, which is exactly the situation in which Knight Capital's technician missed a server. Chapter Ninety-Six returns to this with a defect amplification model and why reviews pay.
What testing contributes
The ISTQB syllabus lists what testing contributes to a project's success, and it is more than finding bugs:
- It "provides a cost-effective means of detecting defects", which debugging then removes, so testing contributes indirectly to higher quality.
- It "provides a means of directly evaluating the quality of a test object", and those measures feed decisions such as whether to release.
- It "provides users with indirect representation on the development project", because testers carry the users' needs into a project users cannot join.
- It "may also be required to meet contractual or legal requirements, or to comply with regulatory standards."
The syllabus also separates testing from quality assurance: testing is "a product-oriented, corrective approach" and "a major form of quality control", while QA is "a process-oriented, preventive approach". That distinction is Module 1's third row, and Chapter Twenty-Four, on quality control and quality assurance, gives it in full.
The four failures, side by side
| Ariane 5 (1996) | Patriot (1991) | Therac-25 (1985 to 1987) | Knight Capital (2012) | |
|---|---|---|---|---|
| What failed | Launcher destroyed about 40 seconds into flight | Missed an incoming Scud; 28 killed | Radiation overdoses | Millions of unintended orders in 45 minutes |
| The defect | Unprotected 64-bit to 16-bit conversion in reused code | Time held in 24-bit registers; drift grew with running time | A race on fast editing; a one-byte counter skipping a check every 256th pass | Old code woken by a repurposed flag on one server |
| Test that was missing | Test with Ariane 5's own trajectory | Test of continuous operation for as long as real use | Tests at an expert operator's speed, and of hundreds of setup passes | Regression test after a change; review of the deployment |
| Kind of testing, in this book | System and acceptance testing with realistic data | Stress and endurance testing | Stress and timing tests; safety reviews | Regression testing; reviews |
What it does not mean
Testing does not guarantee a failure-free product. Every one of these systems was tested. Testing reduces risk; the missing test in each case was one that nobody thought to run.
Why Software Must Be Tested
Reused code is not already tested. Ariane 5 reused Ariane 4's software, and Knight reused an old flag. Code tested in one context is untested in a new one.
Testing is not only for rockets and hospitals. NIST's estimate is dominated by ordinary business software. A portal that overcharges students is a failure in the same sense.
More testing is not automatically better. Testing costs money, and the question is always which tests reduce the most risk for the effort. Chapter Four's principle that testing is context dependent says the same.
Quick revision
- Failures cost money, time and reputation, and in extreme cases cause injury or death (ISTQB v4.0.1).
- Ariane 5 (4 June 1996): unprotected 64-bit to 16-bit conversion of BH in reused Ariane 4 code; no test with Ariane 5's trajectory; "Had such a test been performed ... the failure mechanism would have been exposed."
- Patriot (25 February 1991): time converted in 24-bit registers, each tenth of a second held slightly short; after 100 hours, 0.3433 s of drift and a 687 m range-gate shift; 28 killed; the fix arrived the next day.
- Therac-25 (June 1985 to January 1987): six people massively overdosed; software checks replaced hardware interlocks; a race on fast editing, and a one-byte counter that skipped a safety check every 256th pass.
- Knight Capital (1 August 2012): new code missed on one of eight servers, a repurposed flag woke old code never retested; over $460 million lost in 45 minutes.
- NIST 2002: inadequate testing infrastructure cost the US an estimated $59.5 billion a year; $22.2 billion avoidable.
- Boehm and Basili 2001: fixing after delivery is often 100 times more expensive than in requirements and design.
- Testing detects defects, evaluates quality, represents users and meets legal requirements; it is quality control, not quality assurance.
Test yourself
1. Give four reasons why software testing is necessary. Defects are inevitable in human work; defects found late cost far more to fix than those found early; failures in operation cause financial, legal, reputational and human harm; and testing provides the evidence for release decisions and for contractual and regulatory compliance.
2. What caused the Ariane 5 failure, and which test would have exposed it? Software reused from Ariane 4 converted a 64-bit floating-point value, the horizontal bias, into a 16-bit signed integer without protection, and Ariane 5's faster early trajectory made the value too large, raising an exception in both inertial reference units. The inquiry board found that a ground test injecting Ariane 5's own trajectory would have exposed it.
3. Explain how a tiny error made the Patriot miss the Scud at Dhahran. Its clock counted tenths of a second and converted the count to a real number with only 24 bits of precision, so each tick was stored slightly short. The shortfall grew with running time: after about 100 hours of continuous operation it reached about a third of a second, shifting the radar's range gate about 687 metres, so the system looked in the wrong place.
Why Software Must Be Tested
4. What two missing checks allowed the Knight Capital failure? A review of the deployment, which would have noticed that one of eight servers had not received the new code; and a regression test of the old Power Peg code after it was changed in 2005, which was never done.
5. According to Boehm and Basili, how much more does it cost to fix a problem after delivery? Often 100 times more than finding and fixing it during the requirements and design phase.
6. Is testing the same as quality assurance? No. Testing is product-oriented and corrective, a major form of quality control; quality assurance is process-oriented and preventive, working on the basis that a good process produces a good product.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.