munotes®

Improving Software Reliability

Get access to whole semester resourcesSemester Pass

Chapter Ninety-Three

Syllabus topic Module 2, "Software Quality Assurance: ... Software reliability measurement and improvement"

Pages 543 to 547 of 622

In one line

Software becomes more reliable in four ways, which Lyu names fault prevention, fault removal, fault tolerance and fault/failure forecasting; testing to remove faults works but with diminishing returns, the operational profile tells a team which faults users would actually meet, and where failures can cause harm, reliability is not enough and safety must be engineered in its own right.

In the wording a student can write in an examination: reliability growth is the "improvement in reliability that results from correction of faults" (ISO/IEC/IEEE 24765). Lyu's four technical areas are: fault prevention, "To avoid, by construction, fault occurrences"; fault removal, "To detect, by verification and validation, the existence of faults and eliminate them"; fault tolerance, "To provide, by redundancy, service complying with the specification in spite of faults having occurred or occurring"; and fault/failure forecasting, "To estimate, by evaluation, the presence of faults and the occurrence and consequences of failures." The operational profile, "the set of operations that the software can execute along with the probability with which they will occur" (Lyu), directs testing and fixing to what users do most. Safety, the "expectation that a system does not, under defined conditions, lead to a state in which human life, health, property, or the environment is endangered" (ISO/IEC/IEEE 12207:2026), is a different property from reliability.

Four ways to more reliable software

Area (Lyu)What it doesTechniques Lyu namesOn ExamReg
Fault preventionStops faults being madeRefined requirements, good design methods, structured programming, clear code, formal methods, reuseThe shared validation module and requirements checklist item (Chapter Eighty, on using defect data)
Fault removalFinds and removes faults that were madeTesting, formal inspectionReviews, and every test level of Module 1
Fault toleranceKeeps the service right when a fault is triggeredMonitoring, atomicity of actions, decision verification, exception handling; recovery blocks, N-version programmingA payment that either completes or is rolled back whole
Fault/failure forecastingEstimates faults remaining and failures to comeReliability models, failure data, toolsMusa's model on release 2.1 (Chapter Ninety-Two, on measuring reliability)

NIST Special Publication 500-235 describes the effect of multiple versions: "any disagreement between versions can be reported during testing but a majority voting mechanism helps reduce the likelihood of incorrect output after delivery." Lyu's summary of prevention lists "The interactive refinement of the user's system requirement, the engineering of the software specification process, the use of good software design methods, the enforcement of a structured programming discipline, and the encouragement of writing clear code". For removal, besides testing, he names formal inspection, "a rigorous process focused on finding faults, correcting faults, and verifying the corrections" (Chapter Twenty-Nine, on inspection). His conclusion is that no single area is enough: "software reliability engineers must apply a combination of the above methods for the delivery of reliable software systems."

munotes.in543

Improving Software Reliability

Worked example 1: what more testing buys

Reliability growth testing is fault removal guided by forecasting: test, record failures, remove their faults, and watch the failure intensity fall. Chapter Ninety-Two fitted Musa's basic model to release 2.1's system test. Under that model the failure intensity falls exponentially with test time, so it halves every ln 2 / β1 hours. The program refits the model and follows four halvings.

from math import exp, log

def fit_basic(times, total):                     # Musa's basic model, as in Chapter 92
    n, s = len(times), sum(times)
    lo, hi = 1e-9, 1.0
    for _ in range(200):
        mid = (lo + hi) / 2
        g = n / mid - n * total / (exp(mid * total) - 1) - s
        lo, hi = (mid, hi) if g > 0 else (lo, mid)
    b1 = (lo + hi) / 2
    return n / (1 - exp(-b1 * total)), b1

# release 2.1's system test (FINDINGS 5.7): failure times in test hours, 260 hours in all
times, total = [3, 8, 14, 20, 28, 37, 47, 58, 71, 86, 103, 122, 145, 171, 200, 232], 260
b0, b1 = fit_basic(times, total)
halving = log(2) / b1                            # hours for the failure intensity to halve
print(f"each halving of the failure intensity takes {halving:.1f} more test hours")
t = total
for step in range(4):
    found = b0 * (exp(-b1 * t) - exp(-b1 * (t + halving)))     # expected failures in the step
    print(f"   {t:>5.0f} to {t + halving:>5.0f} h: intensity {b0 * b1 * exp(-b1 * t):.4f}"
          f" -> {b0 * b1 * exp(-b1 * (t + halving)):.4f}; failures expected {found:.2f}")
    t += halving
each halving of the failure intensity takes 78.3 more test hours
     260 to   338 h: intensity 0.0158 -> 0.0079; failures expected 0.89
     338 to   417 h: intensity 0.0079 -> 0.0039; failures expected 0.45
     417 to   495 h: intensity 0.0039 -> 0.0020; failures expected 0.22
     495 to   573 h: intensity 0.0020 -> 0.0010; failures expected 0.11

Each halving of the failure intensity costs the same 78.3 hours of testing, but finds half as many failures as the one before: 0.89 expected failures in the first 78 hours, then 0.45, 0.22 and 0.11. Testing buys reliability at a steadily rising price per failure found. That is the quantitative reason for Lyu's "combination": after a point, the next improvement is cheaper through prevention, which stops faults being made, or tolerance, which stops them becoming failures, than through yet more testing.

Worked example 2: fixing what students would meet

Not every fault costs the same reliability. A fault in an operation students use constantly causes more failures than one in an operation they rarely reach, and the operational profile says which is which. Release 2.1's system test ran 2,000 sessions generated from Chapter Eighty-Eight's usage model and logged the operation in which each of its 16 failures occurred. The program combines the failures per use of each operation with how often a student uses it in a session.

munotes.in544

Improving Software Reliability

# release 2.1's system test (FINDINGS 5.7): times each operation ran, and the failures in it
ran = {"log in": 2000, "fill form": 1100, "pay fee": 936, "payment failed": 94, "hall ticket": 1256}
failed = {"log in": 1, "fill form": 6, "pay fee": 5, "payment failed": 1, "hall ticket": 3}
# the operational profile: expected uses of each operation in one student's session (Chapter 88)
uses = {"log in": 1.0, "fill form": 0.55, "pay fee": 0.468, "payment failed": 0.047,
        "hall ticket": 0.628}

per_use = {op: failed[op] / ran[op] for op in ran}
per_session = {op: uses[op] * per_use[op] for op in ran}
total = sum(per_session.values())
print(f"{'operation':<16}{'failures per use':>18}{'share of the failures a student meets':>39}")
for op in sorted(ran, key=lambda o: -per_session[o]):
    print(f"{op:<16}{per_use[op]:>18.4f}{per_session[op] / total:>39.0%}")
print(f"probability that a session meets a failure: {total:.4f}")
operation         failures per use  share of the failures a student meets
fill form                   0.0055                                    38%
pay fee                     0.0053                                    31%
hall ticket                 0.0024                                    19%
log in                      0.0005                                     6%
payment failed              0.0106                                     6%
probability that a session meets a failure: 0.0080

The order of work. Filling the form and paying the fee account for 38 and 31 per cent of the failures a student would meet; improving those two operations improves reliability as students experience it most. The retry after a failed payment has the highest failure rate per use, 0.0106, yet only 6 per cent of the failures students meet, because few sessions reach it. Ranked by failures per use alone, it would come first; ranked by the operational profile, it comes last.

And yet. A failure in the retry could charge a student twice. The operational profile ranks faults by how often they would be met, not by what they would do, and a rare failure with serious consequences needs attention of its own. That is where reliability ends and safety begins.

Reliability is not safety

The standards define safety as the "expectation that a system does not, under defined conditions, lead to a state in which human life, health, property, or the environment is endangered" (ISO/IEC/IEEE 12207:2026), and a hazard as an "intrinsic property or condition that has the potential to cause harm or damage" (IEEE 1012-2024). Reliability counts failures; safety asks what a failure would do.

Leveson and Turner's account of the Therac-25, the radiation therapy machine of Chapter Three, on why software must be tested, names the confusion among its lessons: "This software was highly reliable. It worked tens of thousands of times before overdosing anyone, and occurrences of erroneous behavior were few and far between. AECL assumed that their software was safe because it was reliable, and this led to complacency." A second lesson is about tolerance: "The software did not contain self-checks or other error-detection and error-handling features" that would have caught the inconsistencies and coding errors.

munotes.in545

Improving Software Reliability

Safety is therefore engineered and assured separately. Hazards are identified and analysed; the software that could lead to them is found; defensive design (self-checks, interlocks, fault tolerance) is added where they could occur; and assurance confirms it. NASA's software assurance standard lists among its purposes "Ensuring that the software systems are safe and that the software safety-critical requirements are followed." ExamReg endangers no one's life, but it handles students' money, and the same reasoning applies at its scale: the payment retry is ExamReg's hazard, and it gets a model, tests and an atomic design of its own.

What it does not mean

Reliability is not improved only by testing. Testing removes faults with diminishing returns; prevention and tolerance supply what testing cannot afford.

The highest failure rate is not the first fix. Faults are ranked by the failures users would meet, which combines the rate with how often the operation is used, and then by their consequences.

Reliable is not safe. The Therac-25 worked tens of thousands of times; the rare failure killed people.

Fault tolerance is not an excuse for faults. It keeps the service right when a fault is triggered; the fault is still found and removed.

Quick revision

  • Reliability growth (ISO/IEC/IEEE 24765): improvement in reliability from correcting faults.
  • Four areas (Lyu): fault prevention (avoid by construction), fault removal (detect by V&V and eliminate), fault tolerance (service despite faults, by redundancy), fault/failure forecasting (estimate by evaluation); a combination is needed.
  • Fault tolerance techniques (Lyu): monitoring, atomicity of actions, decision verification, exception handling; recovery blocks and N-version programming by design diversity.
  • Operational profile (Lyu): operations and their probabilities; rank faults by the failures users would meet.
  • Safety (ISO/IEC/IEEE 12207:2026) is not reliability (Therac-25: "safe because it was reliable"); hazards need analysis and defensive design.
  • Worked examples: each halving of release 2.1's failure intensity takes 78.3 test hours and finds 0.89, 0.45, 0.22, 0.11 failures; form and payment give 38 and 31 per cent of the failures a student meets; the payment retry has the highest rate per use (0.0106) but a 6 per cent share; a session meets a failure with probability 0.0080.

Test yourself

1. Name and explain the four technical areas for achieving reliable software. Fault prevention avoids faults by construction, through good requirements, design methods and coding discipline; fault removal detects faults by verification and validation, such as testing and inspection, and eliminates them; fault tolerance uses redundancy and defensive techniques to keep the service correct when a fault is triggered; fault/failure forecasting estimates the faults remaining and the failures to come with reliability models.

munotes.in546

Improving Software Reliability

2. What is reliability growth testing, and why does it have diminishing returns? Testing in which failures are recorded and their faults removed, so that reliability grows and is tracked with a model. Under an exponential model each halving of the failure intensity takes the same test time but finds half as many failures as the previous halving; for release 2.1, 78.3 hours per halving, finding 0.89, then 0.45, then 0.22 failures.

3. How does the operational profile help improve reliability? It gives the probability with which each operation is used, so failures per use can be weighted by use to show which faults users would actually meet. For ExamReg, filling the form and paying caused 38 and 31 per cent of the failures students would meet, while the payment retry, with the highest failure rate per use, caused only 6 per cent.

4. Describe two fault tolerance techniques. In a single version of the software, exception handling and atomic actions (an operation that completes whole or not at all) partially tolerate faults, keeping a triggered fault from becoming a wrong result. With design diversity, functionally equivalent versions are developed independently; with several versions, a majority vote among their outputs reduces the chance of an incorrect output, the idea behind N-version programming (Lyu also names recovery blocks).

5. Distinguish reliability from safety, with an example. Reliability is how rarely a system fails; safety is whether its failures can lead to harm to life, health, property or the environment. The Therac-25's software worked tens of thousands of times, so it was highly reliable, but its rare failures overdosed patients; its maker assumed it was safe because it was reliable.

6. Why does the payment retry on ExamReg need attention although it contributes few failures? Because its consequence, charging a student twice, is serious even if rare. The operational profile ranks by how often failures would be met, not by their consequences, so hazardous rare functions are given their own analysis, tests and defensive design.

munotes.in547

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!