munotes®

Software Reliability

Get access to whole semester resourcesSemester Pass

Chapter Ninety

Syllabus topic Module 2, "Software Quality Assurance: ... Statistical Quality Assurance and Software Reliability"

Pages 526 to 530 of 622

In one line

Software reliability is how long, and how probably, software runs without failing in the conditions it is actually used in; it is measured by failures over time, summarised as mean time to failure, mean time to repair and availability, and it behaves unlike hardware reliability because software does not wear out but changes.

In the wording a student can write in an examination: software reliability is the "extent to which a system has operated without countable failures in a specified environment for a specified time span" (IEEE 982:2024), classically "the probability of failure-free software operation for a specified period of time in a specified environment" (the ANSI definition, as Lyu gives it). A failure is a "departure of system behavior from system requirements" (IEEE 982:2024). The mean time to failure (MTTF) is the expected time until the next failure; the mean time to repair (MTTR) the "expected or observed duration required to return a malfunctioning system or component to normal operations" (ISO/IEC/IEEE 24765); the mean time between failures (MTBF) the "expected or observed time between consecutive failures in a system or component" (the same standard). Availability is the "ratio of uptime divided by the sum of uptime plus downtime" (IEEE 982:2024), usually computed as MTTF / (MTTF + MTTR).

What reliability is, and is not

Reliability is one of the nine quality characteristics of ISO/IEC 25010:2023, which defines it as the "capability of a product to perform specified functions under specified conditions for a specified period of time without interruptions and failures" (Chapter Twenty, on the quality model today). Three parts of every definition matter.

  • Failures, not faults. Reliability counts what users experience. A fault that is never executed causes no failure and costs no reliability; one fault on the most used path can fail thousands of times.
  • A specified environment. Reliability depends on how the software is used. Lyu defines the operational profile as "the set of operations that the software can execute along with the probability with which they will occur", the same idea as the usage model of Chapter Eighty-Eight, on approaches to SQA. The same program can be highly reliable for one group of users and unreliable for another.
  • A time span. Reliability is always over a period: an hour, a session, a registration window. A number without its period means nothing.

Reliability is not correctness. Correctness is the "degree to which a system or component is free from faults in its specification, design, and implementation" (ISO/IEC/IEEE 24765). A program can be incorrect and still very reliable, if its faults lie where users rarely go: ExamReg's version that charged the wrong late fee at exactly 7 days (Chapter Fifty-Five, on boundary value analysis) failed only for students exactly a week late. And a program proved correct against a wrong specification can fail every day. Correctness asks about the program against its specification; reliability asks about the program in use.

munotes.in526

Software Reliability

Hardware fails by wearing out; software does not

The NIST/SEMATECH e-Handbook describes the failure rate of most manufactured products over their lives as the bathtub curve. It begins with an early failure period, "a high but rapidly decreasing failure rate"; then "the failure rate levels off and remains roughly constant" through the intrinsic failure period, where "most systems spend most of their lifetimes"; and finally the wearout failure period, when "the failure rate begins to increase as materials wear out and degradation failures occur at an ever increasing rate."

Software has no wearout period. In Lyu's words, software "does not wear out, burn out, or deteriorate, i.e., its reliability does not decrease with time." Instead, "software generally enjoys reliability growth during testing and operation since software faults can be detected and removed when software failures occur." And there is a way down as well: "software may experience reliability decrease due to abrupt changes of its operational usage or incorrect modifications to the software." Chapter Twenty-Three, on quality in software development, tabulated these differences; the figure shows them as curves.

Two failure-rate curves over time: hardware's bathtub curve, with early failures, a long flat period and rising wear-out; software's curve, falling as faults are removed and jumping at each change, with no wear-out

Figure 90.1 Hardware wears out; software's failure rate falls as faults are removed and jumps when it is changed

The software curve explains a practical rule. Every change to software, a fix, a new feature, a new kind of user, can move its failure rate up, so reliability figures belong to one version in one environment, and a new release is measured again.

The measures

Lyu defines the MTTF as the expected time until the next failure, adding that it "is also known as MTBF", the name ISO/IEC/IEEE 24765 gives to the "time between consecutive failures". The MTTR is the expected time to repair the system after a failure. From the two, Lyu gives availability, "the probability that a system is available when needed", as

Availability = MTTF / (MTTF + MTTR)

which is the same as IEEE 982:2024's uptime divided by uptime plus downtime, since over a period the failures divide the uptime into MTTF-sized pieces and the downtime into MTTR-sized ones.

Reliability over a time span needs a model of how failures arrive. The simplest is the exponential model, in which the failure rate is constant: the NIST/SEMATECH e-Handbook notes that "The exponential distribution is the only distribution to have a constant failure rate", that its reliability is R(t) = e^(-t/MTTF), and that "another name for the exponential mean is the Mean Time To Fail". It fits the flat part of the bathtub curve, and for software a version whose failure rate is not changing: no fixes, no new usage.

munotes.in527

Software Reliability

When a task needs several components working at once, the handbook's rule for independent components applies: "to calculate the reliability of a system of independent components, multiply the reliability functions of all the components together."

Worked example: release 2.0's first 60 days in use

During release 2.0's first 60 days in use, 1,440 hours, the ExamReg portal stopped serving students six times; the payment gateway it depends on, a third-party service, failed three times in the same days. These are failures, not defects: the log records when the service stopped and for how long, not which fault caused it. The program computes the measures and the reliability over two time spans, assuming a constant failure rate.

from math import exp

HOURS = 60 * 24                            # release 2.0's first 60 days in use (FINDINGS 5.5)
portal = [0.5, 1.0, 0.25, 2.0, 0.75, 1.5]  # hours each outage of the portal lasted
gateway = [0.5, 0.5, 1.0]                  # the payment gateway's outages in the same 60 days

def measures(outages):
    up = HOURS - sum(outages)
    return up / len(outages), sum(outages) / len(outages), up / HOURS   # MTTF, MTTR, availability

for name, outages in [("portal", portal), ("payment gateway", gateway)]:
    mttf, mttr, available = measures(outages)
    print(f"{name}: {len(outages)} failures; MTTF {mttf:.1f} h, MTTR {mttr:.2f} h;"
          f" availability {available:.2%}, and MTTF / (MTTF + MTTR) = {mttf / (mttf + mttr):.2%}")

mttf_portal, mttf_gateway = measures(portal)[0], measures(gateway)[0]
for span, hours in [("a 1-hour session", 1), ("the 15-day registration window", 15 * 24)]:
    r_portal = exp(-hours / mttf_portal)                 # exponential model: constant failure rate
    r_payment = r_portal * exp(-hours / mttf_gateway)    # paying needs both: reliabilities multiply
    print(f"{span}: portal without a failure {r_portal:.4f}; portal and gateway {r_payment:.4f}")
portal: 6 failures; MTTF 239.0 h, MTTR 1.00 h; availability 99.58%, and MTTF / (MTTF + MTTR) = 99.58%
payment gateway: 3 failures; MTTF 479.3 h, MTTR 0.67 h; availability 99.86%, and MTTF / (MTTF + MTTR) = 99.86%
a 1-hour session: portal without a failure 0.9958; portal and gateway 0.9937
the 15-day registration window: portal without a failure 0.2217; portal and gateway 0.1046

The measures. The portal ran 1,434 of the 1,440 hours: an MTTF of 239.0 hours, an MTTR of 1.00 hour and an availability of 99.58 per cent, which the formula MTTF / (MTTF + MTTR) reproduces exactly. The gateway failed half as often and recovered faster, at 99.86 per cent.

The time span decides. For one student in a 1-hour session, the portal gets through without a failure with probability 0.9958, and a payment, which needs the gateway too, with probability 0.9937. Over the whole 15-day registration window, the chance of no portal failure at all is only 0.2217, and of neither failing 0.1046. Both statements describe the same portal. A 99.58 per cent available system will very probably fail at least once during a fortnight, and a reliability requirement must therefore say which span it means: no failure during a student's session and no failure during the registration window are very different promises.

munotes.in528

Software Reliability

The model's limit. The exponential model assumes a constant failure rate, which holds for release 2.0 only as long as nobody changes it. The software curve in the figure is the warning: each fix or new release restarts the measurement. Chapter Ninety-Two, on measuring software reliability, fits models in which the failure rate falls as faults are removed.

What it does not mean

Reliability is not the absence of faults. It is the absence of failures in use; faults in unused code cost no reliability, and one fault on a common path costs a great deal.

Availability is not reliability. A system that fails often but recovers in seconds can be highly available and still unreliable; ExamReg was 99.58 per cent available and yet likely to fail during any fortnight.

Software does not wear out. Its reliability falls when it is changed or used differently, not with age.

MTTF is not a guarantee. It is an average; with a constant failure rate, a system with an MTTF of 239 hours fails within its first 239 hours more often than not.

Quick revision

  • Software reliability (IEEE 982:2024): operation "without countable failures in a specified environment for a specified time span"; (ANSI, via Lyu) "the probability of failure-free software operation for a specified period of time in a specified environment".
  • Failure (IEEE 982:2024): "departure of system behavior from system requirements". Operational profile (Lyu): the operations and their probabilities.
  • Reliability against correctness: correctness is freedom from faults against the specification; reliability is freedom from failures in use.
  • Hardware: the bathtub curve (early failure, intrinsic failure, wearout). Software: no wear-out; reliability grows as faults are removed; it drops with changes in usage or faulty modifications.
  • Measures: MTTF, MTTR, MTBF; availability = MTTF / (MTTF + MTTR) = uptime / (uptime + downtime); exponential model R(t) = e^(-t/MTTF) for a constant failure rate; independent components multiply.
  • Worked example: 6 outages in 1,440 hours: MTTF 239.0 h, MTTR 1.00 h, availability 99.58 per cent; a 1-hour session 0.9958 (0.9937 with the gateway); the 15-day window 0.2217 (0.1046).

Test yourself

1. Define software reliability, and explain why its definition names an environment and a time span. The probability of failure-free operation of software for a specified period of time in a specified environment. The environment matters because reliability depends on how the software is used, its operational profile; the time span matters because the probability of getting through without a failure falls as the period grows.

munotes.in529

Software Reliability

2. Distinguish reliability from correctness. Correctness is the degree to which software is free from faults in its specification, design and implementation; reliability is the degree to which it operates without failures in use. Software with faults in rarely used paths can be very reliable; software correct against a wrong specification can fail constantly.

3. Compare the failure curves of hardware and software. Hardware follows the bathtub curve: a falling early failure rate, a long flat intrinsic period, and a rising wear-out period. Software does not wear out: its failure rate falls as failures reveal faults that are removed, and rises when the software is changed or used in new ways, so its curve falls with jumps at changes.

4. Define MTTF, MTTR and availability, and give the formula connecting them. MTTF is the expected time until the next failure; MTTR the expected time to repair the system after a failure; availability the proportion of time the system is available when needed. Availability = MTTF / (MTTF + MTTR), equivalently uptime / (uptime + downtime).

5. A portal failed 6 times in 1,440 hours, with 6 hours of outages in all. Compute its MTTF, MTTR and availability. Uptime 1,434 hours; MTTF 1,434 / 6 = 239.0 hours; MTTR 6 / 6 = 1.00 hour; availability 1,434 / 1,440, or 239 / (239 + 1), which is 99.58 per cent.

6. Why is a system with 99.58 per cent availability still likely to fail during a two-week period? Because availability measures the share of time it is up, not the chance of an unbroken run. With an MTTF of 239 hours and a constant failure rate, the probability of no failure in 360 hours is e^(-360/239), about 0.22, so a failure during the period is more likely than not.

munotes.in530

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!