Measuring Software Reliability
Chapter Ninety-Two
Syllabus topic Module 2, "Software Quality Assurance: ... Statistical process control techniques: Software reliability measurement and improvement"
Pages 538 to 542 of 622
In one line
Software reliability is measured by recording failures against time, during system test and in use, and fitting a reliability model to the record; the model estimates the reliability reached so far and predicts how much more testing is needed to reach an objective, which turns is it ready? into a question with a number for an answer.
In the wording a student can write in an examination: failure intensity is the "quantitative characterization of in-service reliability as a ratio of the number of observed failures to the duration of an observed span" (IEEE 982:2024); the rate of occurrence of failures (ROCOF) is Lyu's other name for the failure rate function. A failure intensity objective (FIO) is the "failure intensity level to be achieved during pre-release testing", as part of a system's release criteria (IEEE 982:2024). Reliability data come as failure-count data (failures per period) or time-between-failures data. Estimation determines the reliability achieved so far; prediction forecasts it. Musa's basic execution time model assumes the expected number of failures by execution time t is β0(1 - e^(-β1 t)), so the failure intensity β0β1e^(-β1 t) falls exponentially as faults are removed.
What is measured
Chapter Ninety, on software reliability, measured a running system by its outages. Measuring reliability during development needs the same raw material, failures and times, collected while the software is tested. Lyu names the two forms such data take.
- Failure-count data track "the number of failures detected per unit of time": for example, 4 failures in the first 8 hours of testing, 4 in the next 8, 3 in the next.
- Time-between-failures data track "the intervals between consecutive failures": 0.5 hours to the first failure, then 1.2 hours to the second, and so on.
Either can be converted into the other. The time can be calendar time or execution time. Musa's model uses execution time, because, as Lyu reports, "Musa feels that execution time is more reflective of the actual stress induced on the software system than the amount of calendar time that has elapsed."
Failure intensity is the measure that summarises the data: failures per unit of time. Lyu's closely related failure rate function, "also called the rate of occurrence of failures", is the probability of a failure per unit time just after time t, given none before it. For software under test, failure intensity should fall as failures reveal faults that are removed; that falling curve is the reliability growth of Chapter Ninety's figure.
Estimation and prediction
Lyu separates two activities.
- Estimation "determines current software reliability by applying statistical inference techniques to failure data obtained during system test or during system operation": the reliability achieved up to now.
- Prediction "determines future software reliability based upon available software metrics and measures": from failure data once testing has begun, or from product and process measures before it (early prediction).
Measuring Software Reliability
A software reliability model makes prediction possible. It "specifies the general form of the dependence of the failure process on the principal factors that affect it: fault introduction, fault removal, and the operational environment." Its purpose, in Lyu's words, "is twofold: (1) to predict the extra time needed to test the software to achieve a specified objective; (2) to predict the expected reliability of the software when the testing is finished."
Musa's basic execution time model
Lyu calls Musa's basic model the one that "has had the widest distribution among the software reliability models", developed by John Musa of AT&T Bell Laboratories. Its main assumptions, as Lyu lists them:
- The cumulative number of failures by execution time t follows a Poisson process with mean value function μ(t) = β0(1 - e^(-β1 t)): the expected number of failures in a period is proportional to the expected number of faults still undetected. As t grows, μ(t) approaches β0, "the total number of faults that would be detected in the limit".
- The execution times between failures are exponentially distributed piece by piece: "the hazard rate for a single fault is constant."
From these follow the failure intensity λ(t) = β0β1e^(-β1 t), which "decreases exponentially to 0", and the reliability over a further Δt hours, R(Δt | t) = exp(-β0e^(-β1 t)(1 - e^(-β1 Δt))). The two parameters are estimated from the data by maximum likelihood. With n failures at times t1 to tn, and T the total time observed (the last failure time plus any failure-free time after it), the estimates satisfy
- β0 = n / (1 - e^(-β1 T)), and
- n / β1 - nT / (e^(β1 T) - 1) - (t1 + t2 + ... + tn) = 0,
the second equation being solved for β1 first.
Worked example: checking the model, then using it
A program that fits a model should be checked on data with a published answer before it is trusted with new data. Lyu's Example 3.3 gives ten failure times in CPU hours, 15 failure-free hours after the last, and its answer: β0 = 13.6 and β1 = 0.006. The program solves the same equations, checks that it gets Lyu's answer, and then fits ExamReg release 2.1's system test, in which 16 failures were logged over 260 hours of test execution. The release criterion is a failure intensity objective of 0.01 failures per test hour, one failure per 100 hours.
from math import exp, log
def fit_basic(times, total):
"""Musa's basic execution time model: the maximum likelihood estimates of beta0 and beta1
(Lyu, section 3.3.4.4) from the failure times and the total time observed."""
n, s = len(times), sum(times)
def g(b1): # the equation for beta1; it falls as beta1 grows
return n / b1 - n * total / (exp(b1 * total) - 1) - s
lo, hi = 1e-9, 1.0
for _ in range(200): # bisection
mid = (lo + hi) / 2
lo, hi = (mid, hi) if g(mid) > 0 else (lo, mid)
b1 = (lo + hi) / 2
return n / (1 - exp(-b1 * total)), b1
# Lyu's Example 3.3: ten failures, at these CPU hours, then 15 more hours without a failure
b0, b1 = fit_basic([10, 18, 32, 49, 64, 86, 105, 132, 167, 207], 222)
print(f"Lyu's example 3.3: beta0 {b0:.1f}, beta1 {b1:.3f} (the handbook gives 13.6 and 0.006)")
# ExamReg release 2.1's system test (FINDINGS 5.7): failures at these test hours, 260 hours in all
times = [3, 8, 14, 20, 28, 37, 47, 58, 71, 86, 103, 122, 145, 171, 200, 232]
total, objective = 260, 0.01 # objective: 0.01 failures per test hour
b0, b1 = fit_basic(times, total)
now = b0 * b1 * exp(-b1 * total) # failure intensity b0 b1 exp(-b1 t), at t = 260
print(f"ExamReg 2.1: beta0 {b0:.2f} failures in all, beta1 {b1:.5f};"
f" {b0 - len(times):.2f} failures expected still to come")
print(f"failure intensity now {now:.4f} per test hour (MTTF {1 / now:.1f} h)")
r24 = exp(-b0 * exp(-b1 * total) * (1 - exp(-b1 * 24)))
print(f"probability of no failure in the next 24 test hours: {r24:.3f}")
print(f"test hours still needed to reach {objective}: {log(b0 * b1 / objective) / b1 - total:.1f}")Measuring Software Reliability
Lyu's example 3.3: beta0 13.6, beta1 0.006 (the handbook gives 13.6 and 0.006)
ExamReg 2.1: beta0 17.78 failures in all, beta1 0.00885; 1.78 failures expected still to come
failure intensity now 0.0158 per test hour (MTTF 63.4 h)
probability of no failure in the next 24 test hours: 0.711
test hours still needed to reach 0.01: 51.5The check. The program reproduces Lyu's answer, β0 = 13.6 and β1 = 0.006, so its equations and its solver can be trusted with new data.
Estimation. For release 2.1 the model estimates 17.78 failures in all, of which 16 have been seen, so about 1.78 remain to be found by testing of this kind. The failure intensity now is 0.0158 failures per test hour, a mean time to failure of 63.4 hours: the gaps between failures grew from 5 hours at the start to over 30 at the end, and the model has turned that growth into a current rate.
Prediction. The objective is 0.01 failures per test hour, and the current 0.0158 has not reached it. Setting β0β1e^(-β1 t) equal to 0.01 and solving for t gives the answer to Lyu's first purpose: about 51.5 more hours of system testing. For the second, the chance of getting through the next 24 test hours without a failure is 0.711. Instead of the testers feel it is nearly ready, the release decision can say on this model, 52 more test hours reach the objective the college agreed to.
Measuring Software Reliability
The limits. The numbers are only as good as the model's assumptions. Execution time, a constant hazard per fault and testing that resembles real use (the operational profile of Chapter Ninety, on software reliability) must all roughly hold; the likelihood equation for β1 has a solution only when failures are thinning out; and a model should be checked against how well it predicted before its forecasts are relied on. Lyu notes that Musa recommends this model particularly when predicting early reliability, when the program is changing substantially during the data collection, and when studying the effect of a new engineering technique.
Measuring reliability in a project
| Step | What is done | For ExamReg release 2.1 |
|---|---|---|
| Define failure | What counts as a failure, by severity | Any departure from the requirements a user would notice |
| Set an objective | A failure intensity objective for release | 0.01 failures per test hour |
| Collect data | Failure times (or counts) in execution time, during testing that follows the operational profile | 16 failure times over 260 test hours |
| Fit a model | Estimate its parameters; check it on known data first | Musa's basic model; checked on Lyu's Example 3.3 |
| Estimate and predict | Current intensity; further testing needed; reliability ahead | 0.0158 per hour now; about 51.5 more hours |
| Decide | Release when the objective is met, or continue testing | Continue testing |
What it does not mean
A reliability model does not find faults. It summarises failures already found and forecasts the next ones; testing finds faults.
A predicted number is not a promise. It holds only while the model's assumptions hold; a change to the software or to how it is tested restarts the measurement.
Fewer failures in a week is not proof of reliability growth. It may mean less testing was done; that is why time is measured in execution time, not calendar time.
Failures are not defects. One defect can cause many failures, and reliability counts what users would experience, not what the code contains.
Quick revision
- Failure intensity (IEEE 982:2024): observed failures divided by the duration observed. ROCOF: rate of occurrence of failures. Failure intensity objective: the level to be reached in pre-release testing as a release criterion.
- Data (Lyu): failure-count data (failures per period) and time-between-failures data; execution time preferred to calendar time.
- Estimation (reliability achieved so far) and prediction (future reliability; the extra testing needed to reach an objective).
- Musa's basic execution time model (Lyu 3.3.4): μ(t) = β0(1 - e^(-β1 t)); λ(t) = β0β1e^(-β1 t); β0 = total failures in the limit; R(Δt | t) = exp(-β0e^(-β1 t)(1 - e^(-β1 Δt))); maximum likelihood estimates from failure times and total time.
- Worked example: Lyu's Example 3.3 reproduced (13.6, 0.006); release 2.1: 17.78 failures expected in all (1.78 to come), 0.0158 failures per hour now (MTTF 63.4 h), 0.711 probability of 24 failure-free hours, about 51.5 more hours to reach 0.01.
Measuring Software Reliability
Test yourself
1. What data are collected to measure software reliability? Failures and the times at which they occur, either as failure-count data, the number of failures in each period, or as time-between-failures data, the intervals between consecutive failures. Time is preferably execution time, and failures are counted according to a stated definition and severity.
2. Define failure intensity and the failure intensity objective. Failure intensity is the number of observed failures divided by the duration observed, the rate at which failures occur. The failure intensity objective is the level of failure intensity to be reached during pre-release testing as part of the release criteria.
3. Distinguish reliability estimation from reliability prediction. Estimation applies statistical inference to failure data from testing or operation to determine the reliability achieved so far. Prediction determines future reliability, from failure data through a reliability model, or, before testing, from product and process measures (early prediction).
4. State Musa's basic execution time model. The cumulative number of failures by execution time t is a Poisson process with mean value function β0(1 - e^(-β1 t)), where β0 is the total number of failures that would be seen in the limit; the failure intensity β0β1e^(-β1 t) falls exponentially as faults are removed, and the time between failures is exponential with a constant hazard for each fault.
5. How does a reliability model support the release decision? By estimating the current failure intensity from the failure data and predicting the further testing needed to reach the failure intensity objective. For ExamReg release 2.1 the intensity was 0.0158 failures per test hour against an objective of 0.01, and the model predicted about 51.5 more hours of testing.
6. Why was the program run on Lyu's example first? To check that its equations and its numerical solution are right: it reproduced the published answer, β0 = 13.6 and β1 = 0.006, before being trusted with release 2.1's data, whose answer nobody knows in advance.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.