munotes®

Statistical Software Quality Assurance and Six Sigma

Get access to whole semester resourcesSemester Pass

Chapter Eighty-Nine

Syllabus topic Module 2, "Software Quality Assurance: ... Statistical Quality Assurance and Software Reliability"

Pages 520 to 525 of 622

In one line

Statistical software quality assurance makes quality decisions from defect data instead of impressions: collect and categorise the defects, trace each to its underlying cause, rank the causes to find the vital few, remove those causes, and measure again; Six Sigma turns the same idea into a disciplined project method, DMAIC, with a demanding target of 3.4 defects per million opportunities.

In the wording a student can write in an examination: statistical quality control is "The application of statistical techniques to control quality" (ASQ), and statistical process control the "statistically based analysis of a process and measures of process performance, which identify common and special causes of variation in process performance and maintain process performance within limits" (ISO/IEC/IEEE 24765). Applied to software, statistical SQA collects and categorises defect data, traces each defect to its cause, uses the Pareto principle ("most effects come from relatively few causes") to isolate the vital few causes, corrects them, and measures the result. Six Sigma "is a method that provides organizations tools to improve the capability of their business processes"; its projects follow DMAIC (Define, Measure, Analyze, Improve, Control), and "Six Sigma quality performance means 3.4 defects per million opportunities (accounting for a 1.5-sigma shift in the mean)" (ASQ).

Statistics instead of impressions

Every project has opinions about why its software fails: the testers blame rushed code, the developers blame changing requirements, the managers blame the testers. Statistical quality assurance replaces the opinions with counts. ISO's sixth quality management principle states the reason: "Decisions based on the analysis and evaluation of data and information are more likely to produce desired results."

For software the method has five steps, each already met in Module 2.

  1. Collect and categorise. Record every defect with its attributes: type, origin, severity, where it was found (Chapter Seventy-Three, on what a defect is).
  2. Trace each defect to its underlying cause. Not the symptom, not the person: the cause the process can change (Chapter Eighty, on using defect data to improve the process).
  3. Rank the causes. ASQ's glossary gives the principle: the Pareto chart, "named after 19th century economist Vilfredo Pareto", suggests "that most effects come from relatively few causes".
  4. Correct the vital few. Improvement effort goes to the few causes that account for most of the defects, Juran's "few vital breakthroughs" (Chapter Eighty-Two, on Shewhart, Deming and Juran).
  5. Measure again. Compare the next release with this one, against the causes that were not targeted, as Chapter Eighty did.

Worked example 1: release 2.0's defects by cause

When release 2.0's causal analysis traced each of its 200 defects to an underlying cause, nine kinds of cause emerged (the last row gathers eight rarer ones). The program ranks them.

# release 2.0's 200 defects by the underlying cause that causal analysis recorded (FINDINGS 5.2.7)
causes = {"requirement silent or ambiguous": 52, "rules misunderstood by developer": 30,
          "no shared validation": 28, "interface changed without notice": 24, "coding slip": 22,
          "screens never reviewed with users": 16, "test data unlike real records": 10,
          "documents not updated": 10, "eight other causes": 8}
total = sum(causes.values())

print(f"{'cause':<36}{'defects':>8}{'share':>7}{'cumulative':>12}")
running, vital = 0, []
for cause, n in sorted(causes.items(), key=lambda item: -item[1]):
    running += n
    if running - n < total / 2:                  # causes needed to account for half the defects
        vital.append(cause)
    print(f"{cause:<36}{n:>8}{n / total:>7.0%}{running / total:>12.0%}")
share = sum(causes[c] for c in vital) / total
print(f"the vital few: {len(vital)} of {len(causes)} causes, {share:.0%} of the defects")
munotes.in520

Statistical Software Quality Assurance and Six Sigma

cause                                defects  share  cumulative
requirement silent or ambiguous           52    26%         26%
rules misunderstood by developer          30    15%         41%
no shared validation                      28    14%         55%
interface changed without notice          24    12%         67%
coding slip                               22    11%         78%
screens never reviewed with users         16     8%         86%
test data unlike real records             10     5%         91%
documents not updated                     10     5%         96%
eight other causes                         8     4%        100%
the vital few: 3 of 9 causes, 55% of the defects

The vital few. Three causes of the nine account for 55 per cent of the defects, and all three concern the fee and eligibility rules: requirements silent or ambiguous about a rule, developers misunderstanding a rule, and no shared validation of the rules' inputs. The distribution is concentrated, though less than the 80/20 rule suggests: five causes are needed to pass three quarters. As Chapter Eighty-Two, on Shewhart, Deming and Juran, found, the rule describes a tendency, and the data decides.

What the ranking changes. Without it, a natural reaction to 200 defects is to test harder. The ranking points elsewhere: the largest cause lies in the requirements, before any code is written, and the top three share a single remedy, stating the rules once, precisely, and reviewing them with the exam cell. The actions of Chapter Eighty, on using defect data to improve the process (a shared validation module, a requirements review checklist item on invalid input), attacked part of this, and the release 2.1 figures showed the targeted defects falling. The ranking counts defects equally; Chapter One Hundred Four, on Pareto diagrams, weights them by cost and draws the chart.

Six Sigma

Six Sigma began at Motorola: in ASQ's history, "a methodology developed by Motorola to improve its business processes by minimizing defects", which "evolved into an organizational approach that achieved breakthroughs and significant bottom-line results." ASQ lists the threads its definitions share: teams assigned well-defined projects with a direct effect on the organisation; training in statistical thinking at every level, with specialists, called Black Belts, trained in advanced statistics and project management; the DMAIC approach to solving problems; and management support for the whole as a business strategy.

munotes.in521

Statistical Software Quality Assurance and Six Sigma

The name. In ASQ's glossary a sigma is "One standard deviation in a normally distributed process", and Six Sigma quality is "A term generally used to indicate process capability in terms of process spread measured by standard deviations in a normally distributed process." ASQ's Six Sigma page makes the picture concrete: a process at Six Sigma quality keeps its own variation within its control limits, three standard deviations from the centre line, while the requirement's tolerance limits lie six standard deviations away. Its output almost never falls outside the tolerance.

The number. Six Sigma's measure is defects per million opportunities (DPMO): the defects counted, divided by the opportunities for a defect, times a million, the same normalisation as ASQ's parts per million, "the number of defects normalized to a population of one million for ease of comparison." The target is ASQ's: "Six Sigma quality performance means 3.4 defects per million opportunities (accounting for a 1.5-sigma shift in the mean)". The shift is a convention of the calculation: a six sigma process is taken to operate with its mean 1.5 standard deviations off centre, leaving 4.5 standard deviations to the nearest limit.

DMAIC

ASQ describes DMAIC as "a data-driven quality strategy used to improve processes." Its five phases, with the tools ASQ lists for each:

PhaseWhat it does (ASQ)Typical tools (ASQ)
DefineThe problem, the improvement activity, the project goals, and the customer's requirementsProject charter; voice of the customer; value stream map
Measure"Measure process performance."Process map; capability analysis; Pareto chart
Analyze"Analyze the process to determine root causes of variation, poor performance (defects)."Root cause analysis; failure mode and effects analysis; multi-vari chart
Improve"Improve process performance by addressing and eliminating the root causes."Design of experiments; kaizen event
Control"Control the improved process and future process performance."Control plan; statistical process control; 5S; mistake proofing (poka-yoke)

For a new product or process, rather than an existing one, Six Sigma uses DMADV: Define, Measure, Analyze, Design and Verify. And ASQ notes a lineage: "The DMAIC process easily lends itself to the project approach to quality improvement encouraged and promoted by Juran."

Chapter Eighty's causal analysis of input validation defects was, in effect, a DMAIC project, and reading it phase by phase shows the method on software.

PhaseExamReg's input validation project (Chapter Eighty)
DefineThe problem: input validation, the largest defect type, 58 of release 2.0's 200 defects; the customer's requirement: every invalid input refused
Measure2.90 input validation defects per KLOC in release 2.0
AnalyzeThe five whys: no shared validation; requirements silent on invalid input
ImproveA shared validation module; a requirements review checklist item; unit test templates from partitions and boundaries
ControlThe checklist item and templates made standard; the rate watched release by release (1.12 per KLOC in release 2.1)
munotes.in522

Statistical Software Quality Assurance and Six Sigma

Worked example 2: sigma levels, and ExamReg's fees

The program first computes the defects per million opportunities at each sigma level, with the 1.5-sigma shift, from the normal distribution itself; the last row should reproduce ASQ's 3.4. It then measures release 2.0's first month of fees: the exam cell reconciled 18,000 fees, each made of three parts (the form fee, the backlog fee and the late fee), and found 9 parts wrong, on 9 different forms.

from statistics import NormalDist

def dpmo(sigma_level, shift=1.5):
    """Defects per million opportunities at a sigma level, with the conventional 1.5-sigma shift."""
    return 1e6 * (1 - NormalDist().cdf(sigma_level - shift))

def sigma_level(defects_per_million, shift=1.5):
    return NormalDist().inv_cdf(1 - defects_per_million / 1e6) + shift

for k in range(1, 7):
    print(f"{k} sigma: {dpmo(k):>11,.1f} defects per million opportunities")

forms, parts, wrong = 18_000, 3, 9        # release 2.0's first month of fees (FINDINGS 5.2.7)
for name, opportunities in [("every fee part", forms * parts), ("every form", forms)]:
    d = 1e6 * wrong / opportunities
    print(f"opportunity = {name}: {opportunities:,} opportunities, {d:.1f} DPMO,"
          f" sigma level {sigma_level(d):.2f}")
1 sigma:   691,462.5 defects per million opportunities
2 sigma:   308,537.5 defects per million opportunities
3 sigma:    66,807.2 defects per million opportunities
4 sigma:     6,209.7 defects per million opportunities
5 sigma:       232.6 defects per million opportunities
6 sigma:         3.4 defects per million opportunities
opportunity = every fee part: 54,000 opportunities, 166.7 DPMO, sigma level 5.09
opportunity = every form: 18,000 opportunities, 500.0 DPMO, sigma level 4.79

The scale. The sixth row reproduces ASQ's 3.4 exactly, which confirms how the number is built: the chance of a value beyond 4.5 standard deviations of a normal distribution, six sigma less the 1.5-sigma shift. The scale is steep. Moving from four to five sigma cuts the rate from 6,209.7 to 232.6 per million, and from five to six, to 3.4.

The month of fees. Counting every fee part as an opportunity, 9 wrong parts in 54,000 make 166.7 DPMO, a sigma level of 5.09. Counting every form as one opportunity, the same 9 errors make 500.0 DPMO and 4.79 sigma. Nothing about the portal changed between the two lines; only the definition of an opportunity did. A sigma level is therefore meaningless until the opportunity is defined and held fixed, the same lesson as the four defect counts of Chapter Seventy-Three and the line counts of Chapter Sixty-Seven, on size metrics.

Six Sigma and software. The fee example also shows where the metric fits software best: repeated outputs such as fees computed, forms processed or transactions completed, where each output is a natural opportunity. For defects in code there is no natural unit of opportunity (a line? a function? a requirement?), and a sigma level computed per line of code says more about the choice of unit than about quality. DMAIC, by contrast, is a problem-solving method that ASQ says "can be implemented as a standalone quality improvement procedure", and it carries over to software unchanged, as the input validation project shows.

munotes.in523

Statistical Software Quality Assurance and Six Sigma

Six Sigma and lean

ASQ distinguishes the two approaches often combined as lean Six Sigma: "Lean focuses on waste reduction, whereas Six Sigma emphasizes variation reduction." Lean, and CMMI, are the subject of Chapter One Hundred, on Lean, CMMI and choosing a methodology.

What it does not mean

Statistical SQA is not counting for its own sake. Its purpose is the ranking of causes and the action on the vital few; a count that changes no decision is not worth collecting.

The vital few are not always 20 per cent. The Pareto principle describes a concentration whose size the data decides; in release 2.0, 3 causes of 9 held 55 per cent.

A sigma level is not comparable without its definition. The same month of fees is 5.09 or 4.79 sigma depending on what counts as an opportunity.

Six Sigma is not the number six. It is a project method with trained people, management support and a measure; the 3.4 per million is its target, not its content.

Quick revision

  • Statistical quality control (ASQ): "The application of statistical techniques to control quality"; it includes acceptance sampling, which statistical process control does not.
  • Statistical SQA: collect and categorise defects; trace each to its underlying cause; rank the causes (Pareto: "most effects come from relatively few causes"); correct the vital few; measure again.
  • Six Sigma (Motorola; ASQ): teams on well-defined projects, Black Belts, DMAIC, management support; a sigma is one standard deviation; the target is 3.4 defects per million opportunities with the 1.5-sigma shift.
  • DPMO = defects divided by opportunities, times a million.
  • DMAIC: Define, Measure, Analyze, Improve, Control; DMADV (Define, Measure, Analyze, Design, Verify) for new products.
  • Lean against Six Sigma (ASQ): waste reduction against variation reduction.
  • Worked examples: 3 of 9 causes gave 55 per cent of release 2.0's defects, all about the fee rules; sigma table 6,209.7 (4), 232.6 (5), 3.4 (6); the month of fees 166.7 DPMO and 5.09 sigma per fee part, or 500.0 and 4.79 per form.

Test yourself

1. What is statistical software quality assurance? List its steps. The use of statistics on defect data to decide where to improve. Information about defects is collected and categorised; each defect is traced to its underlying cause; the causes are ranked, using the Pareto principle, to isolate the vital few; the vital few are corrected; and the effect is measured in the next release.

munotes.in524

Statistical Software Quality Assurance and Six Sigma

2. In the worked example, what did the ranking of causes show, and what action follows? That three of nine causes, all concerning the fee and eligibility rules (silent or ambiguous requirements, misunderstood rules, and no shared validation), accounted for 55 per cent of the defects. The action is to state the rules once, precisely, review them with the exam cell and validate them in one shared place, rather than simply to test harder.

3. What is Six Sigma, and what does "3.4 defects per million opportunities" mean? A method, begun at Motorola, that improves process capability through well-defined projects led by statistically trained staff, following DMAIC, with management support. The figure is its target: the rate of defects when the tolerance limits are six standard deviations from the process centre and, by convention, the mean is taken as shifted by 1.5 standard deviations, leaving 4.5 standard deviations to the nearest limit.

4. Explain the DMAIC phases, with a software example. Define the problem, goals and customer requirements; Measure process performance; Analyze to find the root causes of defects; Improve by eliminating the root causes; Control the improved process. For ExamReg: define the input validation problem (58 of 200 defects), measure 2.90 per KLOC, analyze with the five whys, improve with a shared validation module and a review checklist item, and control by making them standard and watching the rate (1.12 per KLOC in the next release).

5. How is DPMO computed, and why must the opportunity be defined? Defects divided by the number of opportunities for a defect, multiplied by a million. The opportunity must be defined and held fixed because the same data gives different results under different definitions: ExamReg's 9 wrong fee parts in 18,000 fees are 166.7 DPMO (5.09 sigma) per fee part but 500.0 DPMO (4.79 sigma) per form.

6. Distinguish Six Sigma from lean. Lean focuses on reducing waste, work that adds no value, and on standard work and flow; Six Sigma focuses on reducing variation, using statistical analysis. They are often combined as lean Six Sigma.

munotes.in525

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!