munotes®

Inspection

Get access to whole semester resourcesSemester Pass

Chapter Twenty-Nine

Syllabus topic Module 1, "Verification and Validation (V&V): Inspection"

Pages 158 to 164 of 622

In one line

An inspection is the most formal kind of review: a trained moderator leads a small team with defined roles through five planned operations, the meeting's only aim is to find errors, and the errors found are classified, counted and used to improve the process that made them.

In the wording a student can write in an examination: an inspection is a "formal review of a work product to identify issues, which uses defined team roles and measurement to improve the review process" (ISO/IEC 20246:2017). The method is M. E. Fagan's, published in the IBM Systems Journal in 1976. Its five operations are overview, preparation, inspection, rework and follow-up; its four roles are moderator, designer, coder/implementor and tester, with one participant acting as reader. Fagan reported that inspections found 82 per cent of the errors in one application before testing, raised coding productivity by 23 per cent in one systems programming study, and produced 38 per cent fewer errors than a comparable walk-through sample.

Where inspection came from

Fagan's paper opens from the cost of late errors: "The cost of reworking errors in programs becomes higher the later they are reworked in the process, so every attempt should be made to find and fix errors as early in the process as possible." His answer was a process with inspections at fixed points, each with exit criteria, and he described inspections in three adjectives: "Inspections are a formal, efficient, and economical method of finding errors in design and code."

Fagan's programming process placed inspections after each major work product:

  • I0, after the internal (module) specifications;
  • I1, the design-complete inspection, after the logic specifications;
  • I2, the code inspection, once the code reaches its first clean compilation;
  • IT1 and IT2, inspections of the test plan and of the test cases;
  • PI0, PI1 and PI2, inspections of the publications, because poor documentation "can mislead the user, causing him to make errors quite as important as errors in the program."

The same method serves all of them. The paper describes it through I1 and I2, and says the others "retain the same essential properties" but differ "in materials inspected, number of participants, and some other minor points."

The five operations

Fagan's Table 3 lists the operations, the objective of each, and the rate at which each goes for systems programming.

OperationWhoObjective (Fagan's Table 3)Design I1 rateCode I2 rate
1. OverviewWhole teamCommunication, education500 lines per hourNot necessary
2. PreparationEach person aloneEducation100 lines per hour125 lines per hour
3. InspectionWhole teamFind errors130 lines per hour150 lines per hour
4. ReworkDesigner or coderRework and resolve errors found by inspection20 hours per K.NCSS16 hours per K.NCSS
5. Follow-upModeratorSee that all errors, problems and concerns have been resolvedNone givenNone given
munotes.in158

Inspection

K.NCSS is a thousand non-commentary source statements, roughly a thousand lines of code without comments. Each operation has its own rules.

  1. Overview. "The designer first describes the overall area being addressed and then the specific area he has designed in detail", and the design documents are handed out. A code inspection needs no overview, because the same people inspected the design.
  2. Preparation. Participants "literally do their homework to try to understand the design, its intent and logic." They also study the ranked distributions of error types found by recent inspections, and checklists of clues, so that they look where errors are most likely.
  3. Inspection. A reader chosen by the moderator, usually the coder, paraphrases the design or code. "Every piece of logic is covered at least once, and every branch is taken at least once." Each error found is noted, its type classified and its severity (major or minor) recorded, and the reading moves on. No one designs solutions at the meeting: "The inspection is not intended to redesign, evaluate alternate design solutions, or to find solutions to errors; it is intended just to find errors!" Within one day the moderator writes the inspection report.
  4. Rework. "All errors or problems noted in the inspection report are resolved by the designer or coder/implementor."
  5. Follow-up. The moderator checks that every issue is resolved. "If more than five percent of the material has been reworked, the team should reconvene and carry out a 100 percent reinspection."

Two practical rules come from Fagan's experience. The meeting's error detection falls off after two hours, so "it is advisable to schedule inspection sessions of no more than two hours at a time. Two two-hour sessions per day are acceptable." And the time for inspections and rework "must be scheduled and managed with the same attention as other important project activities", because under pressure inspections are the first thing a project drops.

Fagan also defined what the meeting is looking for: "an error is defined as any condition that causes malfunction or that precludes the attainment of expected or previously specified results."

The roles

"The inspection team is best served when its members play their particular roles", Fagan wrote, and he named four.

RoleFagan's description
Moderator"The key person in a successful inspection." A competent programmer, but not necessarily an expert on the program; best from an unrelated project, to preserve objectivity. Manages the team; in Fagan's words, "he is the coach". Schedules the meetings, reports within one day, follows up the rework, and should be specially trained
Designer"The programmer responsible for producing the program design."
Coder/implementor"The programmer responsible for translating the design into code."
Tester"The programmer responsible for writing and/or executing test cases or otherwise testing the product of the designer and coder."
munotes.in159

Inspection

The reader is not a fifth person: it is a job the moderator gives, usually to the coder. When one person has done two jobs, the roles are refilled from outside: if the same person designed and coded the work, they take the designer's role and "a coder from some related or similar program will perform the role of the coder." On size, "Four people constitute a good-sized inspection team", and the team "should not be artificially increased over four" unless the code touches several interfaces whose owners should be present.

The modern syllabus keeps the essential rule in a sentence: "In inspections, the author cannot act as the review leader or scribe." Fagan's moderator is today's review leader and moderator; his hand-written notes are today's scribe's log.

Looking where the errors are: checklists and error types

Fagan observed that finding errors has to be taught: "it is one thing to direct people to find errors in design or code. It is quite another problem for them to find errors." His answer had two parts. Inspectors study the ranked distribution of error types from recent inspections, so they concentrate on the most common and costly kinds; and they use checklists of clues for each type. His Figure 5 shows part of the design checklist for logic, with questions such as "Are All Constants Defined?" and "Are All Increment Counts Properly Initialized (0 or 1)?" Every error found is also classified as missing, wrong or extra.

The program below recomputes the error distributions from Fagan's Figures 3 and 4, read row by row off the page, and adds the arithmetic of his Tables 1 and 3.

import math

# Fagan 1976, Table 1: the Aetna application, errors found per K.NCSS
by_inspection, by_test, after = 38, 8, 0
total = by_inspection + by_test + after
print(f"Table 1: inspections found {by_inspection} of {total} errors per K.NCSS,"
      f" {by_inspection / total:.1%}")

# Figures 3 and 4: errors by type as (missing, wrong, extra), read off the page
design = {"logic": (126, 57, 24), "prologue/prose": (44, 38, 7), "CB usage": (18, 17, 1),
          "other": (15, 10, 10), "more detail": (24, 6, 2), "interconnect calls": (18, 9, 0),
          "test and branch": (12, 7, 2), "CB definition": (16, 2, 0),
          "maintainability": (8, 5, 3), "return code/msg": (5, 7, 2),
          "interconnect reqts": (4, 5, 2), "performance": (1, 2, 3),
          "register usage": (1, 2, 0), "higher level docu": (1, 0, 1), "FPFS": (1, 0, 0),
          "mod attributes": (1, 0, 0), "pass data areas": (0, 1, 0)}
code = {"logic": (33, 49, 10), "design error": (31, 32, 14), "prologue/prose": (25, 24, 3),
        "CB usage": (3, 21, 1), "code comments": (5, 17, 1), "interconnect calls": (7, 9, 3),
        "maintainability": (5, 7, 2), "PL/S or BAL use": (4, 9, 1), "performance": (3, 2, 5),
        "F1": (0, 8, 0), "test and branch": (2, 5, 0), "register usage": (4, 2, 0),
        "storage usage": (1, 0, 0)}
for name, table in [("Design (Figure 3)", design), ("Code (Figure 4)", code)]:
    errors = sum(sum(row) for row in table.values())
    missing, wrong, extra = (sum(row[i] for row in table.values()) for i in range(3))
    print(f"{name}: {errors} errors; missing {missing / errors:.0%},"
          f" wrong {wrong / errors:.0%}, extra {extra / errors:.0%}")
    ranked = sorted(table, key=lambda t: -sum(table[t]))[:3]
    print("  most frequent:", ", ".join(f"{t} {sum(table[t])} ({sum(table[t]) / errors:.1%})"
                                         for t in ranked))

# Table 3: rates of progress for systems programming, per person, for 1,000 lines
people, lines = 4, 1000
for name, overview, preparation, meeting, rework in [("design I1", 500, 100, 130, 20),
                                                     ("code I2", None, 125, 150, 16)]:
    hours = {"overview": people * lines / overview if overview else 0,
             "preparation": people * lines / preparation,
             "inspection": people * lines / meeting,
             "rework": rework}
    sessions = math.ceil(lines / meeting / 2)            # no session longer than two hours
    print(f"{name}: " + ", ".join(f"{k} {v:.1f}" for k, v in hours.items() if v)
          + f"; total {sum(hours.values()):.1f} people-hours; {sessions} two-hour meetings")
munotes.in160

Inspection

Table 1: inspections found 38 of 46 errors per K.NCSS, 82.6%
Design (Figure 3): 520 errors; missing 57%, wrong 32%, extra 11%
  most frequent: logic 207 (39.8%), prologue/prose 89 (17.1%), CB usage 36 (6.9%)
Code (Figure 4): 348 errors; missing 35%, wrong 53%, extra 11%
  most frequent: logic 92 (26.4%), design error 77 (22.1%), prologue/prose 52 (14.9%)
design I1: overview 8.0, preparation 40.0, inspection 30.8, rework 20.0; total 98.8 people-hours; 4 two-hour meetings
code I2: preparation 32.0, inspection 26.7, rework 16.0; total 74.7 people-hours; 4 two-hour meetings

The recomputed totals match what Fagan printed: 520 design errors split 57, 32 and 11 per cent, logic at 39.8 per cent, and 348 code errors with logic at 26.4 per cent. Three lessons come out of the numbers.

  • Design inspections find mostly what is missing; code inspections mostly what is wrong. More than half the design errors were missing items, while more than half the code errors were wrong ones. A design checklist should therefore ask what has been left out?, and a code checklist what is incorrect?
  • Logic dominates both, and documentation is close behind. Prologue and prose errors, in the comments and descriptions, are the second commonest design error and the third commonest code error; Fagan inspected them because the next programmer relies on them.
  • The rates turn into a plan. Inspecting 1,000 lines of design with four people costs about 99 people-hours by Table 3's rates, and the meeting alone needs four sessions of at most two hours. Fagan's own estimate for the whole process, "overview through follow-up", was "about 90 to 100 people-hours for systems programming". Code, needing no overview and read faster, costs about 75.
munotes.in161

Inspection

On ExamReg, the same arithmetic tells the test lead what a code inspection of the fee module will cost before the first meeting is booked, which is exactly what Fagan meant by managing inspections "with the same attention as other important project activities".

What Fagan reported

Fagan gave results from two settings, and warned in the paper that they "cannot be considered representative of every situation".

A systems programming study at IBM. A piece of an operating system component, designed by three programmers and coded by 13, went through I1 and I2 inspections for the first time. The net saving "translated into a 23 percent increase in the productivity of the coding operation alone." A control sample, taken once the inspections were routine, differed by only 0.9 per cent, so the gain was not a novelty effect. The net savings were 94 programmer hours per K.NCSS from I1 and 51 from I2, while a third inspection after unit test, I3, cost 20 hours more than it saved, and "As a consequence, I3 is no longer in effect." In testing after unit test, the inspected sample had "38 percent less errors" than a comparable piece built with walk-throughs.

An application at Aetna Life and Casualty. A COBOL program of 4,439 non-commentary statements in eight modules, written by two programmers with inspections as the only change to their process, was estimated to need 62 programmer days and took 46.5, including inspection meetings: "The resulting saving in programmer resources was 25 percent." Table 1 records where its errors were found: 38 per K.NCSS by the design and code inspections, 8 by unit and preparation for acceptance testing, and none in acceptance testing or in six months of use. That is Fagan's error detection efficiency of 82 per cent (38 ÷ 46 is 0.826, and the table prints 82), where

error detection efficiency = errors found by an inspection ÷ total errors in the product before inspection.

The reason inspections pay, in Fagan's words, is where they find errors: rework at the early levels "is 10 to 100 times less expensive than if it is done in the last half of the process." Chapter Ninety-Six, on why reviews pay, returns to that cost.

munotes.in162

Inspection

Inspection results are not for appraising people

Fagan was emphatic that inspection data belongs to the programmer: the results "should not under any circumstances be used for programmer performance appraisal." The reason is practical. An inspection works only if authors bring work early and errors are reported freely; the moment error counts are used against the people who made them, both stop. It is the same rule as the ISTQB success factor in Chapter Twenty-Eight, on software reviews: evaluation of participants should never be an objective.

What makes an inspection formal

FeatureIn an inspection
Entry and exit criteriaEach inspection has a defined point, such as first clean compilation for code, and cannot be claimed complete until its rework is done
A trained moderatorLeads, schedules, reports within a day and follows up; never the author
Defined rolesModerator, designer, coder/implementor, tester, and a reader
PreparationEvery participant studies the material, the error-type distributions and the checklists beforehand
A single objective in the meetingFind errors; no design, no solutions
ClassificationEvery error by type, as missing, wrong or extra, and as major or minor
Written recordsThe error list, the module summary and the inspection summary report
Verified reworkFollow-up by the moderator; full reinspection when more than 5 per cent is reworked
MeasurementError rates and types analysed to improve both the product and the process

What it does not mean

An inspection is not a meeting to fix the code. Solutions are noted if obvious and worked out afterwards; the meeting only finds errors.

An inspection does not replace testing. In Fagan's Aetna data inspections found 82 per cent of the errors; testing found the rest.

The moderator is not the author's manager or the author. The moderator should come from an unrelated project, and the author cannot lead or record.

Fagan's rates are not universal. They were measured for systems programming and are, by his note, conservative; application code went four to six times faster.

Quick revision

  • Inspection (ISO/IEC 20246): a formal review that uses defined team roles and measurement to improve the review process.
  • Fagan (IBM Systems Journal, 1976): inspection points I0, I1 (design complete), I2 (code), IT1, IT2 (test plan, test cases), PI (publications).
  • Operations: overview (communication, education), preparation (education), inspection (find errors), rework, follow-up.
  • Roles: moderator ("the key person", trained, from an unrelated project), designer, coder/implementor, tester; a reader paraphrases; four is a good size.
  • Rules: at most two hours per session; report within one day; no solution hunting; reinspect if more than 5 per cent reworked; never use results to appraise programmers.
  • Error detection efficiency = errors found by the inspection ÷ total errors before inspection; 82 per cent at Aetna.
  • Results: coding productivity up 23 per cent; 38 per cent fewer errors than walk-throughs; Aetna 25 per cent fewer programmer days.
  • Design errors are mostly missing (57 per cent); code errors mostly wrong (53 per cent); logic tops both.
munotes.in163

Inspection

Test yourself

1. Describe the five operations of a Fagan inspection. Overview: the designer presents the design to the whole team and hands out the documents. Preparation: each participant studies the material, the common error types and the checklists. Inspection: a reader paraphrases the work, every piece of logic and every branch is covered, and errors are noted and classified without seeking solutions; the moderator reports within a day. Rework: the designer or coder resolves every error. Follow-up: the moderator verifies the rework and calls a full reinspection if more than 5 per cent was reworked.

2. What are the roles in an inspection, and why must the moderator be trained and independent? Moderator, designer, coder/implementor and tester, with a reader chosen by the moderator. The moderator runs the whole process and the meeting, so needs training in leading it; coming from an unrelated project preserves objectivity, and the author may never lead or record.

3. Define error detection efficiency and compute it for Fagan's Aetna data. It is the errors found by an inspection divided by the total errors in the product before inspection. At Aetna, inspections found 38 errors per K.NCSS out of 46 found in total, about 82 per cent.

4. What results did Fagan report for inspections? In a systems programming study, a 23 per cent increase in coding productivity and 38 per cent fewer errors than a walk-through sample; in an application at Aetna, 25 per cent fewer programmer days than estimated, 82 per cent of errors found by inspection, and no errors in acceptance testing or six months of use.

5. What makes an inspection more formal than other reviews? Entry and exit criteria, a trained moderator who is not the author, defined roles, required preparation with checklists, a meeting with the single objective of finding errors, classification and written records of every error, verified rework, and measurement used to improve the process.

munotes.in164

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!