munotes®

Checklist-Based Testing

Get access to whole semester resourcesSemester Pass

Chapter Sixty-Four

Syllabus topic Module 2, "Experience-based: ... Checklist-based testing"

Pages 366 to 371 of 622

In one line

Checklist-based testing tests a product against a list of questions drawn from experience of what matters and what goes wrong; it gives experience-based testing consistency and a record, and a checklist stays useful only if it is pruned, sharpened and extended as the defects it is meant to catch change.

In the wording a student can write in an examination: in the ISTQB syllabus's words, "In checklist-based testing, a tester designs, implements, and executes tests to cover test conditions from a checklist. Checklists can be built based on experience, knowledge about what is important for the user, or an understanding of why and how software fails." Items "are often phrased in the form of a question. It should be possible to check each item separately and directly." Checklists "should not contain items that can be checked automatically, items better suited as entry criteria, exit criteria, or items that are too general". They "should be regularly updated based on defect analysis", but "care should be taken to avoid letting the checklist become too long". The same idea applied to documents is checklist-based reviewing, a "review technique guided by a list of questions or required attributes" (ISO/IEC 20246).

Where checklists come from

A checklist is experience written down. The syllabus names three sources: experience, "knowledge about what is important for the user", and "an understanding of why and how software fails". In practice they come from:

  • Defect history. Every defect that escaped is a candidate question: if a double click once charged a student twice, does a double click act only once? belongs on the list.
  • Standards and models. The quality characteristics of ISO/IEC 25010 (Chapter Twenty, on the quality model today) make a checklist for non-functional testing: is it secure, is it compatible, can a first-time user interact with it? The syllabus notes that checklists support "various test types, including functional and non-functional testing", giving as an example usability heuristics.
  • The users. What matters to the people who use the product: for ExamReg, that a student on a phone can finish the form before the deadline.

The same list serves before the software runs. Chapter Twenty-Eight's review of ExamReg's fee requirement used a checklist; that is checklist-based reviewing, and a team often keeps one list for reviewing requirements and another for testing the running product.

What makes a good checklist item

The syllabus's guidance turns into four tests of an item.

  1. It is a question with a definite answer. Does the form keep what the student typed after an error? can be answered yes or no by trying it.
  2. It can be checked on its own, directly. It does not depend on another item's answer, or on a long investigation.
  3. It is not better done by a machine. Is the page's HTML valid? is a question a validator answers in a second, every build; putting it on a person's list wastes the person and delays the answer. It belongs in an automated check.
  4. It is not too general. Is the form easy to use? cannot be checked directly; it has to be broken into questions that can.
munotes.in366

Checklist-Based Testing

And the list must not be a disguise for something else: an item such as has the build passed its unit tests? is an entry criterion for testing, not a test.

A checklist for ExamReg's web forms

ExamReg's testers keep a checklist for every form on the portal, and they record, release by release, how many defects each item found. Before release 3.0 it held ten items:

IdQuestion
C1Does every required field say it is required before the form is sent?
C2Is each error shown beside its field, in words a first-year student understands?
C3Does the form keep what the student typed after an error?
C4Does a double click on Submit or Pay act only once?
C5Are dates shown and accepted in one format, DD-MM-YYYY?
C6Does the form accept names with an apostrophe, a dot or a hyphen?
C7Can the whole form be filled in with the keyboard alone?
C8Does every page load the analytics script?
C9Is the form easy to use?
C10Is the page's HTML valid?

C6 is new in release 3.0, added after the names Chapter Sixty-One's experience-based tests found refused.

Keeping a checklist alive

The syllabus explains why a checklist decays: "Some checklist entries may gradually become less effective over time because the developers will learn to avoid making the same errors. New entries may also need to be added to reflect newly found high severity defects. Therefore, checklists should be regularly updated based on defect analysis." And it warns against the opposite failure, a list that only grows.

Worked example: the release 3.0 review

The program applies the syllabus's guidance to ExamReg's checklist and its history. An automatable item moves to an automated check; a too-general item is rewritten; an item that has found nothing in the last three releases is retired; everything else stays. Two new questions come from the high-severity defects of release 3.0 that no item covered: the payment sent twice by Refresh, found in Chapter Sixty-Three's exploratory sessions, and the days late miscounted across the end of a month and a year, found by Chapter Sixty-Two's error guessing.

RELEASES = ["1.0", "2.0", "2.1", "3.0"]

# (id, question, defects the item found in each release (None: not yet on the list), note)
checklist = [
    ("C1", "Does every required field say it is required before the form is sent?", [3, 1, 0, 0], ""),
    ("C2", "Is each error shown beside its field, in words a first-year student understands?",
                                                                                     [2, 2, 1, 1], ""),
    ("C3", "Does the form keep what the student typed after an error?",              [1, 0, 0, 0], ""),
    ("C4", "Does a double click on Submit or Pay act only once?",                    [0, 1, 1, 2], ""),
    ("C5", "Are dates shown and accepted in one format, DD-MM-YYYY?",                [2, 0, 0, 0], ""),
    ("C6", "Does the form accept names with an apostrophe, a dot or a hyphen?",      [None, None, None, 3], ""),
    ("C7", "Can the whole form be filled in with the keyboard alone?",               [0, 1, 0, 1], ""),
    ("C8", "Does every page load the analytics script?",                             [0, 0, 1, 0], "automatable"),
    ("C9", "Is the form easy to use?",                                               [1, 0, 1, 0], "too general"),
    ("C10", "Is the page's HTML valid?",                                             [1, 1, 0, 0], "automatable"),
]
# high-severity defects of release 3.0 that no item covered: each becomes a new question
new_items = ["Does pressing Back or Refresh after paying never send a second payment?",
             "Are days late counted correctly across the end of a month and of a year?"]

def verdict(found, note, quiet=3):
    if note == "automatable":
        return "move to an automated check"
    if note == "too general":
        return "rewrite as specific questions"
    recent = [n for n in found[-quiet:] if n is not None]
    if len(recent) == quiet and sum(recent) == 0:
        return f"retire: nothing found in the last {quiet} releases"
    return "keep"

counts = {}
for cid, question, found, note in checklist:
    v = verdict(found, note)
    counts[v.split(":")[0]] = counts.get(v.split(":")[0], 0) + 1
    shown = " ".join("-" if n is None else str(n) for n in found)
    print(f"{cid:<4} {shown:<8} {v}")
for q in new_items:
    print(f"new  from a 3.0 high-severity defect: {q}")

kept = counts.get("keep", 0)
print(f"{len(checklist)} items reviewed: " + ", ".join(f"{n} {v}" for v, n in counts.items()))
print(f"the working checklist: {kept} kept + {len(new_items)} new = {kept + len(new_items)} items")
munotes.in367

Checklist-Based Testing

C1   3 1 0 0  keep
C2   2 2 1 1  keep
C3   1 0 0 0  retire: nothing found in the last 3 releases
C4   0 1 1 2  keep
C5   2 0 0 0  retire: nothing found in the last 3 releases
C6   - - - 3  keep
C7   0 1 0 1  keep
C8   0 0 1 0  move to an automated check
C9   1 0 1 0  rewrite as specific questions
C10  1 1 0 0  move to an automated check
new  from a 3.0 high-severity defect: Does pressing Back or Refresh after paying never send a second payment?
new  from a 3.0 high-severity defect: Are days late counted correctly across the end of a month and of a year?
10 items reviewed: 5 keep, 2 retire, 2 move to an automated check, 1 rewrite as specific questions
the working checklist: 5 kept + 2 new = 7 items
munotes.in368

Checklist-Based Testing

Read each verdict against its history.

  • C3 and C5 are retired. Each found defects early and nothing in the last three releases: the developers learnt to keep a form's contents after an error and to use one date format, exactly the decay the syllabus describes. Retiring them is not saying the questions were wrong; it is making room. If a defect of either kind reappears, the item comes back.
  • C1 stays, although it found nothing in the last two releases, because the rule looks back three releases and C1 found a defect in 2.0. Where to draw that line is the team's judgement; the program only applies it consistently.
  • C4 stays and matters more each release: it found 0, 1, 1 and then 2 defects. A double-click defect costs a student money, so it would stay even if its count fell.
  • C8 and C10 move to automated checks. Whether every page loads its analytics script, and whether the HTML is valid, are questions a program answers on every build; a person checking them by hand, a few times a release, is both slower and less reliable.
  • C9 is rewritten. Is the form easy to use? found defects, but a tester cannot check it directly; it becomes several specific questions (can a first-time user find the Pay button without help? does every button say what it does?).
  • Two new items, each from a high-severity defect that no item would have caught.

The working checklist ends with 7 questions (plus whatever specific questions replace C9): shorter than it began, and aimed at the failures that are happening now.

High-level or detailed

A checklist can say check the error messages or it can say is each error shown beside its field, in words a first-year student understands?. The syllabus names the trade-off: "If the checklists are high-level, some variability in the actual testing is likely to occur, resulting in potentially greater coverage but less repeatability." A detailed item is tested the same way by every tester; a high-level one is interpreted, and different testers look at different things. And where detailed test cases do not exist, "checklist-based testing can provide guidelines and some degree of consistency for the testing."

Strengths and limits

Strengths. A checklist carries experience from one tester to the next, so a new tester starts with the team's knowledge. It gives experience-based testing consistency and a record: each item was checked, with a result. It is quick to apply and easy to review. And it sits comfortably with the other techniques: an item such as C6 can be tested with equivalence partitions of names.

munotes.in369

Checklist-Based Testing

Limits. It finds only what it asks about. It decays if nobody maintains it, and it bloats if nobody prunes it. A high-level item gives variable testing; a detailed one gives narrow testing. And ticking an item is not the same as testing it well: yes to C7 after pressing Tab twice is not a check of the whole form.

What it does not mean

A checklist is not a test case. An item names a test condition; the tester still decides how to check it, with what data, and what counts as a pass.

A longer checklist is not a better one. Items that no longer find defects, that a machine could check, or that cannot be checked directly make the list slower to use and easier to skim.

Retiring an item is not forgetting it. The defect records keep it, and it returns if its kind of defect does.

Checklists are not only for testing software that runs. The same idea guides reviews of requirements, designs and code.

Quick revision

  • Checklist-based testing (ISTQB): tests designed, implemented and executed to cover test conditions from a checklist.
  • Sources: experience, what matters to users, why and how software fails; defect history, standards and models, users.
  • Good items: questions, each checkable separately and directly; not automatable, not entry or exit criteria, not too general.
  • For functional and non-functional testing alike; checklist-based reviewing (ISO/IEC 20246) applies it to documents.
  • Keep it alive: entries lose effectiveness as developers learn; add entries for new high-severity defects; do not let it grow too long.
  • Worked example: 10 items reviewed: 5 kept, 2 retired (nothing in 3 releases), 2 automated, 1 rewritten; 2 added; 7 working items.
  • High-level checklists: more coverage, less repeatability.

Test yourself

1. What is checklist-based testing? Where do checklists come from? A technique in which the tester designs and executes tests to cover the test conditions on a checklist. Checklists are built from experience, from knowledge of what matters to users, and from an understanding of why and how software fails: in practice, from defect history, standards and quality models, and the users' priorities.

2. What should a checklist not contain, and why? Items that can be checked automatically, because a program checks them faster and on every build; items that are really entry or exit criteria, because they are conditions for testing, not tests; and items that are too general, because they cannot be checked directly.

3. Write five checklist items for testing a college's online examination form. Does every required field say it is required before the form is sent? Is each error shown beside its field in plain words? Does the form keep what the student typed after an error? Does a double click on Submit or Pay act only once? Does the form accept names with an apostrophe, a dot or a hyphen?

munotes.in370

Checklist-Based Testing

4. Why must a checklist be updated, and how? Because items lose their effectiveness as developers learn to avoid the errors they target, while new kinds of high-severity defects appear. It is updated from defect analysis: items that have stopped finding defects are retired or reviewed, new items are added for escaped high-severity defects, and the list is kept short.

5. What is the trade-off between high-level and detailed checklists? High-level items leave the tester to interpret them, so testing varies between testers, potentially covering more but with less repeatability; detailed items are tested the same way each time, more repeatably but more narrowly.

6. Compare checklist-based testing with exploratory testing. Both are experience-based. A checklist fixes in advance which conditions are checked, giving consistency and a record of each item; exploratory testing decides what to test during the session, guided by a charter and by what each test reveals, giving flexibility and discovery at the cost of repeatability.

munotes.in371

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!