Statement Testing and Statement Coverage
Chapter Fifty-Nine
Syllabus topic Module 2, "White Box: Statement testing"
Pages 337 to 342 of 622
In one line
Statement testing designs tests to execute every executable statement of the code at least once, and statement coverage is the percentage of statements the tests have executed; 100 per cent guarantees that no statement has gone completely untested, and it is the weakest of the structural criteria, because it says nothing about the choices the code makes or the values it meets.
In the wording a student can write in an examination: statement testing is a "structure-based test case design technique based on exercising executable statements in the source code of the test item" (ISO/IEC/IEEE 29119-4). In the ISTQB syllabus's words, "the coverage items are executable statements", and "Coverage is measured as the number of statements exercised by the test cases divided by the total number of executable statements in the code, and is expressed as a percentage." At 100 per cent, "each statement with a defect will be executed, which may cause a failure demonstrating the presence of the defect". But statement coverage "may not detect defects that are data dependent (e.g., a division by zero that only fails when a denominator is set to zero)", and "100% statement coverage does not ensure that all the decision logic has been tested".
What counts as a statement
The measure divides by "the total number of executable statements", so the first question is what to count. A statement, in ISO/IEC/IEEE 24765's general definition, is "a meaningful expression that defines data, specifies program actions, or directs the assembler or compiler". Statement coverage counts only the executable ones, the ones that do something when the program runs.
The tracer this book uses (Chapter Fifty-Eight, on structural testing, built it) counts every statement in a function's body except its docstring. That includes the headers of compound statements (an if, a for), because evaluating the condition is itself executed, and it excludes the lines that are only punctuation of the language: the def line, a bare else:, comments and blank lines.
Different tools count differently: some count lines, some statements, some the compiled instructions underneath. The same tests can therefore show different percentages in different tools. A statement coverage figure means something only together with the rule that counted it, and figures from two tools should never be compared directly.
Designing tests for statement coverage
Statement testing works from the control flow graph: a node is a statement, so the task is to find a few paths from the entry that between them pass through every node, and inputs that make the program take them. The syllabus sets the aim as "to design test cases that exercise statements in the code until an acceptable level of coverage is achieved". Usually that level is 100 per cent. When it is not reachable, the unreached statement is worth a question of its own: code that no input can execute is dead code, and dead code is a finding.
Statement Testing and Statement Coverage
Here are two functions from ExamReg. The first computes the late fee, and the second summarises a paper's result for the exam cell: a pass percentage, and a flag that sends a paper with fewer than 40 per cent passing to review. The specification adds a rule for a paper nobody appeared for: it has no percentage, and it goes to review.
def late_fee(days_late):
"""The late fee, in rupees, for a form days_late days after the last date (0 to 15)."""
if days_late > 0:
fee = 100
if days_late > 7:
fee = 500
return fee
def result_summary(passed, appeared):
"""A paper's pass percentage, and whether it goes to the exam cell for review."""
percent = passed / appeared * 100
review = False
if percent < 40:
review = True
return round(percent, 1), reviewRead late_fee as a graph. Its five statements lie on one path if both conditions are true, so one test with more than 7 days late executes all of them. result_summary is the same: one test with a pass rate below 40 per cent runs every statement. So statement testing asks for one test of each, designed like this: 10 days late (expected Rs 500), and 10 passes out of 40 (expected 25.0 per cent, sent to review).
The tracer, as Chapter Fifty-Eight wrote it for structural testing; each chapter's programs run on their own, so the file is repeated here unchanged:
"""A small coverage tracer: which statements and which branches of one function ran.
It watches one function with sys.settrace, which calls back on every new line executed.
A decision is an if, elif, while or for; its True branch is taken when the next line
executed is inside its body, its False branch when it is anywhere else (or the function
returns). Limits: one statement per line, and no recursion; that is all it was built for.
"""
import ast
import inspect
import sys
def structure(func):
"""Statements {line: text} and decisions {line: (first, last line of the body)}."""
first = func.__code__.co_firstlineno
source = inspect.getsource(func).splitlines()
tree = ast.parse("\n".join(source)).body[0]
statements, decisions = {}, {}
for node in ast.walk(tree):
docstring = isinstance(node, ast.Expr) and isinstance(node.value, ast.Constant)
if isinstance(node, ast.stmt) and node is not tree and not docstring:
statements[node.lineno + first - 1] = source[node.lineno - 1].strip()
if isinstance(node, (ast.If, ast.While, ast.For)):
decisions[node.lineno + first - 1] = (node.body[0].lineno + first - 1,
node.body[-1].end_lineno + first - 1)
return statements, decisions
def trace(func, args):
"""Call func(*args); return the line numbers it executed, in order."""
lines = []
def on_line(frame, event, arg):
if event == "line":
lines.append(frame.f_lineno)
return on_line
sys.settrace(lambda frame, event, arg: on_line if frame.f_code is func.__code__ else None)
try:
func(*args)
except Exception:
pass # a test's verdict is not the tracer's business
finally:
sys.settrace(None)
return lines
def measure(func, tests):
"""Statement and branch coverage of func over tests (a list of argument tuples)."""
statements, decisions = structure(func)
ran, taken = set(), set()
for args in tests:
lines = trace(func, args)
ran.update(lines)
for i, line in enumerate(lines):
if line in decisions:
low, high = decisions[line]
after = lines[i + 1] if i + 1 < len(lines) else None
taken.add((line, after is not None and low <= after <= high))
branches = {(d, outcome) for d in decisions for outcome in (True, False)}
return {"statements": (len(ran & statements.keys()), len(statements)),
"branches": (len(taken), len(branches)),
"statements missed": [statements[n] for n in sorted(statements.keys() - ran)],
"branches missed": [f"{statements[d]} -> {outcome}"
for d, outcome in sorted(branches - taken)]}Statement Testing and Statement Coverage
Worked example, part 1: 100 per cent, and every test passes
from coverage_tracer import measure
from examreg_checks import late_fee, result_summary
def call(func, args):
return f"{func.__name__}({', '.join(map(repr, args))})"
# (function, tests designed to execute every statement, with results expected from the specification)
plan = [(late_fee, [((10,), 500)]),
(result_summary, [((10, 40), (25.0, True))])]
for func, tests in plan:
for args, expected in tests:
got = func(*args)
print(f"{call(func, args)}: expected {expected}, got {got},"
f" {'pass' if got == expected else 'FAIL'}")
s = measure(func, [args for args, _ in tests])["statements"]
print(f" statement coverage {s[0]} of {s[1]} = {100 * s[0] / s[1]:.0f}%")late_fee(10): expected 500, got 500, pass
statement coverage 5 of 5 = 100%
result_summary(10, 40): expected (25.0, True), got (25.0, True), pass
statement coverage 5 of 5 = 100%Both functions: every statement executed, 100 per cent statement coverage, every test passed. By the criterion of this chapter, both are fully tested. Both are defective.
Worked example, part 2: what 100 per cent missed
Two more inputs, one for each function, both entirely ordinary: a form submitted on the last date itself, and a paper for which every registered student was absent.
from coverage_tracer import measure
from examreg_checks import late_fee, result_summary
def call(func, args):
return f"{func.__name__}({', '.join(map(repr, args))})"
for func, args, expected in [(late_fee, (0,), 0), (result_summary, (0, 0), (None, True))]:
try:
got = func(*args)
except Exception as problem:
got = f"crash: {type(problem).__name__}"
print(f"{call(func, args)}: expected {expected}, got {got}")
c = measure(late_fee, [(10,)])
print(f"late_fee's one test: statements {c['statements'][0]} of {c['statements'][1]},"
f" branches {c['branches'][0]} of {c['branches'][1]}")
for text in c["branches missed"]:
print(" branch never taken:", text)late_fee(0): expected 0, got crash: UnboundLocalError
result_summary(0, 0): expected (None, True), got crash: ZeroDivisionError
late_fee's one test: statements 5 of 5, branches 2 of 4
branch never taken: if days_late > 0: -> False
branch never taken: if days_late > 7: -> FalseStatement Testing and Statement Coverage
Both crash, one for each of the two weaknesses the ISTQB syllabus names.
The decision logic that was never tested. A form on time should pay nothing. But late_fee gives fee a value only inside the two if statements, and when both conditions are false no statement assigns it, so the return finds no value and Python stops with UnboundLocalError. Every statement of the function had been executed; what had never happened was the absence of the first assignment. The path on which the defect lies, where days_late > 0 is false, contains no statement of its own, so statement coverage never asks for it. The last lines of the output say so in the tracer's words: the one test covered 5 of 5 statements but only 2 of the 4 branches, and the two it never took are both False. This is the syllabus's second weakness, that full statement coverage "may not exercise all the branches". The next chapter makes those branches its target.
The value that was never tried. result_summary divides by the number of students who appeared, and when none did the division fails with ZeroDivisionError. The statement had run, with 40 as the divisor; the defect lives not in a path but in a value. This is the syllabus's own example of a data-dependent defect, "a division by zero that only fails when a denominator is set to zero". No structural criterion, branch coverage included, would have required a test with zero: there is no branch here to cover. A boundary value analysis of students appeared would, because 0 is the smallest value the input can take (Chapter Fifty-Five, on boundary value analysis), and so would the tester's habit of trying zero, which Chapter Sixty-Two, on error guessing, turns into a technique.
So what is statement coverage worth?
A great deal, at the bottom end. The syllabus states the guarantee: at 100 per cent, "each statement with a defect will be executed, which may cause a failure demonstrating the presence of the defect". The converse is the practical use. A statement that no test has executed has not been tested at all, and whatever defect it holds will be met first by a user. In Chapter Fifty-Eight's structural testing of check_form, the first test left all three of its problem-reporting statements unexecuted; a measurement of statement coverage points at them by name.
What it cannot do is certify anything. Reaching every statement once proves that no statement was left out of the testing, not that the testing of each was adequate, and the two crashes above came from functions at 100 per cent.
Statement Testing and Statement Coverage
Statement coverage in numbers
Exam questions on statement coverage ask one of three things.
- Compute the coverage. Statement coverage = statements executed ÷ executable statements × 100. If a program has 20 executable statements and the tests execute 17, statement coverage is 85 per cent.
- Find the fewest tests for 100 per cent. Trace paths through the control flow graph so that together they pass through every statement node.
late_feeneeds one test; a function with anifand anelse, each containing a statement, needs at least two, because no single run executes both. - Show that 100 per cent is not enough. Give a test set with full statement coverage and an input it misses that fails: exactly the worked example above.
What it does not mean
100 per cent statement coverage does not mean the code is tested. It means each statement ran at least once, with whatever values reached it.
A path with no statements is still a path. The False side of an if with no else has nothing to execute, and statement coverage never visits it.
Coverage is not the same percentage in every tool. It depends on what the tool counts as a statement.
Statement testing does not choose the expected results. The test at 10 days late was designed to reach statements; its expected Rs 500 came from the fee rule.
Quick revision
- Statement testing (ISO/IEC/IEEE 29119-4): "exercising executable statements in the source code of the test item".
- Statement coverage = statements exercised ÷ executable statements, as a percentage (ISTQB); what counts as executable depends on the tool.
- Design from the control flow graph: paths that together pass through every statement node.
- 100 per cent guarantees that every defective statement was executed at least once, and so may have failed.
- Weaknesses (ISTQB): data-dependent defects, such as division by zero; branches not exercised.
- Worked example: one test gave 100 per cent statement coverage of each function;
late_fee(0)crashed (no value assigned on the untested False path: 2 of 4 branches), andresult_summary(0, 0)crashed (division by zero).
Test yourself
1. Define statement testing and statement coverage. Statement testing is a white-box technique that designs test cases to execute the executable statements of the code. Statement coverage is the number of executable statements exercised by the tests divided by the total number of executable statements, expressed as a percentage.
2. A module has 40 executable statements, and the tests execute 34. What is the statement coverage? 85 per cent: 34 of the 40 statements.
3. What does 100 per cent statement coverage guarantee, and what does it not? It guarantees that every executable statement has been executed at least once, so a defect in any statement has had a chance to cause a failure. It does not guarantee that every decision outcome has been exercised, so a defect on a path with no statements (the false side of an if without an else) can be missed, nor that the statements met the values that make them fail, such as a zero divisor.
Statement Testing and Statement Coverage
4. Give a test set with 100 per cent statement coverage of late_fee, and show a defect it misses. One test, 10 days late, executes all five statements and passes, returning Rs 500. The input 0 days late, which should give Rs 0, makes the function crash, because neither assignment to fee is executed and the return has no value.
5. Why can structural coverage not find a division by zero in result_summary? Because the defect depends on a value, not on a path: the division statement is executed by any test, and no branch or statement requires the divisor to be zero. Only a test with zero students appeared finds it, which boundary value analysis or error guessing would choose.
6. Why may different tools report different statement coverage for the same tests? Because they count different things as executable statements: lines, language statements or compiled instructions, with or without compound statement headers. A coverage figure is meaningful only with its counting rule.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.