munotes®

Unit Testing Best Practices

Get access to whole semester resourcesSemester Pass

Chapter Thirty-Six

Syllabus topic Module 1, "Software Testing Strategies: Unit Testing: purpose, techniques, and best practices"

Pages 198 to 203 of 622

In one line

Good unit tests are small, fast, independent and repeatable, each checks one behaviour in the arrange, act, assert shape and is named for it, and they are written before or with the code; coverage tells you what the tests have not reached, not how good they are.

In the wording a student can write in an examination: unit testing best practices are (1) structure each test as arrange, act, assert (Wake); (2) test one behaviour per test and name the test for it; (3) keep tests independent of each other and repeatable, controlling time and randomness; (4) keep them fast, with no databases, networks or files in unit tests; (5) test the edges and invalid input; (6) keep logic out of tests; (7) write tests first, in test-driven development, where "Tests are written first, then the code is written to satisfy the tests, and then the tests and code are refactored" (ISTQB); and (8) treat coverage as a guide, since even full statement coverage "will not detect defects in all cases".

Structure: arrange, act, assert

Bill Wake named the pattern in 2001 (by his own account on the page cited) and describes its three steps:

  • "Arrange: Set up the object to be tested. We may need to surround the object with collaborators. For testing purposes, those collaborators might be test objects (mocks, fakes, etc.) or the real thing."
  • "Act: Act on the object (through some mutator). You may need to give it parameters (again, possibly test objects)."
  • "Assert: Make claims about the object, its collaborators, its parameters, and possibly (rarely!!) global state."

The shape makes a test readable at a glance: what was set up, what was done, what should be true. Wake's warning is about the alternative, a long test that arranges, acts and asserts over and over: "To understand a test like that, you have to track state over a series of activities." His advice is that "Such multi-step unit tests are usually better off being split into several tests."

He also offers a way to start: write the assert first, asking "Suppose it worked; how would I be able to tell?" The answer is the test's reason for existing.

One behaviour per test, named for it

A test should check one behaviour, so that when it fails there is one thing to look at. That is not the same as one assert line. Wake does not follow a strict one-assert rule: when one action changes several aspects of an object, he checks them together, because "the various assertions each explore a different 'dimension' of the object." The rule is one behaviour: "on-time students pay no late fee" is one behaviour, even if checking it takes two lines.

munotes.in198

Unit Testing Best Practices

Name the test for that behaviour. test_receipt_states_the_late_fee_separately reads, when it fails, as the sentence it is: the receipt does not state the late fee separately. test_3 tells the reader nothing.

Independent and repeatable

Independent. Each test should set up everything it needs and depend on no other test having run first. A framework's fixture, rebuilt before every test as in Chapter Thirty-Five, on writing unit tests with a framework, is the tool. Tests that share state pass or fail depending on their order, and a test run that is sometimes red and sometimes green teaches the team to ignore red.

Repeatable. A test should give the same result every time, on every machine. Anything that varies, today's date, a random number, the network, must be controlled. The stub clock of Chapter Thirty-Four, on drivers, stubs and test doubles, is the pattern: the unit asks a clock the test controls, so "10 days late" is 10 days late whenever the test runs.

Fast, and kept to the unit

Unit tests are run many times a day, so they must take seconds, not minutes. Wake draws the boundary: "Unit tests (for the bulk of the system) don't talk to external systems, databases, files, etc., and Arrange-Act-Assert is a pattern for unit tests." That is why the test pyramid from Chapter Seventeen, on agile, Scrum and DevOps, puts many small, fast unit tests at the base and few slow end-to-end tests at the top: a suite that is slow is a suite that is not run.

Test the edges

Most defects in a unit sit at its edges: the boundaries of a range, empty inputs, the first and last items, invalid data. Chapter Thirty-Three's checklist named them, and its worked example found the unit accepting -2 backlog papers. Each boundary and each kind of invalid input deserves its own named test; Chapter Fifty-Five, on boundary value analysis, turns this into a method.

Keep logic out of tests

A test should be straight-line code: arrange, act, assert, with no loops or conditions of its own. A test that computes its expected answer with the same kind of logic as the code it tests can share the code's mistake and pass; a test with an if in it can quietly skip its own assert. Expected values should be written out as constants, worked out by hand from the requirement: Rs 1450, not FORM_FEE + BACKLOG_FEE + 500. Where many similar cases are needed, a table of inputs and expected results, run by the framework, keeps each case visible; Chapter Fifty, on data-driven testing, shows how.

munotes.in199

Unit Testing Best Practices

Test first: test-driven development

The ISTQB syllabus describes test-driven development (TDD) as a practice that "Directs the coding through test cases (instead of extensive software design)", in which "Tests are written first, then the code is written to satisfy the tests, and then the tests and code are refactored." This book labels the three steps red (write a test for the next behaviour and watch it fail), green (write the least code that makes it pass) and refactor (improve the code's structure while every test stays green). The syllabus adds that TDD, like ATDD and BDD, "implements the principle of early testing", since "the tests are defined before the code is written".

Worked example: one TDD cycle, and a coverage trap

Chapter Thirty-Five ended with a failing test: the exam cell now wants each receipt to show the late fee separately. That failing test is the red step of a TDD cycle. The program runs the same three tests, written in the arrange, act, assert shape and named for their behaviours, against the receipt code at each step of the cycle.

Its second part shows why coverage is a guide and not a goal. The exam cell also asks for the fee per backlog paper on the receipt. The new function has two statements and no branch, so either of its two tests executes every statement in it: 100 per cent statement coverage.

import io
import unittest

PARTS = {"form fee": 800, "backlog fee": 150, "late fee": 500}      # student 2026CS014's bill

def receipt_before(parts, number):             # the code Chapter Thirty-Five's test rejected
    return f"Paid Rs {sum(parts.values())}; receipt {number}"

def receipt_green(parts, number):              # GREEN: the least code that passes
    return f"Paid Rs {sum(parts.values())}; late fee Rs {parts['late fee']}; receipt {number}"

def rupees(amount):
    return f"Rs {amount}"

def receipt_refactored(parts, number):         # REFACTOR: the same output, one place per format
    pieces = [f"Paid {rupees(sum(parts.values()))}",
              f"late fee {rupees(parts['late fee'])}",
              f"receipt {number}"]
    return "; ".join(pieces)

class ReceiptTests(unittest.TestCase):
    make_receipt = None                        # set to the version under test before each run

    def test_receipt_states_the_total_paid(self):
        parts, number = PARTS, "R-1042"                    # arrange
        text = self.make_receipt(parts, number)            # act
        self.assertIn("Paid Rs 1450", text)                # assert

    def test_receipt_states_the_late_fee_separately(self):
        parts, number = PARTS, "R-1042"
        text = self.make_receipt(parts, number)
        self.assertIn("late fee Rs 500", text)

    def test_receipt_quotes_the_receipt_number(self):
        parts, number = PARTS, "R-1042"
        text = self.make_receipt(parts, number)
        self.assertIn("receipt R-1042", text)

for stage, version in [("red, before any change", receipt_before),
                       ("green, the least code that passes", receipt_green),
                       ("refactored, same behaviour", receipt_refactored)]:
    ReceiptTests.make_receipt = staticmethod(version)
    report = io.StringIO()
    unittest.TextTestRunner(stream=report).run(
        unittest.defaultTestLoader.loadTestsFromTestCase(ReceiptTests))
    lines = report.getvalue().splitlines()
    print(f"{stage:<34} {lines[0]:<4} {lines[-1]}")

def fee_per_backlog_paper(backlog_fee, papers):    # two statements, no branch
    per_paper = backlog_fee / papers
    return round(per_paper)

print("per-paper tests pass:", fee_per_backlog_paper(300, 2) == 150,
      fee_per_backlog_paper(150, 1) == 150)
try:
    fee_per_backlog_paper(0, 0)
except ZeroDivisionError as problem:
    print("a student with no backlog papers:", type(problem).__name__)
munotes.in200

Unit Testing Best Practices

red, before any change             .F.  FAILED (failures=1)
green, the least code that passes  ...  OK
refactored, same behaviour         ...  OK
per-paper tests pass: True True
a student with no backlog papers: ZeroDivisionError

The cycle reads straight off the first three lines. Red: against the old receipt, the late-fee test fails (the F among the dots) while the other two pass. Green: one changed line makes all three pass, and no more code is written than the test demands. Refactor: the receipt is rebuilt from named pieces with one helper for amounts, a better shape for the next change, and the three tests prove the output is unchanged. The tests that drove the change stay behind as regression tests.

The last two lines are the coverage trap. Both tests pass, and together, or either alone, they execute every statement of fee_per_backlog_paper. Yet a student with no backlog papers, an entirely ordinary case, crashes the receipt with a division by zero. It is precisely the ISTQB syllabus's own example of a defect that full statement coverage can miss: exercising a statement "will not detect defects in all cases. For example, it may not detect defects that are data dependent (e.g., a division by zero that only fails when a denominator is set to zero)." Coverage told the truth, every statement ran; it just answered a different question from is the code right? An edge test, zero papers, was what was missing.

Coverage: a guide, not a goal

Coverage measures what the tests reached, and is useful exactly for that: a line or branch no test has reached is a line or branch no test has checked. What coverage cannot measure is whether the tests check the right things. A team given a coverage target can reach it with tests that execute everything and assert almost nothing. Use coverage to find the untested parts, then design tests for them with the techniques of Module 2; statement and branch coverage have their own chapters there (Chapters Fifty-Nine and Sixty).

A checklist

PracticeThe reason
Arrange, act, assertA reader sees the setup, the action and the claim at a glance
One behaviour per test, named for itA failure points at one behaviour and reads as a sentence
Independent tests, fresh fixture each timeOrder cannot change the result
Repeatable: control time, randomness, external servicesSame result on every run and machine
Fast, with no databases, networks or filesThe suite is run often enough to matter
Test edges and invalid inputThat is where defects cluster in a unit
No logic in tests; expected values as constantsA test cannot share the code's mistake
Test first (TDD): red, green, refactorEvery line of code is wanted by a test, and the tests stay as regression tests
Coverage as a guideIt finds untested code; it does not prove tested code right
munotes.in201

Unit Testing Best Practices

What it does not mean

One behaviour per test does not mean one assert line. Several asserts about one behaviour belong together.

TDD is not writing all the tests first. It is one small test at a time, each followed by the code that satisfies it and a refactoring.

Refactoring is not adding features. It changes the structure and keeps the behaviour, which the tests confirm.

High coverage is not high quality. A suite can execute everything and check little; the worked example reached every statement and missed a crash.

Quick revision

  • Arrange, act, assert (Wake, 2001): set up, act, make claims; split long multi-step tests.
  • One behaviour per test, named for it; several asserts about one behaviour are fine.
  • Independent (fresh fixture, no order dependence) and repeatable (control time, randomness, external services).
  • Fast: unit tests "don't talk to external systems, databases, files" (Wake).
  • Edges and invalid input; no logic in tests; expected values as constants.
  • TDD (ISTQB): tests first, then code to satisfy them, then refactor; implements early testing. Red, green, refactor.
  • Coverage is a guide: 100 per cent statement coverage missed a division by zero (ISTQB's own example).

Test yourself

1. What is the arrange, act, assert pattern? A structure for a unit test in three parts: arrange sets up the object under test and its collaborators, act performs the one action being tested, and assert makes claims about the result. It makes a test readable at a glance, and long tests that repeat the cycle should usually be split.

2. Why should unit tests be independent and repeatable? How is each achieved? So that a result depends only on the code under test, not on the order of tests or on the day, machine or network. Independence comes from building a fresh fixture for each test and sharing no state; repeatability from controlling time, randomness and external services with test doubles.

3. Describe one cycle of test-driven development. Write a test for the next small behaviour and run it to see it fail; write the least code that makes it pass; then refactor the code, improving its structure while all the tests stay green. The tests remain as regression tests.

4. Why is code coverage a guide and not a goal? Because it measures which code the tests executed, not whether they checked the right results. Even 100 per cent statement coverage can miss data-dependent defects, such as a division by zero that fails only when the divisor is zero; coverage should be used to find untested code, which is then tested with proper test design.

munotes.in202

Unit Testing Best Practices

5. Why should a test not contain its own logic? Because a test that computes its expected result with logic like the code's can repeat the code's mistake and pass, and conditions in a test can skip its assertions; expected values should be worked out from the requirement and written as constants.

munotes.in203

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!