Test Automation: What to Automate and What Not To
Chapter Forty-Eight
Syllabus topic Practical, "Creation of Test Suite Using Selenium IDE"
Pages 262 to 266 of 622
In one line
Test automation lets a machine run tests again and again at almost no cost per run, but writing and maintaining the automated tests costs a great deal, so automation pays only for tests that are run often, are stable, and can be checked mechanically; everything that needs judgement stays with people.
In the wording a student can write in an examination: test automation is the use of tools to execute tests, compare results and report them. The ISTQB syllabus lists its potential benefits, among them "Time saved by reducing repetitive manual work", "Prevention of simple human errors through greater consistency and repeatability", "Reduced test execution times to provide earlier defect detection, faster feedback and faster time to market" and "More time for testers to design new, deeper and more effective tests"; and its risks, among them "Unrealistic expectations about the benefits of a tool" and "Using a test tool when manual testing is more appropriate." Record and playback, as in Selenium IDE, captures a tester's actions in a browser as a script that can be replayed; it is quick to start but brittle, and serious suites are exported and restructured. Automation follows the test pyramid: many automated unit tests, fewer service tests, few end-to-end UI tests.
What automation buys
The ISTQB syllabus's full list of potential benefits is short and worth knowing:
- "Time saved by reducing repetitive manual work (e.g., execute regression tests, re-enter the same test data, compare expected results vs actual results, and check against coding standards)"
- "Prevention of simple human errors through greater consistency and repeatability"
- "More objective assessment (e.g., coverage) and providing measures that are too complicated for humans to determine"
- "Easier access to information about testing to support test management and test reporting"
- "Reduced test execution times to provide earlier defect detection, faster feedback and faster time to market"
- "More time for testers to design new, deeper and more effective tests"
Every item is about repetition. A tool does the same thing the same way every time, which is exactly what regression testing needs: the syllabus calls regression suites "a strong candidate for automation" because they are run after every change and grow with every release (Chapter Thirty-Two, on test levels and test types).
What automation costs
The same section opens with a warning: "Simply acquiring a tool does not guarantee success. Each new tool will require effort to achieve real and lasting benefits (e.g., for tool introduction, maintenance and training)." Among the risks it lists:
- "Unrealistic expectations about the benefits of a tool (including functionality and ease of use)."
- "Inaccurate estimations of time, costs, effort required to introduce a tool, maintain test scripts and change the existing manual test process."
- "Using a test tool when manual testing is more appropriate."
- "Relying on a tool too much, e.g., ignoring the need of human critical thinking."
- Dependence on a vendor, or on open-source software that may be abandoned, and a tool "not compatible with the development platform".
Test Automation: What to Automate and What Not To
The cost that surprises teams most is maintenance. An automated test is code. When the application's pages change, the tests that drive them break even though the application works, and someone must repair them before the suite means anything again.
Worked example: when does automating ExamReg's regression suite pay?
ExamReg's regression suite has 120 test cases. Run by hand, at about three minutes each, one run takes about 6 hours. Automating it takes, say, 60 hours to write and debug once, and each automated run then costs a quarter of an hour to start and read, plus whatever it takes to repair the scripts the application's changes have broken since the last run. These hours are this book's illustration; the program shows how the answer depends on the one number teams most often forget, the upkeep.
import math
manual_per_run = 6.0 # hours: 120 regression cases by hand, about 3 minutes each
build = 60.0 # hours: writing and debugging the automated suite, once
run_per_run = 0.25 # hours: starting the automated run and reading its report
def break_even(upkeep_per_run):
"""Runs after which automation has cost no more than doing the same runs by hand."""
saving = manual_per_run - run_per_run - upkeep_per_run
return math.ceil(build / saving) if saving > 0 else None
for label, upkeep in [("well-structured scripts", 0.75), ("recorded, brittle scripts", 4.0),
("scripts repaired almost every run", 5.9)]:
n = break_even(upkeep)
print(f"{label:<34} upkeep {upkeep:>4} h/run: "
+ (f"pays for itself after {n} runs" if n else "never pays for itself"))
print("runs by hand automated (well-structured)")
for n in (1, 5, 12, 20, 40):
print(f"{n:>4} {n * manual_per_run:>9.1f} {build + n * (run_per_run + 0.75):>11.1f}")well-structured scripts upkeep 0.75 h/run: pays for itself after 12 runs
recorded, brittle scripts upkeep 4.0 h/run: pays for itself after 35 runs
scripts repaired almost every run upkeep 5.9 h/run: never pays for itself
runs by hand automated (well-structured)
1 6.0 61.0
5 30.0 65.0
12 72.0 72.0
20 120.0 80.0
40 240.0 100.0With well-structured scripts, which need about three quarters of an hour of repair per run, automation costs more for the first eleven runs and breaks even at the twelfth; by the fortieth run it has cost 100 hours against 240 by hand. If ExamReg's suite runs on every build through continuous integration (Chapter Thirty-Nine), twelve runs come in a few days, and the investment is repaid almost at once.
With brittle scripts the picture changes completely. Four hours of repair every run pushes break-even out to the thirty-fifth run, and if the scripts need repairing nearly as long as the manual run itself, automation never pays at all. That is the ISTQB syllabus's risk of "Inaccurate estimations of time, costs, effort required to ... maintain test scripts" in numbers, and it is why the next three chapters, on waits, data-driven testing and the page object model, are about making scripts that survive change.
Test Automation: What to Automate and What Not To
Record and playback: Selenium IDE
The practical begins with Selenium IDE, which its project describes as "Open source record and playback test automation for the web". The tester works through the application in the browser; the IDE records each click and keystroke as a command; replaying the recording repeats the actions and checks what the tester asked it to check. Its page lists the features that make recorded tests less fragile: "Selenium IDE records multiple locators for each element it interacts with. If one locator fails during playback, the others will be tried until one is successful"; a recorded test can "re-use one test case inside of another" with the run command, "allowing you to re-use your login logic in multiple places throughout a suite"; and its tests can run "on any browser/OS combination in parallel" through a command-line runner.
Record and playback is the fastest way to start automating, and a good way to learn what a browser test does. Its limits are the reasons the practical then asks students to "export the automation script in WebDriver format":
- Brittleness. A recording is tied to the page as it was on the day it was recorded; a renamed button or a moved field can break it.
- Timing. A recording replays at its own pace; a page that loads more slowly than it did during recording can make a step fail. Chapter Forty-Nine's waits deal with this.
- Hard-wired data. A recording repeats the same inputs; testing many inputs means many recordings, or the data-driven testing of Chapter Fifty.
- Duplication. Every recording that logs in repeats the login steps; when the login page changes, every one breaks. The page object model of Chapter Fifty-One puts each page's details in one place.
What to automate, and what not to
| Automate | Keep manual |
|---|---|
| Regression suites run after every change | Exploratory testing, which designs the next test from the last result (Chapter Sixty-Three) |
| Smoke tests run on every build | Usability: whether a first-year student understands the form |
| Unit and service tests, the base of the pyramid | Tests of features that are still changing every week |
| Data-heavy checks: every fee for every day and concession | One-off checks that will never be repeated |
| Performance and load tests, which need many simultaneous users | Judgements about look and feel, wording and helpfulness |
| Checks across many browsers and platforms | Anything whose expected result a machine cannot decide |
Test Automation: What to Automate and What Not To
The rule under the table is the ISTQB syllabus's warning against "Relying on a tool too much, e.g., ignoring the need of human critical thinking." A tool checks what it was told to check. Deciding what to check, and noticing what nobody thought to check, remains the tester's work.
The test automation pyramid
Chapter Seventeen, on agile, Scrum and DevOps, introduced the test pyramid: many small, fast unit tests at the base, fewer service tests in the middle, few broad UI tests at the top. It is also the shape of a good automation portfolio. Unit tests are cheap to write, fast to run and rarely break for reasons unrelated to a real defect; end-to-end UI tests through a browser are slow, and they break whenever a page changes. A suite built mostly from recorded UI tests is the inverted pyramid, and the break-even arithmetic above explains why it disappoints.
Figure 48.1 The test pyramid: automate mostly at the base, where tests are small, fast and stable
What it does not mean
Automation does not replace testers. It replaces repetition; the design of tests, exploration and judgement remain human work.
Automated tests are not free after they are written. They must be maintained as the application changes, and that upkeep decides whether automation pays.
Record and playback is not a finished automation suite. It is a starting point, to be exported and restructured for anything that must last.
More automated UI tests are not always better. A suite should be mostly unit and service tests, with a few end-to-end tests at the top.
Quick revision
- Benefits (ISTQB v4.0.1, 6.2): less repetitive work, fewer simple human errors, objective measures, easier reporting, faster feedback, more time for deeper tests.
- Risks: unrealistic expectations, underestimated costs of introduction and maintenance, automating what should be manual, over-reliance, vendor or open-source dependence, platform incompatibility.
- Worked example: 60 hours to build, 6 hours by hand per run; break-even after 12 runs with well-structured scripts, 35 with brittle ones, never if every run needs repair.
- Selenium IDE: record and playback; multiple locators per element; reusable test cases; export to WebDriver.
- Record and playback's limits: brittleness, timing, hard-wired data, duplication.
- Automate the repetitive and mechanical; keep exploration, usability and judgement manual; follow the test pyramid.
Test yourself
1. State four benefits and four risks of test automation. Benefits: time saved on repetitive work such as regression tests; fewer simple human errors through consistency; faster execution, giving earlier defect detection and feedback; and more time for testers to design deeper tests. Risks: unrealistic expectations of the tool; underestimated effort to introduce it and maintain scripts; automating tests that should be manual; and over-reliance on the tool at the expense of critical thinking.
Test Automation: What to Automate and What Not To
2. Why does the maintenance of automated tests decide whether automation pays? Because automation costs a large amount once and a small amount per run, and saves the cost of a manual run each time. Maintenance is part of the cost per run; if scripts break often and take long to repair, the saving per run shrinks, and the number of runs needed to recover the building cost grows, possibly without limit.
3. What is record and playback, and what are its limits? A way of automating browser tests by recording a tester's actions as a script and replaying them, as Selenium IDE does. Its limits are brittleness when pages change, failures when timing differs, inputs hard-wired into each recording, and steps such as login duplicated across many recordings.
4. Which tests should be automated, and which should not? Automate tests that are run often and checked mechanically: regression and smoke suites, unit and service tests, data-heavy checks, load tests and cross-browser runs. Keep manual the tests that need judgement or change constantly: exploratory testing, usability, look and feel, one-off checks, and features still being redesigned.
5. How does the test pyramid guide automation? It says to automate mostly at the base, with many small, fast, stable unit tests; fewer tests at the service level; and only a few end-to-end tests through the user interface, which are slow and break whenever pages change.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.