munotes®

Test Execution

Get access to whole semester resourcesSemester Pass

Chapter Nine

Syllabus topic Module 1, "Software Testing Fundamentals: Test execution"; the paired practical, "Screenshot Capture and Logging Mechanism"

Pages 51 to 56 of 622

In one line

Test execution is actually running the tests: following each procedure, writing down exactly what the software did, comparing it with what it should have done, and keeping the evidence.

In the wording a student can write in an examination: test execution is the "process of running a test on the test item, producing actual results" (ISO/IEC/IEEE 29119-2:2021). It includes running the test procedures according to the test execution schedule, comparing actual results with expected results, recording the outcome in a test log, analysing anomalies and reporting failures as incidents, and then retesting fixes and running regression tests.

Why execution needs discipline

Execution is the part of testing everyone imagines, and it is easy to do badly. A tester who runs a test, sees something odd, shrugs and moves on has spent the time and kept none of the value. A tester who runs the right test on the wrong build, or with yesterday's data, produces a result nobody can use. And a tester who reports "the fee page is broken" without saying which test, which build, which data and what exactly was seen, gives the developer a puzzle instead of a defect.

Everything in the earlier activities, the plan, the conditions, the cases and procedures, was prepared so that execution could be quick, repeatable and evidential. This chapter is about making it so.

Before the first test runs

Three things are checked before execution starts, and each has its own name in the standards.

The entry criteria are met. The build has been delivered and installs; the test cases are approved; the people are available (Chapter Six, on the software testing life cycle).

The environment and data are ready. ISO/IEC/IEEE 29119-2 even names the documents that say so: a test environment readiness report, "document that describes the status of the test environment", and a test data readiness report, "document describing the status of each test data requirement". For ExamReg that means the test server has the release 2.0 build, its clock can be set, the payment gateway's test sandbox is reachable, and the six test student accounts exist with the right papers.

A smoke test passes. A short, broad check that the build's main functions respond at all (can a student log in, open the form, reach the fee page?) before hours are spent on detailed tests. A build that fails its smoke test goes back unused. Smoke testing has more to it and returns in Chapter Thirty-Nine, on regression testing, smoke testing and continuous integration.

Running a test and judging it

The tester follows the test procedure exactly as written, supplies the specified inputs, and records the actual results: the "set of behaviors or conditions of a test item, or set of conditions of associated data or the test environment, observed as a result of test execution" (ISO/IEC/IEEE 29119-2). Note that the definition includes the data and the environment: a fee shown correctly on screen but stored wrongly in the database is an actual result too, and a tester who checks only the screen misses it.

munotes.in51

Test Execution

Then comes the verdict. ISO/IEC/IEEE 29119-2 defines a test result as the "indication of whether a specific test case has passed or failed, i.e. if the actual results correspond to the expected results or if deviations were observed". In practice four states are recorded:

StatusMeaningExamReg example
PassActual results match the expected results7 days late, fee shown Rs 100
FailActual results differ from the expected results0 days late, fee shown Rs 100, expected nothing
BlockedThe test cannot be run, because a precondition cannot be metThe payment test cannot run because the gateway sandbox is down
Not runNot yet executed in this cycleHall ticket tests, scheduled for tomorrow

A blocked test is neither a pass nor a fail, and it must never be quietly counted as either. A report that shows 95 per cent passing because the blocked tests were left out of the total is misleading.

When the result is unexpected: anomalies and incidents

A failed comparison is not yet a defect in the software. The ISTQB syllabus describes what happens next: "Anomalies are analyzed to identify their likely causes. This analysis allows us to report the anomalies based on the failures observed." A failure can come from the software, from the test itself (a wrong expected result, the false positive of Chapter Two on errors, faults and failures), from the environment (a misconfigured server) or from the test data.

ISO/IEC/IEEE 29119-3:2021 calls the event a test incident: an "event occurring during the execution of a test that requires investigation". The tester records it in an incident report, "documentation of the occurrence, nature, and status of an incident" (ISO/IEC/IEEE 29119-2), which becomes a defect report once the cause is confirmed to be a defect. Writing a defect report has a chapter of its own in Module 2 (Chapter Seventy-Six).

The test log

Everything that happens during execution goes into the test log, which ISO/IEC/IEEE 24765:2017 defines as a "chronological record of relevant details about the execution of tests". ISO/IEC/IEEE 29119-3 calls the document a test execution log, one that "records details of the execution of one or more test procedures". A useful log entry answers: which test, on which build, in which environment, with which data, when, by whom, with what result, and where the evidence is.

munotes.in52

Test Execution

Evidence: logs and screenshots

The paired practical sets an exercise called "Screenshot Capture and Logging Mechanism", in which an automated test takes a screenshot when it fails and writes its progress to a log. The idea behind it is older than any tool. A failure that cannot be shown did not, for practical purposes, happen: the developer who receives the report will try to reproduce it, and if they cannot, the report is closed. A screenshot shows what the user saw at the moment of failure. A log shows what the program did on the way there. Together they turn "it went wrong" into evidence.

Logs are written at levels, so that a quiet run and a detailed investigation can use the same code. Apache Log4j 2's manual describes levels "from debug to fatal"; Python's standard logging module has DEBUG, INFO, WARNING, ERROR and CRITICAL. A test run typically logs each test's start at INFO, each failure at ERROR with its evidence, and detailed steps at DEBUG, switched on only when chasing a problem.

One currency point for the practical. Apache's own page for Log4j 1 records that "On August 5, 2015 the Logging Services Project Management Committee announced that Log4j 1.x had reached end of life", and it lists vulnerabilities in Log4j 1 of which it says "none of the issues listed will be fixed". A practical written today should use Log4j 2, as Apache recommends.

Worked example: an execution run with its log and evidence

The program below runs four ExamReg fee tests against release 2.0's function, which still has the on-time defect found in Chapter One, on what software testing is. It logs every step with Python's logging module, marks a test blocked when its precondition fails (the payment sandbox is down), and on a failure saves a text snapshot of what the fee page showed, the stand-in here for a screenshot. The log format leaves out the time so that the output is the same on every run; a real log starts each line with a timestamp.

import logging, os, sys

logging.basicConfig(stream=sys.stdout, level=logging.INFO,
                    format="%(levelname)-7s %(message)s")
log = logging.getLogger("examreg-tests")

def late_fee(days_late):                     # release 2.0, with the on-time defect
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    return 100 if days_late <= 7 else 500

def fee_page(days_late):                     # what the student would see
    return f"ExamReg fee page\nLate fee: Rs {late_fee(days_late)}\n"

sandbox_up = False                           # the payment gateway's test sandbox
tests = [("TC-FEE-01", 0, 0, None), ("TC-FEE-03", 7, 100, None),
         ("TC-FEE-04", 8, 500, None), ("TC-PAY-21", 3, 100, "sandbox")]
os.makedirs("evidence", exist_ok=True)
tally = {"pass": 0, "fail": 0, "blocked": 0}
log.info("build 2.0.1, environment TEST-2, cycle 1")
for case_id, days, expected, needs in tests:
    if needs == "sandbox" and not sandbox_up:
        log.warning("%s blocked: payment sandbox unreachable", case_id)
        tally["blocked"] += 1
        continue
    actual = late_fee(days)
    if actual == expected:
        log.info("%s pass: %d days late, fee Rs %d", case_id, days, actual)
        tally["pass"] += 1
    else:
        evidence = os.path.join("evidence", case_id + ".txt")
        with open(evidence, "w") as page:
            page.write(fee_page(days))
        log.error("%s FAIL: %d days late, expected Rs %d, got Rs %d; evidence %s",
                  case_id, days, expected, actual, evidence)
        tally["fail"] += 1
log.info("cycle 1 totals: %s", tally)
print(open("evidence/TC-FEE-01.txt").read(), end="")
munotes.in53

Test Execution

INFO    build 2.0.1, environment TEST-2, cycle 1
ERROR   TC-FEE-01 FAIL: 0 days late, expected Rs 0, got Rs 100; evidence evidence/TC-FEE-01.txt
INFO    TC-FEE-03 pass: 7 days late, fee Rs 100
INFO    TC-FEE-04 pass: 8 days late, fee Rs 500
WARNING TC-PAY-21 blocked: payment sandbox unreachable
INFO    cycle 1 totals: {'pass': 2, 'fail': 1, 'blocked': 1}
ExamReg fee page
Late fee: Rs 100

Read the log as a developer would. The build and environment come first, so there is no argument about which version was tested. The failure names the test, the input, the expected and actual results and where the evidence is, and the evidence file shows exactly what a student would have seen. The blocked test is reported as blocked, with its reason, and counted separately: it is a problem to fix in the environment, not a defect in the fee rule.

After the fix: retesting and regression testing

When the developer fixes the defect and delivers a new build, two different things are tested, and they have different names.

Retesting, also called confirmation testing, is "testing performed to check that modifications made to correct a fault have successfully removed the fault" (ISO/IEC/IEEE 29119-2). Here: run TC-FEE-01 again and see Rs 0.

Regression testing is "testing performed following modifications to a test item or to its operational environment, to identify whether failures in unmodified parts of the test item occur" (ISO/IEC/IEEE 29119-1). Here: rerun the other fee tests, and the form and payment tests near the changed code, to show the fix broke nothing else.

Retesting (confirmation testing)Regression testing
QuestionIs this defect really fixed?Did the change break anything that used to work?
Tests runThe test that failedTests that previously passed
ScopeNarrow: the fixed defectBroad: the areas around, or the whole product
WhenAfter a fix is deliveredAfter any change: a fix, a new feature, an environment change
Automated?Often manualThe first candidate for automation, because it is repeated every build

Stopping and restarting: suspension and resumption

Sometimes testing must stop. ISO/IEC/IEEE 29119-1 defines suspension criteria as "criteria used to (temporarily) stop all or a portion of the testing activities". A test plan states them in advance, together with the resumption requirements that must be met to restart. For ExamReg, for example: suspend if the build fails its smoke test, or if more than a quarter of the scheduled tests are blocked; resume when a build passes the smoke test and the blocking problem is fixed. Deciding this in advance stops testers wasting days on a build that was never going to be usable, and stops a manager pushing testing on regardless.

munotes.in54

Test Execution

What it does not mean

A failed test is not automatically a defect in the software. The test, the data or the environment may be wrong; anomalies are analysed first.

Blocked is not failed, and not passed. It means the test could not run, and it is counted separately.

Retesting and regression testing are not the same. One checks the fix; the other checks everything else around it.

Execution is not only for automated tests. Manual execution follows the same rules: exact procedures, recorded actual results, a log and evidence.

Quick revision

  • Test execution: "process of running a test on the test item, producing actual results" (ISO/IEC/IEEE 29119-2).
  • Before it: entry criteria, environment and data readiness, a passing smoke test.
  • Actual results include data and environment, not only the screen. Statuses: pass, fail, blocked, not run.
  • Anomalies are analysed: the cause may be the software, the test, the environment or the data. Test incident: an event during execution that needs investigation; recorded in an incident report.
  • Test log: "chronological record of relevant details about the execution of tests" (ISO/IEC/IEEE 24765).
  • Evidence: logs (by level) and screenshots on failure. Log4j 1 reached end of life on 5 August 2015; use Log4j 2.
  • Retesting (confirmation): checks a fix. Regression testing: checks unmodified parts after a change.
  • Suspension criteria stop testing; resumption requirements restart it.

Test yourself

1. What is test execution? List its main tasks. Running tests on the test item to produce actual results: executing procedures as scheduled, comparing actual with expected results, logging outcomes, analysing anomalies and reporting incidents, then retesting fixes and running regression tests.

2. Distinguish retesting from regression testing. Retesting reruns the test that failed, to confirm a fix removed the fault. Regression testing reruns tests that previously passed, after any change, to check that failures have not appeared in unmodified parts.

3. What does "blocked" mean as a test status, and why must it be reported separately? The test could not be run because a precondition could not be met, such as a gateway sandbox being down. It is neither a pass nor a fail, and hiding blocked tests in either count distorts the picture of quality.

4. Why is evidence such as a log or a screenshot attached to a failure? So the failure can be shown and reproduced: it records what the user saw and what the program did, with the build, environment and data, which lets the developer find the defect instead of arguing about whether it exists.

munotes.in55

Test Execution

5. What are suspension criteria? Give one example. Criteria, decided in advance, for temporarily stopping all or part of testing; for example, suspending when a build fails its smoke test, and resuming when a build passes it.

6. Which Log4j should a practical use today, and why? Log4j 2. Apache announced on 5 August 2015 that Log4j 1.x had reached end of life, and the vulnerabilities found in it since will not be fixed.

munotes.in56

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!