munotes®

What Software Testing Is

Get access to whole semester resourcesSemester Pass

Chapter One

Syllabus topic Module 1, "Software Testing Fundamentals"

Pages 1 to 7 of 622

In one line

Software testing is checking a piece of software, by running it or by examining it, to find out where it goes wrong and how good it is, before its users find out for you.

In the wording a student can write in an examination: software testing is a set of activities that evaluates the quality of a software product and its work products and discovers defects in them. It may be dynamic, in which the software is executed under specified conditions and its results are observed and evaluated, or static, in which a work product such as a requirement, a design or the code is examined without being executed.

Why testing exists at all

Every program is written by people, and people make mistakes. A programmer misreads a requirement, forgets a case, types < where <= was meant, or builds exactly what was asked for when what was asked for was wrong. None of these mistakes announces itself. The program compiles, it runs, and on most inputs it may even give the right answer.

The mistake shows itself only when the wrong input arrives, and by then the software may be in front of the people who depend on it. Testing exists to move that moment earlier: to find the problem while it is still cheap to fix and nobody has been harmed by it. The next two chapters take the two halves of that sentence separately: what errors, faults and failures actually are (Chapter Two), and why software must be tested, given what a failure costs when it is not found in time (Chapter Three).

What the standards say testing is

Three definitions are worth knowing, because each stresses something different.

The software testing standard, ISO/IEC/IEEE 29119-2:2021, defines testing as a "set of activities conducted to facilitate discovery or evaluation of properties of one or more test items". A test item is simply the thing being tested; the same standard calls it a "work product to be tested", which covers a requirements document as much as a program.

IEEE 730-2014, the standard for software quality assurance processes, defines software testing more narrowly, as an "activity in which a system or component is executed under specified conditions, the results are observed or recorded, and an evaluation is made of some aspect of the system or component". That is dynamic testing: something is run.

The ISTQB Foundation syllabus (version 4.0.1, 2024), the most widely used syllabus for testers, puts it in one line: "Software testing is a set of activities to discover defects and evaluate the quality of software work products."

The definition broken down

Take IEEE 730's sentence apart and each phrase turns out to be doing work.

munotes.in1

What Software Testing Is

  1. A system or component. Testing can be aimed at a single function, a module, a whole application, or a system made of several applications. What is being tested at any moment is the test item, and choosing it is the first decision in any test.
  2. Executed under specified conditions. A test is not "try it and see". The inputs, the starting state of the data, the environment and the order of steps are decided in advance and written down, so that the test can be repeated by someone else and give the same result.
  3. The results are observed or recorded. What the software actually did is captured: the value returned, the page shown, the message printed, the record written to the database.
  4. An evaluation is made. The observed result is compared with the expected result, which was also decided in advance. A test without an expected result is only a demonstration: nobody can say whether it passed. Chapter Seven, on what a test case is, gives the expected result the attention it needs.

The ISO definition adds two things IEEE 730's does not. It speaks of discovery or evaluation, so testing is not only hunting for bugs but also measuring properties such as speed, ease of use or security. And it speaks of test items, not programs, which is what makes a review of a requirements document a form of testing too.

Two kinds of testing: dynamic and static

The standards name the two kinds precisely. ISO/IEC/IEEE 29119-2 defines dynamic testing as "testing in which a test item is evaluated by executing it", and static testing as "testing in which a test item is examined against a set of quality or other criteria without the test item being executed".

Dynamic testingStatic testing
Is the software run?YesNo
What it examinesThe behaviour of running codeRequirements, designs, code, test plans, any document
What it findsFailures, from which defects are tracedDefects directly
Typical methodsUnit, integration, system and acceptance testsReviews, inspections, walkthroughs, static analysis tools
When it can startOnce some code runsAs soon as the first document exists
Covered inModule 1 on strategies, Module 2 on techniquesModule 1 on reviews, inspection and walkthrough

The difference in the "what it finds" row matters. Dynamic testing sees a failure, the software doing the wrong thing, and someone must then work back to the defect that caused it. Static testing looks straight at the work product and sees the defect itself. Chapter Two, on errors, faults and failures, explains those words exactly.

The objectives of testing

Testing is done for more than one reason, and a test plan states which ones apply. The ISTQB syllabus lists the typical objectives as follows; the right-hand column says what each means in practice.

munotes.in2

What Software Testing Is

Objective (ISTQB v4.0.1, 1.1.1)What it means in practice
Evaluating work products such as requirements, user stories, designs, and codeReviewing documents before a line of code exists
Causing failures and finding defectsThe objective most people think of first
Ensuring required coverage of a test objectMaking sure every requirement, or every branch of the code, has been exercised
Reducing the risk level of inadequate software qualityTesting hardest where a failure would hurt most
Verifying whether specified requirements have been fulfilledChecking the product against its specification
Verifying that a test object complies with contractual, legal, and regulatory requirementsFor example, data protection or accessibility rules
Providing information to stakeholders to allow them to make informed decisionsTelling managers whether the product is ready to release
Building confidence in the quality of the test objectEvidence that it works, not only that it fails
Validating whether the test object is complete and works as expected by the stakeholdersChecking it does what users actually need

The syllabus adds that objectives vary with context: the work product, the test level, the risks, the development life cycle in use, and business factors such as time to market. A test of a hospital's dosage calculator and a test of a college canteen's menu page do not have the same objectives.

Worked example: the first test in this book

This book follows one invented system throughout, ExamReg, a college's online examination registration portal. Students log in, fill in the semester examination form, pay the fee and download the hall ticket. Every rule of ExamReg is this book's own example, not a rule of the University.

One of its rules is the late fee, and it is the whole specification the tester is given:

Form submittedLate fee
On or before the last datenone
1 to 7 days lateRs 100
8 to 15 days lateRs 500
More than 15 days lateform not accepted

A programmer writes the function below. A tester, who has not seen the code, writes five test cases from the rule alone, each with an input and an expected result, and runs them.

def late_fee(days_late):
    """Late fee in rupees for a form submitted days_late days after the last date."""
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    if days_late <= 7:
        return 100
    return 500

tests = [            # (days late, expected fee), written from the rule alone
    (0, 0),
    (3, 100),
    (7, 100),
    (8, 500),
    (15, 500),
]
for days, expected in tests:
    actual = late_fee(days)
    verdict = "pass" if actual == expected else "FAIL"
    print(f"{days:>2} days late: expected {expected:>3}, got {actual:>3}  {verdict}")
munotes.in3

What Software Testing Is

 0 days late: expected   0, got 100  FAIL
 3 days late: expected 100, got 100  pass
 7 days late: expected 100, got 100  pass
 8 days late: expected 500, got 500  pass
15 days late: expected 500, got 500  pass

Four tests pass and one fails. Look at what the failure says: a student who submits the form on time is charged Rs 100. The program runs, it gives a sensible-looking number for every input, and on four of the five inputs it is right. Nothing about it looks broken until a test asks the one question it gets wrong.

Notice also where the failing input came from. The tester chose 0 because the rule has a boundary there, between "on time" and "late". Chapter Fifty-Five, on boundary value analysis, turns that instinct into a method.

Testing and debugging are different activities

The failing test has done its job. It has shown that a defect exists. It has not said where the defect is, and finding it is a different activity with a different name. The ISTQB syllabus describes debugging after a dynamic test fails as three steps: reproduction of the failure, diagnosis (finding the defect), and fixing the defect.

Here the three steps are short. Reproduce: call late_fee(0) again and get 100 again. Diagnose: the condition days_late <= 7 is true for 0, so an on-time form falls into the Rs 100 branch; the rule's first case, "on or before the last date", was never written into the code. Fix: handle that case first.

Then two more kinds of testing follow the fix. Confirmation testing reruns the test that failed, to confirm the fix works. Regression testing reruns the tests that passed, to confirm the fix has not broken anything else. Both are run below.

def late_fee(days_late):
    """Late fee in rupees for a form submitted days_late days after the last date."""
    if days_late > 15:
        raise ValueError("form not accepted more than 15 days late")
    if days_late <= 0:                  # on or before the last date: the missing case
        return 0
    if days_late <= 7:
        return 100
    return 500

tests = [(0, 0), (3, 100), (7, 100), (8, 500), (15, 500)]
print("confirmation test:", "pass" if late_fee(0) == 0 else "FAIL")
failed = [days for days, expected in tests if late_fee(days) != expected]
print("regression run:", len(tests) - len(failed), "of", len(tests), "pass")
confirmation test: pass
regression run: 5 of 5 pass
TestingDebugging
PurposeTo show that failures occur, or to find defects in a work productTo find the defect behind a failure and remove it
Starts fromA requirement and an expected resultA failure that has been observed
Usually done byTesters, and developers testing their own unitsThe developer who owns the code
Needs knowledge of the code?Not for black-box testingAlways
Ends withA pass or fail verdict, and a defect reportA corrected program, then confirmation and regression tests
munotes.in4

What Software Testing Is

Debugging has a chapter of its own later in Module 1 (Chapter Forty-Seven), with its techniques.

Dijkstra's warning

The fixed function passes all five tests. Is it now correct? The five tests cannot say. They checked five inputs out of every integer the function might receive, and a defect could be sitting on an input nobody tried.

Edsger Dijkstra put this limit in one sentence in his 1972 ACM Turing Lecture, "The Humble Programmer": "program testing can be a very effective way to show the presence of bugs, but is hopelessly inadequate for showing their absence."

The late-fee function shows both halves. The first test run showed the presence of a bug, effectively and cheaply. The second run showed that five tests pass, which is not the same as showing that no bug remains. What happens, for instance, if days_late is negative? The rule does not say, and the tests did not ask. Chapter Four turns this limit into the first of the seven principles of testing.

A static test of the same rule

Go back to the rule itself and read it as a tester, without running anything. It says what happens on time, 1 to 7 days late, 8 to 15 days late and more than 15 days late. It does not say what the portal should do if the date on a form is before the day the form was opened, or if the number of days is missing. Those are gaps in the requirement.

Noticing them is static testing: a work product was examined against a criterion (is every possible input covered?) without executing anything. It found a defect in the requirement, and it found it before a single line of code depended on the missing answer. The same questions asked after release would arrive as complaints.

What testing does not mean

Testing is not only running the program. Reviewing a requirement, a design or the code is testing too: static testing. The ISTQB syllabus calls the belief that testing only consists of executing tests "a common misconception".

A test that passes does not prove the program correct. It proves the program gave the expected result for that input, in that environment, on that day. Dijkstra's sentence is the reason.

Testing is not debugging. Testing shows that something is wrong; debugging finds out why and removes it. They are often done by different people, and they need different information.

munotes.in5

What Software Testing Is

Finding bugs is not the only objective. Measuring quality, checking compliance, building confidence and giving managers the information to decide on a release are objectives too. A test run in which nothing fails is still useful: it is evidence.

A test without an expected result is not a test. Running the program and looking at what comes out, with nothing decided in advance to compare it with, tells nobody whether it passed.

Where testing stops

Testing samples. It runs some inputs out of an enormous number, so it can reduce the risk that a defect reaches users but can never remove it. It also depends on someone knowing the right answer: if the expected result is itself wrong, a correct program fails the test and a wrong one passes. And testing costs time and money, so every project decides how much testing is enough, a question Module 1 returns to in its strategic approach to testing, which asks when testing is complete (Chapter Thirty-One).

Quick revision

  • Software testing: a set of activities that evaluates quality and discovers defects in software and its work products.
  • ISO/IEC/IEEE 29119-2: testing is a "set of activities conducted to facilitate discovery or evaluation of properties of one or more test items".
  • Dynamic testing executes the test item; static testing examines it without executing it (reviews, static analysis).
  • A test needs an input, specified conditions and an expected result; the actual result is compared with it.
  • ISTQB objectives include evaluating work products, finding defects, coverage, reducing risk, verifying requirements, compliance, informing decisions, building confidence and validating.
  • Testing shows failures; debugging reproduces, diagnoses and fixes; then confirmation testing reruns the failed test and regression testing reruns the rest.
  • Dijkstra (1972): testing can show the presence of bugs, never their absence.
  • Worked example: ExamReg's late_fee(0) returned 100 instead of 0; one test in five found it.

Test yourself

1. Define software testing and name its two kinds. Software testing is a set of activities that evaluates the quality of software and its work products and discovers defects in them. Dynamic testing executes the test item under specified conditions and evaluates the results; static testing examines a work product, such as a requirement or the code, without executing it.

2. Why must a test case have an expected result decided in advance? Because a test is an evaluation: the actual result is compared with the expected one. Without an expected result nobody can say whether the test passed, and a wrong output would be accepted simply because it looked reasonable.

3. Differentiate testing from debugging. Testing shows that failures occur, or finds defects directly in a work product; it starts from a requirement and ends with a pass or fail verdict. Debugging starts from an observed failure, reproduces it, diagnoses the defect and fixes it; it is followed by confirmation testing and regression testing.

munotes.in6

What Software Testing Is

4. In the late-fee example, which test failed, and what did the fix require besides changing the code? The test with 0 days late failed: the function returned Rs 100 instead of 0. After the fix, confirmation testing reran that test and regression testing reran the four tests that had passed, to show the fix broke nothing else.

5. Explain Dijkstra's statement about testing, with an example. Testing can show that bugs are present but cannot show that they are absent, because it checks only the inputs actually tried. In the late-fee example, five passing tests say nothing about a negative number of days, which no test tried and the rule does not cover.

6. Is reading a requirements document and finding a gap in it testing? Justify. Yes. It is static testing: a work product is examined against a criterion, here completeness, without anything being executed. It finds the defect directly and earlier than any dynamic test could.

7. State any four objectives of testing. Any four of: evaluating work products; causing failures and finding defects; ensuring required coverage; reducing the risk of inadequate quality; verifying that requirements are fulfilled; verifying compliance with contractual, legal and regulatory requirements; providing information for decisions; building confidence; validating that the product works as stakeholders expect.

munotes.in7

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!