Load Testing: Users, Ramp-Up, Throughput and Error Rate
Chapter Forty-Five
Syllabus topic Practical, "Load Testing Using Apache JMeter"
Pages 247 to 251 of 622
In one line
A load test drives a system with a planned number of simulated users and measures how it responds: in JMeter, a thread group sets the users, how quickly they arrive (the ramp-up) and how often each repeats (the loop count), and the results are read as response times, percentiles, throughput and the percentage of errors.
In the wording a student can write in an examination: load testing evaluates "the behavior of a test item under anticipated conditions of varying load, usually between anticipated conditions of low, typical, and peak usage" (ISO/IEC/IEEE 29119-1:2022). In Apache JMeter, the practical's tool, a thread group "controls the number of threads JMeter will use to execute your test": each thread is one simulated user; the ramp-up period "tells JMeter how long to take to 'ramp-up' to the full number of threads chosen"; and the loop count sets how many times each thread repeats the test. The results are read from reports such as the Aggregate Report: the average, median and 90% line of response time, the error % ("Percent of requests with errors") and the throughput, where "Throughput = (number of requests) / (total time)".
The practical's task
The practical asks students to "Design and execute load testing scenarios for a web application. Configure thread groups, ramp-up time, and loop count. Analyze response time, throughput, and error percentage using graphical reports." This chapter explains each of those terms from JMeter's own manual and computes each measure from a results file, so that the numbers JMeter shows in the lab can be read with understanding.
Load, stress, endurance and spikes
Chapter Forty-Four placed load testing in the performance family. For ExamReg the differences look like this:
- Load test: 50, 200 and 500 simultaneous students, the range the requirements anticipate, against the target of the fee page in 2 seconds.
- Stress test: 800 and 1,200 students, beyond capacity, to see how the portal fails and whether it recovers.
- Endurance (soak) test: a typical load kept up for many hours, to find slow leaks and the long-running faults of Chapter Three, on why software must be tested.
- A spike: a sudden jump in load. JMeter's manual uses the word for a load problem a badly configured test can create by accident: with 100 threads and a ramp-up of 0, "all the threads would start at the same time, and it would produce an unwanted spike of the load." For ExamReg a spike is also real: the minute hall tickets are released, many students ask for them at once, and a test should reproduce that on purpose.
The thread group: threads, ramp-up and loop count
JMeter's manual describes the thread group as the start of every test: "Thread group elements are the beginning points of any test plan", and its controls let a tester "Set the number of threads", "Set the ramp-up period" and "Set the number of times to execute the test".
Load Testing: Users, Ramp-Up, Throughput and Error Rate
- Threads. "Each thread will execute the test plan in its entirety and completely independently of other test threads. Multiple threads are used to simulate concurrent connections to your server application." A thread is one simulated user.
- Ramp-up period. The manual's own example: "If 10 threads are used, and the ramp-up period is 100 seconds, then JMeter will take 100 seconds to get all 10 threads up and running. Each thread will start 10 (100/10) seconds after the previous thread was begun." Its advice cuts both ways: "Ramp-up needs to be long enough to avoid too large a work-load at the start of a test, and short enough that the last threads start running before the first ones finish (unless one wants that to happen)." Its rule of thumb: "Start with Ramp-up = number of threads and adjust up or down as needed."
- Loop count. "By default, the thread group is configured to loop once through its elements." A loop count of 5 makes each thread run the test five times, so the plan sends threads × loops requests for each sampler.
Reading the results
JMeter's glossary defines the measurements precisely.
- Elapsed time: "JMeter measures the elapsed time from just before sending the request to just after the last response has been received." This is the response time the reports use.
- Latency: measured "from just before sending the request to just after the first response has been received".
- Median: "a number which divides the samples into two equal halves".
- 90% Line: "the value below which 90% of the samples fall". The Aggregate Report puts it in words a student can use: "90 % of the samples took no more than this time. The remaining samples took at least as long as this." It also reports the 95% and 99% lines.
- Error %: "Percent of requests with errors".
- Throughput: "Throughput is calculated as requests/unit of time. The time is calculated from the start of the first sample to the end of the last sample." So "Throughput = (number of requests) / (total time)".
The percentiles matter more than the average. An average can look healthy while a tenth of the students wait far too long; the 90% line says how long nine students in ten waited at most, which is closer to what a requirement like the fee page within 2 seconds means.
Worked example: reading one test's results
A tester runs a first load test on ExamReg's fee page with a thread group of 20 threads, a ramp-up of 20 seconds and a loop count of 2, and saves the results. JMeter can write its results to a CSV file; the one below is this book's illustration of such a file, reduced to the five columns the report needs: when each request started (in milliseconds from the start of the test), its elapsed time, its label, the server's response code and whether it succeeded.
Load Testing: Users, Ramp-Up, Throughput and Error Rate
start_ms,elapsed_ms,label,response_code,success
0,380,fee page,200,true
380,410,fee page,200,true
1000,417,fee page,200,true
1417,463,fee page,200,true
2000,454,fee page,200,true
2454,516,fee page,200,true
3000,491,fee page,200,true
3491,569,fee page,200,true
4000,528,fee page,200,true
4528,622,fee page,200,true
5000,565,fee page,200,true
5565,675,fee page,200,true
6000,602,fee page,200,true
6602,728,fee page,200,true
7000,639,fee page,200,true
7639,431,fee page,200,true
8000,676,fee page,200,true
8676,484,fee page,200,true
9000,413,fee page,200,true
9413,537,fee page,200,true
10000,450,fee page,200,true
10450,590,fee page,200,true
11000,487,fee page,200,true
11487,643,fee page,200,true
12000,524,fee page,200,true
12524,696,fee page,200,true
13000,561,fee page,200,true
13561,749,fee page,200,true
14000,598,fee page,200,true
14598,452,fee page,200,true
15000,635,fee page,200,true
15635,505,fee page,200,true
16000,672,fee page,200,true
16672,558,fee page,200,true
17000,409,fee page,200,true
17409,120,fee page,503,false
18000,446,fee page,200,true
18446,664,fee page,200,true
19000,95,fee page,503,false
19095,717,fee page,200,trueThe program works out what the thread group means, then computes an Aggregate Report row from the file by the definitions above, and finally counts how many requests were ever in flight at the same moment.
import csv
import math
threads, ramp_up_s, loops = 20, 20, 2 # the thread group, as configured
print(f"thread group: {threads} threads, ramp-up {ramp_up_s} s, loop count {loops}:"
f" a thread starts every {ramp_up_s / threads:.1f} s, {threads * loops} requests planned")
with open("fee-page-results.csv") as f:
samples = [(int(r["start_ms"]), int(r["elapsed_ms"]), r["success"] == "true")
for r in csv.DictReader(f)]
times = sorted(e for _, e, _ in samples)
n = len(times)
def line(pct): # "pct % of the samples took no more than this"
return times[math.ceil(pct / 100 * n) - 1]
errors = sum(1 for _, _, ok in samples if not ok)
first_start = min(s for s, _, _ in samples)
last_end = max(s + e for s, e, _ in samples)
throughput = n / ((last_end - first_start) / 1000) # requests / total time
print(f"samples {n}, average {sum(times) / n:.0f} ms, median {line(50)} ms,"
f" 90% line {line(90)} ms, 95% line {line(95)} ms, 99% line {line(99)} ms")
print(f"min {times[0]} ms, max {times[-1]} ms, error {errors / n:.1%},"
f" throughput {throughput:.2f} requests/s")
events = sorted([(s, 1) for s, _, _ in samples] + [(s + e, -1) for s, e, _ in samples])
in_flight = peak = 0
for _, change in events:
in_flight += change
peak = max(peak, in_flight)
print(f"most requests in flight at once: {peak} (the plan had {threads} threads)")thread group: 20 threads, ramp-up 20 s, loop count 2: a thread starts every 1.0 s, 40 requests planned
samples 40, average 529 ms, median 528 ms, 90% line 676 ms, 95% line 717 ms, 99% line 749 ms
min 95 ms, max 749 ms, error 5.0%, throughput 2.02 requests/s
most requests in flight at once: 2 (the plan had 20 threads)Load Testing: Users, Ramp-Up, Throughput and Error Rate
Most of the report is what a student expects to read. Forty samples; an average and a median near 530 milliseconds; a 90% line of 676 ms, so nine requests in ten took no more than that; two failures with response code 503, an error rate of 5 per cent; and a throughput of about 2 requests a second.
The last line is the real finding. The plan had 20 threads, but at no moment were more than 2 requests in flight, because each thread finished its two requests in about a second while the next thread started a second later. The manual's warning describes exactly this: ramp-up must be "short enough that the last threads start running before the first ones finish". This test never came near 20 simultaneous students, so it says nothing about the portal's behaviour with 20, let alone the 500 of the requirement. The next run should shorten the ramp-up, raise the loop count, or set a duration, and the concurrency should be checked again before anyone reads the response times as an answer.
Two other cautions follow from the manual. The 503 responses may be the server refusing work under load, or a fault that has nothing to do with load; the report counts them, and a tester must look at them. And throughput depends on how the test was built: the Aggregate Report notes that "If other samplers and timers are in the same thread, these will increase the total time, and therefore reduce the throughput value."
A load test plan for ExamReg
| Question | ExamReg's answer |
|---|---|
| What load is anticipated? | Up to 500 simultaneous students on the last date |
| Which pages? | Login, exam form, fee page, payment hand-off, hall ticket |
| Threads and ramp-up | 50, 200 and 500 threads, with ramp-ups short enough that all are running together |
| Duration | Long enough at each level for response times to settle, checked by the in-flight count |
| Targets | 90% line of the fee page within 2 seconds; error % below 1 per cent |
| Environment | A copy of production, with the payment provider's sandbox |
| Reading | Aggregate Report per page; response time against active threads, graphed |
What it does not mean
The number of threads is not the number of simultaneous users. It is the most that can be active; the ramp-up, loop count and response times decide how many actually overlap, as the worked example showed.
The average response time is not the students' experience. Percentiles, the 90% and 95% lines, say how long most students waited at worst.
Load Testing: Users, Ramp-Up, Throughput and Error Rate
A load test is not a stress test. Load testing stays within anticipated load; stress testing goes beyond it on purpose.
A high error percentage is not always the server's fault. A test that sends malformed requests, or runs out of test data, produces errors too; each kind must be looked at.
Quick revision
- Load testing: behaviour under anticipated load, from low to peak (ISO/IEC/IEEE 29119-1:2022).
- Thread group: threads (simulated users), ramp-up period (time to start them all; delay = ramp-up ÷ threads), loop count (repetitions per thread).
- Ramp-up "long enough to avoid too large a work-load at the start", "short enough that the last threads start running before the first ones finish"; start with ramp-up = threads.
- Elapsed time, latency, median, 90% line (90 per cent of samples took no more), error %, throughput = requests ÷ total time.
- Worked example: 40 samples, median 528 ms, 90% line 676 ms, 5 per cent errors, about 2 requests/s, and never more than 2 requests in flight, so the test did not load the portal as planned.
Test yourself
1. Explain the thread group's settings: number of threads, ramp-up period and loop count. The number of threads is the number of simulated users, each running the test plan independently. The ramp-up period is how long JMeter takes to start all the threads, so each starts ramp-up divided by threads seconds after the previous one. The loop count is how many times each thread repeats the test.
2. If 30 threads are configured with a ramp-up period of 120 seconds, when does each thread start? Each successive thread starts 4 seconds after the previous one, the 120 seconds shared among 30 threads, so the last starts 116 seconds into the test.
3. What are the 90% line, error % and throughput in a JMeter report? The 90% line is the response time that 90 per cent of samples did not exceed. Error % is the percentage of requests that failed. Throughput is the number of requests divided by the total time from the start of the first sample to the end of the last.
4. Why is the 90% line often more useful than the average? Because an average can look acceptable while a tenth of users wait far too long; the 90% line states the worst wait for nine users in ten, which matches requirements such as a page appearing within 2 seconds.
5. A test with 20 threads shows good response times, but no more than 2 requests were ever in flight. What went wrong, and what should be changed? The ramp-up was too long for requests so short: each thread finished before the next few started, so the server never carried 20 users at once. The ramp-up should be shortened, the loop count raised or a duration set, and the concurrency checked before the response times are trusted.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.