Run Charts
Chapter One Hundred Seven
Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Run charts"
Pages 612 to 617 of 622
In one line
A run chart plots data in the order it happened against its own median, and four simple rules, about shifts, trends, the number of runs and standout points, turn that picture into an objective check for patterns too small to trust by eye but too real to be chance.
In the wording a student can write in an examination: a run chart is "a graphical display of data plotted in some type of order", usually time, against a median centreline rather than calculated limits (Perla, Provost and Murray). Its advantage, in the authors' words, is that "it preserves the time order of the data", unlike a significance test that only compares separate, already-aggregated summaries: the same mean and standard deviation can come from a process that improved and held, one that had already improved before a change, or one that improved and slid back, and only the time-ordered picture tells them apart. Four rules find non-random patterns against the median: Rule 1, shift, "Six or more consecutive points either all above or all below the median"; Rule 2, trend, "Five or more consecutive points all going up or all going down"; Rule 3, runs, too few or too many crossings of the median line, judged against tabled critical values; and Rule 4, astronomical point, a point "obviously, even blatantly, different from the rest", which is "subjective" where the first three are "probability based".
Why a median, and why these four rules
The median, not a calculated limit, is the run chart's centreline, for two reasons Perla, Provost and Murray give: "it provides the point at which half the observations are expected to be above and below the centreline" and "the median is not influenced by extreme values in the data." That second reason matters directly for ExamReg: a control chart's limits move if the data used to set them contains an outlier, which is why Chapter Ninety-One had to remove the two 503 responses before trusting its limits at all. A median barely moves.
The three probability-based rules (shift, trend, runs) are built for exactly this kind of small-sample question: is a pattern real, or could ordinary chance have produced it, at about a 5 per cent risk of a false alarm. Rule 4 is different by design, a place for judgement where the other three are silent; the authors note it should not be confused with a chart's highest or lowest point, which every run chart has whether or not anything is wrong.
Worked example: the fee page's 38 served times, run-ordered
Chapter Ninety-One's control chart found the fee page's 38 served times in control: every point inside 254.2 to 848.7 ms, no moving range too large, no Western Electric rule fired, and it promised that this chapter would put the same data to a different, more sensitive test. The program runs that test: the same 38 times Chapter Forty-Five's load testing recorded, in the order they were sent, against their own median.
Run Charts
from statistics import median
# Chapter Forty-Five's fee page load test, all forty requests (elapsed ms, response code);
# the two 503s are the special causes Chapter Ninety-One's control chart already removed
FEE_PAGE = [(380, 200), (410, 200), (417, 200), (463, 200), (454, 200), (516, 200), (491, 200),
(569, 200), (528, 200), (622, 200), (565, 200), (675, 200), (602, 200), (728, 200),
(639, 200), (431, 200), (676, 200), (484, 200), (413, 200), (537, 200), (450, 200),
(590, 200), (487, 200), (643, 200), (524, 200), (696, 200), (561, 200), (749, 200),
(598, 200), (452, 200), (635, 200), (505, 200), (672, 200), (558, 200), (409, 200),
(120, 503), (446, 200), (664, 200), (95, 503), (717, 200)]
times = [ms for ms, code in FEE_PAGE if code == 200]
m = median(times)
print(f"n = {len(times)}, median = {m}")
# Rule 1, shift: 6 or more consecutive points all above, or all below, the median
def longest_shift(values, m):
best, side, run = 0, None, 0
for v in values:
if v == m:
continue # on the median: skip, neither breaks nor extends
s = v > m
if s == side:
run += 1
else:
side, run = s, 1
best = max(best, run)
return best
shift = longest_shift(times, m)
print(f"rule 1, shift: longest run on one side of the median = {shift} "
f"({'SIGNAL' if shift >= 6 else 'no signal'})")
# Rule 2, trend: 5 or more consecutive points all rising, or all falling (repeats do not count)
def longest_trend(values):
best, direction, run = 0, None, 0
prev = None
for v in values:
if prev is not None and v != prev:
d = v > prev
if d == direction:
run += 1
else:
direction, run = d, 2 # the pair that started the new direction
best = max(best, run)
prev = v
return best
trend = longest_trend(times)
print(f"rule 2, trend: longest run rising or falling = {trend} "
f"({'SIGNAL' if trend >= 5 else 'no signal'})")
# Rule 3, runs: count crossings of the median, plus one; compare with table 1 (n = 38: 14 to 26)
def count_runs(values, m):
sides = [v > m for v in values if v != m]
return 1 + sum(a != b for a, b in zip(sides, sides[1:]))
runs = count_runs(times, m)
lower, upper = 14, 26
verdict = "too few" if runs < lower else "too many" if runs > upper else "within range"
print(f"rule 3, runs: {runs} runs (table 1 for n=38: {lower} to {upper}, {verdict})")
# Rule 4, astronomical point: subjective; none of the 38 stands out as obviously different
print("rule 4, astronomical point: none of the 38 is obviously unlike the rest, by eye")
first_seven = times[:7]
print(f"\nfirst seven times: {first_seven}")
print(f"all seven below the median: {all(t < m for t in first_seven)}")Run Charts
n = 38, median = 547.5
rule 1, shift: longest run on one side of the median = 7 (SIGNAL)
rule 2, trend: longest run rising or falling = 4 (no signal)
rule 3, runs: 18 runs (table 1 for n=38: 14 to 26, within range)
rule 4, astronomical point: none of the 38 is obviously unlike the rest, by eye
first seven times: [380, 410, 417, 463, 454, 516, 491]
all seven below the median: TrueFigure 107.1 The fee page's 38 served times, run-ordered against their median: the first seven are a shift the control chart did not flag
Rule 1 fires. The first seven served times are 380, 410, 417, 463, 454, 516 and 491 ms, all seven below the median of 547.5. Seven consecutive points on one side clears Rule 1's threshold of six, a signal at about the 5 per cent risk level the rule is built for.
Rules 2 and 3 do not. The longest rising or falling run is four points, one short of Rule 2's five. The 38 times cross the median 18 times, inside table 1's range of 14 to 26 for a sample this size, so Rule 3 finds nothing unusual about how often the line changes sides. Rule 4 finds no single point that stands out from the rest by eye; the two extreme low values are the 503s, already removed as special causes, not something the run chart itself had to catch.
A signal the control chart missed. Chapter Ninety-One's individuals chart judged the same 38 points in control: every one inside its limits, no rule of its own fired. A control chart's rules are built around distance from the centre and the size of successive differences; a run that stays on one side of the median without ever straying far from it can pass every one of those tests and still not be random. That is what "a different, more sensitive test" meant: not a better chart, a chart built to catch a different shape of pattern.
A plausible reading, not a proven one. A run of low times at the very start of a load test is a familiar shape in practice: the first few requests can benefit from a warm cache, a freshly opened connection, or a server not yet carrying its full test load, before the run settles into whatever the rest of the test looks like. Rule 1 only says the first seven are unlikely to be chance; it does not say why, and this book's data gives no independent check, the way Chapter One Hundred Six's scatter diagram checked the seven basic quality tools' open question of Chapter One Hundred Three with a second dataset. A real investigation would rerun the test and watch whether the first few requests are low again.
Run Charts
The run chart against the control chart
| Control chart (individuals) | Run chart | |
|---|---|---|
| Centreline | Mean of the data | Median of the data |
| What it needs | Enough data for a stable mean and moving range; sensitive to outliers while computing limits | Usable from a handful of points; the median resists outliers |
| What it catches | Points too far from the centre, or patterns in how far successive points differ | Long runs on one side of the median, long rises or falls, too few or too many crossings, one obviously odd point |
| What it does not catch well | A run that stays close to the centre but consistently to one side | The size of a shift, or how far outside a normal range a point falls |
| This chapter's data | 38 points in control: every one inside 254.2 to 848.7 ms | The same 38 points: a shift, the first seven all below the median |
Neither chart is the better one in general; they are built to catch different shapes of non-random pattern, and Perla, Provost and Murray note that a run chart is often the first, simplest display drawn before a more demanding chart like this book's individuals chart is built at all.
What it does not mean
A shift is not a magnitude. Rule 1 says seven points are unlikely to be on one side of the median by chance; it says nothing about how far below the median they are, which is a question for the control chart's limits, not the run chart's rules.
Passing three rules is not proof of nothing wrong. Rule 4 is deliberately subjective; a chart can pass every counted rule and still show something a reader's judgement catches that no rule was built to count.
The median rule needs enough points. Perla, Provost and Murray state plainly that "The shift and run rules require more than 10 points before they are applicable"; a run chart with five or six points can still be drawn and read by eye, but Rules 1 and 3 should not be applied to it.
A run chart signal is a reason to look, not a finished explanation. This chapter's shift is real by the rule's own test; the warm-start explanation offered for it is plausible, not established, exactly the caution Chapter One Hundred Six gave its own scatter diagram's association.
Run Charts
Quick revision
- Run chart (Perla, Provost and Murray, 2011): data plotted in time order against the median; preserves order that a summary statistic destroys.
- Median as centreline: half the points expected above and below it; unlike a mean, not pulled by extreme values.
- Rule 1, shift: 6 or more consecutive points all on one side of the median.
- Rule 2, trend: 5 or more consecutive points all rising or all falling; repeats do not count.
- Rule 3, runs: too few or too many crossings of the median, against table 1's critical values for the sample size.
- Rule 4, astronomical point: a point obviously unlike the rest; subjective, unlike the first three.
- Worked example: the fee page's 38 served times, in control by Chapter Ninety-One's control chart, fail Rule 1: the first seven are all below the median of 547.5, a shift the control chart's own rules did not raise.
Test yourself
1. What is a run chart, and what is its main advantage over a summary statistic? Data plotted in the order it occurred, usually against a median centreline. Its advantage is that it preserves time order: the same mean and standard deviation can arise from a sustained improvement, an improvement that had already started, or one that did not hold, and only the time-ordered chart tells these apart.
2. Why does a run chart use the median rather than the mean as its centreline? Because the median is the point half the observations are expected to fall above and below, and, unlike the mean, it is not pulled by extreme values in the data, so a single outlier does not shift where "normal" is drawn.
3. State Rules 1 and 2 and what each detects. Rule 1, shift: six or more consecutive points all above or all below the median signals a shift in the process. Rule 2, trend: five or more consecutive points all rising or all falling signals a trend; repeated equal values count once and do not break the run.
4. In the worked example, which rule fired, and on what evidence? Rule 1, shift. The first seven of the fee page's 38 served times, 380 to 491 ms, all fall below the median of 547.5 ms, clearing Rule 1's threshold of six consecutive points on one side.
5. Why could a control chart judge the same data in control while the run chart found a signal? A control chart's rules are built around distance from the centreline and the size of successive differences; a run that stays close to the centre but consistently on one side of it can satisfy those rules while still failing a rule built specifically to detect runs relative to the median.
Run Charts
6. Why is the warm-start explanation for the shift described as plausible rather than proven? Because Rule 1 only shows that seven points in a row below the median is unlikely to be chance; it gives no reason on its own. The warm-cache or fresh-connection explanation fits common experience of load tests, but nothing in this book's data independently confirms it, so it remains a reading of the signal, not a demonstrated cause.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.