munotes®

Scatter Diagrams

Get access to whole semester resourcesSemester Pass

Chapter One Hundred Six

Syllabus topic Module 2, "Software Reviews & Quality Improvement Techniques: ... Scatter Diagrams"

Pages 608 to 611 of 622

In one line

A scatter diagram plots one variable against another to show whether they move together, letting a team see a relationship, and often its shape and its exceptions, before running any statistic; the diagram can only show that two things are associated, never that one causes the other.

In the wording a student can write in an examination: a scatter diagram (also called a scatter plot or X-Y graph) "graphs pairs of numerical data, with one variable on each axis, to look for a relationship between them. If the variables are correlated, the points will fall along a line or curve. The better the correlation, the tighter the points will hug the line" (ASQ). The NIST/SEMATECH handbook states its purpose as checking "for Relationship": a scatter plot "reveals relationships or association between two variables", and can show whether two variables are related, whether that relation is linear or not, whether the spread of one changes with the other, and whether there are outliers. Both sources give the same warning in different words: NIST states plainly that "There is no statistical procedure, the scatter plot included, that proves cause-and-effect", and ASQ warns "do not assume that one variable caused the other. Both may be influenced by a third variable."

Building one, and testing what it shows

ASQ's procedure: collect paired data where a relationship is suspected; draw the independent variable on the horizontal axis and the dependent variable on the vertical; plot a point, or a touching pair of points, for every pair of values; then look at the pattern. If a line or curve is clear, ASQ says a team "may stop because variables are correlated" and "may wish to use regression or correlation analysis now." If it is not clear, ASQ gives a manual significance test: split the points into four quadrants at their medians, add the diagonally opposite quadrants' counts to get two sums, take the smaller (Q) and the total (N), and compare Q against a trend-test table for that N; a small enough Q says the pattern is unlikely to be chance. This chapter takes the first path ASQ names, computing a correlation coefficient by program, because release 2.0's data is exactly the "clear pattern, check it numerically" case.

Worked example: does the day 6 spike track logins, or something else?

Chapter One Hundred Three, on the seven basic quality tools, left a question open: the help desk's complaints jumped on day 6 of the first registration week, and it could not say from the check sheet alone whether the portal simply failed more under that day's load, or whether more students were merely using it that day. A scatter diagram of that day's logins against that day's complaints, across all six days, answers the first half; fitting the first five days alone and checking what they would have predicted for day 6 answers the second.

munotes.in608

Scatter Diagrams

from statistics import correlation

# release 2.0's first registration week: daily logins (FINDINGS 5.14) and help desk
# complaints (FINDINGS 5.11), by day; day 6 is the last day before the late fee starts
days = [1, 2, 3, 4, 5, 6]
logins =     [430, 390, 470, 400, 460, 900]
complaints = [8,   8,   7,   11,  12,  32]

print(f"{'day':>3}{'logins':>8}{'complaints':>12}{'per 100 logins':>16}")
for d, lg, cp in zip(days, logins, complaints):
    print(f"{d:>3}{lg:>8}{cp:>12}{100 * cp / lg:>16.2f}")

r_all = correlation(logins, complaints)
r_five = correlation(logins[:5], complaints[:5])
print(f"\nPearson's r, all six days: {r_all:.3f}")
print(f"Pearson's r, days 1 to 5 only: {r_five:.3f}")

# the least-squares line through days 1 to 5, extended to day 6's logins
n = 5
mean_x = sum(logins[:5]) / n
mean_y = sum(complaints[:5]) / n
slope = (sum((lg - mean_x) * (cp - mean_y) for lg, cp in zip(logins[:5], complaints[:5]))
         / sum((lg - mean_x) ** 2 for lg in logins[:5]))
intercept = mean_y - slope * mean_x
predicted_day6 = intercept + slope * logins[5]
print(f"\ndays 1 to 5 alone predict {predicted_day6:.1f} complaints at day 6's {logins[5]} logins;"
      f" day 6 actually had {complaints[5]}")
day  logins  complaints  per 100 logins
  1     430           8            1.86
  2     390           8            2.05
  3     470           7            1.49
  4     400          11            2.75
  5     460          12            2.61
  6     900          32            3.56

Pearson's r, all six days: 0.965
Pearson's r, days 1 to 5 only: -0.033

days 1 to 5 alone predict 8.3 complaints at day 6's 900 logins; day 6 actually had 32
A scatter diagram of daily complaints against daily logins: five ordinary days clustered near 400 logins, a dashed near-flat line fitted through them, and day 6 far above and to the right, well clear of the point that line predicts at its login count

Figure 106.1 Release 2.0's first registration week: complaints against logins, and what days 1 to 5 alone would have predicted for day 6

All six days look strongly related. Pearson's r across all six is 0.965, close to the ASQ description of points that "hug the line" tightly. A team that stopped there could read it as confirmation that complaints simply track traffic.

Without day 6, there is nothing to see. The same statistic on days 1 to 5 alone is minus 0.033, no relationship at all. Those five days' logins move within a narrow band, 390 to 470, which is exactly the case ASQ's considerations warn about: "consider whether the independent (x-axis) variable has been varied widely. Sometimes a relationship is not apparent because the data do not cover a wide enough range." The high overall r is not five ordinary days confirming a pattern; it is one distant point, high on both axes, dominating a statistic computed from only six.

What days 1 to 5 would have predicted, against what happened. Fitting days 1 to 5 alone and reading off their line at day 6's login count gives 8.3 complaints. Day 6 had 32, nearly four times that. If day 6 were only a bigger version of an ordinary day, this is roughly what it would have produced instead.

munotes.in609

Scatter Diagrams

Association, and the explanation a scientist supplies. NIST is explicit that a scatter plot can never prove cause and effect, and that "it is ultimately only the researcher (relying on the underlying science/engineering) who can conclude that causality actually exists." The statistic here only shows that day 6 breaks the pattern the other five days set; it does not say why. Chapter One Hundred Five, on cause-effect diagrams, already supplied a mechanism that fits: the gateway's confirmation callback racing the form's session timer, a race more likely to be lost as concurrent load rises near the deadline. The complaint rate of 3.56 per 100 logins on day 6, against 1.49 to 2.75 across the other five, is consistent with a rate that worsens under load, not merely a count that grows with it; the scatter diagram supplies the association, that chapter's fishbone and five whys supply the engineering reason a reader can trust it.

What it does not mean

A strong overall correlation is not evidence by itself. With only six points, one extreme pair can produce a high r even when the rest show nothing; always check what the statistic looks like with that point removed, as Chapter Ninety-One's control chart did with its own 72-hour outlier.

No correlation in a narrow range does not mean no relationship exists. Days 1 to 5's flat line reflects a login count that barely changed across those days; it says nothing about what would happen at day 6's volume, which is exactly why day 6 is worth plotting rather than assumed away.

Correlation is not causation, however strong it looks. Both logins and complaints could rise together on day 6 because a third factor, the approaching deadline, drives both independently; the scatter diagram cannot rule that out by itself.

A relationship found here does not transfer to every release. Release 2.1's fix for the gateway callback (were one made) would need its own week of data before the same chart could be trusted again.

Quick revision

  • Scatter diagram (ASQ; also scatter plot, X-Y graph): plots paired data, one variable per axis, to look for a relationship; a tighter hug to a line or curve means a stronger correlation.
  • Procedure (ASQ): plot the pairs; if a pattern is clear, use regression or correlation analysis; otherwise split into quadrants at the medians and compare Q, the smaller diagonal sum, against N on a trend-test table.
  • What it answers (NIST/SEMATECH): whether two variables are related, linearly or not, whether the spread of one depends on the other, and whether there are outliers.
  • Not proof of cause: NIST states directly that no statistical procedure, the scatter plot included, proves cause and effect; ASQ warns that a third variable may drive both.
  • Worked example: all six days, r = 0.965; days 1 to 5 alone, r = minus 0.033; their line predicts 8.3 complaints at day 6's login count, against 32 actual, consistent with the gateway-callback mechanism Chapter One Hundred Five's cause-effect diagram found worsening under load.
munotes.in610

Scatter Diagrams

Test yourself

1. What is a scatter diagram, and what does a tighter clustering of points along a line mean? A plot of paired data, one variable on each axis, used to look for a relationship between them. Points that hug a line or curve tightly indicate a stronger correlation; a wide scatter indicates a weak one.

2. Outline ASQ's procedure for testing whether a scatter diagram's pattern is real. Plot the pairs. If a line or curve is obvious, correlation or regression analysis may be used directly. Otherwise, divide the points into four quadrants at their medians, sum the counts in diagonally opposite quadrant pairs, take the smaller sum (Q) and the total (N), and compare Q against a trend-test table for that N to judge whether the pattern could be chance.

3. Why can a scatter diagram never prove that one variable causes another? Because association only shows that two variables move together; a scatter diagram cannot rule out a third variable driving both, or the causation running the other way. NIST states this directly: no statistical procedure, the scatter plot included, proves cause and effect.

4. In the worked example, why did the correlation change so much between all six days and the first five alone? Across all six days, one point, day 6, is far higher on both axes than the rest and dominates the statistic, giving r = 0.965. Among the first five days alone, logins vary little and show no relationship to complaints, r = minus 0.033; the apparent strong correlation came from a single point, not from five days confirming a pattern.

5. What did fitting days 1 to 5 alone predict for day 6, and what actually happened? Their least-squares line predicted about 8.3 complaints at day 6's login count. Day 6 actually had 32 complaints, nearly four times the prediction, showing day 6 was not simply a larger version of an ordinary day.

6. How does Chapter One Hundred Five's cause-and-effect work relate to this chapter's finding? The scatter diagram shows only that day 6 breaks the pattern the other days set, an association; it cannot say why. The five whys of Chapter One Hundred Five identified a specific mechanism, the gateway's callback racing the session timer under load, that plausibly explains why the complaint rate, not just the count, rises on the highest-traffic day.

munotes.in611

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!