← Home

Statistics

Grade 12
Step 1 of 35
INTRODUCTION
1 / 35
Ready? Try the Quiz →

Full lesson notes

Everything covered in this lesson, in one place - useful for revision or printing.

\[ \hat{y} = a + bx \qquad -1 \le r \le 1 \]

Grade 12 Statistics is about bivariate data — two variables measured on the same individuals — and whether one can be used to predict the other.

In this lesson you will learn to:

1. Summarising Data

\[ \bar{x} = \frac{\sum x}{n} \]

MeasureWhat it isWhen to use it
Mean \(\bar{x}\)The arithmetic averageData is roughly symmetric with no outliers
MedianThe middle value once orderedData is skewed or has outliers
ModeThe most frequent valueCategorical data, or the most common item
Only the mean uses every value — which is exactly why one extreme value can drag it away from the typical value.

\[ \text{min} \;\; Q_1 \;\; Q_2 \;\; Q_3 \;\; \text{max} \]

For the ordered data \(12;\ 15;\ 17;\ 20;\ 22;\ 25;\ 30\):

Minimum \(=12\)
\(Q_1\) = median of the lower half \(\{12;15;17\} = 15\)
\(Q_2\) = median \(= 20\)
\(Q_3\) = median of the upper half \(\{22;25;30\} = 25\)
Maximum \(=30\)

\[ Q_1 \text{ at } \tfrac{1}{4}(n+1) \qquad Q_2 \text{ at } \tfrac{1}{2}(n+1) \qquad Q_3 \text{ at } \tfrac{3}{4}(n+1) \]

With \(n=7\) values:

\(Q_1\) position \(= \tfrac{1}{4}(8) = 2\)nd value \(= 15\)
\(Q_2\) position \(= \tfrac{1}{2}(8) = 4\)th value \(= 20\)
\(Q_3\) position \(= \tfrac{3}{4}(8) = 6\)th value \(= 25\)
If the position is not a whole number, take the average of the two values on either side.

\[ \text{order the data FIRST — always} \]

Key idea: every quartile calculation assumes the data is in ascending order. Sorting first is not optional.

2. Dispersion

\[ \text{same mean} \;\ne\; \text{same data} \]

Two classes can both average 60% while one is tightly bunched and the other is all over the place. Dispersion measures that spread.

MeasureFormulaResistant to outliers?
Range\(\text{max} - \text{min}\)No — it uses only the extremes
IQR\(Q_3 - Q_1\)Yes — it uses the middle half
Standard deviation\(\sigma=\sqrt{\dfrac{\sum(x-\bar{x})^2}{n}}\)No — it uses every value

\[ \sigma = \sqrt{\frac{\sum (x-\bar{x})^{2}}{n}} \]

For the data \(4;\ 8;\ 10;\ 12;\ 16\):

\(x\)\(x-\bar{x}\)\((x-\bar{x})^2\)
4\(-6\)36
8\(-2\)4
1000
1224
16636
\(\bar{x}=10\)\(\sum = 80\)
Variance \(= \dfrac{80}{5} = 16\)
Standard deviation \(= \sqrt{16} = 4\)

\[ \bar{x} \pm \sigma \;=\; 10 \pm 4 \]

The standard deviation is roughly the typical distance of a value from the mean.

\(\bar{x}-\sigma = 6\) and \(\bar{x}+\sigma = 14\)
Three of the five values (8, 10, 12) lie within one standard deviation of the mean.
A SMALL standard deviation means consistent, tightly clustered data. A LARGE one means the values are widely scattered. Always interpret it in context — the examiner asks for a sentence, not just a number.

\[ \text{centre} + \text{spread} = \text{a complete description} \]

Key idea: never report a mean without a measure of spread. Two data sets with the same mean can tell completely different stories.

3. Box & Whisker

\[ \text{min} \;\rule[2pt]{18pt}{1pt}\; Q_1 \;\blacksquare\; Q_2 \;\blacksquare\; Q_3 \;\rule[2pt]{18pt}{1pt}\; \text{max} \]

A box-and-whisker diagram is the five-number summary drawn to scale.

\[ \text{long tail} \;\Rightarrow\; \text{skewed in that direction} \]

ShapeMeaning
Median in the middle, whiskers equalSymmetric
Median nearer \(Q_1\), long right whiskerSkewed to the RIGHT (positively skewed)
Median nearer \(Q_3\), long left whiskerSkewed to the LEFT (negatively skewed)
Skewed right: mean > median. Skewed left: mean < median. The extreme tail drags the mean toward it.

\[ \text{outlier if } x Q_3 + 1{,}5\,\text{IQR} \]

For the data \(21;\ 23;\ 24;\ 25;\ 26;\ 28;\ 45\) with \(Q_1=23\) and \(Q_3=28\):

\(\text{IQR} = 28-23 = 5\)
Lower fence \(= 23 - 1{,}5(5) = 15{,}5\)
Upper fence \(= 28 + 1{,}5(5) = 35{,}5\)
\(45 > 35{,}5\), so 45 is an outlier

\[ 1{,}5 \times \text{IQR} \text{ is the test} \]

Key idea: an outlier is not simply “a big number”. Use the \(1{,}5\times\text{IQR}\) rule and show the fence calculation — that is where the marks are.

4. Ogives

\[ \text{cumulative frequency} = \text{a running total} \]

An ogive is a cumulative frequency curve. Build the table by adding as you go:

IntervalFrequencyCumulative frequency
\(0 \le x < 20\)44
\(20 \le x < 40\)913
\(40 \le x < 60\)1427
\(60 \le x < 80\)835
\(80 \le x \le 100\)540
Total40

\[ \text{plot at the UPPER boundary of each interval} \]

This is the detail most learners lose marks on.

So plot \((20;4)\), \((40;13)\), \((60;27)\), \((80;35)\), \((100;40)\) — starting from \((0;0)\).

\[ Q_1 \text{ at } \tfrac{n}{4} \qquad Q_2 \text{ at } \tfrac{n}{2} \qquad Q_3 \text{ at } \tfrac{3n}{4} \]

With \(n=40\):

\(Q_1\): read across from cumulative frequency \(10\)
\(Q_2\) (median): read across from \(20\)
\(Q_3\): read across from \(30\)
Note the difference: for a LIST you use position \(\tfrac{1}{4}(n+1)\); reading off an OGIVE you use \(\tfrac{n}{4}\). Both are accepted in their own context.

\[ \text{cumulative frequency} \;\rightarrow\; \text{quartiles and percentiles} \]

Key idea: an ogive turns grouped data back into quartiles, percentiles and “how many scored more than…” questions.

5. Scatter Plots

2 4 6 8 10 12 14 20 40 60 80 100 Hours studied Test mark (%)

Bivariate data records TWO variables for each individual. Here, eight learners' hours studied and test mark:

Hours \(x\)2356891112
Mark \(y\)3041455860727585
The independent variable goes on the horizontal axis, the dependent variable on the vertical axis.

\[ \text{direction} \;+\; \text{form} \;+\; \text{strength} \]

A full description needs all three:

FeatureOptions
DirectionPositive (rises) / negative (falls) / none
FormLinear / quadratic / exponential / no pattern
StrengthStrong / moderate / weak — how tightly the points hug the pattern
For our data: a strong positive linear relationship.

\[ \text{always plot the data before calculating anything} \]

Key idea: the scatter plot tells you whether a straight line is even appropriate. A regression line fitted to curved data is meaningless.

6. Correlation

\[ -1 \;\le\; r \;\le\; 1 \]

The correlation coefficient \(r\) puts a number on the direction and strength of a linear relationship.

Value of \(r\)Interpretation
\(r = 1\)Perfect positive linear
\(0{,}9 \le r < 1\)Very strong positive
\(0{,}7 \le r < 0{,}9\)Strong positive
\(0{,}4 \le r < 0{,}7\)Moderate positive
\(0 < r < 0{,}4\)Weak positive
\(r = 0\)No linear relationship
Negative values mirror this exactly: \(r=-0{,}95\) is a very strong NEGATIVE linear relationship.

\[ r \approx 0{,}98 \]

For the hours-versus-mark data, the calculator gives \(r = 0{,}98\).

Interpretation: there is a very strong positive linear relationship between the number of hours studied and the test mark.
Write the interpretation as a SENTENCE naming both variables. “\(r=0{,}98\)” on its own does not earn the interpretation mark.

\[ \text{correlation} \;\ne\; \text{causation} \]

A strong \(r\) does not prove that one variable causes the other.

Ice-cream sales and drowning rates correlate strongly — but ice cream does not cause drowning. Both are driven by hot weather.
The third factor is called a lurking variable. This is a standard exam question, and it expects you to name an alternative explanation.

\[ r \text{ measures LINEAR strength only} \]

Key idea: \(r\) close to 0 means no linear relationship — there could still be a strong curved one. That is why you plot first.

7. Regression

\[ \hat{y} = a + bx \]

The least-squares regression line is the single straight line that makes the total of the squared vertical distances from the points to the line as small as possible.

\(b = \dfrac{\sum (x-\bar{x})(y-\bar{y})}{\sum (x-\bar{x})^{2}}\)
\(a = \bar{y} - b\bar{x}\)
The hat on \(\hat{y}\) matters: it is a PREDICTED value, not an observed one.

\(x\) \(y\) \(x-\bar{x}\) \(y-\bar{y}\) \((x-\bar{x})(y-\bar{y})\) \((x-\bar{x})^2\) 2 30 \(-5\) \(-28{,}25\) 141,25 25 3 41 \(-4\) \(-17{,}25\) 69,00 16 5 45 \(-2\) \(-13{,}25\) 26,50 4 6 58 \(-1\) \(-0{,}25\) 0,25 1 8 60 1 1,75 1,75 1 9 72 2 13,75 27,50 4 11 75 4 16,75 67,00 16 12 85 5 26,75 133,75 25 \(\bar{x}=7\) \(\bar{y}=58{,}25\) \(\sum = 467\) \(\sum = 92\)

Every regression calculation reduces to these two column totals.

\[ b = \frac{467}{92} = 5{,}08 \qquad a = 58{,}25 - 5{,}08(7) = 22{,}72 \]

\(b = \dfrac{\sum(x-\bar{x})(y-\bar{y})}{\sum(x-\bar{x})^{2}} = \dfrac{467}{92} = 5{,}0761\ldots \approx 5{,}08\)
\(a = \bar{y} - b\bar{x} = 58{,}25 - (5{,}0761)(7) = 22{,}7174\ldots \approx 22{,}72\)
\(\hat{y} = 22{,}72 + 5{,}08x\)

2 4 6 8 10 12 14 20 40 60 80 100 Hours studied Test mark (%)

\(a = 22{,}72\) is the y-intercept: the predicted mark for \(0\) hours of study.
\(b = 5{,}08\) is the gradient: each extra hour of study is associated with about \(5\) more percentage points.
The line passes through \((\bar{x};\bar{y}) = (7;\ 58{,}25)\) — a useful check on your work.

\[ \hat{y} = 22{,}72 + 5{,}08x \]

Key idea: \(b\) is the rate of change and \(a\) is the starting value. Interpreting them in context is worth as many marks as calculating them.

8. Predictions

\[ \hat{y} = 22{,}72 + 5{,}08(10) = 73{,}5 \]

To predict, substitute the \(x\)-value into the equation.

10 hours: \(\hat{y} = 22{,}72 + 5{,}08(10) = 73{,}5\%\)
4 hours: \(\hat{y} = 22{,}72 + 5{,}08(4) = 43{,}0\%\)
Round sensibly and state the units. A mark is a percentage, so \(73{,}5\%\) — not \(73{,}52\).

\[ 2 \le x \le 12 \;\Rightarrow\; \text{reliable} \]

TermMeaningReliable?
InterpolationPredicting INSIDE the range of the data (\(2 \le x \le 12\))Yes
ExtrapolationPredicting OUTSIDE the rangeNo — use with great caution
25 hours: \(\hat{y} = 22{,}72 + 5{,}08(25) = 149{,}6\%\)
A mark of \(149{,}6\%\) is impossible. The model is only valid over the range of the data collected — this is the classic exam illustration of why extrapolation fails.

\[ \text{one stray point can tilt the whole line} \]

Because least squares minimises squared distances, a single far-away point has an outsized effect.

Never simply delete an outlier because it is inconvenient. Investigate it, and say what you did.

\[ \text{number} + \text{variable names} + \text{context} \]

Exam answers must be sentences. Compare:

Weak: “\(r = 0{,}98\), strong.”
Full marks: “\(r = 0{,}98\) indicates a very strong positive linear relationship between the number of hours studied and the test mark achieved.”
Full marks: “For every additional hour studied, the predicted mark increases by approximately 5 percentage points.”

\[ \text{predict inside the data, explain in context} \]

Key idea: the calculation earns some marks; the interpretation earns the rest. Always answer in a full sentence naming both variables.

\[ \sigma=\sqrt{\frac{\sum(x-\bar{x})^{2}}{n}} \qquad \hat{y}=a+bx \qquad b=\frac{\sum(x-\bar{x})(y-\bar{y})}{\sum(x-\bar{x})^{2}} \]

You can now:

In the examination, most lost marks are interpretation marks, not calculation marks. Practise writing the sentence every time.