← Quizzes

Statistics: Ungrouped Data

Grade 11
Step 1
INTRODUCTION
1 / 30
Ready? Try the Quiz →

Full lesson notes

Everything covered in this lesson, in one place - useful for revision or printing.

\( \bar{x} \quad\; Q_2 \quad\; \sigma \)

Ungrouped data is data that has not been classified or grouped into intervals — every individual value is simply listed. In this chapter we learn how to summarise and describe such a data set.

You will learn to:
• find the mean, median and mode (central tendency)
• measure the spread: range, quartiles and the IQR
• build a five-number summary and a box-and-whisker diagram
• calculate variance and standard deviation
• describe skewness and spot outliers

1. Central Tendency

\[ \bar{x} = \dfrac{\text{sum of all values}}{\text{number of values}} = \dfrac{\sum x}{n} \]

A measure of central tendency is a single value that represents a whole data set. There are three: the mean, the median and the mode.

The mean \( (\bar{x}) \) is the average: add up every value, then divide by how many there are \( (n) \).

\[ \bar{x} = \dfrac{42}{7} = 6 \]

Find the mean of this data set:

4  7  8  5  12  2  4
Add them: \(4+7+8+5+12+2+4 = 42\)
There are \(n = 7\) values, so \(\bar{x} = \dfrac{42}{7} = 6\).

\[ 2\;\;4\;\;4\;\;\underset{\uparrow}{5}\;\;7\;\;8\;\;12 \]

The median is the middle value once the data is placed in order.

Order the same set: 2  4  4  5  7  8  12. With \(n = 7\) (odd), the middle is the 4th value, so the median = 5.

Even \(n\): there are two middle values — average them. E.g. for 8 values, average the 4th and 5th.

\[ \text{mode} = 4 \]

The mode is the value that appears most often (highest frequency).

In 2 4 4 5 7 8 12, the value 4 appears twice and everything else once — so the mode is 4.

A set can have no mode (all values different) or more than one mode.

\[ \bar{x}=6 \qquad \text{median}=5 \qquad \text{mode}=4 \]

For the set 4 7 8 5 12 2 4 the three measures are mean 6, median 5, mode 4.

Each summarises the "centre" of the data in a different way. Which one is most reliable depends on the shape of the data — you will see why in the last topic.

2. Range & Quartiles

\[ \text{Range} = \text{max} - \text{min} \]

Measures of dispersion tell us how spread out the data is. A measure of central tendency represents the data better when the spread is small.

The simplest measure is the range — the largest value minus the smallest.

\[ Q_1 \;\mid\; Q_2 \;\mid\; Q_3 \]

Quartiles divide an ordered data set into four equal parts.

• \(Q_1\) (lower quartile) — a quarter of the data lies below it
• \(Q_2\) (the median) — the halfway mark
• \(Q_3\) (upper quartile) — three quarters of the data lies below it

\[ Q_1 = \dfrac{48+50}{2}=49 \qquad Q_3=\dfrac{74+76}{2}=75 \]

Order the data, find the median \((Q_2)\), then find the median of each half.

18 40 48 50 53 63  64  65 68 74 76 81 82
Middle value \(Q_2 = 64\). Lower half is 18 40 48 50 53 63, whose median is \(Q_1=\tfrac{48+50}{2}=49\). Upper half is 65 68 74 76 81 82, whose median is \(Q_3=\tfrac{74+76}{2}=75\).

\[ \text{IQR}=Q_3-Q_1 \qquad \text{Semi-IQR}=\tfrac{1}{2}(Q_3-Q_1) \]

The interquartile range (IQR) is the spread of the middle 50% of the data. It ignores the extreme values, so it is not affected by outliers.

Using \(Q_1=49\) and \(Q_3=75\):   \(\text{IQR}=75-49=26\).
The semi-interquartile range is just half of that: \(\tfrac{1}{2}(26)=13\).

\[ \text{Range}=\text{max}-\text{min} \qquad \text{IQR}=Q_3-Q_1 \]

Range uses only the two extremes; the IQR uses the quartiles to measure the spread of the central half of the data. Both are found from just two values, which makes them quick but limited — they ignore everything in between.

3. Five-Number Summary

\[ \text{Min} \;<\; Q_1 \;<\; Q_2 \;<\; Q_3 \;<\; \text{Max} \]

The five-number summary captures the whole shape of a data set with just five values:

minimum, lower quartile \(Q_1\), median \(Q_2\), upper quartile \(Q_3\), and maximum.

These five numbers are exactly what we need to draw a box-and-whisker diagram.

\[ 18\;\;\;49\;\;\;64\;\;\;75\;\;\;82 \]

Speeds (km/h) of motorists past a school:

48 65 82 68 74 53 18 63 64 76 81 50 40
Ordered: 18 40 48 50 53 63 64 65 68 74 76 81 82.
Min \(=18\),   \(Q_1=49\),   \(Q_2=64\),   \(Q_3=75\),   Max \(=82\).
Notice the max (82) is much closer to \(Q_3\) (75) than the min (18) is to \(Q_1\) (49) — a clue about the shape.

\[ [\;\text{Min},\; Q_1,\; Q_2,\; Q_3,\; \text{Max}\;] \]

The five-number summary is the bridge between raw data and a visual. Once you have these five values in order, you can immediately read off the range (Max − Min), the IQR \((Q_3-Q_1)\), and draw the diagram in the next topic.

4. Box & Whisker

A box-and-whisker diagram is a picture of the five-number summary drawn against a number line.

• the box stretches from \(Q_1\) to \(Q_3\) (the middle 50%)
• the line inside the box is the median \(Q_2\)
• the two whiskers reach out to the min and the max

Each section of the diagram contains about 25% of the data:

left whisker 18–49  |  left of box 49–64  |  right of box 64–75  |  right whisker 75–82
A wide section means the data there is spread out; a narrow section means it is bunched together.

The position of the median inside the box, and the whisker lengths, reveal the skewness.

Here the median (64) sits to the right of centre and the left whisker is much longer. The data is bunched at the high end with a tail to the low end — it is negatively skewed (skewed to the left).

A box-and-whisker diagram turns five numbers into an instant picture of centre, spread and shape. Longer whisker or box on one side tells you which way the data is skewed — no calculation needed.

5. Variance & Std Dev

\[ \sigma \;=\; \text{typical distance from the mean} \]

Range and IQR use only two values. Standard deviation \((\sigma)\) uses every value — it measures how far a typical data point sits from the mean.

A small \(\sigma\) means the data huddles close to the mean; a large \(\sigma\) means it is widely spread. An observation far from the mean (many standard deviations away) is unusual.

\[ \sigma^2 = \dfrac{\sum (x-\bar{x})^2}{n} \qquad \sigma = \sqrt{\sigma^2} \]

The variance \((\sigma^2)\) is the average of the squared distances from the mean. The standard deviation \((\sigma)\) is the square root of the variance, bringing us back to the original units.

In practice you calculate \(\sigma\) directly on your calculator in STAT / SD mode — but you must understand what it measures.

\[ \sigma^2=\dfrac{66}{7}\approx 9{,}43 \qquad \sigma=\sqrt{9{,}43}\approx 3{,}07 \]

Data: 4 7 8 5 12 2 4, with mean \(\bar{x}=6\). Square each distance from the mean and add:

(4−6)²=4   (7−6)²=1   (8−6)²=4   (5−6)²=1
(12−6)²=36   (2−6)²=16   (4−6)²=4
Sum \(=66\), so \(\sigma^2=\tfrac{66}{7}\approx 9{,}43\) and \(\sigma=\sqrt{9{,}43}\approx 3{,}07\).

\[ [\,\bar{x}-\sigma\,;\,\bar{x}+\sigma\,] \]

We often ask how many values lie within one (or two) standard deviations of the mean — the interval \([\bar{x}-\sigma\,;\,\bar{x}+\sigma]\).

If \(\bar{x}=10\) and \(\sigma=2{,}5\), then within one standard deviation is \([7{,}5\,;\,12{,}5]\), and within two is \([5\,;\,15]\).
This tells us how typical or extreme a value is.

\[ \sigma^2 = \dfrac{\sum (x-\bar{x})^2}{n} \qquad \sigma = \sqrt{\sigma^2} \]

Variance is the mean squared deviation; standard deviation is its square root and shares the data's units. Because they use every value, they describe spread more fully than the range or IQR — and they are the foundation for more advanced statistics.

6. Skewness & Outliers

\[ \begin{aligned} \text{Symmetrical:}\;\; & \bar{x} = \text{median} \\ \text{Positively skewed (right):}\;\; & \bar{x} > \text{median} \\ \text{Negatively skewed (left):}\;\; & \bar{x} < \text{median} \end{aligned} \]

The shape of the data falls into three patterns:

Symmetrical — evenly spread, \(\bar{x}=\text{median}\)
Positively skewed (tail to the right) — \(\bar{x}>\text{median}\)
Negatively skewed (tail to the left) — \(\bar{x}<\text{median}\)

\[ \text{skew pulls } \bar{x} \text{ toward the tail} \]

The mean is dragged in the direction of the tail, so skewness makes the mean less reliable.

If the data is positively skewed, the mean is too high; if negatively skewed, it is too low. The median stays reliable whatever the skew, because it depends only on position.

Test marks: 40 41 41 42 44 48 54 62 72 90 give mean 53,4 but median 46 — positively skewed, so the median represents the class better.

\[ [\,Q_1-1{,}5\times\text{IQR}\;;\; Q_3+1{,}5\times\text{IQR}\,] \]

An outlier is a value that lies exceptionally far from the rest of the data (often from a recording or measurement error). A value is an outlier if it falls outside the interval above.

• \(Q_1-1{,}5\times\text{IQR}\) is the lower fence
• \(Q_3+1{,}5\times\text{IQR}\) is the upper fence

\[ [\,-10{,}5\;;\;25{,}5\,] \;\Rightarrow\; 30 \text{ is an outlier} \]

Data: 2 3 3 5 8 9 9 11 12 14 30, with \(Q_1=3\) and \(Q_3=12\).

\(\text{IQR}=12-3=9\)
Fences: \([\,3-1{,}5(9)\;;\;12+1{,}5(9)\,]=[-10{,}5\;;\;25{,}5]\)
The value 30 lies outside this interval, so it is an outlier.

\[ \text{median: robust} \;\;\;\; \bar{x}\text{: sensitive} \]

Putting it together:

• The median is usually best — it is not affected by skewness or outliers.
• The mean is only reliable when the data is roughly symmetrical with no outliers.
• The mode only matters when it has a genuinely high frequency.

\[ \bar{x} \text{median} \]

Skewness compares the mean and median; outliers are found with the 1,5 × IQR fences. Both warn you when the mean is being distorted — and both point to the median as the safer summary of the centre.

\[ \begin{aligned} \bar{x} &= \dfrac{\sum x}{n} & \text{IQR} &= Q_3-Q_1 \\[4pt] \sigma^2 &= \dfrac{\sum (x-\bar{x})^2}{n} & \sigma &= \sqrt{\sigma^2} \end{aligned} \]

You now know how to describe any ungrouped data set:

Centre: mean, median, mode
Spread: range, quartiles, IQR, semi-IQR, standard deviation
Five-number summarybox-and-whisker diagram
Shape: symmetrical vs positively/negatively skewed
Outliers via the \(1{,}5\times\text{IQR}\) fences, and why the median is the most reliable centre.

Ready to test yourself?