← Home

Data Handling

Grade 9 Mathematics
Step 1 of 25
INTRODUCTION
1 / 25
Ready? Try the Quiz →

Full lesson notes

Everything covered in this lesson, in one place - useful for revision or printing.

\( \text{collect} \rightarrow \text{organise} \rightarrow \text{summarise} \rightarrow \text{represent} \rightarrow \text{interpret} \)

Data Handling is Content Area 4 of the CAPS Senior Phase — roughly 15% of your Grade 9 paper. It is the full cycle: from asking a question to judging whether the answer is fair.

Across this lesson you will learn to:

1. Collecting

\( \text{varies between people} \;\Rightarrow\; \text{statistical question} \)

Statistics begins with a question — but not every question is a statistical one.

A statistical question is answered by collecting data that varies: "What is the most popular sport at our school?"

A fixed-answer question has one correct answer and produces no data set: "How many learners are in Grade 9 at our school?"

\( \text{sample} \subset \text{population} \)

The population is the entire group being studied — say all 480 Grade 9 learners at a school.

A sample is a smaller group chosen from that population. We use a sample when asking everyone is impractical.

The whole art is choosing a sample that behaves like the population.

\( \text{random} \;\;|\;\; \text{systematic} \;\;|\;\; \text{convenience} \)


Convenience sampling is fast but usually biased: the people you can reach easily are rarely typical of the whole population.

\( \dfrac{480}{60} = 8 \)

A teacher wants to know which sport 480 Grade 9 learners would like introduced. Asking all 480 in one lesson is not practical.

Population: all 480 learners.
Sample: 60 learners, chosen systematically.
Interval: \( 480 \div 60 = 8 \) — so take every 8th name on the register.

Then ask ONE closed question so the answers can be tallied: "Which sport would you most like the school to introduce? (netball / chess / athletics / swimming)"

\( \text{interval} = \dfrac{\text{population}}{\text{sample size}} \)

Collecting, in one card. Ask a question whose answers vary. Decide on your population, then choose a sample and say which method you used.

In the exam, "name the sampling method" and "explain why the results may be biased" are two separate marks. Answer both.

2. Organising

\( \text{raw data} \rightarrow \text{tally} \rightarrow \text{frequency} \)

Raw data tells you nothing until it is sorted.

A tally table counts with marks grouped in fives — four uprights, then a diagonal stroke across them. The frequency is the total count for each value.

Always total your frequencies and check the total matches how many items you started with.

\( 3 + 7 + 6 + 3 + 1 = 20 \)

Shoe sizes of 20 learners: 5, 6, 7, 6, 8, 5, 7, 7, 6, 9, 6, 8, 7, 5, 6, 7, 8, 6, 7, 6

Size 5 → 3  ·  Size 6 → 7  ·   Size 7 → 6  ·  Size 8 → 3  ·   Size 9 → 1

The frequencies total 20, which matches the 20 learners — so nothing was missed or double-counted.

\( 0\text{–}9,\; 10\text{–}19,\; 20\text{–}29,\; 30\text{–}39,\; 40\text{–}49 \)

When data has many different values — test marks out of 50 for 24 learners — a row per mark is unreadable. Group it into class intervals of equal width instead.

You trade detail for readability. You can no longer see the individual marks, which is exactly why the mean of grouped data can only be estimated.

\( \text{equal width} \;\;|\;\; \text{no overlap} \;\;|\;\; \text{no gaps} \)

\( \sum \text{frequency} = \text{sample size} \)

Organising, in one card. Tally, total, and check. If your frequencies do not add up to the number of items you started with, something is wrong — fix it before you calculate anything.

Group only when the range is too wide to list, and keep the intervals equal.

3. Summarising

\( \bar{x} = \dfrac{\sum x}{n} \)

The mean adds every value and divides by how many there are.

For 4, 7, 7, 9, 12, 15, 7, 10:
\( \bar{x} = \dfrac{71}{8} = 8{,}875 \)

The mean uses all the data — which is its strength, and also its weakness: one extreme value drags it.

\( \text{median} = \dfrac{7+9}{2} = 8 \)

The median is the middle value once the data is ordered.

Ordered: 4, 7, 7, 7, 9, 10, 12, 15

Eight values means an even count, so there are two middles — average them: \( (7+9) \div 2 = 8 \).

Order the data first, every time. The middle of the list as written is not the median.

\( \text{range} = 15 - 4 = 11 \)

The mode is the value occurring most often — here 7, which appears three times. A set can have no mode, one mode, or several. Two values tied means the set is bimodal; that is not the same as having no mode.

The range is highest minus lowest. Note it measures spread, not centre.

\( \text{estimated mean} = \dfrac{\sum(\text{midpoint} \times \text{frequency})}{\sum \text{frequency}} \)

Grouped data hides the individual values, so the mean can only be estimated. We assume every value in an interval sits at that interval's midpoint.

For 0–9 the midpoint is \( (0+9) \div 2 = 4{,}5 \).

Multiply each midpoint by its frequency, add those products, then divide by the total frequency.

\( \dfrac{410}{20} = 20{,}5 \)

Midpoints 4,5  14,5  24,5  34,5  44,5 with frequencies 3, 7, 6, 3, 1.

Products: 13,5 + 101,5 + 147 + 103,5 + 44,5 = 410
Total frequency: 3 + 7 + 6 + 3 + 1 = 20

Estimated mean = \( 410 \div 20 = 20{,}5 \)

\( \text{mean} \;\;|\;\; \text{median} \;\;|\;\; \text{mode} \;\;|\;\; \text{range} \)

Summarising, in one card. Three measures of centre and one of spread.

Mean — add and divide. Median — order, then take the middle. Mode — most frequent, ties allowed. Range — highest minus lowest.

For grouped data you estimate the mean using midpoints, and you report a modal class.

4. Representing

\( \text{the graph must match the data} \)

Choosing correctly is itself a marked skill.

\( \text{gaps} = \text{categories} \qquad \text{touching} = \text{continuous} \)

This is the most commonly dropped mark in the whole topic.

A bar graph shows separate groups — number of siblings, favourite subject — so the bars stand apart.

A histogram shows a continuous scale cut into intervals — test marks — so the bars touch.

The spacing is the information. Get it wrong and the graph says something false.

\( \text{angle} = \dfrac{\text{frequency}}{\text{total}} \times 360^\circ \)

40 learners named a favourite subject: Maths 12, Science 8, English 10, Life Orientation 6, Other 4.

Maths: \( \dfrac{12}{40} \times 360^\circ = 108^\circ \)
Science: 72°  ·  English: 90°  ·   Life Orientation: 54°  ·  Other: 36°

\( 108 + 72 + 90 + 54 + 36 = 360^\circ \)

Before you pick up the protractor, add your angles. If they do not total exactly 360°, you have made an arithmetic error — and drawing it will only waste the time you have left.

The same habit applies everywhere in this topic: frequencies must total your sample size, and percentages must total 100.

\( \text{time} \rightarrow x\text{-axis} \)

A broken-line graph shows how a value changes over time — daily temperatures across a week. Plot the points and join them with straight segments. Time always goes on the x-axis.

A double bar graph compares two sets across the same categories — boys versus girls per sport. Two bars per category, shaded differently, and a key is compulsory: without it the graph cannot be read.

\( \text{title} + \text{axis labels} + \text{units} + \text{scale from } 0 \)

Representing, in one card. Pick the graph that matches the data, then label it completely.

Every graph needs a title, both axis labels with units, and a scale starting at zero. Missing labels cost marks even when the drawing is accurate.

"Name the most appropriate graph and give a reason" is two marks — always give the reason.

5. Interpreting

\( \text{is this fair?} \)

This is the critical-thinking stage, and where most marks are lost.

\( \dfrac{58 - 52}{52} \times 100 \approx 11{,}5\% \)

A newspaper compares two companies: A sold 52 units, B sold 58. The graph's y-axis starts at 50, not 0.

Visually B's bar looks about three times taller. The real difference is 6 units — about 11,5%.

Because the axis is cut, a small real difference has been drawn to look overwhelming. Name the problem and explain its effect on the reader — that is two marks.

\( \bar{x} = \dfrac{9 \times 12\,000 + 142\,000}{10} = \text{R}25\,000 \)

A company says: "On average, our staff earn R25 000 a month."

But nine of the ten staff earn R12 000, and the owner earns R142 000.

The mean is technically correct and completely misleading — not one employee earns R25 000. The median is R12 000, which is what a typical staff member actually takes home.

\( \text{neutral wording} + \text{balanced options} \)

Leading: "Most learners agree the tuck shop is too expensive — do you agree?" It tells the respondent what to think.

Fair: "Do you think the tuck shop prices are reasonable? (Yes / No / Unsure)"

A fair question uses neutral wording and offers balanced options — including a way to say neither.

\( \text{name the problem} + \text{explain the effect} \)

Interpreting, in one card. Never accept a graph at face value.

Check where the axis starts, how the sample was chosen, whether the question was neutral, and whether an outlier is distorting the average.

When asked to criticise, always do both things: name the problem and say what effect it has on the reader.

\[ \begin{aligned} \bar{x} &= \frac{\sum x}{n} \\[4pt] \text{est. mean} &= \frac{\sum(\text{mid} \times f)}{\sum f} \\[4pt] \text{range} &= \text{max} - \text{min} \\[4pt] \text{angle} &= \frac{f}{\text{total}} \times 360^\circ \end{aligned} \]

The whole topic, end to end.


Now test yourself on the quiz.