← Home

Data Handling

Grade 12 Mathematical Literacy
Step 1 of 25
INTRODUCTION
1 / 25
Ready? Try the Quiz →

Full lesson notes

Everything covered in this lesson, in one place - useful for revision or printing.

\( \text{collect} \rightarrow \text{organise} \rightarrow \text{summarise} \rightarrow \text{represent} \rightarrow \text{interpret} \)

Data handling is the full cycle of working with information — from asking a question to explaining what the numbers mean.

In this lesson you will learn to:

1. Collecting Data

\( \text{population} \supset \text{sample} \)

The population is the whole group you want to know about. A sample is the smaller group you actually collect data from.

Question: “How long do Grade 11s at your school spend on homework?”
Population: all Grade 11s at the school  |  Sample: the 40 learners you survey

\( \text{a good sample is random and representative} \)

A sample must represent the whole population fairly.

Surveying only the soccer team about sport time gives biased data — they are not typical of all learners.
Random selection (e.g. every 5th name on the register) reduces bias.

\( \text{survey} \quad \text{questionnaire} \quad \text{observation} \quad \text{records} \)

Data can come from:

Questionnaires — written questions (keep them short and clear)
Interviews — asking people directly
Observation — counting or measuring yourself
Existing records — municipal bills, school marks, weather data

\( \text{clear question} + \text{fair sample} = \text{useful data} \)

Key idea: decide the question first, then choose a sample that fairly represents the population. Biased collection ruins everything that follows.

2. Classifying Data

\( \text{categorical (words)} \quad \text{vs} \quad \text{numerical (numbers)} \)

Categorical — groups or labels: favourite subject, taxi route, yes/no
Numerical — counts or measurements: age, height, electricity used
You can calculate a mean of numerical data, but not of categorical data.

\( \text{discrete: counted} \qquad \text{continuous: measured} \)

Numerical data splits further:

Discrete — counted in whole steps: number of children, cars, goals
Continuous — measured on a scale: height, mass, time, rainfall
Continuous data can take ANY value in a range, like 63.7 mm of rain.

\( \text{words or numbers? counted or measured?} \)

Key idea: classify data BEFORE choosing a graph or calculation — pie charts suit categorical data; histograms suit continuous numerical data.

3. Organising Data

Transport Frequency Walk 12 Taxi 18 Bus 6 Car 4

A frequency table counts how often each value appears. The frequency column must add up to the total number of data values:

\( 12+18+6+4 = 40 \text{ learners} \)

\( 0\text{–}9, \quad 10\text{–}19, \quad 20\text{–}29, \ \ldots \)

Continuous or wide-ranging data is grouped into class intervals.

Test marks out of 100 might use intervals \(0\text{–}29,\ 30\text{–}49,\ 50\text{–}69,\ 70\text{–}100\).
Intervals must not overlap, and every value must fit into exactly one interval.

\( \text{raw list} \rightarrow \text{frequency table} \)

Key idea: a frequency table turns a messy list into something you can graph and summarise. Always check the frequencies total the sample size.

4. Measures of Centre

\( \text{mean} = \dfrac{\text{sum of values}}{\text{number of values}} \)

The mean is the everyday “average”.

Marks: 55, 70, 62, 48, 65
\( \text{mean} = \dfrac{55+70+62+48+65}{5} = \dfrac{300}{5} = 60 \)

\( 48,\ 55,\ \underline{62},\ 65,\ 70 \)

The median is the middle value once the data is in order.

48, 55, 62, 65, 70 → median \(= 62\)
With an EVEN count, average the two middle values: for 48, 55, 62, 70 the median is \( \tfrac{55+62}{2} = 58.5 \)

\( \text{mode} = \text{the value that appears most} \)

Shoe sizes: 5, 6, 6, 7, 6, 8, 5 → mode \(= 6\)
Data can have two modes (bimodal) or none. The mode is the ONLY measure that works for categorical data — e.g. the modal transport is “taxi”.

\( \text{outliers pull the mean, not the median} \)

Salaries: R8 000, R9 000, R9 500, R10 000, R60 000.

Mean \(= \text{R}19\,300\) — pulled up by the one high salary
Median \(= \text{R}9\,500\) — describes the typical worker far better
When data has outliers, prefer the median.

\( \text{mean} \quad \text{median} \quad \text{mode} \)

Key idea: all three describe the “centre” — mean uses every value, median resists outliers, mode shows the most common. Choose the one that fits the data.

5. Measures of Spread

\( \text{range} = \text{maximum} - \text{minimum} \)

The range measures how spread out the data is.

Marks 48 to 70: range \(= 70-48 = 22\)
Two classes can share a mean of 60 while one has range 10 (consistent) and the other 50 (all over the place).

\( Q_1 \quad Q_2(\text{median}) \quad Q_3 \)

Quartiles split ordered data into four equal parts.

Data: 12, 15, 17, 20, 22, 25, 30
\( Q_2 \) (median) \(= 20\)
\( Q_1 \) = median of the lower half \(\{12, 15, 17\} = 15\)
\( Q_3 \) = median of the upper half \(\{22, 25, 30\} = 25\)

\( \text{IQR} = Q_3 - Q_1 \)

The IQR is the spread of the middle half of the data:

\( \text{IQR} = 25 - 15 = 10 \)
Unlike the range, the IQR ignores extreme values — one freak result cannot distort it.

\( \text{range} = \max - \min \qquad \text{IQR} = Q_3 - Q_1 \)

Key idea: centre tells you the typical value; spread tells you how consistent the data is. Report both when comparing two data sets.

6. Reading Graphs

\( \text{bar: compare categories} \qquad \text{pie: parts of a whole} \)

Bar graph — compares categories; bars have gaps
Compound bar — compares two groups side by side (e.g. boys vs girls)
Pie chart — shows each category as a share of \(360^\circ\)
In a pie chart, a category with 25% of the data gets \( 0.25 \times 360^\circ = 90^\circ \).

\( \text{histogram: grouped continuous data} \)

Histogram — like a bar graph but with NO gaps, for grouped continuous data (e.g. rainfall intervals)
Broken-line graph — shows change over time (e.g. monthly electricity usage)

\( \text{check the scale before you trust the picture} \)

Graphs can mislead:

• A vertical axis that starts at 50 instead of 0 makes small differences look huge
• Uneven intervals distort trends
• A pie chart of a tiny sample can exaggerate one category
Always read the axis labels, units and totals first.

\( \text{right graph for the right data} \)

Key idea: categorical data → bar or pie; grouped continuous data → histogram; change over time → line graph. And always inspect the scale.

\[ \text{mean} = \dfrac{\text{sum}}{n} \qquad \text{IQR} = Q_3 - Q_1 \]

You now know how to: