Data Handling: Collecting, Organising & Summarising

Grade 8 Mathematics · Interactive Lesson
← Quizzes
INTRODUCTION
Ready? Try the Quiz →

Full lesson notes

Everything covered in this lesson, in one place - useful for revision or printing.

\[ \text{Question} \rightarrow \text{collect} \rightarrow \text{organise} \rightarrow \text{summarise} \]

Data handling is a cycle. Before we can draw graphs or make conclusions, we must first ask the right questions, collect the data, organise it neatly, and summarise it using statistics. Let us begin with the first four stages of the cycle.

1. Data Cycle & Questions

\[ \text{How do learners travel to school?} \]

A good statistical question is clear and unbiased. It must be answerable with data. Notice how this question simply asks for a fact without suggesting what the "correct" answer should be.

\[ \text{Avoid: "Don't you agree that responsible learners walk?"} \]

Leading questions suggest an answer or use emotional language. The word "responsible" makes learners feel guilty if they do not walk, which biases the data. Always check your questions for hidden bias.

\[ \text{Walk, Taxi, Bus, Car, Train, Other} \]

Providing specific multiple-choice options makes data much easier to organise and summarise later. It ensures everyone answers in the same format, creating clean categorical data.

\[ \text{Clear questions} \rightarrow \text{Unbiased data} \]

The foundation of good data handling is a clear, unbiased question and a well-designed questionnaire. If your starting question is flawed, your final conclusions will be flawed too.

2. Population vs Sample

\[ \text{Population} = \text{The whole group being studied} \]

If we want to know how all Grade 8 learners at a school travel to school, the population is every single Grade 8 learner at that school. It is the entire group we care about.

\[ \text{Sample} = \text{A smaller selected group} \]

Asking every single person in a population takes too much time and money. Instead, we select a smaller group to survey. The data we collect from this sample helps us estimate what the whole population would say.

\[ \text{Representative} \neq \text{Convenience} \]

A sample must represent the population fairly. If you only ask learners who arrive early to school, you miss the latecomers who might rely on different transport. This is called a convenience sample, and it introduces bias.

\[ \text{Population} \rightarrow \text{Sample} \rightarrow \text{Data} \]

Always define your population clearly, and ensure your sample is large enough and diverse enough to represent that population fairly without bias.

3. Categorical Data

\[ \text{Walk, Bus, Taxi, Walk, Car...} \]

Raw data is information in its original, unorganised form. It is very difficult to spot patterns or draw conclusions from a messy list of raw responses. We must organise it.

\[ \text{|||| } \rightarrow 5 \]

Tally marks are a quick way to count frequencies as you read through raw data. We group them in fives (four vertical lines crossed by a diagonal line) to make the final counting faster and more accurate.

\[ \text{Frequency} = \text{Total count per category} \]

The frequency is the number of times a specific category occurs. In a frequency table, we replace the messy tally marks with clean, single numbers representing the totals.

\[ 7 + 5 + 4 + 3 + 1 = 20 \]

A crucial data quality check: the sum of all frequencies must exactly equal the total sample size. If they do not match, you have missed a data point or made a counting error.

\[ \text{Raw} \rightarrow \text{Tally} \rightarrow \text{Frequency} \]

Organising categorical data into frequency tables turns chaos into clarity. Always verify that your frequencies sum to your total sample size before moving on to graphs.

4. Numerical Data

\[ 42, 35, 47 \dots \rightarrow 35, 38, 41 \dots \]

Before summarising numerical data, always order it from smallest to largest. This immediately reveals the extreme values (the minimum and maximum) and makes finding the median much easier.

\[ \text{Stem } 4 \mid \text{Leaf } 2 = 42 \]

A stem-and-leaf display groups data by place value while keeping the original numbers visible. The stem represents the tens digit, and the leaf represents the units digit.

\[ 3 \mid 5 \ 8 \]

The leaves on each stem must be arranged in ascending order. Always include a key (like 3|5 = 35) so the reader understands the place value. This display is excellent for spotting the shape of the data.

\[ 10-19, \ 20-29, \ 30-39 \]

For very large datasets, stem-and-leaf displays become too crowded. Instead, we group data into intervals. Intervals must be non-overlapping, and every value must fit into exactly one interval.

\[ \text{Order} \rightarrow \text{Stem-and-leaf} \rightarrow \text{Intervals} \]

Organising numerical data reveals patterns, clusters, and extremes that are completely hidden in raw lists of numbers.

5. Central Tendency

\[ \bar{x} = \frac{\sum x}{n} = \frac{68}{8} = 8.5 \]

The mean (average) is calculated by adding all the values together and dividing by the number of values. Note that the mean (8.5) does not have to be a number that actually appears in your data set.

\[ \frac{8+8}{2} = 8 \]

The median is the middle value of the ordered data. If there is an even number of values, there is no single middle number. Instead, you calculate the mean of the two middle values.

\[ \text{Mode} = 8 \text{ (most frequent)} \]

The mode is simply the value that appears most often in the data set. A data set can have one mode, multiple modes, or no mode at all if all values appear equally.

\[ \text{Mean, Median, Mode} \]

These three measures of central tendency tell us where the "centre" of the data lies. The mean uses all data, the median resists extremes, and the mode shows the most popular value.

6. Spread & Quality

\[ \text{Range} = \text{Max} - \text{Min} = 12 - 6 = 6 \]

The range measures the spread or dispersion of the data. It is calculated by subtracting the smallest value from the largest value. A larger range means the data is more spread out.

\[ \text{Extremes} = 6 \text{ and } 12 \]

The extremes are simply the minimum and maximum values in your ordered dataset. Identifying them is the first step in calculating the range and checking for outliers.

\[ \sum f = n \]

If your frequency table sums to 23, but your sample size is 25, you have an error. Two responses are missing or recorded incorrectly. Always check your tables before drawing graphs or calculating means.

\[ \text{Central Tendency} + \text{Spread} = \text{Full Picture} \]

Averages alone do not tell the whole story. You must always report a measure of spread (like the range) alongside a measure of central tendency to give a complete summary of the data.

\[ \text{Collect} \rightarrow \text{Organise} \rightarrow \text{Summarise} \]

Excellent work. You have mastered the first half of the data cycle: asking unbiased questions, selecting representative samples, organising data into tables and stem-and-leaf displays, and summarising it using mean, median, mode, and range. You are now ready to test your knowledge.