Cambridge Lower Secondary CheckpointStage 8

Statistics and Probability (8Ss)

Mathematics Stage 8 Chapter Notes

What this chapter covers

Statistics and Probability - Statistics
ShareWhatsAppPost
Statistics and Probability (8Ss) notes

Unable to load PDF

The notes viewer could not load. Please refresh the page.

Read online free. Download a watermarked copy with a free account.

Read the notes

The full Statistics and Probability (8Ss) notes as text: skim, search, and jump between subtopics.

~18 min read

1. Frequency Tables and Bar Charts

To make sense of raw data, we first need to organise it. A frequency table is a simple way to do this. We use a tally column to count data points, then write the total count in the frequency column. For data with two categories, like 'boy/girl' and 'likes/dislikes sport', we use a two-way table. Once data is organised, we can draw a bar chart to represent it visually. Bar charts use bars of equal width to show the frequency of each category. The height of the bar represents the frequency. It's crucial to have gaps between the bars for categorical data and to label both axes clearly.

Key term

Frequency: The number of times a particular data value or category occurs in a set of data.

Examiner insight

Examiners award marks for clearly constructed tables and charts. Always double-check that your totals in a two-way table add up correctly both horizontally and vertically.

Common pitfall

When drawing bar charts, students often forget to label the axes or make the bars different widths, which is incorrect for standard bar charts.

Worked example 14 marks

The colours of 20 cars in a car park are recorded: Red, Blue, Silver, Red, Black, Silver, Silver, White, Blue, Red, Black, Silver, Red, Blue, White, Silver, Silver, Red, Black, Silver.a) Create a frequency table for this data.b) What is the modal colour?

  1. 1

    a) First, set up a table with columns for Colour, Tally, and Frequency.

  2. 2

    Go through the list and add a tally mark for each colour.

  3. 3

    Colour | Tally | Frequency

  4. 4

    Red | |||| | 5

  5. 5

    Blue | ||| | 3

  6. 6

    Silver | |||| || | 7

  7. 7

    Black | ||| | 3

  8. 8

    White | || | 2

  9. 9

    Total | | 20

  10. 10

    b) The modal colour is the one with the highest frequency. Looking at the table, the highest frequency is 7, which corresponds to Silver. So, the modal colour is Silver.

Worked example 23 marks

A school surveyed 80 students about whether they walk to school. The results are partly shown in the two-way table. Complete the table.

  1. 1

    Start with the given table:

    WalkDon't WalkTotal
    Boy2145
    Girl17
    Total80
  2. 2

    Calculate the number of boys who don't walk: Total Boys - Boys who walk = 45 - 21 = 24.

  3. 3

    Calculate the total number of students who walk: Total Students - Total who don't walk. We need another value first.

  4. 4

    Calculate the total number of girls: Total Students - Total Boys = 80 - 45 = 35.

  5. 5

    Calculate the number of girls who walk: Total Girls - Girls who don't walk = 35 - 17 = 18.

  6. 6

    Calculate the total who walk: Boys who walk + Girls who walk = 21 + 18 = 39.

  7. 7

    Calculate the total who don't walk: Boys who don't walk + Girls who don't walk = 24 + 17 = 41. (Check: 39 + 41 = 80. Correct).

  8. 8

    Final Table:

    WalkDon't WalkTotal
    Boy212445
    Girl181735
    Total394180

Recap

  • Use a tally chart to count raw data before finding the total frequency.
  • A two-way table organises data across two different categories.
  • Bar charts must have labelled axes, bars of equal width, and gaps between bars.
  • The height of each bar shows the frequency for that category.
  • The mode is the category with the highest frequency.

Quick check

  1. A bar chart shows the frequency of pets. The bar for 'Dog' has a height of 12. What does this mean?1 mark
  2. In a two-way table, a row total is 50 and a column total is 30. The grand total is 100. Is this possible?1 mark

2. Stem-and-Leaf Diagrams

A stem-and-leaf diagram is a clever way to show numerical data. It groups the data by place value (the 'stem') and shows each individual data point (the 'leaf'). The main advantage is that it displays the shape of the distribution without losing the original data values. To create one, you split each number into a 'stem' (the first digit or digits) and a 'leaf' (the last digit). Always remember to include a key to explain how to read your diagram. For comparing two sets of data, a back-to-back stem-and-leaf diagram is extremely useful.

Key term

Key: An essential part of a stem-and-leaf diagram that explains what the stem and leaf values represent.

Examiner insight

Examiners specifically look for a key with every stem-and-leaf diagram. Without a key, the diagram is meaningless and will score zero marks.

Fun fact

Stem-and-leaf diagrams were invented by the influential statistician John Tukey in the 1970s as a quick tool for data analysis before computers became widespread.

Worked example 14 marks

The ages of 15 people at a party are: 21, 34, 18, 25, 27, 34, 41, 19, 25, 21, 30, 18, 23, 29, 44.a) Draw an ordered stem-and-leaf diagram.b) Find the median age.

  1. 1

    a) First, identify the stems. The data ranges from 18 to 44, so the stems will be 1, 2, 3, and 4.

  2. 2

    Draw the unordered diagram first by going through the list:

    1 | 8 9 8 2 | 1 5 7 5 1 3 9 3 | 4 4 0 4 | 1 4

  3. 3

    Now, create the ordered diagram by putting the leaves for each stem in numerical order. Don't forget the key.

  4. 4

    Ordered Diagram:

    1 | 8 8 9 2 | 1 1 3 5 5 7 9 3 | 0 4 4 4 | 1 4 Key: 1 | 8 means 18

  5. 5

    b) To find the median, find the middle value. There are 15 people. The middle position is (15 + 1) / 2 = 8th value.

  6. 6

    Count through the ordered leaves to the 8th value. The 8th value is in the '2' stem row. It is 25.

  7. 7

    The median age is 25.

Worked example 24 marks

The stem-and-leaf diagram shows the heights (in cm) of boys and girls in a class. Compare the heights of the boys and girls.

Girls | | Boys 9, 8 | 14 | 7, 5, 2 | 15 | 1, 4 6, 1 | 16 | 2, 5, 8 | 17 | 0, 3 Key: For boys, 15 | 1 means 151 cm. For girls, 8 | 14 means 148 cm.

  1. 1

    First, calculate a measure of average for both groups. The median is best here.

  2. 2

    Girls' heights: 148, 149, 152, 155, 157, 161, 166. There are 7 girls. Median is the (7+1)/2 = 4th value, which is 155 cm.

  3. 3

    Boys' heights: 151, 154, 162, 165, 168, 170, 173. There are 7 boys. Median is the 4th value, which is 165 cm.

  4. 4

    Next, calculate a measure of spread for both groups. The range is suitable.

  5. 5

    Girls' range: 166 - 148 = 18 cm.

  6. 6

    Boys' range: 173 - 151 = 22 cm.

  7. 7

    Now, write two comparative sentences using the context. One for the average, one for the spread.

  8. 8

    Comparison 1 (Average): On average, the boys are taller than the girls (median of 165 cm vs 155 cm).

  9. 9

    Comparison 2 (Spread): The heights of the girls are more consistent than the boys' heights because their range is smaller (18 cm vs 22 cm).

Recap

  • A stem-and-leaf diagram shows the shape of the data while keeping all original values.
  • Always include a key to explain how to read the diagram.
  • An ordered diagram is needed to easily find the median and range.
  • Back-to-back diagrams are used to compare two datasets.
  • The median is the middle value in an ordered set of data.

Quick check

  1. In a stem-and-leaf diagram, the row '5 | 0 1 1 7' appears. What are the four data values on this row?1 mark
  2. Why is an ordered stem-and-leaf diagram more useful than an unordered one?1 mark

3. Averages and Spread

To summarise a dataset, we use measures of 'average' or 'central tendency'. The three main types are the Mean (the sum of all values divided by the number of values), the Median (the middle value when the data is in order), and the Mode (the most frequent value). To describe how spread out the data is, we use a 'measure of spread'. The simplest is the Range, which is the difference between the highest and lowest values. A small range means the data is consistent, while a large range means it is varied.

Mean = (Sum of all values) / (Number of values)

Mean from frequency table, Mean = Σfx / Σf

Range = Highest value – Lowest value

Key term

Outlier: A data point that is significantly different from the other data points in a set.

Examiner insight

When asked to compare distributions, a two-part answer is required for full marks. You must make one statement comparing the averages (e.g., 'Group A has a higher median score') and a second statement comparing the spread (e.g., 'Group B's scores are more consistent as the range is smaller'), both linked to the context of the data.

Common pitfall

The most common mistake is forgetting to order the data before finding the median. For an even number of data points, students often forget to find the mean of the two middle values.

Worked example 14 marks

Find the mean, median, mode, and range of this set of numbers: 5, 8, 2, 5, 9, 1, 6.

  1. 1

    Order the data first to find the median and range: 1, 2, 5, 5, 6, 8, 9.

  2. 2

    Mean: (1 + 2 + 5 + 5 + 6 + 8 + 9) / 7 = 36 / 7 ≈ 5.14 (to 2 d.p.).

  3. 3

    Median: The middle value in the ordered list. There are 7 values, so the median is the (7+1)/2 = 4th value, which is 5.

  4. 4

    Mode: The most frequent value. The number 5 appears twice, more than any other number. So the mode is 5.

  5. 5

    Range: Highest value - Lowest value = 9 - 1 = 8.

Worked example 23 marks

A survey asks 20 people how many pets they have. The results are in the frequency table. Calculate the mean number of pets.

Pets (x)Frequency (f)
04
19
25
32
  1. 1

    To find the mean from a frequency table, we need a new column, 'fx', which is the value multiplied by its frequency.

  2. 2
    Pets (x)Frequency (f)fx
    040 × 4 = 0
    191 × 9 = 9
    252 × 5 = 10
    323 × 2 = 6
  3. 3

    Now find the totals (Σ) of the 'f' and 'fx' columns.

  4. 4

    Σf = 4 + 9 + 5 + 2 = 20 (This is the total number of people, as given in the question).

  5. 5

    Σfx = 0 + 9 + 10 + 6 = 25 (This is the total number of pets).

  6. 6

    Use the formula: Mean = Σfx / Σf.

  7. 7

    Mean = 25 / 20 = 1.25. The mean number of pets is 1.25.

Recap

  • The mean is the sum of values divided by the count of values.
  • The median is the middle value of an ordered dataset.
  • The mode is the most frequently occurring value.
  • The range measures the spread of data (Highest - Lowest).
  • To compare two datasets, you must compare both an average and the spread.

Quick check

  1. What is the median of the numbers 10, 4, 8, 2?2 marks
  2. Which measure of average is most affected by an extreme outlier?1 mark

4. Basic Probability and Complementary Events

Probability measures the likelihood of an event happening, on a scale from 0 (impossible) to 1 (certain). It can be written as a fraction, decimal, or percentage. A key concept is that of complementary events. If you have an event 'A', its complement, written as A', is the event 'not A'. Because it either rains or it doesn't, the event 'rain' and the event 'no rain' are complementary. The probabilities of complementary events always add up to 1.

P(Event) = (Number of favourable outcomes) / (Total number of possible outcomes)

P(A') = 1 - P(A)

Key term

Complementary Events: Two events are complementary if they are the only two possible outcomes of an experiment, and their probabilities sum to 1.

Common pitfall

Mixing up fractions, decimals, and percentages. If a question gives a probability as a fraction, it's usually best to give your answer as a fraction too unless instructed otherwise.

Worked example 11 mark

The probability that a train is on time is 0.85. What is the probability that the train is not on time?

  1. 1

    The events 'on time' and 'not on time' are complementary events.

  2. 2

    Their probabilities must add up to 1.

  3. 3

    Let P(on time) = 0.85.

  4. 4

    P(not on time) = 1 - P(on time)

  5. 5

    P(not on time) = 1 - 0.85 = 0.15.

Worked example 23 marks

A bag contains 5 red, 3 blue, and 2 green counters. A counter is picked at random.a) What is the probability of picking a blue counter?b) What is the probability of not picking a blue counter?

  1. 1

    a) First, find the total number of counters: 5 + 3 + 2 = 10 counters.

  2. 2

    The number of favourable outcomes (picking blue) is 3.

  3. 3

    P(blue) = (Number of blue counters) / (Total number of counters) = 3/10.

  4. 4

    b) The event 'not picking blue' is the complement of 'picking blue'.

  5. 5

    We can use the formula: P(not blue) = 1 - P(blue).

  6. 6

    P(not blue) = 1 - 3/10 = 7/10.

  7. 7

    Alternatively, 'not blue' means 'red or green'. There are 5 + 2 = 7 such counters. So, P(not blue) = 7/10.

Recap

  • Probability is a measure of chance between 0 (impossible) and 1 (certain).
  • The sum of probabilities of all possible outcomes is 1.
  • Complementary events are opposites, like 'win' and 'not win'.
  • The probability of an event not happening is 1 minus the probability that it does happen.

Quick check

  1. The probability of winning a game is 2/7. What is the probability of not winning?1 mark
  2. A spinner has sections labelled A, B, C, D. P(A)=0.1, P(B)=0.4, P(C)=0.2. What is P(D)?2 marks

5. Experimental Probability

Sometimes we can't calculate probability based on theory. For example, what is the probability a piece of toast lands butter-side down? We can only find out by doing an experiment. This is called experimental probability, or relative frequency. You perform an experiment or simulation many times (called trials) and count how many times the desired event occurs. The more trials you do, the more reliable your experimental probability becomes as an estimate for the true, theoretical probability.

Experimental Probability = (Number of times the event occurred) / (Total number of trials)

Key term

Trial: A single performance of a probability experiment, such as flipping a coin once or rolling a die once.

Common pitfall

Stating that an experimental probability is an exact value. It's always an estimate, so use words like 'estimate' or 'about' when making predictions.

Fun fact

Insurance companies are built on experimental probability. They use vast amounts of data (trials) on accidents, illness, and life expectancy to calculate the probability of events and set their premiums.

Worked example 14 marks

A biased spinner is spun 200 times. It lands on red 68 times.a) Calculate the experimental probability of the spinner landing on red.b) If the spinner is spun 500 times, estimate the number of times it will land on red.

  1. 1

    a) Use the formula for experimental probability.

  2. 2

    Number of times event occurred = 68. Total number of trials = 200.

  3. 3

    Experimental P(Red) = 68 / 200.

  4. 4

    Simplify the fraction: 68/200 = 34/100 = 17/50. It can also be written as 0.34.

  5. 5

    b) To estimate the number of red outcomes in 500 spins, multiply the number of spins by the probability of landing on red.

  6. 6

    Estimated number of reds = P(Red) × Number of spins

  7. 7

    Estimated number of reds = 0.34 × 500 = 170.

  8. 8

    So, we would expect it to land on red about 170 times.

Recap

  • Experimental probability is an estimate based on the results of an experiment.
  • It is calculated by dividing the number of successful outcomes by the total number of trials.
  • Theoretical probability is what should happen in theory (e.g., P(Heads) = 1/2).
  • Experimental probability gets closer to theoretical probability as the number of trials increases.
  • You can use experimental probability to predict future outcomes.

Quick check

  1. A drawing pin is dropped 100 times. It lands 'point up' 35 times. What is the experimental probability of it landing 'point up'?1 mark
  2. Why is an experimental probability based on 1000 trials more reliable than one based on 10 trials?1 mark

6. Combined Events and Sample Spaces

Often, we are interested in the probability of two or more events happening, called combined events. Examples include rolling two dice or flipping three coins. To handle these, we need a systematic way to list all possible outcomes. This complete set of outcomes is called the sample space. You can represent a sample space using an organised list, a two-way table (often called a sample space diagram), or a tree diagram. Once you have the sample space, you can find the probability of any event by counting the number of favourable outcomes and dividing by the total number of outcomes.

Key term

Sample Space: The set of all possible outcomes of a probability experiment.

Examiner insight

For combined events, examiners expect to see a systematic method. A clearly drawn sample space or tree diagram will often earn method marks, even if your final calculation is wrong. Don't do it in your head.

Fun fact

The study of probability was kick-started in the 17th century by mathematicians Blaise Pascal and Pierre de Fermat, who were trying to solve a problem about a gambling game.

Worked example 14 marks

Two fair six-sided dice are rolled, and their scores are added together.a) Draw a sample space diagram to show all possible outcomes.b) Find the probability that the total score is 7.

  1. 1

    a) A sample space diagram is a table showing the outcomes of the first die along one axis and the second die along the other. The cells show the sum of the scores.

  2. 2
    Die 1 / Die 2123456
    1234567
    2345678
    3456789
    45678910
    567891011
    6789101112
  3. 3

    b) First, find the total number of possible outcomes. The table is 6x6, so there are 36 outcomes in total.

  4. 4

    Next, count the number of favourable outcomes (where the total is 7). Looking at the diagonal in the table, we can see there are 6 ways to get a total of 7: (1,6), (2,5), (3,4), (4,3), (5,2), (6,1).

  5. 5

    P(Total = 7) = (Number of ways to get 7) / (Total number of outcomes) = 6/36.

  6. 6

    Always simplify the fraction if possible: P(Total = 7) = 1/6.

Worked example 23 marks

A bag contains 3 red balls and 2 blue balls. A ball is picked, its colour noted, and it is put back (replaced). A second ball is then picked. Draw a tree diagram and find the probability of picking two red balls.

  1. 1

    Start the tree diagram with the first pick. The two branches are Red (R) and Blue (B).

  2. 2

    P(R) = 3/5 and P(B) = 2/5. Write these probabilities on the branches.

  3. 3

    From the end of each first branch, draw two more branches for the second pick. Since the ball is replaced, the probabilities remain the same.

  4. 4

    The complete diagram will have four final outcomes: RR, RB, BR, BB.

  5. 5

    To find the probability of an outcome, multiply the probabilities along the branches leading to it.

  6. 6

    We want the probability of picking two red balls, which is the 'RR' path.

  7. 7

    P(R and R) = P(R on first pick) × P(R on second pick)

  8. 8

    P(RR) = (3/5) × (3/5) = 9/25.

Recap

  • A sample space is the list of all possible outcomes.
  • Sample space diagrams (tables) are great for showing the outcomes of two combined events.
  • Tree diagrams are useful for sequences of events, showing the probabilities on each branch.
  • To find the probability of a sequence of events on a tree diagram, multiply the probabilities along the branches.
  • The total number of outcomes is the denominator of your probability fraction.

Quick check

  1. You flip a coin twice. List the full sample space.1 mark
  2. On a tree diagram, the probability of event A is 0.5 and event B is 0.2. What is P(A and then B)?1 mark

End-of-chapter exercise

Test yourself on the whole chapter. Work through these before moving on.

  1. The ages of 11 members of a chess club are: 24, 15, 39, 42, 15, 28, 50, 33, 29, 15, 21. Find the mode, median, and range of their ages.4 marks
  2. The probability of a biased coin landing on heads is 0.44. The coin is flipped 350 times. Estimate the number of times it will land on tails.3 marks
  3. The table shows the number of goals scored by a football team in 40 matches. Calculate the mean number of goals scored per match. Goals | 0 | 1 | 2 | 3 | 4 ---|---|---|---|---|--- Frequency | 8 | 15 | 11 | 5 | 13 marks
  4. A spinner can land on Red, Blue, or Green. The probability it lands on Red is 1/4 and the probability it lands on Blue is 1/2. What is the probability it lands on Green?2 marks
  5. A survey of 100 people asked if they preferred tea or coffee. 40 of the 60 men preferred tea. 25 women preferred coffee. Construct and complete a two-way table to show this information.4 marks
  6. The heights of two groups of plants were measured after a month. Group A: Mean height = 25 cm, Range = 8 cm. Group B: Mean height = 28 cm, Range = 5 cm. Make two comparisons between the heights of the plants in Group A and Group B.2 marks
  7. A restaurant menu has 3 starters and 4 main courses. A customer chooses one starter and one main course. Draw a sample space diagram to find the total number of different two-course meal combinations.3 marks
  8. A bag contains 5 red counters and 3 blue counters. A counter is taken, its colour is noted, and it is NOT replaced. A second counter is then taken. Draw a tree diagram to represent this and find the probability that both counters are the same colour.5 marks
  9. The times, in minutes, taken by 15 runners to complete a race are shown. Draw an ordered stem-and-leaf diagram for this data. Remember to include a key. 23, 41, 35, 28, 35, 42, 29, 30, 36, 23, 35, 49, 28, 31, 263 marks
  10. A survey of 180 students' favourite subjects showed 60 chose Maths, 45 chose Science, 30 chose English, and 45 chose other subjects. Calculate the angle you would use for each subject on a pie chart.4 marks

Go deeper

Practise and revise with member-only material for this chapter.

Free notes are just the start.

Unlock every Workbook and Chapter at a Glance, and generate your own worksheets and predicted papers.

Explore plans

Related chapters