← Mathematics

Probability and Statistics

Syllabus topics 8 and 9 · the two topics where careful reading earns most of the marks

What the check found — and it is good news

On the Foundations Check, statistics, probability and data came back secure to step 6 and broke at step 7. Step 7 is the hardest Extended rung. Nothing below it broke. Read that plainly: this is one of your stronger areas, and it is one of only two strands where you got all the way to the top rung before anything went wrong.

So this guide is not a rebuild. It is a finish. The basics get a quick, honest pass — you can already do them, and skipping them entirely would be a mistake since they are worth easy marks — but most of the page is aimed at the three places step 7 actually lives:

One more thing from the check lands here. Sign errors recurred across several strands. In statistics they have exactly two homes: range = largest − smallest and IQR = Q3 − Q1. Both are subtractions where the order is fixed and a reversed order gives a negative answer that is always wrong. A spread can never be negative. That is your check.

What this guide covers

CodeSub-topicExpected difficulty for you
E8.1Introduction to probabilityMedium — notation, not ideas
E8.2Relative and expected frequenciesMedium
E8.3Probability of combined eventsHigh — with vs without replacement
E8.4Conditional probabilityHigh — the step-7 gap
E9.1Classifying statistical dataLow
E9.2Interpreting statistical dataLow — but two comparisons, not one
E9.3Averages and measures of spreadHigh — estimated mean by hand
E9.4Statistical charts and diagramsMedium
E9.5Scatter diagramsLow
E9.6Cumulative frequency diagramsLow, once the plotting point is right
E9.7HistogramsMedium — frequency density is the whole topic

Paper 2 is non-calculator: 2 hours, 100 marks, 50% of your grade. Every number on this page is chosen so it can be done by hand, and every method shown is a by-hand method.

The thing that is already working

Number sense and the four operations: solid on all seven rungs. Not one gap, right up to the Extended-hard level. That matters more in this topic than in any other, because probability and statistics are almost entirely arithmetic wearing a hat. Fractions multiplied along tree branches, fractions added between them, a column of f × x products summed, a division at the end — that is the whole of it, and you can already do all of those.

So when a statistics question goes wrong for you, it will not be the arithmetic. It will be one of two things: reading the question (with or without replacement, given or not given, frequency or frequency density), or the plotting rule (upper class boundary, midpoint, class width). Both are things you can decide before you calculate anything. That is why this guide spends so much time on the sentence before the sum.

How this guide works

Every idea appears four times, with less help each time.

  1. Worked in full — every step, with a reason on each line.
  2. Last step is yours — work out the last line before you press the button.
  3. Last two are yours — the same, harder.
  4. All yours — type an answer and check it.

Section 12 is a mixed set: fifteen questions, unlabelled and out of order, because in an exam nobody tells you whether you are looking at a histogram question or a conditional probability question.

E8.1 · medium1 · Introduction to probability — the scale, the notation, and the total of 1▼
▶  Watch: E8.1 Introduction to probability
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

Probability is a number, and the number always lives between 0 and 1. Never 1 in 4, never 25 out of 100 written as a ratio, never a percentage unless the question asks for one — a fraction, a decimal or a percentage, but a single number on this scale.

0 impossible 0.25 unlikely 0.5 even chance 0.75 likely 1 certain a fair coin lands heads no probability lives left of here or right of here

The notation Cambridge uses

WrittenMeansRule
P(A)the probability that A happens0 ≤ P(A) ≤ 1
P(A′) or P(not A)the probability A does not happenP(A′) = 1 − P(A)
a full outcome listevery possible result, once eachall the probabilities add to 1

For one event with equally likely outcomes:

P(event) = number of outcomes you want ÷ total number of outcomes

The word equally likely is doing real work there. If the outcomes are not equally likely — a biased spinner, a weighted dice — you cannot count, you must be given or must estimate the probabilities. That is section 2.

The habit that saves marks: write the total down first, before you count what you want. Most wrong single-event probabilities are wrong in the denominator, not the numerator — someone counted the reds and forgot that the reds are part of the total too.
Worked in full
A bag holds 5 red counters, 3 blue counters and 4 green counters. One counter is taken at random. Work out P(blue).
1
Total = 5 + 3 + 4 = 12
Count everything first. This is the denominator and it is where the marks leak.
2
Blue counters = 3
This is the numerator: the outcomes you want.
3
P(blue) = 3/12
Wanted over total. Leave it here for a second — do not simplify before you have checked the count.
4
P(blue) = 1/4
Divide top and bottom by 3. A fraction in lowest terms is expected unless the question says otherwise.
5
Check: 5/12 + 3/12 + 4/12 = 12/12 = 1
The three probabilities add to 1, so nothing has been miscounted.
Last step is yours
Using the same bag: work out P(not green).
1
P(green) = 4/12 = 1/3
Four green out of twelve.
2
P(not green) = 1 − P(green)
The complement rule. Everything that is not green.
3
1 − 1/3 = 2/3
Write 1 as 3/3 first, then subtract: 3/3 − 1/3 = 2/3. Check it against the bag — 5 red + 3 blue = 8 out of 12 = 2/3. It agrees.
1 2 3 4 5 6 7 8 sectors 1, 2 and 3 are shaded
Last two are yours
The eight-sector spinner above is fair. Sectors 1, 2 and 3 are shaded. Work out P(the spinner lands on an unshaded sector).
1
Total sectors = 8
Fair spinner, so all eight sectors are equally likely.
2
Shaded = 3, so P(shaded) = 3/8
Count the shaded ones.
3
P(unshaded) = 1 − 3/8 = 5/8
Write 1 as 8/8: 8/8 − 3/8 = 5/8.
4
Check: 3/8 + 5/8 = 8/8 = 1
Shaded and unshaded together are everything, so they must add to 1.
The mistake: answering “3 out of 8” or “3 : 5”. A probability of 3/8 is not a ratio. Writing 3 : 5 answers a different question and gets nothing, even though the person clearly understood the spinner.
All yours
A fair dice is rolled once. Work out P(the score is 5 or more). Give your answer as a fraction in its lowest terms.
A spinner has three outcomes: A, B and C. P(A) = 0.35 and P(B) = 0.2. Work out P(C).
The probability that it rains tomorrow is 7/10. Work out the probability that it does not rain.
Non-calculator: to turn a probability fraction into a percentage by hand, aim for a denominator of 100. 3/8 has denominator 8, and 8 does not divide 100 — so instead use 3 ÷ 8 as a short division: 3.000 ÷ 8 = 0.375, then × 100 = 37.5%. Short division of a small number by 8 is three steps and always terminates.
E8.2 · medium2 · Relative and expected frequencies — backwards and forwards▼
▶  Watch: E8.2 Relative and expected frequencies
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

Everything in section 1 assumed the outcomes were equally likely. Real objects are not always fair — a drawing pin, a bent coin, a plastic cone — and for those you cannot count outcomes. You have to experiment and estimate.

relative frequency = number of successes ÷ number of trials

Relative frequency is an estimate of probability. Cambridge is strict about the word: if the question says “estimate the probability”, it wants a relative frequency, and if it says “the spinner is fair”, it wants a counted probability. The more trials, the better the estimate — that sentence is itself a mark in a lot of papers.

expected number of occurrences = probability × number of trials
Read the two words: relative frequency looks backwards at an experiment that has already happened. Expected frequency looks forwards at trials that have not happened yet. One is a division, the other is a multiplication. If you can tell which direction the question points, you have already chosen the operation.
Worked in full
A drawing pin is thrown 200 times. It lands point up 130 times. (a) Estimate the probability that it lands point up. (b) Estimate how many times it would land point up in 500 throws.
1
(a) relative frequency = 130 ÷ 200
Successes over trials. Do not simplify yet.
2
130/200 = 13/20
Divide top and bottom by 10.
3
13/20 = 0.65
13 ÷ 20: double both to get 26/40, halve to 13/20… quicker route is 13/20 = 65/100 = 0.65, since 20 × 5 = 100 and 13 × 5 = 65.
4
(b) expected = 0.65 × 500
Probability times the new number of trials.
5
0.65 × 500 = 0.65 × 5 × 100 = 3.25 × 100 = 325
Split the 500 into 5 × 100. Multiplying by 5 then by 100 is easier by hand than 0.65 × 500 in one go.
6
Answer: 0.65 and 325 times
Both parts stated. Note the second answer is a whole number of throws.
Last step is yours
A biased dice is rolled 400 times and shows a six 92 times. Estimate the number of sixes in 1000 rolls of the same dice.
1
P(six) is estimated as 92/400
Relative frequency from the experiment.
2
92/400 = 23/100 = 0.23
Divide top and bottom by 4. 400 ÷ 4 = 100 and 92 ÷ 4 = 23, so the decimal falls out at once.
3
expected = 0.23 × 1000 = 230
Multiplying a decimal by 1000 moves the point three places: 0.23 becomes 230.
Last two are yours
A fair coin is flipped 60 times. (a) How many heads are expected? (b) The coin actually gave 34 heads. Work out the relative frequency of heads, as a decimal.
1
(a) P(head) = 1/2
The coin is fair, so this one is counted, not estimated.
2
expected = 1/2 × 60 = 30
Half of 60.
3
(b) relative frequency = 34/60 = 17/30 ≈ 0.57
Divide top and bottom by 2 to get 17/30. As a decimal 17 ÷ 30 = 0.5666…, which is 0.57 to 2 decimal places. It is close to 0.5, which is what you would expect from a fair coin over only 60 flips.
The mistake: saying the coin must be biased because 34 heads is not 30. An expected frequency is not a prediction of what will happen — it is an average over many repeats. Cambridge wants “the relative frequency is close to 0.5, so the coin appears fair”, and it wants you to mention that 60 trials is a small number.
All yours
A spinner is spun 250 times and lands on red 40 times. Estimate the number of reds in 1000 spins.
The probability that a seed germinates is 0.8. A gardener plants 350 seeds. How many are expected to germinate?
In 500 trials an event happened 175 times. Write the relative frequency as a decimal.

Fair, biased and random

wordwhat it means
fairevery outcome is equally likely: a fair four-sided spinner lands on each number with probability ¼
biasedsome outcomes are more likely than others. You can only detect bias from many trials, by comparing relative frequencies with the fair probabilities
randomchosen with no pattern, so that every item has the same chance of being picked
Worked in full
A four-sided spinner numbered 1 to 4 is spun 200 times. It lands on 1 38 times, on 2 41 times, on 3 79 times and on 4 42 times. Is the spinner fair?
1
relative frequency of 3 = 79 ÷ 200 = 0.395
Successes over trials.
2
a fair spinner would give about ¼ = 0.25 for each number
Expected 0.25 × 200 = 50 threes.
3
0.395 is far above 0.25, and 200 is a lot of spins: the spinner looks biased towards 3
Quote both numbers in the reason. With only 8 spins, a result like this could easily be chance.
All yours
For the spinner above, work out the relative frequency of landing on 2. Give a decimal.
A fair coin is thrown 300 times. How many heads would you expect?
Non-calculator: to turn a/b into a decimal by hand, scale the denominator to 10, 100 or 1000 rather than doing long division. 40/250 → multiply by 4 → 160/1000 = 0.16. 175/500 → multiply by 2 → 350/1000 = 0.35. It works whenever the denominator is built from 2s and 5s, which in exam questions it almost always is.
E8.3 · high3 · Combined events — sample spaces, Venn diagrams and trees▼
▶  Watch: E8.3 Probability of combined events
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

Two events, one question. There are only three tools, and choosing the right one takes about five seconds:

ToolUse it when
Sample space diagramTwo things happen at once and every outcome is equally likely — two dice, a coin and a spinner. Small and countable.
Venn diagramYou are told how many are in each category and how many overlap. The words both, either, neither.
Tree diagramTwo or more stages, one after the other. The words then, followed by, a second counter.
The two rules that run this whole section:
Multiply ALONG the branches (a path through the tree = this AND then that).
Add BETWEEN the paths (this path OR that path).
Everything else is detail.

Sample space diagrams

second die first die 1 2 3 4 5 6 1 2 3 4 5 6 2 3 4 5 6 7 3 4 5 6 7 8 4 5 6 7 8 9 5 6 7 8 9 10 6 7 8 9 10 11 7 8 9 10 11 12 36 equally likely outcomes. The six shaded cells give a total of 7.
Worked in full
Two fair dice are rolled and their scores are added. Work out P(total = 7).
1
Total outcomes = 6 × 6 = 36
Six faces on the first die, six on the second, and each pairing is a separate outcome.
2
Totals of 7: 1+6, 2+5, 3+4, 4+3, 5+2, 6+1
Six of them. Note 1+6 and 6+1 are different outcomes — they are different cells in the grid.
3
P(total 7) = 6/36
Wanted over total.
4
P(total 7) = 1/6
Divide top and bottom by 6. Seven is the most likely total on two dice, which is a useful fact to carry.
Using the same two dice, work out P(the total is 10 or more). Give a fraction in its lowest terms.

Venn diagrams

A Venn diagram is a counting picture. The rectangle is everyone, each circle is one category, and the overlap belongs to both. Fill the overlap first — that single habit fixes most Venn errors, because the numbers you are given usually include the overlap and the numbers you write must not double-count it.

E = 30 students H — hockey N — netball 12 6 8 4
Worked in full
In a class of 30 students, 18 play hockey, 14 play netball and 6 play both. (a) Complete the Venn diagram. (b) Work out P(a student chosen at random plays neither sport).
1
Both = 6, so write 6 in the overlap
Fill the middle first. Always.
2
Hockey only = 18 − 6 = 12
The 18 includes the 6 who play both. Subtract them out or you count six people twice.
3
Netball only = 14 − 6 = 8
Same reasoning on the other side.
4
In at least one sport = 12 + 6 + 8 = 26
Add the three regions inside the circles.
5
Neither = 30 − 26 = 4
Everyone else. This goes in the rectangle but outside both circles.
6
P(neither) = 4/30 = 2/15
Wanted over total, then divide top and bottom by 2.
The mistake: writing 18 in the hockey-only region. The sentence “18 play hockey” describes the whole circle, not the part of it outside the overlap. If you write 18 and 14 and 6, your regions add to 38 in a class of 30 — and a total larger than the class is the signal that this is what has happened.
Using the Venn diagram above, work out P(a student plays hockey only). Give a fraction in its lowest terms.

Tree diagrams — with replacement

A tree is for stages. Each branch carries a probability; each complete path from left to right is one outcome. With replacement means the first object goes back before the second is taken, so the bag is identical at stage 2 — the second set of branches is a copy of the first.

WITH replacement — the counter goes back, so stage 2 looks exactly like stage 1 R B 3/8 5/8 R 3/8 RR = 3/8 × 3/8 = 9/64 B 5/8 RB = 3/8 × 5/8 = 15/64 R 3/8 BR = 5/8 × 3/8 = 15/64 B 5/8 BB = 5/8 × 5/8 = 25/64
Worked in full
A bag holds 3 red and 5 blue counters. One is taken, its colour recorded, and it is put back. A second is then taken. Work out (a) P(both red) and (b) P(one of each colour).
1
Total = 3 + 5 = 8, so P(R) = 3/8 and P(B) = 5/8
Stage 1. These two must add to 1: 3/8 + 5/8 = 8/8. They do.
2
The counter goes back, so stage 2 is identical: P(R) = 3/8, P(B) = 5/8
This is what “with replacement” buys you. Nothing changes.
3
(a) P(RR) = 3/8 × 3/8
Multiply along the top path.
4
3/8 × 3/8 = 9/64
Tops times tops, bottoms times bottoms. 3 × 3 = 9, 8 × 8 = 64.
5
(b) one of each = RB or BR
Two different paths, so two products and then an addition.
6
P(RB) = 3/8 × 5/8 = 15/64 and P(BR) = 5/8 × 3/8 = 15/64
The same product both ways round, but they are genuinely two separate outcomes and both must be counted.
7
P(one of each) = 15/64 + 15/64 = 30/64
Same denominator, so just add the tops. This is where “add between the branches” happens.
8
30/64 = 15/32
Divide top and bottom by 2.
Non-calculator: when you add along the bottom of a tree, the denominators are already equal, because every path came from the same two stages. So the addition is just tops: 15/64 + 15/64 = 30/64. Only simplify at the very end — simplifying each path first destroys the common denominator and forces you to rebuild it.

Tree diagrams — without replacement

This is the distinction that gets examined more than anything else in probability. Without replacement means the first object is kept, so at stage 2 there is one fewer object in the bag — and if the first was red, there is also one fewer red.

Two things change at stage 2, and both must change: the denominator drops by one on every branch, and the numerator drops by one only on the branch matching what you already took.

WITHOUT replacement — one counter has gone, so every stage-2 denominator is 7 R B 3/8 5/8 R 2/7 RR = 3/8 × 2/7 = 6/56 B 5/7 RB = 3/8 × 5/7 = 15/56 R 3/7 BR = 5/8 × 3/7 = 15/56 B 4/7 BB = 5/8 × 4/7 = 20/56
Worked in full
Same bag: 3 red and 5 blue. Two counters are taken without replacement. Work out (a) P(both red) and (b) P(one of each colour).
1
Stage 1 is unchanged: P(R) = 3/8, P(B) = 5/8
Nothing has been removed yet.
2
Stage 2 denominators are all 7
One counter has gone, whatever colour it was. 8 − 1 = 7. Write the 7s in before you think about the tops.
3
If the first was red: 2 reds and 5 blues remain, so P(R) = 2/7, P(B) = 5/7
The red numerator dropped from 3 to 2. Check: 2/7 + 5/7 = 7/7. Correct.
4
If the first was blue: 3 reds and 4 blues remain, so P(R) = 3/7, P(B) = 4/7
This time the blue numerator dropped. Check: 3/7 + 4/7 = 7/7.
5
(a) P(RR) = 3/8 × 2/7 = 6/56
Multiply along. 3 × 2 = 6 and 8 × 7 = 56.
6
6/56 = 3/28
Divide top and bottom by 2.
7
(b) P(RB) = 3/8 × 5/7 = 15/56 and P(BR) = 5/8 × 3/7 = 15/56
Both paths, both products.
8
P(one of each) = 15/56 + 15/56 = 30/56 = 15/28
Add tops, then divide top and bottom by 2.
The mistake: changing the denominator but forgetting the numerator, so writing P(R then R) = 3/8 × 3/7. The check that catches it every time: each pair of stage-2 branches must add to 1. 3/7 + 5/7 = 8/7, which is bigger than 1 and therefore impossible. Two seconds of checking, and the whole question is saved.
Last step is yours
A box has 4 mints and 6 toffees. Two sweets are taken without replacement. Work out P(both toffees).
1
Total = 4 + 6 = 10, so P(toffee) = 6/10 first
Stage 1.
2
One sweet has gone, so stage 2 denominators are 9
10 − 1.
3
The first was a toffee, so 5 toffees remain: P(toffee) = 5/9
Both the numerator and the denominator dropped by one, because the sweet removed was a toffee.
4
P(TT) = 6/10 × 5/9 = 30/90 = 1/3
Multiply along: 6 × 5 = 30, 10 × 9 = 90. Then 30/90 = 1/3. Cancelling before multiplying is faster still: the 5 and the 10 cancel to 1 and 2, leaving 6/2 × 1/9 = 3/9 × ... but the safest route by hand is multiply, then simplify once.
Last two are yours
A bag holds 2 green and 5 yellow beads. Two are taken without replacement. Work out P(at least one green).
1
Total = 7. “At least one green” means GG, GY or YG — three paths.
Three paths is a lot of adding. There is a shorter way.
2
The opposite of “at least one green” is “no green at all”, which is YY — one path
Use the complement. One product instead of three.
3
P(YY) = 5/7 × 4/6 = 20/42 = 10/21
Stage 2: one bead gone so denominator 6, and it was yellow so 5 yellows become 4.
4
P(at least one green) = 1 − 10/21 = 11/21
Write 1 as 21/21, then 21/21 − 10/21 = 11/21.
Whenever you see “at least one”, reach for the complement. P(at least one) = 1 − P(none). It turns three or more products into one, and in a non-calculator paper that is the difference between a clean answer and a page of fractions.
All yours
A bag holds 4 white and 3 black counters. Two are taken without replacement. Work out P(both white). Give a fraction in its lowest terms.
The same bag (4 white, 3 black), two taken without replacement. Work out P(both black). Fraction in lowest terms.
A coin is flipped twice. Work out P(exactly one head). Fraction in lowest terms.
A bag holds 5 red and 4 blue. One counter is taken WITH replacement, then another. Work out P(both blue). Fraction in lowest terms.
E8.4 · high4 · Conditional probability — the word “given” shrinks the total▼
▶  Watch: E8.4 Conditional probability
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

This is where the check said step 7 broke, and it is the one genuinely new idea in the whole probability topic. Everything else is counting. This is about what the word “given” does to the total.

Conditional probability answers: we already know B happened — now what is the chance of A? The notation is P(A | B), read as “the probability of A given B”. It is a handy shorthand for your own working, but 0580 does not require it: in an answer you can just as well write P(French, given a girl).

The one sentence to remember: being told that B has happened does not change the numerator — it shrinks the denominator. The population is no longer everybody. It is only the people in B. So: find the new total first, then count inside it.
P(A | B) = number in both A and B ÷ number in B

From a two-way table — the easiest place to see it

French Spanish Total Boys 12 8 20 Girls 15 5 20 Total 27 13 40 Given the student is a girl, the world shrinks to the shaded row: 20 people, 15 of them French.
Worked in full
40 students each study one language. The two-way table above shows the numbers. A student is chosen at random. Work out (a) P(the student studies French) and (b) P(the student studies French, given that the student is a girl).
1
(a) Total students = 40
The bottom-right corner of a two-way table is always the grand total. Use it.
2
French students = 27
The French column total.
3
P(French) = 27/40
No word “given”, so the denominator is everybody. 27/40 will not simplify — 27 is odd.
4
(b) “given that the student is a girl” → use only the Girls row
The world has shrunk. Cover the Boys row with your hand.
5
New total = 20
The Girls row total, not 40.
6
Girls doing French = 15
Reading across the Girls row to the French column.
7
P(French | girl) = 15/20
New numerator over new total.
8
15/20 = 3/4
Divide top and bottom by 5. Compare with part (a): 27/40 is about 0.675, but 3/4 is 0.75. Knowing the student is a girl changed the answer, which is the whole point.
The mistake: answering 15/40. That is P(French AND girl), not P(French GIVEN girl). Both are legitimate questions and Cambridge asks both, sometimes in the same part — so the word to hunt for in the sentence is given, or if we know that, or of the girls. All three mean the same thing: change the denominator.

From a Venn diagram

Same idea, different picture. Go back to the 30 students: 12 hockey only, 6 both, 8 netball only, 4 neither.

Last step is yours
A student is chosen at random from the 30. Work out P(the student plays netball, given that the student plays hockey).
1
“Given hockey” → use only the hockey circle
Ignore the netball-only region and the outside entirely.
2
Number who play hockey = 12 + 6 = 18
The whole circle, both of its regions. This is the new denominator.
3
Of those, the ones who also play netball = 6
Only the overlap. This is the new numerator.
4
P(netball | hockey) = 6/18 = 1/3
Divide top and bottom by 6. One in three hockey players also plays netball.

From a tree diagram

A tree is already conditional. Every probability on the second set of branches is a conditional probability — that is exactly what “without replacement” was doing in section 3. When a question asks for a conditional probability from a tree, you are usually being asked to read a stage-2 branch directly, or to divide one path by a group of paths.

P(A | B) = P(A and B) ÷ P(B) (extra, not required for 0580: use the counts, the table or the tree)
Last two are yours
A bag holds 3 red and 5 blue counters. Two are taken without replacement. Given that the first counter was red, work out the probability that the second is also red. Then work out P(the second is red), not given anything.
1
Given the first was red → read the top stage-2 branch
No calculation at all. The tree already carries this number.
2
P(second red | first red) = 2/7
Two reds left out of seven counters. That is a conditional probability, sitting in plain sight.
3
For P(second red) with no condition, add the two paths ending in red: RR and BR
The second counter can be red either way round.
4
P(RR) = 3/8 × 2/7 = 6/56 and P(BR) = 5/8 × 3/7 = 15/56
Multiply along each path.
5
P(second red) = 6/56 + 15/56 = 21/56 = 3/8
Add tops over the common denominator: 21/56, and 21 ÷ 7 = 3 with 56 ÷ 7 = 8, so 3/8. It equals P(first red) — that is not a coincidence, and it is worth noticing.
Non-calculator: conditional questions in Paper 2 almost always come from a table or a Venn diagram rather than from the formula, precisely because the arithmetic then is a single division of small whole numbers. If you find yourself dividing one fraction by another, stop and check whether the question actually gave you counts you could have used instead.
All yours
From the two-way table above (Boys: 12 French, 8 Spanish; Girls: 15 French, 5 Spanish): work out P(the student is a boy, given that the student studies Spanish). Give a fraction in its lowest terms.
From the same table: work out P(the student studies Spanish, given that the student is a boy). Fraction in lowest terms.
30 students: 12 hockey only, 6 both, 8 netball only, 4 neither. Work out P(the student plays hockey, given that they play netball). Fraction in lowest terms.
40 students: 15 boys and 8 girls play a sport; 7 boys and 10 girls do not. A girl is chosen at random. Work out the probability that she plays a sport. Fraction in lowest terms.
A bag has 2 green and 6 yellow beads. Two are taken without replacement. Given the first was yellow, what is P(the second is green)? Fraction in lowest terms.
E9.1 · low5 · Classifying statistical data — tally tables and two-way tables▼
▶  Watch: E9.1 Classifying statistical data
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

This sub-topic is worth few marks and takes little time, but the marks are free and Cambridge does ask for the words. Two jobs: classify the data, and organise it into a table.

The four words

WordMeansExample
Qualitativea description, not a numbereye colour, favourite sport
Quantitativea numberheight, number of pets
Discretecounted — only certain values possiblenumber of pets: 0, 1, 2, never 1.4
Continuousmeasured — any value in a rangeheight: 162.37 cm is perfectly possible

Discrete and continuous are both kinds of quantitative data. The test for continuous is simple: could a more accurate instrument give you another decimal place? Height, mass, time and length always can. Shoe size cannot — there is no shoe size 7.318.

The mistake: calling shoe size continuous because sizes go up in halves. Halves are still a fixed list. And calling age discrete because people say “I am 15” — age is continuous, it is just usually rounded down.

Tally tables

A tally table turns a list into a frequency table. The fifth mark goes diagonally across the previous four, so the marks read in fives and can be counted at a glance.

ColourTallyFrequency
Red||||  |||8
Blue||||  ||||10
Green||||4
Yellow||||  ||||  ||12
Total34

The total row is not decoration. It is the check: if the frequencies do not add to the number of items you were given, something has been miscounted, and finding that out now costs nothing.

Worked in full
A survey records the number of pets owned by 25 students: 0, 2, 1, 0, 3, 1, 1, 2, 0, 4, 1, 2, 0, 1, 1, 3, 2, 0, 1, 2, 5, 1, 0, 2, 1. Complete a frequency table and state the type of data.
1
Type: quantitative and discrete
It is a number, and you cannot own 1.5 pets.
2
List the possible values first: 0, 1, 2, 3, 4, 5
Set up every row before counting anything. This stops you from inventing a category halfway through.
3
Tally 0: positions 1, 4, 9, 13, 18, 23 → frequency 6
Work through the list once, in order, marking as you go. Never scan back and forth.
4
1 appears 9 times, 2 appears 6 times, 3 appears 2 times, 4 once, 5 once
Same method for each value.
5
Check: 6 + 9 + 6 + 2 + 1 + 1 = 25
It matches the 25 students, so the tally is complete.

Two-way tables

A two-way table classifies by two things at once — that is what makes it the natural home for conditional probability, as in section 4. Fill it by using the totals: every row adds across, every column adds down, and the corner checks both.

Last step is yours
Complete this two-way table. 60 people were asked whether they walk or drive to work. 35 are women. 12 men walk. 22 people in total drive.
1
Grand total = 60
Put it in the corner first.
2
Women = 35, so men = 60 − 35 = 25
Column totals must add to the grand total.
3
Drive total = 22, so walk total = 60 − 22 = 38
Row totals must too.
4
Men who walk = 12, so men who drive = 25 − 12 = 13
Work inside the men column now that its total is known.
5
Women who walk = 38 − 12 = 26
The walk row totals 38 and 12 of them are men, so the rest are women. Check the last cell two ways: women who drive = 35 − 26 = 9, and also 22 − 13 = 9. Both give 9, so the table is right.
Last two are yours
From that completed table (men: 12 walk, 13 drive; women: 26 walk, 9 drive), a person is chosen at random. Work out (a) P(they walk) and (b) P(they are a woman, given that they drive).
1
(a) walkers = 12 + 26 = 38 out of 60
No condition, so the denominator is everybody.
2
P(walk) = 38/60 = 19/30
Divide top and bottom by 2.
3
(b) Given they drive → denominator is 22
The drive row total.
4
P(woman | drives) = 9/22 = 9/22
Nine women drive out of 22 drivers. It will not simplify — 9 and 22 share no factors.
All yours
Classify this data: the time taken by each student to run 100 metres. Answer with one word: discrete or continuous.
In a tally table the frequencies are 7, 11, 5 and 9. How many items were surveyed in total?
A two-way table has 80 people in total. 46 are adults. Of the adults, 19 chose tea. In total 35 chose tea. How many children chose tea?
E9.2 · low6 · Interpreting statistical data — always TWO comparisons▼
▶  Watch: E9.2 Interpreting statistical data
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

Reading a chart is easy. The mark that gets dropped is the comparison, and it is dropped in the same way by almost everybody: they give one comparison when Cambridge wants two.

The rule for every “compare the two sets” question: one sentence about an average (usually the median or the mean) and one sentence about the spread (usually the range or the interquartile range). Two sentences, two marks. One sentence, one mark, every time.
And put both sentences in context — not “the median is higher” but “the median mark in Class 1 is higher, so Class 1 did better on average”.
0 2 4 6 8 10 12 14 grade A grade B grade C grade D Class 1 Class 2 frequency A dual bar chart. The bars sit side by side, never stacked on top of each other.
Worked in full
The dual bar chart above shows grades A to D for two classes. Class 1: A 4, B 9, C 12, D 5. Class 2: A 10, B 11, C 6, D 3. Compare the performance of the two classes.
1
Class 1 total = 4 + 9 + 12 + 5 = 30 and Class 2 total = 10 + 11 + 6 + 3 = 30
Check the totals first. They are equal here, which makes the comparison fair — if they were not, you would have to talk in proportions.
2
Class 2 has more of the top grades: 10 + 11 = 21 at A or B, against 4 + 9 = 13
Group the grades before comparing them. Comparing bar by bar produces four sentences and no argument.
3
Sentence 1 (average): Class 2 performed better on average, since 21 of its 30 students got A or B compared with 13 of 30 in Class 1.
An average-type statement, in context, with numbers quoted from the chart.
4
Sentence 2 (spread): Class 1 is more spread out across the grades, with its students split fairly evenly over all four, while Class 2 is concentrated in A and B.
A spread statement. This is the sentence that most students never write.
5
Both sentences mention grades and classes, not just numbers
Context is a marking requirement, not politeness.
The mistake: “Class 2 is better.” That is a conclusion with no evidence and it earns nothing. Every comparison sentence needs a number from the diagram in it. The fastest fix is to write the number first and the opinion second: “21 out of 30 got A or B in Class 2 compared with 13 out of 30, so Class 2 did better.”
Last step is yours
Two machines fill bottles. Machine A: median 501 ml, IQR 4 ml. Machine B: median 500 ml, IQR 11 ml. Which machine would a factory prefer, and why?
1
Average: the medians are 501 and 500, almost identical
Both machines put in about the right amount, so the average does not separate them.
2
Spread: IQR 4 ml against IQR 11 ml
A much smaller IQR for Machine A.
3
A smaller IQR means the middle half of the bottles are closer to each other, so the machine is more consistent
This is what spread means in context. Say it in the language of the question.
4
Conclusion: Machine A, because its median is essentially the same but its IQR is far smaller, so it fills bottles far more consistently.
Both measures mentioned, both in context, and a decision at the end. That is the full three marks.
Last two are yours
Two football teams score these goals in 7 matches. Team X: 0, 1, 1, 2, 2, 3, 12. Team Y: 2, 3, 3, 3, 4, 4, 5. Compare them, and say which average is the fairer one to use for Team X.
1
Team X mean = (0+1+1+2+2+3+12) ÷ 7 = 21 ÷ 7 = 3
Add in pairs: 0+1 = 1, 1+2 = 3, 2+3 = 5, and 1 + 3 + 5 + 12 = 21.
2
Team Y mean = (2+3+3+3+4+4+5) ÷ 7 = 24 ÷ 7 ≈ 3.4
A slightly higher mean.
3
Team X median = 2 (the 4th value), Team Y median = 3
Both lists are already in order. With 7 values the median is the 4th.
4
Team X range = 12 − 0 = 12, Team Y range = 5 − 2 = 3
Largest minus smallest, in that order. A range is never negative — if yours is, you subtracted the wrong way round.
5
The 12 is an outlier. It drags Team X mean up to 3 while five of the seven matches produced 2 goals or fewer, so the median of 2 describes Team X more honestly.
That is the standard Cambridge answer: the mean is affected by extreme values, the median is not.
Sign-error watch. range = largest − smallest. Written the other way round, 0 − 12 = −12, and a negative spread is impossible. Same for IQR = Q3 − Q1. Before you write either answer down, glance at it: is it positive? That single glance catches the error you make most often, in the one place in statistics where it can happen.
All yours
Class P has a median test mark of 62 and a range of 40. Class Q has a median of 62 and a range of 15. Which class was more consistent? Answer with one letter.
A data set has smallest value 7 and largest value 31. Work out the range.
For a set of marks, Q1 = 38 and Q3 = 57. Work out the interquartile range.
Team A has a mean of 6 goals and a range of 2. Team B has a mean of 6 goals and a range of 9. How many measures do you need to quote to compare the two teams fully? Answer with a digit.

What the data cannot tell you

Before you write a conclusion, ask five questions. Each one is a common exam trap:

questionwhy it matters
Is the sample big enough, and was it chosen at random?a small or biased sample may not represent everyone
Are the totals equal?if not, compare proportions, not counts
Does the vertical axis start at 0?a broken axis makes small differences look large
Does a correlation show that one thing causes the other?no: both may be caused by something else
Is the prediction outside the range of the data?extrapolating beyond the data is unreliable
Worked in full
At school X, 30 of 60 students walk to school. At school Y, 45 of 150 students walk. Lena says: “More students walk at Y, so walking is more popular at Y.” Is she right?
1
the totals are different (60 and 150), so compare proportions
Counts alone are unfair when the groups are different sizes.
2
X: 30/60 = 1/2; Y: 45/150 = 3/10
Half against three tenths.
3
no: walking is more popular at X
More students walk at Y only because Y is a bigger school.
All yours
At school Y, 45 of 150 students walk. What fraction walk? Give it in its lowest terms.
A bar chart’s vertical axis starts at 90. Bar A reaches 100 and bar B reaches 95. The drawn bar for B is what fraction of the drawn length of bar A?
E9.3 · high7 · Averages and spread — including the estimated mean▼

The biggest sub-topic in the whole of statistics, and the one where step 7 sits. Four averages and two spreads, and then the piece that separates grades: estimating the mean from a grouped frequency table.

MeasureHowWatch out for
Meanadd them all, divide by how manydragged by extreme values
Medianput in order, then take the middleforgetting to order first
Modethe most common valuethere can be two, or none
Rangelargest − smallestuses only two values, so one outlier ruins it
Q1, Q3medians of the lower and upper halveswhether to include the median itself
IQRQ3 − Q1the middle half only — ignores outliers, which is the point
Finding the median position by hand: with n values in order, the median is the value in position (n + 1) ÷ 2. If n = 9 that is position 5, a real value. If n = 8 that is position 4.5, meaning halfway between the 4th and 5th — so add those two and halve. Quartiles: Q1 is the middle of the lower half and Q3 the middle of the upper half. With an odd number of values, leave the median out of both halves; with an even number, the list splits exactly in two. Some books use positions instead, Q1 at (n + 1) ÷ 4 and Q3 at 3(n + 1) ÷ 4: with an odd number of values this gives the same answer, with an even number it can differ a little. Both answers are accepted here; use the one your teacher uses.
Worked in full
Find the median, quartiles, range and interquartile range of: 11, 4, 16, 7, 20, 6, 14, 9.
1
In order: 4, 6, 7, 9, 11, 14, 16, 20
Order first, always. Most median errors are made before any arithmetic starts.
2
n = 8, so the median is at position (8 + 1) ÷ 2 = 4.5
Between the 4th and 5th values.
3
median = (9 + 11) ÷ 2 = 20 ÷ 2 = 10
Add the two middle values and halve.
4
Lower half: 4, 6, 7, 9. Q1 = (6 + 7) ÷ 2 = 6.5
With an even n, the halves split cleanly and no value is shared.
5
Upper half: 11, 14, 16, 20. Q3 = (14 + 16) ÷ 2 = 15
Same method on the top half.
6
range = 20 − 4 = 16
Largest minus smallest.
7
IQR = Q3 − Q1 = 15 − 6.5 = 8.5
Q3 first. Both spreads came out positive, which is the check.

Estimated mean from a grouped frequency table — the step-7 skill

When data is grouped you no longer know the individual values, only the class each one fell in. So you cannot find the true mean. You estimate it, by assuming every value in a class sits at the class midpoint. The word “estimate” in the question is not politeness; it is telling you which method to use.

midpoint = (lower boundary + upper boundary) ÷ 2     estimated mean = Σfx ÷ Σf
Worked in full
40 students had their heights measured. 140 < h ≤ 150: 4 students. 150 < h ≤ 160: 11. 160 < h ≤ 170: 18. 170 < h ≤ 180: 7. Estimate the mean height.
1
Midpoints: 145, 155, 165, 175
Halfway across each class: (140 + 150) ÷ 2 = 145, and so on. They go up in 10s because the classes do.
2
Σf = 4 + 11 + 18 + 7 = 40
Check it against the 40 students in the question before going further.
3
4 × 145 = 580
Do each f × x on its own line. 4 × 145: 4 × 100 = 400, 4 × 45 = 180, total 580.
4
11 × 155 = 1705
11 × 155 = 10 × 155 + 155 = 1550 + 155 = 1705. Multiplying by 11 is always “times ten, then add one more”.
5
18 × 165 = 2970
18 × 165 = 20 × 165 − 2 × 165 = 3300 − 330 = 2970. Rounding up to 20 and subtracting back is far easier by hand than long multiplication.
6
7 × 175 = 1225
7 × 175 = 7 × 100 + 7 × 75 = 700 + 525 = 1225.
7
Σfx = 580 + 1705 + 2970 + 1225 = 6480
Add in pairs: 580 + 1705 = 2285, and 2970 + 1225 = 4195. Then 2285 + 4195 = 6480.
8
estimated mean = 6480 ÷ 40 = 162
Divide by 4 then by 10: 6480 ÷ 4 = 1620, and 1620 ÷ 10 = 162 cm.
9
Sense check: 162 lies inside 140 to 180 and near the biggest class, 160 to 170
An estimated mean must land inside the data. If it does not, an f × x has gone wrong.
Non-calculator shortcut worth having — the assumed mean. Pick the midpoint of the biggest class as a base, here 165, and work with how far each midpoint is from it: −20, −10, 0, +10.
Σf × d = 4(−20) + 11(−10) + 18(0) + 7(+10) = −80 − 110 + 0 + 70 = −120.
mean = 165 + (−120 ÷ 40) = 165 − 3 = 162. Same answer, and the largest number you handled was 120 instead of 6480.
This is exactly where your sign errors would bite. The deviations below the base are negative and the ones above are positive, and they must keep their signs all the way to the end. Write the minus signs large, add the negatives together first (−80 − 110 = −190), then bring in the positives (−190 + 70 = −120). Doing all the negatives before any positives removes most of the risk.
Last step is yours
30 people recorded the time t minutes they waited. 0 < t ≤ 10: 6 people. 10 < t ≤ 20: 14. 20 < t ≤ 30: 10. Estimate the mean waiting time.
1
Midpoints: 5, 15, 25
Halfway across each class.
2
Σf = 6 + 14 + 10 = 30
Matches the 30 people.
3
6 × 5 = 30, 14 × 15 = 210, 10 × 25 = 250
14 × 15 = 14 × 10 + 14 × 5 = 140 + 70 = 210.
4
Σfx = 30 + 210 + 250 = 490
Straight addition.
5
estimated mean = 490 ÷ 30 = 16.3 minutes (3 s.f.)
490 ÷ 30 = 49 ÷ 3 = 16.333…, so 16.3 minutes to 3 significant figures. It sits inside 10 to 20, where most of the people are, so it is believable.
Last two are yours
The table shows goals scored in 20 matches. 0 goals: 5 matches. 1 goal: 7. 2 goals: 5. 3 goals: 3. Work out the mean, the median and the modal number of goals.
1
This data is not grouped — the values are exact, so the mean is exact too
Ungrouped discrete data. No midpoints needed.
2
Σfx = 5(0) + 7(1) + 5(2) + 3(3) = 0 + 7 + 10 + 9 = 26
Multiply each value by its frequency, then add.
3
mean = 26 ÷ 20 = 1.3 goals
26 ÷ 20 = 13 ÷ 10 = 1.3.
4
median position = (20 + 1) ÷ 2 = 10.5, so halfway between the 10th and 11th values
Running totals: the first 5 values are 0, values 6 to 12 are 1. So both the 10th and the 11th are 1.
5
median = 1 goal
Both middle values are 1, so the median is 1.
6
mode = 1 goal, since 7 matches is the highest frequency
The mode is the value with the biggest frequency, not the frequency itself. Answering 7 is the classic slip — 7 is how often it happened, 1 is what happened.
The mistake: giving the modal frequency instead of the modal value. And in a grouped table, the answer is a modal class — a whole interval such as 160 < h ≤ 170, not a single number.
All yours
40 plants had their heights measured. 0 < h ≤ 20: 7 plants. 20 < h ≤ 40: 18. 40 < h ≤ 60: 15. Estimate the mean height in cm.
Find the median of: 3, 8, 5, 12, 9, 4, 11.
Find the interquartile range of: 2, 5, 6, 8, 11, 13, 15, 20.
A grouped table has classes 0 to 10, 10 to 20, 20 to 30 with frequencies 3, 9, 8. What is the midpoint of the middle class?
Five numbers have a mean of 12. Four of them are 8, 10, 14 and 15. Work out the fifth.
For a set of data, the smallest value is 12, Q1 = 25, the median is 34, Q3 = 46 and the largest value is 58. Work out the interquartile range.

The modal class

In a grouped table with equal class widths, the modal class is the class with the highest frequency. The answer is the whole interval, such as 160 < h ≤ 170: not its frequency, and not a single number.

height h (cm)140 < h ≤ 150150 < h ≤ 160160 < h ≤ 170170 < h ≤ 180
frequency411187
Last step is yours
Write down the modal class for the heights in the table.
1
the highest frequency is 18
Find the biggest frequency first.
2
modal class: 160 < h ≤ 170
Answer with the interval. Writing 18 scores nothing.
All yours
Times: 0 < t ≤ 10: 6; 10 < t ≤ 20: 14; 20 < t ≤ 30: 12; 30 < t ≤ 40: 8. Write down the modal class, like 0<t<=10.
E9.4 · medium8 · Charts and diagrams — bar, pie, pictogram, stem-and-leaf▼
▶  Watch: E9.4 Statistical charts and diagrams
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

Cambridge does not usually ask you to draw a whole chart from nothing. It asks you to complete one, read one, or calculate one number needed to draw one. The calculations are short; the marks go to accuracy and labelling.

Bar charts — simple, dual and stacked

Bars have equal widths and gaps between them, and the height is the frequency. That last sentence is only true for bar charts. In a histogram the height is not the frequency, and confusing the two is section 11.

0 5 10 15 20 walk 6 bus 9 car 5 Class 1 top of the bar = 20 = the total A stacked bar. Each block is read from its own height, not from the axis.

In a stacked bar the blocks sit on top of each other, so the top of the bar is the total and each block is read as a difference between two positions, not as a height off the axis. In a dual bar (section 6) they sit side by side and each is read off the axis directly.

The mistake: reading the middle block of a stacked bar straight off the axis. In the chart above the bus block runs from 6 up to 15, so it is 15 − 6 = 9, not 15.

Pie charts

A pie chart shows proportion, not amount. The whole circle is 360° and the whole data set, so:

angle = (frequency ÷ total) × 360°
100° 80° 60° 120° Maths — 10 students, 100° Science — 8 students, 80° English — 6 students, 60° Art — 12 students, 120° 36 students, so each student is worth 360 ÷ 36 = 10°.
Worked in full
36 students chose a favourite subject: Maths 10, Science 8, English 6, Art 12. Work out the angle for each sector.
1
Total = 10 + 8 + 6 + 12 = 36
Check it against the 36 in the question.
2
360 ÷ 36 = 10° per student
Find the angle for ONE item first. Every angle is then a single multiplication, and by hand that is far quicker than four separate fraction sums.
3
Maths: 10 × 10 = 100°
Ten students at 10° each.
4
Science: 8 × 10 = 80°
Same method.
5
English: 6 × 10 = 60° and Art: 12 × 10 = 120°
Same again.
6
Check: 100 + 80 + 60 + 120 = 360°
The angles must total 360°. If they do not, one is wrong, and this check finds it before the drawing does.
Non-calculator: always compute 360 ÷ total first and see whether it is a whole number. Exam totals are chosen so that it usually is: 36 gives 10°, 30 gives 12°, 40 gives 9°, 24 gives 15°, 60 gives 6°, 90 gives 4°, 120 gives 3°. If it is not whole — say a total of 50, giving 7.2° — multiply first and divide last: frequency × 360 ÷ 50, keeping whole numbers as long as possible.
Last step is yours
A pie chart shows how 90 people travel to work. The sector for “train” has an angle of 64°. How many people travel by train?
1
This is the reverse direction: angle known, frequency wanted
Rearranging the same relationship.
2
360 ÷ 90 = 4° per person
One person is worth 4°.
3
people = 64 ÷ 4 = 16
Sixty-four degrees at four degrees each. Check: 16 × 4 = 64. Correct.

Pictograms

A pictogram uses a symbol to stand for a number of items, and part-symbols for the remainder. The key is the whole question. If one circle stands for 8 people, then half a circle is 4 and a quarter is 2. Reading a pictogram without checking the key is how a correct count becomes a wrong answer.

Stem-and-leaf diagrams

A stem-and-leaf keeps every original value while showing the shape of the data, which is why you can find the exact median and range from one. The stem is the tens digit, the leaves are the units, and the leaves must be in order and evenly spaced.

Stem Leaf 2 3 5 7 3 1 4 4 8 4 2 5 5 5 5 1 Key: 3 | 4 means 34. Without the key the diagram earns nothing.
Last two are yours
From the stem-and-leaf above (key 3 | 4 = 34), work out the number of values, the range, the mode and the median.
1
Count the leaves: 3 + 4 + 4 + 1 = 12 values
Count leaves, never stems.
2
Smallest = 23, largest = 51
First leaf on the first row, last leaf on the last row. The diagram is already ordered, which is its main advantage.
3
range = 51 − 23 = 28
Largest minus smallest. Positive, so the order was right.
4
mode = 45
The leaf 5 appears three times on the stem 4, so 45 occurs three times — more than any other value.
5
median position = (12 + 1) ÷ 2 = 6.5, so the 6th and 7th values are 34 and 38, giving median = (34 + 38) ÷ 2 = 36
Count along the leaves in order: 23, 25, 27, 31, 34, 34, 38, ... The 6th is 34 and the 7th is 38, so the median is 36.
All yours
A pie chart represents 24 people. Work out the angle of the sector representing 5 people.
A pie chart represents 60 students. One sector has an angle of 54°. How many students does it represent?
In a pictogram, one square stands for 12 cars. A row shows 3 and a half squares. How many cars is that?
In the stacked bar above, walk = 6 and the bus block runs from 6 up to 15. How many people took the bus?

Pictograms

Read the key first. If one symbol stands for 8 books, half a symbol is 4 books and a quarter is 2. To draw one: frequency ÷ the key value = the number of symbols, using part-symbols for what is left over.

MonTueWedThuFri= 8 books
Worked in full
The pictogram shows books borrowed from a library. How many were borrowed on Tuesday and on Thursday?
1
Tuesday: 3½ symbols = 3.5 × 8 = 28 books
Three whole books and a half.
2
Thursday: 4¼ symbols = 4.25 × 8 = 34 books
A quarter of a symbol is 2 books.

Simple frequency distributions

A frequency table lists each value with how often it occurs. Read the wording carefully: “at least 7” includes 7; “more than 7” does not. It is drawn as a bar chart with the values along the bottom.

score456789
frequency359643
Worked in full
Use the table. How many students scored at least 7? What is the mode, and how many students are there?
1
at least 7: 6 + 4 + 3 = 13
At least 7 includes 7 itself.
2
mode = 6
The score with the highest frequency, 9. The mode is the score, not the 9.
3
total = 3 + 5 + 9 + 6 + 4 + 3 = 30
Add the frequencies, not the scores.
All yours
On the pictogram, one symbol represents 8 books. How many symbols show 18 books? Give a decimal.
How many books were borrowed on Friday?
In the score table, how many students scored more than 7?
E9.5 · low9 · Scatter diagrams — correlation and the line of best fit▼
▶  Watch: E9.5 Scatter diagrams
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

A scatter diagram plots two measurements for each individual and asks whether they are related. There are only three jobs: plot, describe, draw a line of best fit and use it.

positive negative zero / no correlation Describe strength as well as direction: strong, moderate or weak.

Describing correlation takes two words, not one: a direction and a strength.

DirectionMeans
Positiveas one goes up, so does the other
Negativeas one goes up, the other goes down
Zero / noneno pattern — the points are scattered

Strength is strong if the points sit close to a straight line, weak if they are loosely grouped around it. And say what it means: “strong positive correlation — students who revised for longer tended to score higher”.

0 0 2 2 4 4 6 6 8 8 10 10 12 12 hours revised test score Strong positive correlation. The line of best fit follows the trend with roughly as many points above as below.
Drawing a line of best fit by hand: it must follow the trend with roughly as many points above it as below, and it does not have to pass through the origin. It should pass through the mean point if you know it. Use a ruler, draw it right across the plotted range, and do not join the dots.
Worked in full
The scatter diagram above shows hours revised against test score for ten students. (a) Describe the correlation. (b) A student revised for 7 hours but their score was not recorded. Estimate it.
1
(a) As hours revised increases, score increases
Read the direction off the picture: the points climb from left to right.
2
The points lie close to a straight line, so the correlation is strong
Strength as well as direction.
3
Answer: strong positive correlation — students who revised longer tended to score higher
Direction, strength and context. All three are needed.
4
(b) Find 7 on the horizontal axis and go straight up to the line of best fit
Not up to a data point — the line is what you predict from.
5
Read across to the vertical axis: about 8
An estimate from a hand-drawn line, so any sensible nearby value earns the mark.
6
Answer: about 8 marks
State it as an estimate, since that is exactly what it is.
The mistake: saying correlation proves cause. Ice cream sales and sunburn cases correlate strongly, and neither causes the other — hot weather causes both. Cambridge asks this, and the answer is that correlation shows an association, not a cause.
Interpolation and extrapolation. Reading inside the range of the data is interpolation and is reliable. Reading outside it — predicting a score for 25 hours of revision — is extrapolation and is not reliable, because there is no evidence the pattern continues. If a question asks whether a prediction is reliable, that is the answer it wants.
Last step is yours
A scatter diagram of car age against value shows points falling steadily from left to right, lying close to a straight line. (a) Describe the correlation. (b) Explain what it means.
1
The points fall from left to right, so the direction is negative
As age increases, value decreases.
2
They lie close to a line, so the strength is strong
Both parts of the description.
3
(a) strong negative correlation
Two words, both needed.
4
(b) Older cars tend to be worth less
The interpretation, in the language of the question. Say “tend to” — a correlation describes a tendency, not a rule.
Last two are yours
Eight students have a mean of 5 hours revision and a mean score of 6.5. A line of best fit is drawn through (0, 3) and passes through the mean point. Use it to estimate the score of a student who revised for 8 hours.
1
The line goes through (0, 3) and (5, 6.5)
The mean point is a point the line must pass through.
2
gradient = (6.5 − 3) ÷ (5 − 0) = 3.5 ÷ 5 = 0.7
Change in score over change in hours. Watch the subtraction order — top and bottom must both run from the first point to the second.
3
equation: score = 3 + 0.7 × hours
Starting value plus gradient times hours.
4
at 8 hours: 3 + 0.7 × 8 = 3 + 5.6
0.7 × 8 = 5.6.
5
score ≈ 8.6
3 + 5.6 = 8.6. Eight hours is just outside the plotted range of the data, so this is extrapolation and should be called an estimate that may not be reliable.
All yours
A scatter diagram shows temperature against number of hot drinks sold. The points fall from left to right and lie close to a straight line. Describe the correlation in two words.
Points on a scatter diagram are spread all over the grid with no pattern. What is the correlation? Answer in one word.
A line of best fit passes through (2, 5) and (8, 17). Work out its gradient.
Using that line (gradient 2, through the point (2, 5)), estimate the value of y when x = 5.
E9.6 · low10 · Cumulative frequency — plot at the upper class boundary▼
▶  Watch: E9.6 Cumulative frequency diagrams
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

A cumulative frequency diagram answers “how many were less than or equal to this value?”. It is the only sensible way to get a median and quartiles out of grouped data, because grouped data has lost the individual values.

Building the table

Cumulative frequency is a running total. Each entry is the one before it plus the next frequency, and the last entry must equal n.

Time t (minutes)FrequencyCumulative frequencyPlot at
0 < t ≤ 1044(10, 4)
10 < t ≤ 20104 + 10 = 14(20, 14)
20 < t ≤ 301614 + 16 = 30(30, 30)
30 < t ≤ 401430 + 14 = 44(40, 44)
40 < t ≤ 50644 + 6 = 50(50, 50)
The single rule that decides these questions: plot at the UPPER class boundary, never the midpoint. The value 14 means “14 people took 20 minutes or less”, and that statement is only true at 20. Plotting at 15 would claim 14 people finished in 15 minutes or less, which nothing in the table says.
0 0 10 10 20 20 30 30 40 40 50 50 Q1 median Q3 time (minutes) — plotted at the UPPER class boundary cumulative frequency n = 50, so read across at 12.5, 25 and 37.5 — not at 12, 25 and 38.

Join the points with a smooth curve, and start it at (0, 0) — nobody took less than zero minutes, so the cumulative frequency there is 0. The curve always rises and never falls, since a running total cannot go down.

Worked in full
Use the diagram above (n = 50) to estimate the median, the quartiles and the interquartile range.
1
n = 50
Read it from the top of the curve, or from the last cumulative frequency.
2
median is at cumulative frequency 50 ÷ 2 = 25
For a cumulative frequency curve you use n ÷ 2, not (n + 1) ÷ 2. The curve is continuous, so the plus one is not needed.
3
Read across from 25 to the curve, then down: median ≈ 27 minutes
Across first, then down. Draw both lines on the diagram — Cambridge gives a mark for showing the read-off.
4
Q1 is at 50 ÷ 4 = 12.5
A quarter of the way up.
5
Read across from 12.5: Q1 ≈ 18.5 minutes
It falls in the 10 to 20 class, which agrees with the table since the running total reaches 14 at t = 20.
6
Q3 is at 3 × 50 ÷ 4 = 37.5
Three quarters of the way up.
7
Read across from 37.5: Q3 ≈ 35 minutes
In the 30 to 40 class, again matching the table.
8
IQR = Q3 − Q1 = 35 − 18.5 = 16.5 minutes
Q3 first, Q1 second. Positive, so the order was right.
The mistake: reading the median at 25 on the horizontal axis. The median position is a cumulative frequency, so it lives on the vertical axis. You go across from the vertical axis, then down to the horizontal. Getting this the wrong way round produces an answer with the wrong units, which is the tell.
Last step is yours
From the same curve, estimate how many people took longer than 35 minutes.
1
The curve gives “less than or equal to”, so start there
Go up from 35 on the time axis to the curve.
2
At t = 35 the cumulative frequency is about 37
That is the number who took 35 minutes or less.
3
longer than 35 = 50 − 37 = about 13 people
Subtract from the total. Any question phrased “more than” needs this last subtraction, and forgetting it is the most common single error in the topic.
Last two are yours
A cumulative frequency curve for 80 pupils shows a median mark of 54, Q1 = 41 and Q3 = 66. Estimate (a) the interquartile range and (b) the number of pupils scoring more than 66.
1
(a) IQR = Q3 − Q1
Definition, and the order is fixed.
2
66 − 41 = 25 marks
Positive, as a spread must be.
3
(b) Q3 = 66 means three quarters of the pupils scored 66 or less
That is what the upper quartile means, and it saves you a read-off.
4
three quarters of 80 = 60, so more than 66 is 80 − 60 = 20 pupils
A quarter of the data always lies above Q3, and a quarter of 80 is 20.
Non-calculator: the three positions you need are n ÷ 2, n ÷ 4 and 3n ÷ 4. Halve, halve again, then add the two: for n = 50, half is 25, half of that is 12.5, and 25 + 12.5 = 37.5. No long division anywhere.
All yours
A cumulative frequency curve is drawn for 120 people. At what cumulative frequency should you read across to estimate the upper quartile?
Frequencies for four classes are 5, 12, 9 and 4. What is the third cumulative frequency?
A class is 20 < t ≤ 30. At which value of t should this class be plotted on a cumulative frequency curve?
For a cumulative frequency curve with n = 200, Q1 = 34 and Q3 = 58. Work out the interquartile range.
From a curve for 60 students, 45 students scored 70 marks or less. How many scored more than 70?

Percentiles

The kth percentile is the value with k% of the data at or below it. On a cumulative frequency curve, read across from k/100 × n, then down. The median is the 50th percentile, Q1 the 25th and Q3 the 75th. “The top 10%” starts at the 90th percentile.

010203040500102030405090th ≈ 4110th ≈ 1145 = 90% of 505 = 10% of 50time t (minutes)cumulative frequency
Worked in full
Use the cumulative frequency curve above (n = 50) to estimate the 90th percentile and the 10th percentile.
1
90th percentile: 90% of 50 = 45. Read across at 45, then down
The curve reaches 45 at about 41 minutes.
2
90th percentile ≈ 41 minutes
The slowest 10% of the 50 people took more than about 41 minutes.
3
10th percentile: 10% of 50 = 5. Read across at 5: about 11 minutes
Same method at the bottom of the curve.
All yours
A cumulative frequency curve is drawn for 120 people. At what cumulative frequency do you read across for the 80th percentile?
For the curve above (n = 50), at what cumulative frequency do you read across for the 30th percentile?
E9.7 · medium11 · Histograms — unequal widths and frequency density▼
▶  Watch: E9.7 Histograms
Short hand-worked explanations for exactly this sub-topic. Opens on YouTube in a new tab. Watch one, then come straight back and try the questions below — watching without testing yourself feels like learning but is not.

Here is the one. More marks are thrown away on histograms than on anything else in the statistics topic, and the reason is a single sentence that most students never take seriously:

In a histogram, the HEIGHT of a bar is not the frequency.
The AREA of the bar is the frequency.

When every class has the same width this does not matter, because area and height are then proportional and the picture looks the same. Cambridge therefore sets histogram questions with unequal class widths, every time, precisely because that is when the distinction bites.

frequency density = frequency ÷ class width
frequency = frequency density × class width
Get the class width right first. For a class 20 < t ≤ 40 the width is 40 − 20 = 20. Not 19, not 21. For a class written as “20 to 40” on continuous data it is still 20. Where classes are written with discrete-looking boundaries such as 10–19 and 20–29, the real boundaries are 9.5 to 19.5, so the width is 10.
0 10 20 30 40 50 60 70 0 0.5 1 1.5 2 2.5 3 fd = 1.5 area = 15 fd = 2.5 area = 25 fd = 1.5 area = 30 fd = 0.5 area = 15 time t (minutes) — note the classes are NOT the same width frequency density Height is frequency density. The AREA of each bar is the frequency: 15, 25, 30, 15 — total 85.
Worked in full
85 people were timed. 0 < t ≤ 10: 15 people. 10 < t ≤ 20: 25. 20 < t ≤ 40: 30. 40 < t ≤ 70: 15. Work out the frequency density for each class.
1
Widths: 10, 10, 20, 30
Upper boundary minus lower boundary for each class. Write this column before anything else — it is the column people skip.
2
0 to 10: fd = 15 ÷ 10 = 1.5
Frequency over width.
3
10 to 20: fd = 25 ÷ 10 = 2.5
Same width, so this bar really is the tallest.
4
20 to 40: fd = 30 ÷ 20 = 1.5
Thirty people, but spread over twice the width, so the same height as the first bar. The bar with the most people is not the tallest bar.
5
40 to 70: fd = 15 ÷ 30 = 0.5
Fifteen people over a width of 30. 15 ÷ 30 = 1/2.
6
Check: areas are 1.5 × 10 = 15, 2.5 × 10 = 25, 1.5 × 20 = 30, 0.5 × 30 = 15
Each area gives back its frequency, and 15 + 25 + 30 + 15 = 85, which matches. This check takes fifteen seconds and it catches every width error.

Compare the correct histogram above with what happens if the frequencies are used as heights instead:

0 10 20 30 40 50 60 70 WRONG: frequency used as the height. The 20–40 bar now looks twice as important as it is.
The mistake: plotting frequency up the vertical axis. In the wrong diagram the 20–40 bar towers over everything because it is both tall and wide, so it claims an area of 600 people out of 85. Frequency density exists to stop exactly this. If a histogram question hands you a vertical axis already labelled “frequency density”, that is the examiner telling you which calculation is expected.
Last step is yours
A histogram has a bar covering 30 < x ≤ 45 with a frequency density of 4. How many items are in that class?
1
width = 45 − 30 = 15
Upper minus lower.
2
frequency = frequency density × class width
The formula run backwards. This is the direction Cambridge asks most often.
3
frequency = 4 × 15 = 60 items
Four per unit, over fifteen units.
Last two are yours
A histogram shows two bars. Bar A covers 0 < x ≤ 20 with height 3. Bar B covers 20 < x ≤ 50 with height 2. Work out the total frequency, and state which class contains more items.
1
Bar A width = 20, height 3
Height means frequency density.
2
Bar A frequency = 3 × 20 = 60
Area of the bar.
3
Bar B width = 50 − 20 = 30, height 2
Bar B is shorter but wider.
4
Bar B frequency = 2 × 30 = 60
The same frequency, from a shorter bar. This is the whole idea in one line.
5
total = 60 + 60 = 120, and neither class contains more — they are equal
A shorter bar can hold just as many items if it is wide enough. Judging by height alone would have said Bar A, and that is the marked-wrong answer.
Non-calculator: frequency densities are chosen to be friendly. If a division does not come out neatly, check the width before you doubt the arithmetic — a width read as 19 instead of 20 is the usual cause. And when going backwards, multiply the width by the density mentally in parts: 4 × 15 = 4 × 10 + 4 × 5 = 40 + 20 = 60.
All yours
A class 60 < m ≤ 100 has a frequency of 32. Work out the frequency density.
A class 0 < x ≤ 5 has frequency density 7. How many items are in the class?
A class 15 < y ≤ 35 contains 50 items. Work out the frequency density.
In a histogram, bar P has width 10 and height 6, bar Q has width 40 and height 2. Which bar represents more items? Answer P or Q.
A histogram has classes of widths 5, 5 and 10 with frequencies 20, 35 and 30. What is the frequency density of the third class?
mixed12 · Mixed set — no labels, no order▼

Fifteen questions from everything above, shuffled and unlabelled. Deciding which method applies is the part ordinary revision quietly does for you, and it is the part the exam tests. Before you calculate anything, answer two questions in your head: is this with or without replacement? and am I being told something has already happened? Those two decisions settle most of probability, and for statistics the equivalent is: frequency, or frequency density?

nothing answered yet
1. A bag holds 4 red and 5 green counters. Two are taken without replacement. Work out P(both red). Give a fraction in its lowest terms.
2. A class 25 < x ≤ 45 has frequency 70. Work out the frequency density.
3. Find the median of: 14, 9, 21, 6, 15, 11, 8, 19.
4. A pie chart represents 45 people. Work out the angle for a sector representing 6 people.
5. 50 students: 30 study Physics, 22 study Chemistry, 9 study both. Work out P(a student studies Chemistry, given they study Physics). Fraction in lowest terms.
6. A drawing pin is dropped 250 times and lands point down 90 times. Estimate the probability it lands point down, as a decimal.
7. Estimate the mean of this grouped table. 0 < t ≤ 20: 5. 20 < t ≤ 40: 12. 40 < t ≤ 60: 3.
8. Two fair dice are rolled. Work out P(both show the same number). Fraction in lowest terms.
9. A cumulative frequency curve is drawn for 160 people. At what cumulative frequency do you read across for the lower quartile?
10. For a data set, Q1 = 27 and Q3 = 46. Work out the interquartile range.
11. A bag holds 3 white and 7 black beads. One is taken WITH replacement, then another. Work out P(both white). Fraction in lowest terms.
12. A histogram bar covers 12 < x ≤ 20 with a frequency density of 6.5. How many items are in the class?
13. The probability a train is late is 0.15. Out of 320 trains, how many are expected to be late?
14. A bag has 5 yellow and 3 blue counters. Two are taken without replacement. Work out P(at least one blue). Fraction in lowest terms.
15. Class A has a median of 58 and an IQR of 6. Class B has a median of 58 and an IQR of 21. Which class was more consistent? Answer A or B.

Where to go next

If any of questions 1, 11 or 14 went wrong, the problem is the with-or-without-replacement decision, not the fractions — go back to section 3 and redo the two tree diagrams side by side until the difference between the two stage-2 columns is automatic.

If 2 or 12 went wrong, go back to section 11 and write the class-width column first, every time, before touching a division.

If 5 went wrong, section 4. The word to hunt for is given, and the thing it changes is the denominator.

If 7 went wrong, section 7 and the assumed-mean method, which keeps the numbers small enough to do in your head and keeps the sign errors visible.