On the Foundations Check, statistics, probability and data came back secure to step 6 and broke at step 7. Step 7 is the hardest Extended rung. Nothing below it broke. Read that plainly: this is one of your stronger areas, and it is one of only two strands where you got all the way to the top rung before anything went wrong.
So this guide is not a rebuild. It is a finish. The basics get a quick, honest pass — you can already do them, and skipping them entirely would be a mistake since they are worth easy marks — but most of the page is aimed at the three places step 7 actually lives:
One more thing from the check lands here. Sign errors recurred across several strands. In statistics
they have exactly two homes: range = largest − smallest and
IQR = Q3 − Q1. Both are subtractions where the order is fixed and a reversed order gives a
negative answer that is always wrong. A spread can never be negative. That is your check.
| Code | Sub-topic | Expected difficulty for you |
|---|---|---|
| E8.1 | Introduction to probability | Medium — notation, not ideas |
| E8.2 | Relative and expected frequencies | Medium |
| E8.3 | Probability of combined events | High — with vs without replacement |
| E8.4 | Conditional probability | High — the step-7 gap |
| E9.1 | Classifying statistical data | Low |
| E9.2 | Interpreting statistical data | Low — but two comparisons, not one |
| E9.3 | Averages and measures of spread | High — estimated mean by hand |
| E9.4 | Statistical charts and diagrams | Medium |
| E9.5 | Scatter diagrams | Low |
| E9.6 | Cumulative frequency diagrams | Low, once the plotting point is right |
| E9.7 | Histograms | Medium — frequency density is the whole topic |
Paper 2 is non-calculator: 2 hours, 100 marks, 50% of your grade. Every number on this page is chosen so it can be done by hand, and every method shown is a by-hand method.
Number sense and the four operations: solid on all seven rungs. Not one gap, right up to the
Extended-hard level. That matters more in this topic than in any other, because probability and statistics
are almost entirely arithmetic wearing a hat. Fractions multiplied along tree branches, fractions added
between them, a column of f × x products summed, a division at the end — that is the
whole of it, and you can already do all of those.
So when a statistics question goes wrong for you, it will not be the arithmetic. It will be one of two things: reading the question (with or without replacement, given or not given, frequency or frequency density), or the plotting rule (upper class boundary, midpoint, class width). Both are things you can decide before you calculate anything. That is why this guide spends so much time on the sentence before the sum.
Every idea appears four times, with less help each time.
Section 12 is a mixed set: fifteen questions, unlabelled and out of order, because in an exam nobody tells you whether you are looking at a histogram question or a conditional probability question.
Probability is a number, and the number always lives between 0 and 1. Never 1 in 4, never 25 out of 100 written as a ratio, never a percentage unless the question asks for one — a fraction, a decimal or a percentage, but a single number on this scale.
| Written | Means | Rule |
|---|---|---|
P(A) | the probability that A happens | 0 ≤ P(A) ≤ 1 |
P(A′) or P(not A) | the probability A does not happen | P(A′) = 1 − P(A) |
| a full outcome list | every possible result, once each | all the probabilities add to 1 |
For one event with equally likely outcomes:
The word equally likely is doing real work there. If the outcomes are not equally likely — a biased spinner, a weighted dice — you cannot count, you must be given or must estimate the probabilities. That is section 2.
Everything in section 1 assumed the outcomes were equally likely. Real objects are not always fair — a drawing pin, a bent coin, a plastic cone — and for those you cannot count outcomes. You have to experiment and estimate.
Relative frequency is an estimate of probability. Cambridge is strict about the word: if the question says “estimate the probability”, it wants a relative frequency, and if it says “the spinner is fair”, it wants a counted probability. The more trials, the better the estimate — that sentence is itself a mark in a lot of papers.
| word | what it means |
|---|---|
| fair | every outcome is equally likely: a fair four-sided spinner lands on each number with probability ¼ |
| biased | some outcomes are more likely than others. You can only detect bias from many trials, by comparing relative frequencies with the fair probabilities |
| random | chosen with no pattern, so that every item has the same chance of being picked |
a/b into a decimal by hand, scale the
denominator to 10, 100 or 1000 rather than doing long division. 40/250 → multiply by 4 → 160/1000
= 0.16. 175/500 → multiply by 2 → 350/1000 = 0.35. It works whenever the denominator is built from
2s and 5s, which in exam questions it almost always is.Two events, one question. There are only three tools, and choosing the right one takes about five seconds:
| Tool | Use it when |
|---|---|
| Sample space diagram | Two things happen at once and every outcome is equally likely — two dice, a coin and a spinner. Small and countable. |
| Venn diagram | You are told how many are in each category and how many overlap. The words both, either, neither. |
| Tree diagram | Two or more stages, one after the other. The words then, followed by, a second counter. |
A Venn diagram is a counting picture. The rectangle is everyone, each circle is one category, and the overlap belongs to both. Fill the overlap first — that single habit fixes most Venn errors, because the numbers you are given usually include the overlap and the numbers you write must not double-count it.
A tree is for stages. Each branch carries a probability; each complete path from left to right is one outcome. With replacement means the first object goes back before the second is taken, so the bag is identical at stage 2 — the second set of branches is a copy of the first.
This is the distinction that gets examined more than anything else in probability. Without replacement means the first object is kept, so at stage 2 there is one fewer object in the bag — and if the first was red, there is also one fewer red.
Two things change at stage 2, and both must change: the denominator drops by one on every branch, and the numerator drops by one only on the branch matching what you already took.
This is where the check said step 7 broke, and it is the one genuinely new idea in the whole probability topic. Everything else is counting. This is about what the word “given” does to the total.
Conditional probability answers: we already know B happened — now what is the chance of A?
The notation is P(A | B), read as “the probability of A given B”. It is a handy
shorthand for your own working, but 0580 does not require it: in an answer you can just as well write P(French, given a girl).
Same idea, different picture. Go back to the 30 students: 12 hockey only, 6 both, 8 netball only, 4 neither.
A tree is already conditional. Every probability on the second set of branches is a conditional probability — that is exactly what “without replacement” was doing in section 3. When a question asks for a conditional probability from a tree, you are usually being asked to read a stage-2 branch directly, or to divide one path by a group of paths.
This sub-topic is worth few marks and takes little time, but the marks are free and Cambridge does ask for the words. Two jobs: classify the data, and organise it into a table.
| Word | Means | Example |
|---|---|---|
| Qualitative | a description, not a number | eye colour, favourite sport |
| Quantitative | a number | height, number of pets |
| Discrete | counted — only certain values possible | number of pets: 0, 1, 2, never 1.4 |
| Continuous | measured — any value in a range | height: 162.37 cm is perfectly possible |
Discrete and continuous are both kinds of quantitative data. The test for continuous is simple: could a more accurate instrument give you another decimal place? Height, mass, time and length always can. Shoe size cannot — there is no shoe size 7.318.
A tally table turns a list into a frequency table. The fifth mark goes diagonally across the previous four, so the marks read in fives and can be counted at a glance.
| Colour | Tally | Frequency |
|---|---|---|
| Red | |||| ||| | 8 |
| Blue | |||| |||| | 10 |
| Green | |||| | 4 |
| Yellow | |||| |||| || | 12 |
| Total | 34 |
The total row is not decoration. It is the check: if the frequencies do not add to the number of items you were given, something has been miscounted, and finding that out now costs nothing.
A two-way table classifies by two things at once — that is what makes it the natural home for conditional probability, as in section 4. Fill it by using the totals: every row adds across, every column adds down, and the corner checks both.
Reading a chart is easy. The mark that gets dropped is the comparison, and it is dropped in the same way by almost everybody: they give one comparison when Cambridge wants two.
range = largest − smallest. Written the
other way round, 0 − 12 = −12, and a negative spread is impossible. Same for
IQR = Q3 − Q1. Before you write either answer down, glance at it: is it positive?
That single glance catches the error you make most often, in the one place in statistics where it can
happen.Before you write a conclusion, ask five questions. Each one is a common exam trap:
| question | why it matters |
|---|---|
| Is the sample big enough, and was it chosen at random? | a small or biased sample may not represent everyone |
| Are the totals equal? | if not, compare proportions, not counts |
| Does the vertical axis start at 0? | a broken axis makes small differences look large |
| Does a correlation show that one thing causes the other? | no: both may be caused by something else |
| Is the prediction outside the range of the data? | extrapolating beyond the data is unreliable |
The biggest sub-topic in the whole of statistics, and the one where step 7 sits. Four averages and two spreads, and then the piece that separates grades: estimating the mean from a grouped frequency table.
| Measure | How | Watch out for |
|---|---|---|
| Mean | add them all, divide by how many | dragged by extreme values |
| Median | put in order, then take the middle | forgetting to order first |
| Mode | the most common value | there can be two, or none |
| Range | largest − smallest | uses only two values, so one outlier ruins it |
| Q1, Q3 | medians of the lower and upper halves | whether to include the median itself |
| IQR | Q3 − Q1 | the middle half only — ignores outliers, which is the point |
(n + 1) ÷ 2. If n = 9 that is position 5, a real value. If
n = 8 that is position 4.5, meaning halfway between the 4th and 5th — so add those two and halve.
Quartiles: Q1 is the middle of the lower half and Q3 the middle of the upper half. With an odd number of values, leave the median out of both halves; with an even number, the list splits exactly in two. Some books use positions instead, Q1 at (n + 1) ÷ 4 and Q3 at 3(n + 1) ÷ 4: with an odd number of values this gives the same answer, with an even number it can differ a little. Both answers are accepted here; use the one your teacher uses.When data is grouped you no longer know the individual values, only the class each one fell in. So you cannot find the true mean. You estimate it, by assuming every value in a class sits at the class midpoint. The word “estimate” in the question is not politeness; it is telling you which method to use.
In a grouped table with equal class widths, the modal class is the class with the highest frequency. The answer is the whole interval, such as 160 < h ≤ 170: not its frequency, and not a single number.
| height h (cm) | 140 < h ≤ 150 | 150 < h ≤ 160 | 160 < h ≤ 170 | 170 < h ≤ 180 |
|---|---|---|---|---|
| frequency | 4 | 11 | 18 | 7 |
Cambridge does not usually ask you to draw a whole chart from nothing. It asks you to complete one, read one, or calculate one number needed to draw one. The calculations are short; the marks go to accuracy and labelling.
Bars have equal widths and gaps between them, and the height is the frequency. That last sentence is only true for bar charts. In a histogram the height is not the frequency, and confusing the two is section 11.
In a stacked bar the blocks sit on top of each other, so the top of the bar is the total and each block is read as a difference between two positions, not as a height off the axis. In a dual bar (section 6) they sit side by side and each is read off the axis directly.
A pie chart shows proportion, not amount. The whole circle is 360° and the whole data set, so:
360 ÷ total first and
see whether it is a whole number. Exam totals are chosen so that it usually is: 36 gives 10°, 30 gives
12°, 40 gives 9°, 24 gives 15°, 60 gives 6°, 90 gives 4°, 120 gives 3°. If it is not
whole — say a total of 50, giving 7.2° — multiply first and divide last:
frequency × 360 ÷ 50, keeping whole numbers as long as possible.A pictogram uses a symbol to stand for a number of items, and part-symbols for the remainder. The key is the whole question. If one circle stands for 8 people, then half a circle is 4 and a quarter is 2. Reading a pictogram without checking the key is how a correct count becomes a wrong answer.
A stem-and-leaf keeps every original value while showing the shape of the data, which is why you can find the exact median and range from one. The stem is the tens digit, the leaves are the units, and the leaves must be in order and evenly spaced.
Read the key first. If one symbol stands for 8 books, half a symbol is 4 books and a quarter is 2. To draw one: frequency ÷ the key value = the number of symbols, using part-symbols for what is left over.
A frequency table lists each value with how often it occurs. Read the wording carefully: “at least 7” includes 7; “more than 7” does not. It is drawn as a bar chart with the values along the bottom.
| score | 4 | 5 | 6 | 7 | 8 | 9 |
|---|---|---|---|---|---|---|
| frequency | 3 | 5 | 9 | 6 | 4 | 3 |
A scatter diagram plots two measurements for each individual and asks whether they are related. There are only three jobs: plot, describe, draw a line of best fit and use it.
Describing correlation takes two words, not one: a direction and a strength.
| Direction | Means |
|---|---|
| Positive | as one goes up, so does the other |
| Negative | as one goes up, the other goes down |
| Zero / none | no pattern — the points are scattered |
Strength is strong if the points sit close to a straight line, weak if they are loosely grouped around it. And say what it means: “strong positive correlation — students who revised for longer tended to score higher”.
A cumulative frequency diagram answers “how many were less than or equal to this value?”. It is the only sensible way to get a median and quartiles out of grouped data, because grouped data has lost the individual values.
Cumulative frequency is a running total. Each entry is the one before it plus the next frequency, and the last entry must equal n.
| Time t (minutes) | Frequency | Cumulative frequency | Plot at |
|---|---|---|---|
| 0 < t ≤ 10 | 4 | 4 | (10, 4) |
| 10 < t ≤ 20 | 10 | 4 + 10 = 14 | (20, 14) |
| 20 < t ≤ 30 | 16 | 14 + 16 = 30 | (30, 30) |
| 30 < t ≤ 40 | 14 | 30 + 14 = 44 | (40, 44) |
| 40 < t ≤ 50 | 6 | 44 + 6 = 50 | (50, 50) |
Join the points with a smooth curve, and start it at (0, 0) — nobody took less than zero minutes, so the cumulative frequency there is 0. The curve always rises and never falls, since a running total cannot go down.
The kth percentile is the value with k% of the data at or below it. On a cumulative frequency curve, read across from k/100 × n, then down. The median is the 50th percentile, Q1 the 25th and Q3 the 75th. “The top 10%” starts at the 90th percentile.
Here is the one. More marks are thrown away on histograms than on anything else in the statistics topic, and the reason is a single sentence that most students never take seriously:
When every class has the same width this does not matter, because area and height are then proportional and the picture looks the same. Cambridge therefore sets histogram questions with unequal class widths, every time, precisely because that is when the distinction bites.
20 < t ≤ 40 the width is 40 − 20 = 20. Not 19, not 21. For a class written as
“20 to 40” on continuous data it is still 20. Where classes are written with discrete-looking
boundaries such as 10–19 and 20–29, the real boundaries are 9.5 to 19.5, so the width is 10.Compare the correct histogram above with what happens if the frequencies are used as heights instead:
Fifteen questions from everything above, shuffled and unlabelled. Deciding which method applies is the part ordinary revision quietly does for you, and it is the part the exam tests. Before you calculate anything, answer two questions in your head: is this with or without replacement? and am I being told something has already happened? Those two decisions settle most of probability, and for statistics the equivalent is: frequency, or frequency density?
If any of questions 1, 11 or 14 went wrong, the problem is the with-or-without-replacement decision, not the fractions — go back to section 3 and redo the two tree diagrams side by side until the difference between the two stage-2 columns is automatic.
If 2 or 12 went wrong, go back to section 11 and write the class-width column first, every time, before touching a division.
If 5 went wrong, section 4. The word to hunt for is given, and the thing it changes is the denominator.
If 7 went wrong, section 7 and the assumed-mean method, which keeps the numbers small enough to do in your head and keeps the sign errors visible.