Topic 5: Statistics
Three 60-minute lesson plans on measures of central tendency and dispersion, especially standard deviation, pitched at DP AA (SL/HL), MYP Grade 10, and IGCSE Grade 10.
A. DP Mathematics: Analysis & Approaches — SL/HL, Year 2
Caveat: content below is drawn from the pre-publication draft syllabus (first assessment 2029, Topic D: Probability & Statistics, subtopic D1). Verify against the final published guide once released.
Anchor concept: Measures of central tendency and dispersion, including standard deviation, for grouped/continuous data (syllabus ref. D1). Taught in Year 2 per SOW_DP2_AAHL, after Calculus is completed.
Prior knowledge assumed: Mean/median/mode and basic dispersion measures for ungrouped data (MYP5/IGCSE); use of GDC statistical mode.
Learning objectives
- SL+HL: Calculate mean, median, mode, and standard deviation for both discrete and grouped continuous data, using GDC statistical functions and by understanding the underlying formula.
- SL+HL: Identify outliers using standard criteria and interpret the effect of an outlier on the mean vs the median.
- HL only: Derive and apply the variance formula (\sigma^2=\frac{\sum(x-\bar x)^2}{n}) explicitly (rather than reading it off the GDC), and extend to finding the variance of a discrete random variable (previews D3).
Command terms used
Calculate (mean, SD, from raw or grouped data); Determine (whether a value is an outlier); Comment on (the effect of an outlier); Derive (HL — the variance formula).
Assessment focus (2029 draft AOs)
Computation — accurate use of GDC statistical functions for grouped data; Interpretation — explaining what a given standard deviation means about a dataset’s spread in context, and comparing two datasets by their means and SDs.
ATL / skills focus
Thinking (comparing distributions using more than one measure at once — mean alone can mislead); communication (writing a comparative statistical statement, e.g. “Group A had a higher mean but also higher variability than Group B”).
Starter (≈10 min)
Give two small datasets (e.g. two classes’ test scores) with the same mean but visibly different spread. Ask: “These have the same average — does that mean the two classes performed identically? Why or why not?”
Main teaching sequence (≈35 min)
- Recap mean/median/mode for ungrouped data; extend to grouped/continuous data using midpoints and GDC frequency-table input.
- Introduce standard deviation as a measure of spread; use GDC to compute it for the starter’s two datasets and confirm the intuition that “same mean, different spread” is now quantified.
- Outliers: introduce a standard criterion (e.g. more than 1.5×IQR beyond the quartiles, or a given (z)-score threshold) and discuss why the mean is more sensitive to outliers than the median using a worked example (add one extreme value to a dataset and recompute both).
- HL-only extension (last ~10 min): derive the variance formula from the definition of “average squared deviation from the mean,” and apply it to find the variance of a simple discrete random variable (e.g. a fair die), previewing D3. SL students instead complete extended practice comparing multiple grouped datasets using mean and SD together.
Formative assessment / check for understanding
Have students predict, before calculating, which of the starter’s two datasets will have the larger SD, then verify — checks the conceptual link between “visual spread” and “numerical SD” before the mechanics take over.
Plenary / exit ticket (≈10 min)
“Two machines fill bottles with a target volume of 500ml. Machine A: mean 500ml, SD 2ml. Machine B: mean 502ml, SD 8ml. Which machine would you choose for consistent filling, and why?” (Genuine Interpretation task requiring both measures together.)
Resources
GDC statistical mode for grouped-data mean/SD (see Use of GDC in 2026.pdf); Haese AA SL textbook
chapter M09 (Statistics) under IB DP Mathematics/PPT of SL Mathematics/.
Vertical link note
Builds on MYP5’s box-plot/cumulative-frequency work and introductory standard deviation (below); the explicit variance derivation and discrete random variable link (HL) go beyond both MYP and IGCSE, connecting forward to DP’s probability distributions unit.
B. MYP Mathematics — Grade 10 (MYP Year 5)
Anchor concept: Statistical measures and representation for continuous data — box plots, cumulative frequency, and an introduction to standard deviation (Reasoning with Data branch, MYP5 Standard/Extended).
Prior knowledge assumed: Mean/median/mode for discrete data; basic frequency tables and bar charts (MYP3/4).
Learning objectives
- Construct and interpret a box-and-whisker plot from a dataset, including identifying quartiles and the interquartile range (IQR).
- Construct a cumulative frequency diagram for grouped continuous data and use it to estimate the median and quartiles.
- (Extended) Calculate the standard deviation of a dataset using GDC/technology and interpret it as a measure of spread alongside the mean.
Assessment criteria addressed
Criterion A (Knowing and understanding) — correct construction/reading of statistical diagrams; Criterion C (Communicating) — interpreting and comparing distributions using correct terminology (median, IQR, spread).
Command terms used
Construct (a box plot or cumulative frequency diagram); Describe (the shape/spread of a distribution); Compare (retained in MYP even though removed from the DP 2029 glossary); Interpret.
ATL / skills focus
Communication skills (moving between a raw dataset, a diagram, and a written description of what it shows); thinking skills (selecting an appropriate diagram type for a given dataset).
Starter (≈10 min)
Give students a set of exam scores as a simple list. Ask them to find the median and range quickly, then show a box plot of the same data and ask what extra information it reveals that median/range alone didn’t (the quartiles/IQR).
Main teaching sequence (≈35 min)
- Formalise box plots: five-number summary, constructing one from raw data, reading off IQR as a spread measure.
- Move to grouped continuous data (e.g. heights of 30 students in class intervals); construct a cumulative frequency table and diagram, and use it to estimate median and quartiles graphically.
- Compare two datasets side by side using box plots (e.g. two class’s test scores) and have students write 2–3 sentences comparing centre and spread — the C-criterion focus.
- Extended-level addition: introduce standard deviation via GDC for the same datasets, and discuss how it complements (or sometimes disagrees with) the IQR as a spread measure.
Formative assessment / check for understanding
Check students’ five-number summaries before they draw their box plots — an error here propagates through the whole diagram, so it’s worth catching early.
Plenary / exit ticket (≈10 min)
“Here are box plots for two classes’ test scores. Write two sentences comparing their performance, referring to both the median and the spread.” (Direct Criterion C task.)
Resources
Squared paper or GDC/spreadsheet for box plot and cumulative frequency construction; sample grouped datasets.
Vertical link note
This is the direct precursor to the DP D1 work above, which formalises standard deviation (including its formula and variance at HL) rather than treating it as a GDC output only; it extends IGCSE’s grouped-data mean/SD work (below) by adding cumulative frequency and box-plot representation.
C. IGCSE Mathematics — Grade 10 (Extended tier)
Caveat: IGCSE content below follows standard Cambridge (0580) / Edexcel International GCSE (4MA1) Extended-tier syllabus content. No official board specification file was found on this machine — cross-check against the current specification before formal use.
Anchor concept: Mean, median, mode, and standard deviation from grouped and continuous data.
Prior knowledge assumed: Mean/median/mode for discrete/ungrouped data; frequency tables.
Learning objectives
- Calculate an estimate of the mean from a grouped frequency table using midpoints.
- Identify the class containing the median and the modal class for grouped data.
- Calculate the standard deviation of a dataset using a calculator, and interpret it as a measure of spread.
Assessment objectives addressed
AO1 (recall and apply standard techniques — mean/SD calculation); AO3 (solve problems in context — comparing two datasets using calculated statistics).
Command terms used
Calculate, Estimate (mean from grouped data is explicitly an estimate, since exact raw values are unknown), State (modal class, median class).
Starter (≈10 min)
Give a small grouped frequency table (e.g. class test scores in intervals of 10). Ask students to identify the modal class and the class containing the median by inspection, before any calculation.
Main teaching sequence (≈35 min)
- Formalise estimating the mean from grouped data: use interval midpoints × frequency, divided by total frequency — work through the starter’s table as a full worked example.
- Identify modal class and median class formally, distinguishing these from a single modal value or exact median (which grouped data can’t give exactly).
- Introduce standard deviation calculation via calculator statistical mode for the same grouped dataset; interpret the result as “typical distance from the mean.”
- Comparison task: two grouped datasets (e.g. two shops’ daily sales) — students calculate mean and SD for both and write a comparative statement.
Formative assessment / check for understanding
Spot-check midpoint values before students proceed to the full mean calculation — a common error point (using interval boundaries instead of midpoints).
Plenary / exit ticket (≈10 min)
“The table shows the time (in minutes) 40 students took to complete a task, in grouped intervals. Calculate an estimate of the mean and the standard deviation, and state the modal class.”
Resources
Past-paper style grouped-frequency statistics questions (Extended tier); scientific/graphic calculator with statistical mode.
Vertical link note
Provides the grouped-data mean/SD mechanics that MYP5 extends with box plots and cumulative frequency representation, and that DP AA formalises further with the explicit variance formula and (HL) its link to discrete random variables.