Central tendency • grouped data • quartiles • box & whisker
Grade 12 CAPS: bivariate data and scatter plots, the least-squares regression line, the correlation coefficient r, and interpolation vs. extrapolation — the final piece of the 3-year Statistics course, built on everything Grade 10 and 11 already taught.
How two-variable (bivariate) data relates, and how to summarise, quantify and use that relationship. Your calculator does the heavy arithmetic here — your job is knowing what to ask it for, and what the answer actually means.
This is the final stop.
Central tendency • grouped data • quartiles • box & whisker
Histograms • ogives • variance & standard deviation • outliers
Bivariate data • scatter plots • regression & correlation
One week, five connected skills, all about relating two variables.
Represent bivariate data as a scatter plot and judge by eye whether it looks linear, quadratic or exponential.
Use a calculator to find the equation of the least-squares regression line.
Use a calculator to find the correlation coefficient \(r\), and interpret it in terms of strength and direction.
Use the regression equation to interpolate and extrapolate, and discuss which is more reliable.
Revise symmetric and skewed data and be ready to discuss skewness — CAPS's own wording for this topic explicitly names it alongside interpolation and extrapolation.
Two measurements per subject, plotted as one point each.
Everything in Grade 10 and 11 was UNIVARIATE — one variable per subject (just test scores, just heights). Bivariate data measures TWO variables on the SAME set of subjects — e.g. each learner's study hours (\(x\)) AND test score (\(y\)) together. Each learner contributes one point \((x,y)\).
Before calculating anything, look at the overall pattern: does it rise roughly in a straight line (linear), curve up then down or down then up (quadratic), or start slow and rocket upward (exponential)?
Plotting the points yourself first, then classifying three already-drawn patterns.
Five seedlings' daily hours of sunlight (\(x\)) and height after two weeks in cm (\(y\)) are: \((2,8),\ (4,15),\ (6,19),\ (8,25),\ (10,30)\). Plot this as a scatter plot.
For each of the three scatter plots below, state whether the pattern looks linear, quadratic or exponential.
The single straight line that best fits a linear-looking scatter plot.
\(\hat{y}=a+bx\) — \(a\) is the \(y\)-intercept, \(b\) is the gradient. The "hat" on \(\hat y\) signals this is a PREDICTED value, not an actual data point.
Enter every \((x,y)\) pair into your calculator's statistics/regression mode, then read \(a\) and \(b\) directly off the screen — CAPS does not require deriving these values by hand.
Two data sets, using your calculator's regression mode.
Seven learners' weekly study hours (\(x\)) and test scores out of 100 (\(y\)) are: \((1,42),(2,48),(3,55),(4,60),(5,68),(6,74),(7,80)\). Determine the equation of the least-squares regression line.
A different data set gives \((10,50),(15,58),(20,52),(25,65),(30,60),(35,72),(40,68)\). Determine the regression line equation.
The exam only ever requires the calculator method — but seeing the real formula once builds genuine trust in the shortcut.
\(b=\dfrac{n\sum xy-\left(\sum x\right)\left(\sum y\right)}{n\sum x^2-\left(\sum x\right)^2}\)\(\qquad a=\bar{y}-b\bar{x}\)
Find the regression line for \((1,3),(2,5),(3,6),(4,8),(5,10)\) using the formulas above, then confirm it matches your calculator.
Using the study-hours data from earlier, \((1,42),(2,48),(3,55),(4,60),(5,68),(6,74),(7,80)\), with \(b\approx6{,}39\): find \(\sigma_x\) and \(\sigma_y\), then use them to compute \(r\).
A scatter plot's points start nearly flat and then climb increasingly steeply. Which shape does this suggest?
In the regression equation ŷ = a + bx, what does b represent?
One number that says how strong and which direction the linear relationship is.
\(-1\le r\le1\), read directly off your calculator alongside \(a\) and \(b\).
Positive \(r\): as \(x\) increases, \(y\) tends to increase too. Negative \(r\): as \(x\) increases, \(y\) tends to decrease.
\(|r|\) close to 1: a strong linear relationship — the points sit close to the regression line. \(|r|\) close to 0: a weak or no linear relationship — the points are scattered widely around any line.
| \(|r|\) range | Strength | Example |
|---|---|---|
| \(0\) to \(0{,}25\) | Very weak | \(r=0{,}05\) — almost no linear relationship |
| \(0{,}25\) to \(0{,}5\) | Weak | \(r=-0{,}40\) — weak negative |
| \(0{,}5\) to \(0{,}75\) | Moderate | \(r=0{,}55\) — moderate positive |
| \(0{,}75\) to \(0{,}9\) | Strong | \(r=-0{,}82\) — strong negative |
| \(0{,}9\) to \(1\) | Very strong | \(r=0{,}98\) — very strong positive |
A perfect-fit lead-in, then the two regression examples, plus a genuinely negative case.
Find \(r\) for: \((1,10),(2,20),(3,30),(4,40),(5,50)\).
Find and interpret \(r\) for the study-hours/test-score data from the previous slide.
Find and interpret \(r\) for the second (weaker) data set from the previous slide.
A runner's speed (km/h, \(x\)) and time to finish a fixed race (minutes, \(y\)) are: \((5,95),(10,80),(15,68),(20,55),(25,42),(30,30)\). Find and interpret \(r\).
A data set has r = -0,92. What does this tell you?
Which value of r shows the WEAKEST linear relationship?
Using the regression equation to predict — and knowing when to trust the prediction.
Predicting a \(y\)-value for an \(x\) INSIDE the original data's range. Generally reliable, since the pattern is directly supported by real observed data on both sides.
Predicting for an \(x\) OUTSIDE the original range. Less reliable — the relationship might not continue the same way beyond the data you actually collected.
A direct substitution first, then the full interpolation-vs-extrapolation reasoning.
A regression line \(\hat{y}=20+5x\) was fitted to data from \(x=2\) to \(x=10\). Predict \(\hat y\) when \(x=6\).
Using \(\hat{y}=35{,}43+6{,}39x\) (fitted to data from \(x=1\) to \(x=7\)): (a) predict the score for 4,5 study hours. (b) predict the score for 15 study hours. (c) which prediction is more reliable, and why?
A full multi-part Question 1, exactly the shape a real NSC paper uses — a short skewness part on ONE data set, then the regression work on a completely different one.
The recovery times (in days) of a group of patients after a minor procedure are summarised by the five-number summary: minimum 2, \(Q_1=4\), median 5, \(Q_3=9\), maximum 20. (a) Describe the skewness of this data set. Separately, a tutor records 6 learners' number of practice quizzes completed (\(x\)) and their exam mark (\(y\)): \((2,45),(4,52),(6,58),(8,63),(10,71),(12,75)\). (b) Describe the shape of the data. (c) Determine the regression equation. (d) Determine \(r\) and interpret it. (e) Predict the mark for a learner who completes 9 quizzes, and comment on the reliability of this prediction.
A regression line is fitted to data with x ranging from 10 to 50. Predicting y at x = 30 is an example of...
Why is extrapolation generally considered less reliable than interpolation?
Work first. Open one answer only when your own line of working is complete.
\(\hat{y}=12{,}5+2{,}1x\). Answer: gradient \(b=2{,}1\); for every 1-unit increase in \(x\), \(\hat y\) increases by 2,1.
\(r=-0{,}12\). Answer: very weak, almost no linear relationship (the negative sign barely matters when \(|r|\) is this small).
Data range \(x\in[5,25]\), predict at \(x=18\). Answer: interpolation (18 is inside [5,25]).
Data range \(x\in[5,25]\), predict at \(x=40\). Answer: extrapolation (40 is outside [5,25]) — less reliable.
Fitting a straight-line regression to obviously curved (quadratic/exponential) data.
Reporting only the number for \(r\) without saying what it MEANS.
Treating an extrapolated prediction as equally trustworthy as an interpolated one.
Look at the scatter plot's shape first, every time.
Always pair \(r\)'s value with both a strength word and a direction word.
State explicitly whether a given prediction is interpolation or extrapolation.
Use the Mastery Bank to build fluency. Then take the test without notes and use the result to choose the exact slide to revisit.
18 questions by skill, from reading a regression equation through to a full multi-part bivariate investigation. Answers reveal only after your attempt.
Next step 2A short exam-style self-check. Mark it, then revisit the exact slide that matches a missed skill.
Four free, independent videos — not made by Equation Station SA.
Plotting bivariate data and reading its shape.
Khan Academy · Constructing a scatter plot
What "least-squares" actually means, visually.
Eugene O'Loughlin · How To... Perform Simple Linear Regression by Hand
Building intuition for what different r values look like.
The Organic Chemistry Tutor · Correlation Coefficient
When a prediction can be trusted, and when it can't.
Khan Academy · Example estimating from regression line
The core teaching is above. These are the next steps, not a replacement for it.
18 questions by skill, with concise reveal answers and methods.
Use the short exam-style self-check when you want a fast confidence check.
Use its own worked examples for extra explanation and exercises.
Official state-owned learner books and teacher support for Grade 12 Mathematics.
This topic began in Grade 10 and Grade 11.
Short answers for the checks learners make while preparing for the Grade 12 CAPS exam.
Represent bivariate (two-variable) data as a scatter plot and judge whether it looks linear, quadratic or exponential; use a calculator to find the least-squares regression line and the correlation coefficient r; interpret r in terms of the strength and direction of the relationship; and use the regression equation to interpolate and extrapolate, discussing the reliability of each; and revise symmetric and skewed data, ready to describe skewness from a five-number summary or box plot.
No — CAPS Grade 12 expects you to use your calculator's regression (linear regression / LR) mode to find the equation and r directly. You do need to know what the values mean and how to use them, not just how to compute them from scratch.
Interpolation means predicting a y-value for an x-value INSIDE the range of the original data — this is generally reliable. Extrapolation means predicting outside that range, which is far less reliable since the pattern may not continue.
r is always between -1 and 1. Values close to 1 or -1 show a strong linear relationship (positive or negative); values close to 0 show a weak or no linear relationship. The sign tells you the direction: positive means y increases as x increases, negative means y decreases as x increases.
Finish the interactive slides, open the Mastery Bank, then take the short self-test without notes.