18 questions arranged by DBE cognitive level — scatter plots, the least-squares regression line, the correlation coefficient r, interpolation vs. extrapolation, and skewness. Work each one on paper first, then reveal the memo.
18
practice questions
4
cognitive levels
18
worked memos
100%
independently verified
How to use this bank.
Start at Level 1 and move up — don't jump to Level 4 first.
Use your own calculator's regression mode to check every \(a\), \(b\) and \(r\) value.
Always check whether an \(x\)-value is inside or outside the data's range before predicting.
Reveal the memo only after a genuine attempt.
Accuracy note: every question below was independently solved from scratch before publication, cross-checked against its own working rather than assumed correct. These are original "Equation Station SA Practice Question" items, written to match the exact CAPS scope taught in the Summary Notes for this grade.
L1 — Knowledge 4 Qs
L2 — Routine Procedures 6 Qs
L3 — Complex Procedures 4 Qs
L4 — Problem Solving 4 Qs
22%
Level 1 | Knowledge
Direct Recall
Reading a and b from a given equation, interpreting a given r value, classifying interpolation/extrapolation, and explaining why shape matters.
Q1Equation Station SA Practice Question2 marks
Regression
Reading a and b
A regression line has equation \(\hat{y}=15{,}4+2{,}3x\). State the values of \(a\) and \(b\), and explain what \(b\) means.
Memo
✓ Comparing to \(\hat{y}=a+bx\): \(\boxed{a=15{,}4}\), \(\boxed{b=2{,}3}\).✓ \(b\) is the gradient — for every 1-unit increase in \(x\), \(\hat y\) is predicted to increase by 2,3.
Q2Equation Station SA Practice Question2 marks
Correlation
Interpreting a Given r Value
A data set has \(r=0{,}15\). What does this tell you about the linear relationship between the two variables?
Memo
✓ \(r=0{,}15\) is positive but very close to 0.✓ \(\boxed{\text{This shows a very weak (almost no) positive linear relationship}}\) — the points are scattered widely with little linear pattern.
Q3Equation Station SA Practice Question1 mark
Interpolation/Extrapolation
Classifying a Prediction
A regression line is fitted to data with \(x\) ranging from 20 to 60. Is predicting \(\hat y\) at \(x=45\) interpolation or extrapolation?
Memo
✓ \(45\) lies INSIDE the original range \([20,60]\).✓ \(\boxed{\text{This is interpolation}}\), generally the more reliable type of prediction.
Q4Equation Station SA Practice Question2 marks
Scatter Plots
Why Shape Matters First
Explain why you should look at a scatter plot's overall shape before fitting a straight-line regression to it.
Memo
✓ A straight line only makes sense as a model if the data genuinely follows a roughly linear pattern.✓ \(\boxed{\text{Fitting a straight line to data that is actually quadratic or exponential would misrepresent the true relationship}}\), even if the calculator still produces numbers for it.
34%
Level 2 | Routine Procedures
One Established Method
Finding a regression equation, finding and interpreting r, making a straightforward prediction, and classifying skewness from a five-number summary.
Q5Equation Station SA Practice Question3 marks
Regression
Finding a Regression Equation (I)
Six shops' advertising spend (R hundreds, \(x\)) and weekly sales (R hundreds, \(y\)) are: \((3,20),(5,28),(7,33),(9,42),(11,48),(13,55)\). Determine the regression equation.
Memo
✓ Enter all 6 pairs into your calculator's regression mode.✓ Read off \(a\approx9{,}78\), \(b\approx3{,}49\).✓ \(\boxed{\hat{y}=9{,}78+3{,}49x}\)
Q6Equation Station SA Practice Question2 marks
Correlation
Finding and Interpreting r (I)
For the same data as Q5, determine \(r\) and interpret it.
Memo
✓ Calculator regression mode: \(\boxed{r\approx0{,}998}\)✓ Positive and extremely close to 1: \(\boxed{\text{a very strong positive linear relationship}}\) — higher advertising spend is very consistently linked to higher sales.
Q7Equation Station SA Practice Question3 marks
Regression
Finding a Regression Equation (Negative)
Five cars' age in years (\(x\)) and resale value in R thousands (\(y\)) are: \((2,88),(4,75),(6,65),(8,52),(10,40)\). Determine the regression equation.
Memo
✓ Enter all 5 pairs into regression mode.✓ Read off \(a\approx99{,}7\), \(b\approx-5{,}95\).✓ \(\boxed{\hat{y}=99{,}7-5{,}95x}\) — the negative gradient makes sense, since resale value falls as a car ages.
Q8Equation Station SA Practice Question2 marks
Interpolation/Extrapolation
A Straightforward Interpolation
Using \(\hat{y}=9{,}78+3{,}49x\) (fitted to data from \(x=3\) to \(x=13\)), predict \(y\) when \(x=8\).
Memo
✓ \(x=8\) is inside the range \([3,13]\), so this is interpolation.✓ \(\hat{y}=9{,}78+3{,}49(8)=9{,}78+27{,}92\approx\boxed{37{,}7}\)
Q9Equation Station SA Practice Question3 marks
Interpolation/Extrapolation
A Straightforward Extrapolation
Using the same equation \(\hat{y}=9{,}78+3{,}49x\) (fitted to data from \(x=3\) to \(x=13\)), predict \(y\) when \(x=25\), and comment on the reliability of this prediction.
Memo
✓ \(x=25\) is far OUTSIDE the range \([3,13]\), so this is extrapolation.✓ \(\hat{y}=9{,}78+3{,}49(25)=9{,}78+87{,}25\approx\boxed{97{,}0}\)✓ \(\boxed{\text{This prediction is less reliable}}\), since it lies well outside the range the relationship was actually observed to hold — the true pattern at that spend level is unknown.
Q10Equation Station SA Practice Question2 marks
Skewness
Skewness in a Combined Statistics Question
Eleven job applicants' scores (%) on a screening test are summarised by the five-number summary: minimum \(15\), \(Q_1=62\), median \(70\), \(Q_3=75\), maximum \(80\). Describe the skewness of this data set. (This is exactly how real NSC papers test skewness inside the Statistics and Regression question — as a short sub-part on its own separate data set, alongside the bivariate regression work.)
Memo
✓ Median to \(Q_1\): \(70-62=8\). Median to \(Q_3\): \(75-70=5\) — the median sits CLOSER to \(Q_3\).✓ Lower whisker: \(Q_1-\text{min}=62-15=47\). Upper whisker: \(\text{max}-Q_3=80-75=5\) — the lower whisker is far longer.✓ Both checks agree: \(\boxed{\text{the data is negatively (left) skewed}}\) — most applicants scored well, but a few scored much lower, stretching the lower whisker out.
22%
Level 3 | Complex Procedures
Multi-Step Reasoning
Combining regression and correlation in one question, comparing two relationships, and reasoning about shape from a description.
Q11Equation Station SA Practice Question2 marks
Correlation
Finding and Interpreting r (Negative)
For the car resale data in Q7, determine \(r\) and interpret it.
Memo
✓ Calculator regression mode: \(\boxed{r\approx-0{,}999}\)✓ Negative and extremely close to \(-1\): \(\boxed{\text{a very strong negative linear relationship}}\) — resale value decreases almost perfectly consistently as age increases.
Q12Equation Station SA Practice Question5 marks
RegressionCorrelation
Combined Regression and Correlation
A shop records daily temperature (\(^{\circ}\)C, \(x\)) and ice-cream revenue (R hundreds, \(y\)): \((18,120),(22,180),(25,210),(28,260),(30,300),(33,340)\). Determine the regression equation and \(r\), and interpret \(r\).
Memo
✓ Regression mode: \(a\approx-148{,}07\), \(b\approx14{,}73\). \(\boxed{\hat{y}=-148{,}07+14{,}73x}\)✓ \(\boxed{r\approx0{,}997}\) — positive and extremely close to 1: a very strong positive linear relationship between temperature and revenue.
Q13Equation Station SA Practice Question3 marks
Correlation
Comparing Two Relationships
Data set P has \(r=0{,}92\); data set Q has \(r=0{,}48\). Which data set's regression line gives more reliable predictions, and why?
Memo
✓ Both values of \(r\) are positive, but \(|0{,}92|\) is far closer to 1 than \(|0{,}48|\).✓ \(\boxed{\text{Data set P's regression line is more reliable}}\), since its points sit much closer to the fitted line — the linear relationship is stronger and the model fits the data better.
Q14Equation Station SA Practice Question2 marks
Scatter Plots
Judging Shape from a Description
A scatter plot of a ball's height over time starts low, rises to a maximum, then falls back down again. What shape does this data most likely follow, and why would a straight-line regression be inappropriate here?
Memo
✓ Rising to a single maximum then falling is the classic n-shaped pattern: \(\boxed{\text{quadratic}}\).✓ A straight line cannot rise and then fall — it can only ever increase, decrease, or stay flat, so fitting one here would badly misrepresent the actual relationship.
22%
Level 4 | Problem Solving
Full Bivariate Investigations
Complete multi-part questions combining every skill, and reasoning about a flawed extrapolation.
Q15Equation Station SA Practice Question7 marks
RegressionCorrelationInterpolation
Full Investigation (I)
A tutor records 6 learners' hours of extra lessons (\(x\)) and their exam mark (\(y\)): \((1,15),(3,22),(5,30),(7,35),(9,44),(11,50)\). (a) Determine the regression equation. (b) Determine \(r\) and interpret it. (c) Predict the mark for a learner who takes 6 hours of extra lessons, and comment on the reliability of this prediction.
Memo
✓ (a) Regression mode: \(a\approx11{,}58\), \(b\approx3{,}51\). \(\boxed{\hat{y}=11{,}58+3{,}51x}\)✓ (b) \(\boxed{r\approx0{,}998}\) — positive and extremely close to 1: a very strong positive linear relationship between extra lessons and exam mark.✓ (c) \(x=6\) is inside the range \([1,11]\), so this is interpolation. \(\hat{y}=11{,}58+3{,}51(6)=11{,}58+21{,}09\approx\boxed{32{,}7}\). Since this is interpolation within a very strong linear relationship, \(\boxed{\text{the prediction is highly reliable}}\).
Q16Equation Station SA Practice Question7 marks
RegressionCorrelationExtrapolation
Full Investigation (II) — With Extrapolation
A seedling's height (cm, \(y\)) is measured over 12 days (\(x\)): \((2,3),(4,7),(6,12),(8,15),(10,20),(12,24)\). (a) Determine the regression equation. (b) Determine \(r\) and interpret it. (c) Use the equation to predict the height on day 50, and explain why this prediction should not be trusted.
Memo
✓ (a) Regression mode: \(a\approx-1{,}2\), \(b\approx2{,}1\). \(\boxed{\hat{y}=-1{,}2+2{,}1x}\)✓ (b) \(\boxed{r\approx0{,}999}\) — a very strong positive linear relationship over these 12 days.✓ (c) \(x=50\) is far OUTSIDE the range \([2,12]\), so this is extrapolation. \(\hat{y}=-1{,}2+2{,}1(50)=-1{,}2+105=\boxed{103{,}8}\text{ cm}\). A seedling reaching over a metre tall this quickly is biologically unrealistic — \(\boxed{\text{plant growth naturally slows over time, so a linear pattern observed over 12 days cannot be assumed to continue for 50 days}}\).
Q17Equation Station SA Practice Question3 marks
Interpolation/Extrapolation
Spot the Flawed Reasoning
A learner fits a regression line to data on 5 runners' age (18 to 35 years) versus their marathon time, then uses the equation to predict the marathon time of an 80-year-old runner. Explain what is wrong with this approach.
Memo
✓ The original data only covers ages 18 to 35, so predicting at age 80 is extreme extrapolation — more than double the highest age actually observed.✓ \(\boxed{\text{There is no evidence the same linear relationship between age and marathon time continues to hold this far outside the observed range}}\); the true relationship for much older runners could look completely different, making this prediction unreliable.
Q18Equation Station SA Practice Question4 marks
Scatter PlotsCorrelation
Connecting r to the Scatter Plot Itself
Two scatter plots both show a rising trend. Plot X has \(r=0{,}99\); Plot Y has \(r=0{,}60\). Describe, in terms of how close the points sit to the regression line, what the visual difference between the two plots would look like.
Memo
✓ \(r\) close to 1 means the points sit VERY close to the regression line, forming a tight, almost line-like cluster.✓ \(\boxed{\text{Plot X's points would sit tightly clustered around its regression line, while Plot Y's points would be noticeably more scattered above and below its line}}\), even though both plots still show an overall rising trend.