PSAT 10 Data Analysis: Scatterplots, Standard Deviation & Sample Inference

What You Need to Know

You’ll see questions where you must read data displays and make reasonable conclusions (not just compute). This mini-unit hits three big skills:

  • Scatterplots & linear models: Describe relationships, interpret a line of best fit, and use/interpret residuals.
  • Standard deviation: Compare spread, predict how changes to data affect spread, and recognize when data are more/less variable.
  • Sample inference: Decide whether a sample result can be generalized to a population, and interpret estimates + margin of error.

Big idea: You’re often not asked to do heavy calculations; you’re asked to interpret what the numbers/graphs mean and whether a conclusion is justified.

Critical reminder: Correlation/association is not causation. A trend in a scatterplot does not prove one variable causes the other.


Step-by-Step Breakdown

A) Scatterplots: describe, model, interpret
  1. Identify variables

    • xx: explanatory (input)
    • yy: response (output)
  2. Describe the association (SOCS for scatterplots)

    • Strength: strong / moderate / weak (how tightly points cluster)
    • Outliers: points far from the pattern
    • Correlation direction: positive / negative / none
    • Shape/Form: linear vs curved
  3. If the trend is roughly linear, use the line of best fit (often provided)

    • Slope mm: “For each increase of 11 in xx, predicted yy changes by mm.”
    • Intercept bb: predicted yy when x=0x = 0 (only meaningful if x=0x = 0 makes sense in context).
  4. Make a prediction

    • Plug into y^=mx+b\hat{y} = mx + b.
  5. Compute/interpret residuals

    • Residual =y−y^= y - \hat{y}.
    • Positive residual: actual is above the line.
    • Negative residual: actual is below the line.
    • Smaller residual magnitude ∣y−y^∣|y-\hat{y}| means a better prediction.
  6. Check for extrapolation

    • Predicting outside the observed xx-range is risky and often labeled “not supported.”

Mini worked example (residual):

  • Line: y^=2x+5\hat{y} = 2x + 5
  • At x=4x = 4, predicted y^=13\hat{y} = 13.
  • If actual y=16y = 16, residual =16−13=3= 16 - 13 = 3 (point is above the line by 33 units).

B) Standard deviation: compare spread and data transformations
  1. Know what standard deviation measures

    • Standard deviation σ\sigma (or ss) describes the typical distance from the mean μ\mu (or xˉ\bar{x}).
  2. Compare two sets using center + spread

    • Same mean but larger σ\sigma ⇒ more variability.
    • Larger range doesn’t always mean larger σ\sigma, but it’s a clue.
  3. Recognize the effect of transformations

    • Add/subtract a constant: mean shifts, spread stays the same.
    • Multiply/divide by a constant: mean and standard deviation scale.

Mini worked example (transformation):

  • If you convert inches to centimeters: y=2.54xy = 2.54x.
  • Then σy=2.54σx\sigma_y = 2.54\sigma_x and μy=2.54μx\mu_y = 2.54\mu_x.

C) Sample inference: when can you generalize?
  1. Identify the population and the sample

    • Population: the full group you want conclusions about.
    • Sample: the measured subset.
  2. Check sampling method

    • Best: random sample (every member has a chance).
    • Bad: convenience/voluntary response; these create bias.
  3. State what can be inferred

    • If random + representative, you can generalize from sample to population.
    • If biased or non-random, conclusions should be limited to the sample.
  4. Interpret a margin of error (if given)

    • If an estimate is p^\hat{p} with margin of error MOE\text{MOE}, a typical conclusion is:
      plausible population proportion≈p^±MOE\text{plausible population proportion} \approx \hat{p} \pm \text{MOE}
  5. Use sample size logic

    • Larger nn ⇒ less variability in estimates (usually smaller MOE).
    • Small nn ⇒ more variability; conclusions are less precise.

Key nuance: Random assignment (in experiments) supports cause-and-effect. Random sampling supports generalizing to a population. They are different.


Key Formulas, Rules & Facts

Scatterplots & linear models
ItemFormula / RuleWhen to useNotes
Line of best fity^=mx+b\hat{y} = mx + bPredict yy from xxy^\hat{y} is predicted value, not actual
Slope meaningm=Δy^Δxm = \frac{\Delta \hat{y}}{\Delta x}Interpret rate of changeInclude units: “units of y\text{units of }y per unit of x\text{unit of }x”
Residualy−y^y - \hat{y}Check prediction errorSign tells above/below the line
ExtrapolationPredicting outside observed xxDecide if conclusion is supportedUsually not reliable
Standard deviation essentials
ItemFormula / RuleWhen to useNotes
Population SD (definition)σ=1N∑(xi−μ)2\sigma = \sqrt{\frac{1}{N}\sum (x_i-\mu)^2}If entire population givenPSAT often focuses on concept, not computation
Sample SD (definition)s=1n−1∑(xi−xˉ)2s = \sqrt{\frac{1}{n-1}\sum (x_i-\bar{x})^2}If a sample is usedThe n−1n-1 is the “degrees of freedom” correction
Shift dataIf y=x+cy = x + c then σy=σx\sigma_y = \sigma_xAdding/subtracting constantSpread does not change
Scale dataIf y=axy = ax then σy=∣a∣σx\sigma_y = |a|\sigma_xUnit conversions, rescalingSpread scales by factor ∣a∣|a|
OutliersOutliers increase σ\sigmaComparing spreadsσ\sigma is sensitive to extreme values
Sample inference basics
ItemRuleWhen to useNotes
Representative sampleRandom sample best supports generalizationSurvey inferenceConvenience samples can be biased
Estimate of proportionp^=successesn\hat{p} = \frac{\text{successes}}{n}Survey proportion questionsSometimes given directly
Informal intervalp^±MOE\hat{p} \pm \text{MOE}If MOE providedInterpret as “plausible range”
Sample size effectLarger nn ⇒ smaller typical errorCompare two surveysNot linear: doubling nn doesn’t halve error

Examples & Applications

Example 1: Describe a scatterplot (trend + outlier)

A scatterplot of xx = hours studied and yy = test score shows points rising left to right with tight clustering, but one point far below the cluster.

  • Direction: positive association.
  • Form: roughly linear.
  • Strength: strong.
  • Outlier: that low-score point (maybe illness/test anxiety).
    Exam move: Mention the outlier can weaken the model and increase prediction error.
Example 2: Interpret slope and intercept in context

Given y^=1.5x+40\hat{y} = 1.5x + 40 where xx = number of practice problems and yy = quiz score.

  • Slope 1.51.5: each additional practice problem predicts about 1.51.5 more points on the quiz.
  • Intercept 4040: predicts 4040 points when x=0x = 0 (meaningful only if taking 00 practice problems is realistic).
Example 3: Residuals + “best fit” reasoning

Two lines are proposed for the same scatterplot.

  • Line A gives residuals mostly small (like −2-2 to 22).
  • Line B gives residuals larger (like −6-6 to 66).
    Conclusion: Line A fits better because it minimizes typical vertical error ∣y−y^∣|y-\hat{y}|.
Example 4: Sample inference + margin of error

A random sample of 200200 students finds p^=0.60\hat{p} = 0.60 prefer online homework, with MOE=0.05\text{MOE} = 0.05.

  • Plausible population proportion:
    0.60±0.05⇒[0.55, 0.65]0.60 \pm 0.05 \Rightarrow [0.55,\,0.65]
  • Interpretation: It’s reasonable to believe about 55%55\% to 65%65\% of all students in that school prefer online homework.
    Trap watch: You can generalize to that school (the population sampled from), not automatically to all students everywhere.

Common Mistakes & Traps

  1. Mixing up xx and yy

    • Wrong: interpreting slope as “per unit of yy.”
    • Fix: slope is always “change in predicted yy per change in xx.”
  2. Assuming causation from a trend

    • Wrong: “Because xx and yy are correlated, xx causes yy.”
    • Fix: correlation could be coincidence or a lurking variable.
  3. Treating the intercept as automatically meaningful

    • Wrong: using bb when x=0x = 0 is outside the data range or impossible.
    • Fix: only interpret bb if x=0x = 0 is sensible and within/near observed data.
  4. Extrapolating far beyond the data

    • Wrong: predicting at xx much larger/smaller than shown.
    • Fix: predictions are safest within the observed xx interval.
  5. Residual sign confusion

    • Wrong: thinking positive residual means point is below the line.
    • Fix: residual =y−y^= y - \hat{y}; if actual yy is bigger, the point is above the line.
  6. Believing standard deviation measures “average” value

    • Wrong: interpreting σ\sigma like a typical score instead of a typical distance from the mean.
    • Fix: σ\sigma is about spread around μ\mu.
  7. Forgetting how transformations affect σ\sigma

    • Wrong: saying adding 1010 increases σ\sigma.
    • Fix: adding/subtracting constant shifts center only; multiplying scales spread.
  8. Overgeneralizing from a biased sample

    • Wrong: “A poll of volunteers proves the whole school thinks this.”
    • Fix: without random sampling (or at least strong representativeness), inference is weak.

Memory Aids & Quick Tricks

Trick / mnemonicWhat it helps you rememberWhen to use
SOCSStrength, Outliers, Correlation direction, ShapeDescribing scatterplots fast
Residual = Actual − PredictedSign and meaning of residualsAny line-of-best-fit question
Shift vs Scalex+cx+c: SD same; axax: SD scales by ∣a∣|a|Unit conversions / data changes
Random sample ⇒ GeneralizeWhen inference to population is justifiedSurvey questions
Random assignment ⇒ CauseWhen cause-and-effect is justifiedExperiment questions

Quick Review Checklist

  • You can describe a scatterplot using direction, form, strength, outliers.
  • You can interpret mm in y^=mx+b\hat{y} = mx + b as “per 11 unit increase in xx.”
  • You can compute residuals with y−y^y - \hat{y} and interpret the sign.
  • You avoid extrapolation outside the observed xx-values.
  • You know σ\sigma measures typical distance from the mean, and outliers increase σ\sigma.
  • You remember: x+cx+c doesn’t change σ\sigma; axax makes σ\sigma become ∣a∣σ|a|\sigma.
  • You only generalize from sample to population when sampling is random/representative.
  • You can interpret p^±MOE\hat{p} \pm \text{MOE} as a plausible population range (if MOE is provided).

You’ve got this—focus on interpreting what the data says and what it doesn’t allow you to claim.