PSAT 10 Data Analysis: Scatterplots, Standard Deviation & Sample Inference
What You Need to Know
You’ll see questions where you must read data displays and make reasonable conclusions (not just compute). This mini-unit hits three big skills:
- Scatterplots & linear models: Describe relationships, interpret a line of best fit, and use/interpret residuals.
- Standard deviation: Compare spread, predict how changes to data affect spread, and recognize when data are more/less variable.
- Sample inference: Decide whether a sample result can be generalized to a population, and interpret estimates + margin of error.
Big idea: You’re often not asked to do heavy calculations; you’re asked to interpret what the numbers/graphs mean and whether a conclusion is justified.
Critical reminder: Correlation/association is not causation. A trend in a scatterplot does not prove one variable causes the other.
Step-by-Step Breakdown
A) Scatterplots: describe, model, interpret
Identify variables
- : explanatory (input)
- : response (output)
Describe the association (SOCS for scatterplots)
- Strength: strong / moderate / weak (how tightly points cluster)
- Outliers: points far from the pattern
- Correlation direction: positive / negative / none
- Shape/Form: linear vs curved
If the trend is roughly linear, use the line of best fit (often provided)
- Slope : “For each increase of in , predicted changes by .”
- Intercept : predicted when (only meaningful if makes sense in context).
Make a prediction
- Plug into .
Compute/interpret residuals
- Residual .
- Positive residual: actual is above the line.
- Negative residual: actual is below the line.
- Smaller residual magnitude means a better prediction.
Check for extrapolation
- Predicting outside the observed -range is risky and often labeled “not supported.”
Mini worked example (residual):
- Line:
- At , predicted .
- If actual , residual (point is above the line by units).
B) Standard deviation: compare spread and data transformations
Know what standard deviation measures
- Standard deviation (or ) describes the typical distance from the mean (or ).
Compare two sets using center + spread
- Same mean but larger ⇒ more variability.
- Larger range doesn’t always mean larger , but it’s a clue.
Recognize the effect of transformations
- Add/subtract a constant: mean shifts, spread stays the same.
- Multiply/divide by a constant: mean and standard deviation scale.
Mini worked example (transformation):
- If you convert inches to centimeters: .
- Then and .
C) Sample inference: when can you generalize?
Identify the population and the sample
- Population: the full group you want conclusions about.
- Sample: the measured subset.
Check sampling method
- Best: random sample (every member has a chance).
- Bad: convenience/voluntary response; these create bias.
State what can be inferred
- If random + representative, you can generalize from sample to population.
- If biased or non-random, conclusions should be limited to the sample.
Interpret a margin of error (if given)
- If an estimate is with margin of error , a typical conclusion is:
- If an estimate is with margin of error , a typical conclusion is:
Use sample size logic
- Larger ⇒ less variability in estimates (usually smaller MOE).
- Small ⇒ more variability; conclusions are less precise.
Key nuance: Random assignment (in experiments) supports cause-and-effect. Random sampling supports generalizing to a population. They are different.
Key Formulas, Rules & Facts
Scatterplots & linear models
| Item | Formula / Rule | When to use | Notes |
|---|---|---|---|
| Line of best fit | Predict from | is predicted value, not actual | |
| Slope meaning | Interpret rate of change | Include units: “ per ” | |
| Residual | Check prediction error | Sign tells above/below the line | |
| Extrapolation | Predicting outside observed | Decide if conclusion is supported | Usually not reliable |
Standard deviation essentials
| Item | Formula / Rule | When to use | Notes |
|---|---|---|---|
| Population SD (definition) | If entire population given | PSAT often focuses on concept, not computation | |
| Sample SD (definition) | If a sample is used | The is the “degrees of freedom” correction | |
| Shift data | If then | Adding/subtracting constant | Spread does not change |
| Scale data | If then | Unit conversions, rescaling | Spread scales by factor |
| Outliers | Outliers increase | Comparing spreads | is sensitive to extreme values |
Sample inference basics
| Item | Rule | When to use | Notes |
|---|---|---|---|
| Representative sample | Random sample best supports generalization | Survey inference | Convenience samples can be biased |
| Estimate of proportion | Survey proportion questions | Sometimes given directly | |
| Informal interval | If MOE provided | Interpret as “plausible range” | |
| Sample size effect | Larger ⇒ smaller typical error | Compare two surveys | Not linear: doubling doesn’t halve error |
Examples & Applications
Example 1: Describe a scatterplot (trend + outlier)
A scatterplot of = hours studied and = test score shows points rising left to right with tight clustering, but one point far below the cluster.
- Direction: positive association.
- Form: roughly linear.
- Strength: strong.
- Outlier: that low-score point (maybe illness/test anxiety).
Exam move: Mention the outlier can weaken the model and increase prediction error.
Example 2: Interpret slope and intercept in context
Given where = number of practice problems and = quiz score.
- Slope : each additional practice problem predicts about more points on the quiz.
- Intercept : predicts points when (meaningful only if taking practice problems is realistic).
Example 3: Residuals + “best fit” reasoning
Two lines are proposed for the same scatterplot.
- Line A gives residuals mostly small (like to ).
- Line B gives residuals larger (like to ).
Conclusion: Line A fits better because it minimizes typical vertical error .
Example 4: Sample inference + margin of error
A random sample of students finds prefer online homework, with .
- Plausible population proportion:
- Interpretation: It’s reasonable to believe about to of all students in that school prefer online homework.
Trap watch: You can generalize to that school (the population sampled from), not automatically to all students everywhere.
Common Mistakes & Traps
Mixing up and
- Wrong: interpreting slope as “per unit of .”
- Fix: slope is always “change in predicted per change in .”
Assuming causation from a trend
- Wrong: “Because and are correlated, causes .”
- Fix: correlation could be coincidence or a lurking variable.
Treating the intercept as automatically meaningful
- Wrong: using when is outside the data range or impossible.
- Fix: only interpret if is sensible and within/near observed data.
Extrapolating far beyond the data
- Wrong: predicting at much larger/smaller than shown.
- Fix: predictions are safest within the observed interval.
Residual sign confusion
- Wrong: thinking positive residual means point is below the line.
- Fix: residual ; if actual is bigger, the point is above the line.
Believing standard deviation measures “average” value
- Wrong: interpreting like a typical score instead of a typical distance from the mean.
- Fix: is about spread around .
Forgetting how transformations affect
- Wrong: saying adding increases .
- Fix: adding/subtracting constant shifts center only; multiplying scales spread.
Overgeneralizing from a biased sample
- Wrong: “A poll of volunteers proves the whole school thinks this.”
- Fix: without random sampling (or at least strong representativeness), inference is weak.
Memory Aids & Quick Tricks
| Trick / mnemonic | What it helps you remember | When to use |
|---|---|---|
| SOCS | Strength, Outliers, Correlation direction, Shape | Describing scatterplots fast |
| Residual = Actual − Predicted | Sign and meaning of residuals | Any line-of-best-fit question |
| Shift vs Scale | : SD same; : SD scales by | Unit conversions / data changes |
| Random sample ⇒ Generalize | When inference to population is justified | Survey questions |
| Random assignment ⇒ Cause | When cause-and-effect is justified | Experiment questions |
Quick Review Checklist
- You can describe a scatterplot using direction, form, strength, outliers.
- You can interpret in as “per unit increase in .”
- You can compute residuals with and interpret the sign.
- You avoid extrapolation outside the observed -values.
- You know measures typical distance from the mean, and outliers increase .
- You remember: doesn’t change ; makes become .
- You only generalize from sample to population when sampling is random/representative.
- You can interpret as a plausible population range (if MOE is provided).
You’ve got this—focus on interpreting what the data says and what it doesn’t allow you to claim.