1/140
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is the difference between a population and a sample?
Population = the COMPLETE collection of people/objects you want to study (e.g., all US college students). Sample = the subset actually selected and analyzed to learn about the population (e.g., 796 surveyed students).
What is a parameter vs. a statistic? (definition + notation)
Parameter = a number describing a POPULATION (usually unknown; needs every member). Statistic = a number calculated from a SAMPLE, used to ESTIMATE a parameter. Notation (statistic / parameter): proportion p-hat / p; mean x-bar / mu (μ); correlation r / rho (ρ). Memory trick: Statistic–Sample, Parameter–Population.
Example: A survey of 796 college students finds 288 (36.2%) binge drank last month. Identify the population, sample, statistic, parameter.
Population = all US college students. Sample = the 796 surveyed. Statistic = p-hat = 288/796 = 0.362 (36.2%). Parameter = p = the true proportion of ALL college students who binge drink (unknown).
Example: Statistic or parameter? (a) Average age of ALL employees at a company is 35. (b) 80% of students enrolled at a certain college are full time. (c) A poll of 57 respondents shows support for a school band issue. (d) 500 high-school students surveyed: 60% intend to go to college.
(a) Parameter (every employee). (b) Parameter (all enrolled students). (c) Statistic (poll = sample). (d) Statistic (sample survey).
Example: Give the notation. (a) 2010 Census: 50,325,523 of 308,745,538 people identify as Hispanic/Latino. (b) Random sample of 300 Colorado citizens: 62 identify as Hispanic/Latino.
(a) Census = whole population, so it's a PARAMETER: p = 50,325,523 / 308,745,538. (b) Sample, so it's a STATISTIC: p-hat = 62/300 ≈ 0.207.
What is a census? Why do we usually sample instead?
A census is a 'sample' that includes the ENTIRE population (e.g., the US census every 10 years). It's usually too costly/slow/impractical, so we sample and use statistics to make inferences about parameters.
What is a representative sample, and what is a probability sampling plan?
Representative sample = resembles the population across all characteristics, so the statistic accurately reflects the parameter. Probability sampling plan = a method of choosing a sample using some form of RANDOM selection (SRS, systematic, stratified, cluster, multistage).
What are the two main types of variables?
Categorical = responses are grouped into categories/labels (word answers; also numbers that are just labels, like zip code). Quantitative = real numbers where differences/arithmetic are meaningful (income, population size). Test: Is it a word/label (categorical) or a measured/counted number (quantitative)?
Example: Categorical or quantitative? (1) Hours slept last night (2) Favorite color (3) Money spent yesterday (4) Do you take the bus to campus?
(1) Quantitative (2) Categorical (3) Quantitative (4) Categorical (yes/no).
Nominal vs. ordinal categorical variables?
Nominal = categories with NO meaningful order (college major, hair color, blood type, occupation, nationality, state). Ordinal = categories where the ORDER matters (final letter grades, rain level: light/moderate/heavy; poor/fair/good ratings).
Discrete vs. continuous quantitative variables?
Discrete = takes a set/countable number of values (# of home runs, household size, # of students). Continuous = can take ANY value in an interval (waiting time, weight, distance, age, speed, heart rate). Quick test: can you count it (discrete) or measure it (continuous)?
Draw the 'variable family tree.'
Variable → Categorical (Nominal | Ordinal) or Quantitative (Discrete | Continuous).
What are explanatory vs. response variables?
Explanatory variable (x) = the variable that might AFFECT/explain the other. Response variable (y) = the outcome that might change as a result. Example: median household income (explanatory) → county population (response).
Define cases and variables; what is a data matrix?
Data matrix = a table to organize data: each ROW is a case (the entity measured: respondent, subject, experimental unit), each COLUMN is a variable (a characteristic recorded).
What is a Simple Random Sample (SRS)? Steps? Example?
SRS = probability sampling plan where EVERY member of the population has an equal chance of being selected (like a lottery). Steps: define the population → decide sample size → randomly select (random number generator) → collect data. Example: pick 100 of 974 bridges: assign each bridge a number (0–973) and use a random number generator to choose 100 numbers.
What is a Systematic Random Sample? Steps? Example?
Systematic sample = choose every kth member by a predetermined rule. Steps: define population → decide sample size → compute interval k = population size ÷ sample size → pick a random starting point and take every kth member. Example: 100 of 974 bridges → k = 974/100 ≈ 9.7 → use 9; select every 9th bridge (after a random start) until you have 100.
What is a Stratified Random Sample? Example?
Population is first divided into groups of SIMILAR individuals (strata); then a simple random sample is taken from EACH stratum. Example: to estimate support for the Supreme Court, split the US into Democrats, Republicans, Independents, then take an SRS from each group. (Every group is guaranteed representation.)
What is a Cluster Sample? Example?
Population is divided into groups (clusters) that are each a mini-version of the population; take a random sample of CLUSTERS; EVERY member of each chosen cluster is in the sample. Example: to learn what freshmen think of campus food, list all freshman dorms, randomly select a few dorms, survey ALL freshmen in those dorms.
Stratified vs. cluster sampling — what is the difference? (Boston cream pie analogy)
Stratified: groups are DIFFERENT from each other (strata) → sample from EVERY group. Cluster: groups are SIMILAR to each other (each a mini-population) → sample only SOME groups, but take everyone in them. Pie: stratified = take a bite of each layer (cake, custard, frosting); cluster = randomly pick whole slices and eat all layers of each. Rule: sample from every group = stratified; take all of some groups = cluster.
What is multistage sampling?
Sampling that COMBINES several methods (no set recipe). Example: randomly select clusters, then randomly sample individuals within the selected clusters. Helps be representative and avoid bias.
What is bias in sampling?
Any SYSTEMATIC failure of a sampling method to represent its population (some groups over-represented, others under-represented).
What is a voluntary response sample? Examples? Why biased?
Large group invited to respond; only those who CHOOSE to respond are counted (call-in shows, internet polls). Biased because people with strong opinions are far more likely to respond.
What is a convenience sample? Examples?
Sample of whoever is EASIEST to reach (mall surveys, standing outside Starbucks, convenient internet polls). Not representative of the whole population.
What is a bad sampling frame? What is undercoverage?
Sampling frame = the list you sample from. If it's incomplete, individuals are excluded → bias (e.g., a survey from a landlord's resident list only reaches residents who rent from that landlord). Undercoverage = part of the population is not sampled or is under-represented relative to its true share (e.g., a study-hours survey that leaves out commuter students).
What is nonresponse bias?
Bias introduced when a large fraction of those sampled fail to respond (and responders differ from non-responders). Fix: follow up / reduce nonresponse.
What is response bias? Examples?
Anything in the survey design that influences the responses: respondents tailoring answers to please the interviewer; unwillingness to reveal personal/sensitive facts; leading or poorly worded questions.
Open vs. closed survey questions (pros/cons)?
Open = respondent can answer anything (Pro: more insight; Con: slow to answer and analyze). Closed = fixed choices (Pro: easy to answer/analyze; Con: limited, no choice may be perfectly accurate). Example: 'How old are you?' (open, exact age) vs. age ranges 18–24, 25–34… (closed).
Observational study vs. experiment?
Observational study: researchers merely OBSERVE data that arise; no treatment is imposed. Experiment: researchers deliberately IMPOSE treatment(s) and control who gets what. Only a randomized experiment can support causal conclusions.
Prospective vs. retrospective observational studies?
Prospective = recruit subjects NOW and follow them into the future to see outcomes. Pros: less bias, researcher involved in data collection. Cons: time-consuming; outcome may be rare; attrition, subjects change behavior, costly. Retrospective = select subjects, then look BACK at past conditions/behavior. Pros: fast. Cons: relies on memory → response bias; hard to find enough people with the outcome.
Example: Smokers and non-smokers are recruited and followed for 10 years to see who develops lung cancer. Classify and name the variables.
Prospective observational study. Explanatory variable = smoking status (categorical); response variable = lung cancer status (yes/no). Potential problems: subjects don't return (attrition), subjects quit/start smoking mid-study, costly.
Example: People who tested positive for COVID report their vaccination status and symptoms. Classify and give a problem.
Retrospective observational study (looking back at the past). Explanatory = vaccination status; response = symptoms. Problem: misremembering/forgetting symptoms (response bias).
What is association vs. causation?
Association = values of one variable tend to be related to values of the other. Causation = changing one variable actually INFLUENCES the other. Association does NOT imply causation; causation can be inferred only from a randomized experiment.
Example: Association or causation? (1) Taking a practice exam raises exam scores. (2) Families with many cars own many TVs. (3) Low-dose aspirin reduces heart attacks. (4) Goldfish in large ponds are larger. (5) Putting a goldfish in a larger pond makes it grow.
(1) Causation (2) Association (3) Causation (4) Association (observed pattern only) (5) Causation (an intervention/'putting' implies a treatment). Causation language = change/cause/increases because we imposed it.
What is a confounding (lurking) variable? Examples?
A variable related to BOTH the explanatory and response variables that confuses their relationship. Ice cream sales ↔ shark attacks: time of year/hot weather drives both. Firefighters ↔ fire damage: fire severity drives both (more firefighters sent AND more damage). Confounders can be controlled in an experiment, but are hard to avoid in observational studies.
Experimental terminology: factor, levels, treatment, response variable?
Factor = an explanatory variable in an experiment. Levels = the specific values chosen for a factor. Treatment = the specific combination of factor levels a unit receives. Response variable = what's measured. Example: sleep and exercise (2 factors) → grades (response).
Treatment group, control group, placebo?
Treatment group = receives the treatment. Control group = receives no treatment (or baseline). Placebo = an ineffective 'fake' treatment given to the control group so they don't know they aren't getting the real one (controls for psychological effects).
What are the 4 principles of experimental design?
1) Control — hold other variables constant (e.g., everyone drinks 12 oz of water with the pill). 2) Randomization — randomly assign subjects to treatments to even out differences (age, immune system) and prevent accidental bias. 3) Replication — enough subjects/cases to estimate the effect accurately. 4) Blocking — group similar units, then randomize within groups.
What is blocking?
When units are similar in a way that is NOT a factor under study: (a) group them into blocks, (b) randomly assign treatments WITHIN each block. Blocking lets you account for a variable you can't randomize (like sex or risk level). Example: split patients into low-risk and high-risk blocks, then randomize half of each block to treatment.
What is a Completely Randomized Design (CRD)? Example?
Every subject is randomly assigned a treatment with no regard to other characteristics (no blocking). Example: 40 dental patients randomly use fluoride toothpaste or identical non-fluoride toothpaste for 6 months. Factor = fluoride (yes/no); no blocks; response = number of cavities.
What is a Randomized Block Design (RBD)? Example?
Subjects are first divided into blocks by a variable that can't be randomized, then randomly assigned treatments WITHIN each block. Example: 16 Alzheimer's patients (6 male, 10 female); half get physical activity sessions, half routine care. Factor = physical activity; block = sex (male/female); response = Tinetti gait & balance score (0–28).
What is a Factorial Design? Example?
An experiment with MORE THAN ONE factor where all combinations of factor levels are used as treatments. Example: 42 subjects get THC (0, 7.5, 12.5 mg) and CBD (0, 20 mg). 3 × 2 = 6 treatments (the 0 mg/0 mg is essentially the control). Response = number of 5-second conversation pauses.
How to identify the design quickly?
One factor, no blocking → completely randomized. Blocking variable present → randomized block. Two+ factors combined → factorial.
What is blinding? Placebo effect vs. experimenter effect?
Blinding = hiding which treatment a subject is receiving from subjects and/or experimenters. Placebo effect = subjects respond to the IDEA of treatment rather than the treatment itself. Experimenter effect = experimenter's (often subconscious) knowledge of the treatment influences how they treat/measure subjects.
Single-blind vs. double-blind?
Single-blind = EITHER the subjects OR the experimenters are blinded (not both). Double-blind = BOTH subjects and experimenters are blinded.
Example: Toothpaste study — who should be blinded?
Subjects: blind them (identical-looking toothpaste) so they don't change brushing habits. Experimenter (dentist counting cavities): blinding also helps prevent subconscious bias during the exam → ideally double-blind.
Example: Alzheimer's physical-activity study — blinding?
Subjects: NOT possible to blind (they know whether they exercised). Experimenter: can be blinded when scoring the Tinetti test to avoid bias → single-blind.
Example: THC/CBD study — blinding?
Subjects: should be blinded (avoid placebo/psychosomatic effect). Experimenter: blinded so they keep the same conversation pace for everyone → double-blind.
Ethics in experiments — what is required before, during, and after?
Before: Institutional Review Board (IRB) approval; informed consent from all subjects (agree to participate, aware of risks). During/after: data kept confidential/anonymous; subjects have autonomy and the right to withdraw anytime; subjects must be debriefed (esp. if adverse effects or deception). Some experiments are unethical to run at all (e.g., assigning people to smoke or breathe polluted air).
Examples of unethical studies?
Tuskegee Study: Black men with syphilis were denied treatment under the guise of 'free medical care' for ~40 years (no informed consent). Milgram experiment: participants were deceived and told to give increasingly 'lethal' shocks, with pressure from a lab-coated authority and no debriefing.
Count vs. proportion vs. percentage?
Count = number of observations in a category. Proportion = count ÷ total. Percentage = proportion × 100%. Frequency table = counts only; relative frequency table = proportions/percentages.
Example: 20 STAT 1000 students: Freshman 4, Sophomore 6, Junior 7, Senior 3. Make the relative frequency table.
Freshman 4/20 = 0.20 = 20%; Sophomore 6/20 = 0.30 = 30%; Junior 7/20 = 0.35 = 35%; Senior 3/20 = 0.15 = 15%. (Total = 100%.)
Pie chart: when to use it, and what are its weaknesses?
Slice size = proportion; use when categories DON'T overlap and you want to show how a 'whole' is divided. Weaknesses: doesn't show counts, hard to compare close categories, 'Other' category isn't specific, poor with many categories (works best with a few).
Bar chart: what does it show and when is it better?
Height of bar = count (or proportion) per category. Better than a pie chart when categories DO overlap/aren't parts of a single whole (e.g., which continents each person has visited) or when you want to show scale. Order/remove categories (e.g., 'None') for better scale.
Example: Convert a bar chart of traffic fatalities (Pedestrian 4735, Pedalcyclist 743, Other 190) to relative frequencies.
Total = 4735 + 743 + 190 = 5,668. Pedestrian = 4735/5668 = 0.8354 (83.54%); Pedalcyclist = 743/5668 = 0.1311 (13.11%); Other = 190/5668 = 0.0335 (3.35%).
What is a cross-classification (contingency / two-way) table? Which variable goes where?
A table of counts/proportions/percentages comparing TWO categorical variables. Rows = explanatory variable, columns = response variable. Row/column totals are the margins; the overall total is in the corner.
Example (1,529 films by genre and rating): How many were rated R? How many were dramas?
R-rated = 452 (column total). Dramas = 529 (row total).
Example: What percent of comedies were rated R? What percent of R-rated films were comedies?
Comedies rated R: 124/312 = 0.397 = 39.7% (denominator = comedy row total, 312). R-rated films that were comedies: 124/452 = 0.274 = 27.4% (denominator = R column total, 452). Always ask: 'what group am I given?' — that total is the denominator.
Example (forecast vs. actual weather, 365 days; Rain/Rain=27, Forecast Rain/No Rain=63, No-rain forecast/Rain=7, No/No=268): On what % of days did it actually rain? Was rain predicted? Was the forecast correct?
Actually rained: (27+7)/365 = 34/365 = 9.31%. Rain predicted: (27+63)/365 = 90/365 ≈ 24.7%. Forecast correct: (27+268)/365 = 295/365 = 80.8%.
What is a conditional distribution?
The proportion/percentage in each RESPONSE category GIVEN that the case is in a specific explanatory group. Each row (explanatory group) adds to 100%. Example: Job vs. year — Freshman 22/46 = 47.83% have a job; Sophomore 170/312 = 54.49%; Junior 97/134 = 72.39%; Senior 37/53 = 69.81%; overall 326/565 = 57.70%.
Side-by-side vs. segmented (stacked) bar chart vs. mosaic plot?
Side-by-side: a bar for each response category within each explanatory group (shows counts or proportions; compare directly). Segmented/stacked: ONE bar per explanatory group split into segments whose heights = conditional percentages (each bar sums to 100%). Mosaic: like segmented, but bar WIDTH also reflects group size, so area is proportional to count. Titanic example: segmented bars show the % surviving within each class; you can't know overall counts from percentages alone.
What is Simpson's Paradox? Example?
A trend that appears in combined data disappears or REVERSES when the data are separated into groups (a lurking variable is hidden in the aggregate). Hospital example: overall delay rate — large hospital 130/1000 = 13% vs. small 30/300 = 10% (small looks better). But within each surgery type, the large hospital is better: major 15% vs. 20%; minor 5% vs. 8%. The small hospital only looks better because it does mostly low-risk minor surgeries. Always check subgroups.
How do we display a quantitative variable? (histogram, stem-and-leaf, dotplot, density plot)
Histogram: bars show frequency within equal-width bins (height = # of cases in the bin; bins must be same width, non-overlapping). Stem-and-leaf: like a histogram but keeps individual values (stem = leading digits, leaf = last digit; e.g., 8|5 = 88). Dotplot: one dot per case along an axis. Density plot: a smoothed-out histogram curve.
How do you construct a histogram?
1) Divide the range into bins (equal width). 2) Count observations in each bin. 3) Draw bars with height = count; horizontal axis = data values, vertical axis = counts. Limitation: you can only answer questions that line up with bin edges (e.g., '% between $10 and $25' is OK; '% between $12 and $24' can't be answered).
What does SOCS stand for?
Shape, Outliers, Center, Spread — the four things we describe about a single quantitative distribution.
SHAPE: what 3 features do we describe?
1) Symmetry — can you fold it down the middle and have the sides match? 2) Modality — number of peaks (mode): unimodal, bimodal, multimodal, or uniform (no peak; all bars about equal). 3) Skewness — is one tail longer?
Right-skewed vs. left-skewed? Where do the mean and median fall?
Skew is named for the direction of the LONG TAIL. Right-skewed: tail stretches right → mean > median. Left-skewed: tail stretches left → mean < median. (The mean is pulled toward the tail.)
Example: Calculus exam scores (400 students) histogram shows two peaks. Describe the shape.
Not symmetric-skewed; BIMODAL, no skew. (Common for exams: students who come to class vs. those who don't.)
Example: Credit card spending (500 customers) — shape and outliers?
Unimodal, right-skewed, with outlier(s) at the high end. Outliers must be reported.
What is an outlier?
An extreme value that doesn't appear to belong with the rest of the data. Always report/investigate them; they can have big effects on the mean and SD.
CENTER: median vs. mean — definitions and how to compute?
Median = middle value (half above, half below). Sort the data; odd n → the middle value; even n → the average of the two middle values (may not be a data value). Mean = x-bar = (sum of all values) ÷ n (balance point of the histogram). Mean = x-bar (statistic), mu (μ) (parameter).
Spotify example (13 songs): 138 162 178 197 204 209 216 222 231 245 262 273 297. Find the mean and median.
Median = 7th value = 216 sec. Mean = 2834/13 = 218 sec.
Floods example (n = 20): 29, 38, 38, 43, 48, 49, 56, 68, 76, 80, 82, 82, 86, 87, 103, 113, 118, 131, 136, 176. Find the median and mean.
Median = average of 10th and 11th = (80 + 82)/2 = 81. Mean = 1639/20 = 81.95.
Mean vs. median: which is resistant to outliers/skew?
The MEDIAN is resistant (not much influenced by outliers or skewness); the MEAN is pulled toward the tail. Spotify example: replace 297 with 15654 → median stays 216, mean jumps to 1399.3. For skewed data/outliers, report the median.
Example: A clerk types the CEO's $200,000 salary as $2,000,000. Effect on median and mean?
Median: no change (it's based on position/middle). Mean: increases a lot (pulled up by the outlier).
SPREAD: three measures?
Range, interquartile range (IQR), and standard deviation.
How do you calculate the range and IQR? (Spotify data)
Range = max − min = 297 − 138 = 159. IQR = Q3 − Q1. Q1 = median of the lower half; Q3 = median of the upper half (for odd n, do NOT include the overall median in either half). Spotify: lower half 138–209 → Q1 = (178+197)/2 = 187.5; upper half 222–297 → Q3 = (245+262)/2 = 253.5; IQR = 253.5 − 187.5 = 66.
Floods example: find Q1, Q3, and IQR.
Lower half (first 10 values) median = (48+49)/2 = 48.5 = Q1. Upper half median = (103+113)/2 = 108 = Q3. IQR = 108 − 48.5 = 59.5.
Why is the IQR a good measure of spread for skewed data? (credit card quartiles $73.84 and $624.80)
IQR = 624.80 − 73.84 = $550.96. It measures the spread of the middle 50% and is NOT affected by outliers or skewness, unlike the range and standard deviation.
How is standard deviation calculated and what does it mean?
Variance s² = Σ(x − x-bar)² ÷ (n − 1) (sum of squared deviations from the mean divided by count minus 1). Standard deviation s = √variance. Roughly the typical distance of values from the mean. Always positive (≥ 0); bigger = more spread. Notation: s (sample statistic), sigma (σ) (population parameter). Steps: find mean → subtract the mean from each value → square → sum → divide by n − 1 → square root. Example: songs 138, 178, 197, 209, 222, 297 → mean 206.83, sum of squares ≈ 14,030.8, variance ≈ 2,806.2, s ≈ 52.97.
Example: A meteorologist's coldest week was 36°F but he recorded 2° (Celsius) by mistake. Effect on mean, median, range, IQR, SD?
Mean: changes (decreases). Median: stays the same (position-based). Range: changes (increases). IQR: stays the same. Standard deviation: changes (increases).
Why must we know the SHAPE as well as center and spread?
Mean and SD don't uniquely define a distribution — different-shaped datasets can have the same mean and SD.
What is the five-number summary?
Min, Q1, Median, Q3, Max (summarizes quantitative data via position).
Example: Five-number summary of Spotify data.
Min 138, Q1 187.5, Median 216, Q3 253.5, Max 297.
Example: Five-number summary of floods data.
Min 29, Q1 48.5, Median 81, Q3 108, Max 176.
What is a boxplot and how is it built?
A display of the five-number summary: a box from Q1 to Q3 with a line at the median; whiskers extend to the most extreme values that are NOT outliers; outliers are plotted as separate points. Standard outlier cutoffs: below Q1 − 1.5×IQR or above Q3 + 1.5×IQR. (This is why you see 'cutoffs' on boxplots.)
Example: 20 exam scores 33 54 59 62 65 67 69 71 73 74 76 77 79 82 83 85 88 91 94 98 — five-number summary and boxplot.
Min 33, Q1 = 66, Median = 75, Q3 = 84, Max 98; IQR = 18. Fences: 66 − 27 = 39 and 84 + 27 = 111. The 33 is below 39 → OUTLIER (plotted as a dot); lower whisker ends at 54; upper whisker ends at 98.
Example: Floods data — boxplot outlier check.
IQR = 59.5; fences: 48.5 − 89.25 = −40.75 and 108 + 89.25 = 197.25. Min (29) and max (176) are inside the fences → NO outliers; whiskers go to 29 and 176.
How does a boxplot show skewness?
Symmetric: whiskers about equal, median centered. Right-skewed: longer upper whisker/upper tail, median closer to Q1. Left-skewed: longer lower whisker, median closer to Q3. Example: house values boxplot with a long upper whisker → histogram would be right-skewed.
How do you compare groups with graphs?
Use the SAME scale/bins for histograms; use side-by-side boxplots for easy comparison of center (median), spread (IQR/box length), shape, and outliers. Example: wood vs. steel roller coasters — similar/slightly higher median for wood, but steel has greater variability and both high and low outliers.
What is a z-score? Formula? Why use it?
z = (x − mean) ÷ SD = number of standard deviations a value is above (+) or below (−) the mean. Lets you compare values from DIFFERENT distributions/units ('standardizing'). Example: z = 2 → 2 SDs above the mean; z = −1 → 1 SD below.
Interpreting z-scores: how unusual is a value?
|z| < 1: common, not unusual. 1–2: relatively common, slightly unusual. 2–3: unusual but not unheard of. > 3: highly unusual/unlikely.
Example: Bode Miller — slalom 51.01 s (mean 52.67, SD 1.614) vs. downhill 113.91 s (mean 116.26, SD 1.914). Which race was better?
Slalom z = (51.01 − 52.67)/1.614 = −1.03; Downhill z = (113.91 − 116.26)/1.914 = −1.23. Lower time = better, so the more negative z is better → he performed better in the DOWNHILL.
Example: Exam 1 score 90 (mean 88, SD 4); Exam 2 score 80 (mean 75, SD 5). Which gets dropped?
Exam 1 z = (90 − 88)/4 = 0.5; Exam 2 z = (80 − 75)/5 = 1.0. Exam 1 is relatively worse → the 90 on Exam 1 is dropped.
What is shifting (adding/subtracting a constant)? Effect?
Add/subtract c from every value: center (mean, median, quartiles, min/max) shifts by c; spread (SD, range, IQR) does NOT change; shape does NOT change.
What is scaling (multiplying/dividing)? Effect?
Multiply/divide every value by c: center AND spread both change (mean, median × c; SD, IQR, range × |c|); shape does not change. New mean = mean × c; new SD = SD × c.
Example: mean 80, median 85, SD 2.5. Apply: (1) ×2 (2) +5 (3) −10 (4) ÷10.
(1) mean 160, median 170, SD 5. (2) mean 85, median 90, SD 2.5. (3) mean 70, median 75, SD 2.5. (4) mean 8, median 8.5, SD 0.25.
What happens to a distribution when converted to z-scores?
Always mean = 0 and SD = 1 (subtract the mean = shifting; divide by SD = scaling). Shape stays the same.
How do you describe a scatterplot?
Direction (positive, negative, or none), Form (linear, curved, no pattern), Strength (weak, moderate, strong — how tightly clustered), plus any outliers. Explanatory variable on x-axis, response on y-axis. Example: verbal SAT (x) vs. math SAT (y): positive, linear, moderate.
Positive vs. negative association?
Positive: as x increases, y tends to increase (upward trend). Negative: as x increases, y tends to decrease (downward trend).
What is the correlation coefficient? Notation and range?
r (sample statistic), rho (ρ) (population parameter) = a numerical measure of the direction and strength of a LINEAR association between two quantitative variables. Always between −1 and 1; no units; sign = direction; magnitude = strength.