AP Psychology Statistics Practice Notes

Measures of Central Tendency

Central tendency refers to statistical measures that identify a single score as representative of an entire distribution. The three primary measures are the mean, median, and mode.

  • Definitions of Measures:

    • Mean: The arithmetic average of a distribution, calculated by summing all individual scores and dividing by the total number of scores (nn).
    • Median: The middle score in a distribution when all data points are arranged in numerical order. If the distribution contains an even number of scores, the median is the average of the two middle values.
    • Mode: The most frequently occurring score(s) in a distribution.
  • Calculations for Dataset A (IQ Scores):

    • Raw Scores: 75, 80, 100, 55, 80, 105, 142
    • Ordered Dataset: 55, 75, 80, 80, 100, 105, 142
    • Sample Size (nn): 77
    • Mean Calculation:Mean=55+75+80+80+100+105+1427=6377=91\text{Mean} = \frac{55 + 75 + 80 + 80 + 100 + 105 + 142}{7} = \frac{637}{7} = 91
    • Median Calculation: The 4th score in the ordered dataset is 8080
    • Mode Calculation: The value 8080 appears twice, which is more frequent than any other score; therefore, the mode is 8080
  • Calculations for Dataset B (AP Psychology Test Scores):

    • Raw Scores: 2, 3, 3, 3, 3, 4, 4, 4, 5, 5
    • Sample Size (nn): 1010
    • Mean Calculation:Mean=2+3+3+3+3+4+4+4+5+510=3610=3.6\text{Mean} = \frac{2 + 3 + 3 + 3 + 3 + 4 + 4 + 4 + 5 + 5}{10} = \frac{36}{10} = 3.6
    • Median Calculation: The average of the 5th score (33) and 6th score (44):         Median=3+42=3.5\text{Median} = \frac{3 + 4}{2} = 3.5
    • Mode Calculation: The value 33 appears 4 times, making it the mode
  • Calculations for Dataset C (ACT Test Scores):

    • Raw Scores: 20, 22, 23, 25, 26, 26, 33
    • Sample Size (nn): 77
    • Mean Calculation:Mean=20+22+23+25+26+26+337=1757=25\text{Mean} = \frac{20 + 22 + 23 + 25 + 26 + 26 + 33}{7} = \frac{175}{7} = 25
    • Median Calculation: The 4th score in the ordered dataset is 2525
    • Mode Calculation: The value 2626 appears twice, making it the mode

Applications of Central Tendency and Outlier Effects

  • Candy Consumption Analysis (Individual Data):

    • Data Points: Henri = 1010, Sylvia = 1010, Jerry = 88, Tricia = 66, Tahli = 11
    • Total Candy Bars: 10+10+8+6+1=3510 + 10 + 8 + 6 + 1 = 35
    • Number of Individuals (nn): 55
    • Mean Calculation:Mean=355=7\text{Mean} = \frac{35}{5} = 7
    • Correct Choice: c. 7
  • Age Distribution Analysis (Berry Family Children):

    • Data Points: Ages 2, 3, 7, 9, 9
    • Number of Children (nn): 55
    • Median Determination: The middle (3rd) age in ordered sequence is 77
    • Correct Choice: c. 7
  • Sensitivity to Extreme Scores (Outliers):

    • An extreme score (outlier) significantly pulls the mean toward the direction of the skew because every individual value directly contributes to the sum used in computing the arithmetic mean.
    • The median and mode remain resistant (robust) to extreme outliers.
    • Most Affected Measure: Mean
    • Correct Choice: e. mean
  • Raffle Ticket Earnings Analysis (Girls' Club):

    • Raw Individual Earnings: $7, $13, $3, $5, $2, $9, $3
    • Ordered Dataset: $2, $3, $3, $5, $7, $9, $13
    • Mean Calculation:Mean=2+3+3+5+7+9+137=427=6\text{Mean} = \frac{2 + 3 + 3 + 5 + 7 + 9 + 13}{7} = \frac{42}{7} = 6
    • Median Calculation: The 4th score is 55
    • Mode Calculation: The value 33 occurs most frequently
    • Comparison: Mean (66) is greater than the mode (33) and greater than the median (55)
    • Correct Choice: b. greater than; greater than
  • Selecting Median over Mean for Skewed Distributions:

    • The median is a more appropriate measure of central tendency than the mean when a dataset contains strong skewness or notable extreme scores that disproportionately distort the mean.
    • In the dataset 16, 40, 4, 8, 24 (ordered: 4, 8, 16, 24, 40), the presence of higher extreme values skews the distribution, making the median a better central descriptor.
    • Correct Choice: a. 16, 40, 4, 8, 24

Measures of Variation and Descriptive Analogies

  • Conceptual Pairing of Statistical Terms:

    • Variation describes the spread or dispersion of a data set, whereas central tendency describes its central focal value.
    • Range measures variation, and median measures central tendency.
    • Analogy: Variation is to central tendency as range is to median.
    • Correct Choice: a. range; median
  • Indication of Score Dispersion:

    • Standard deviation quantifies the average distance or variation of individual scores relative to the distribution's mean score.
    • Correct Choice: c. standard deviation

Normal Distribution and the Empirical Rule

The normal distribution is a symmetrical, bell-shaped curve where the mean, median, and mode are all located at the center. Spread around the mean is defined by standard deviation units following the Empirical Rule (68−95−99.768-95-99.7 Rule).

  • Empirical Rule Thresholds:

    • Approximately 68%68\% of all scores fall within ±1\pm 1 standard deviation of the mean (μ±1σ\mu \pm 1\sigma).
    • Approximately 95%95\% of all scores fall within ±2\pm 2 standard deviations of the mean (μ±2σ\mu \pm 2\sigma).
    • Approximately 99.7%99.7\% of all scores fall within ±3\pm 3 standard deviations of the mean (μ±3σ\mu \pm 3\sigma).
  • Wechsler IQ Test Parameters:

    • Mean (μ)=100\text{Mean } (\mu) = 100
    • Standard Deviation (σ)=15\text{Standard Deviation } (\sigma) = 15
  • Population Percentage Between 70 and 130:

    • Lower Bound=100−2(15)=70\text{Lower Bound} = 100 - 2(15) = 70
    • Upper Bound=100+2(15)=130\text{Upper Bound} = 100 + 2(15) = 130
    • This interval spans from −2σ-2\sigma to +2σ+2\sigma.
    • Result: Approximately 95%95\% of the population falls within this range.
    • Correct Choice: d. 95%
  • Population Percentage Within One Standard Deviation of the Mean:

    • This interval spans from −1σ-1\sigma to +1σ+1\sigma (100−15=85100 - 15 = 85 to 100+15=115100 + 15 = 115).
    • Result: Approximately 68%68\% of all scores fall within this range.
    • Correct Choice: b. 68
  • Population Percentage Between 100 and 130:

    • Lower Bound=100\text{Lower Bound} = 100 (the mean)
    • Upper Bound=100+2(15)=130\text{Upper Bound} = 100 + 2(15) = 130 (+2σ+2\sigma
    • Since the normal curve is symmetrical, the area from the mean (μ\mu) to +2σ+2\sigma is half of the total 95%95\% area contained within ±2σ\pm 2\sigma:         95%2=47.5%\frac{95\%}{2} = 47.5\%
    • Correct Choice: c. 47.5%

Correlation Coefficients and Effect Size

  • Interpreting Correlation Coefficients (rr):

    • Correlation coefficients range from −1.00-1.00 to +1.00+1.00.
    • The sign (+ or −+\text{ or }-) indicates direction (positive or negative relationship).
    • The numerical magnitude indicates strength (0.000.00 represents no linear relationship; 1.001.00 represents a perfect linear relationship).
    • A very weak positive relationship demonstrating that higher IQ relates slightly to better math word problem-solving ability corresponds to a small positive value near zero, such as r=+0.10r = +0.10
    • Correct Choice: c. +.10
  • Interpreting Effect Size (Cohen's d):

    • Cohen's dd measures the standardized difference between two group means.
    • A Cohen's dd value of 0.200.20 is considered small, 0.500.50 is medium, and 0.800.80 or higher is considered large.
    • In research comparing participants sleeping 8 hours versus 5 hours, a value of d=0.95d = 0.95 reflects a very substantial impact.
    • Conclusion: Sleep has a large impact on stress.
    • Correct Choice: b. Sleep has a large impact on stress

Inferential Statistics, Generalizability, and Experimental Design

  • Statistical Significance and pp-Values:

    • The pp-value indicates the probability that the observed differences between experimental conditions occurred purely by random chance.
    • A standard threshold for statistical significance in psychology is p<0.05p < 0.05
    • Finding p=0.02p = 0.02 means there is only a 2%2\% probability that the observed improvement in recall scores occurred by chance.
    • Conclusion: The impact of the memory training program on recall scores was unlikely due to chance.
    • Correct Choice: b. The impact of the memory training program on recall scores was unlikely due to chance
  • Random Sampling and Generalizability:

    • Random sampling ensures every member of a defined population has an equal chance of selection.
    • Proper random sampling allows research findings to be generalized back to the specific population from which the sample was drawn.
    • Selecting 200 teachers using a random number generator from a school district's complete staff list allows findings to generalize specifically to all teachers in that district.
    • Correct Choice: c. The researcher can conclude that her results apply to teachers in the entire school district
  • Random Assignment and Control of Confounding Variables:

    • Random assignment involves placing participants into experimental or control conditions purely by chance.
    • Random assignment equalizes pre-existing individual differences (such as baseline anxiety levels) across conditions prior to the experimental manipulation.
    • Primary Benefit: It minimizes or reduces the influence of confounding variables across groups.
    • Correct Choice: d. The researcher has reduced the influence of all confounding variable