Comprehensive Study Notes on Descriptive Statistics, Density Curves, Normal Distributions, and Z-Scores


Recap of Descriptive Measures: Center and Spread

  • Measures of Center:

    • Mode: The most frequently occurring value in a dataset.

    • Median: The exact middle value of an ordered distribution when values are sorted from smallest to largest (50th50\text{th} percentile).

    • Mean: The arithmetic average of all data points in the distribution.

  • Measures of Spread: Resistant vs. Non-Resistant Categories:

    • Resistant Spread Measures (Robust to extreme outliers and extreme peripheral values):

      • Five-Number Summary: Minimum, First Quartile (Q1Q_1), Median (Q2Q_2), Third Quartile (Q3Q_3), and Maximum.

      • Interquartile Range (IQRIQR): The distance between the 25th25\text{th} percentile (Q1Q_1) and the 75th75\text{th} percentile (Q3Q_3). Adding an extreme outlier at the fringe of the distribution does not substantially alter the IQRIQR.

    • Non-Resistant Spread Measures (Highly sensitive to outliers):

      • Standard Deviation (ss or σ\sigma) and Variance (s2s^2 or σ2\sigma^2).

      • Calculated by taking the deviation of every single value from the mean and squaring it. Squaring deviations weights larger peripheral distances disproportionately.

Activity Inequality and Public Health Case Study

  • Study Design & Dataset:

    • Analyzed anonymized smartphone step-count data from participants across 111 countries111\,\text{countries} to evaluate worldwide physical activity patterns.

    • Research Problem: Determining why obesity rates vary significantly across nations despite similar national mean daily step counts.

  • Findings on Mean vs. Spread:

    • The mean daily step count is not an effective predictor of national obesity rates.

    • Obesity rates are driven by activity inequality (the spread/distribution of physical activity across a population) rather than the center of the distribution.

    • Obesity concentrates primarily among the extreme lower tail of the physical activity distribution (highly inactive populations taking minimal steps).

  • Statistical Explanatory Power (R2R^2):

    • R2R^2 Definition: A statistical metric representing the proportion of variance in an outcome variable predictable from an independent variable.

    • Activity Inequality R2R^2: 0.640.64 (approximately 64%64\% of variance in national obesity rates is explained by activity inequality).

    • Mean Daily Steps R2R^2: 0.470.47 (only 47%47\% of variance in national obesity rates is explained by mean daily steps).

  • Comparing Two Hypothetical Countries:

    • Country A: High clustering around the mean step count (6,000 steps6{,}000\,\text{steps}). Low standard deviation, tight distribution, very small inactive tail, resulting in low obesity rates.

    • Country B: Identical mean step count (6,000 steps6{,}000\,\text{steps}), but high standard deviation and spread. Features a large population segment in the highly inactive tail alongside a highly active segment, resulting in higher obesity rates.

  • Demographic Stratification & Sampling Biases:

    • Gender Stratification: While overall population distribution is approximately 50%50\% male and 50%50\% female, activity data is often unstratified by gender. Failing to break down data by gender hides gender gaps, risking public health interventions that amplify inequality.

    • Smartphone Sampling Biases:

      • Individual-Level Bias: Smartphone ownership favors younger and wealthier demographics.

      • Population Demographics: Older populations walk less on average due to age patterns. A healthy country with an aging population (e.g., Japan) may exhibit lower mean steps due to demographic composition while maintaining low obesity rates.

  • Core Methodological Principles:

    1. Always report measures of spread alongside the mean.

    2. Extreme values (tails), rather than central values, often drive real-world systemic outcomes.

    3. Stratify descriptive statistics by sub-groups (e.g., gender, age, region) to uncover hidden patterns of inequality.

    4. Evaluate sample characteristics against overall population demographics to spot sampling bias.

Building Blocks of Spread: Deviation to Standard Deviation

  • 1. Deviation:

    • Formula: x−xˉx - \bar{x} (sample) or x−μx - \mu (population).

    • Represents the distance and direction of a single data point from the mean. Can be positive or negative.

  • 2. Squared Deviation:

    • Formula: (x−xˉ)2(x - \bar{x})^2

    • Squares the distance of an observation from the mean. Ensures all values are positive and rescales units into squared terms.

  • 3. Variance:

    • Formula (Population): σ2=∑(x−μ)2N\sigma^2 = \frac{\sum (x - \mu)^2}{N}

    • Formula (Sample): s2=∑(x−xˉ)2n−1s^2 = \frac{\sum (x - \bar{x})^2}{n - 1}

    • The average squared deviation. Measured in squared units (e.g., hours2\text{hours}^2), making direct intuitive interpretation difficult.

  • 4. Standard Deviation:

    • Formula: σ=σ2\sigma = \sqrt{\sigma^2} or s=s2s = \sqrt{s^2}

    • The square root of the variance. Returns the metric to the original measurement units, representing the typical distance of data points from the mean.

Linear Transformations of Data

  • Core Rule: Linear transformations change the center and/or spread of a distribution, but NEVER change the underlying shape of the distribution.

  • Types of Linear Transformations:

    • Shift Transformation (Adding/Subtracting a Constant aa):

      • Formula: x′=x+ax' = x + a

      • Shifts measures of center (mean, median, mode) by aa.

      • Measures of spread (standard deviation, IQRIQR, variance) remain unchanged.

      • Example: Converting temperature offsets (e.g., shifting Fahrenheit to Celsius scales).

    • Rescaling Transformation (Multiplying/Dividing by a Constant bb):

      • Formula: x′=b×xx' = b \times x

      • Rescales measures of center by bb and measures of spread (standard deviation, IQRIQR) by ∣b∣|b|.

      • Example: Converting miles to kilometers, cents to Canadian dollars, or rescaling an exam score out of 1,200 points1{,}200\,\text{points}.

    • Combined Transformation:

      • Formula: x′=a+b×xx' = a + b \times x

      • Rescales spread by ∣b∣|b| and shifts center by aa.

Density Curves and Normal Distributions

  • Density Curve Definition: A smooth mathematical curve that models and approximates the shape of an empirical data histogram (e.g., standard test scores of 947 seventh-graders947\,\text{seventh-graders} on the Iowa Test of Basic Skills).

  • Two Mandatory Rules for Density Curves:

    1. The curve must lie entirely on or above the horizontal axis (f(x)≥0f(x) \ge 0 for all xx). Negative concentration of observations is impossible.

    2. The total area underneath the density curve must equal exactly 11 (100%100\% of all observations).

  • Mean and Median on Density Curves:

    • Symmetric Unimodal Curve: Mean = Median = Mode, located at the central peak.

    • Right-Skewed Curve: Tail extends right; the mean is pulled rightward by extreme values (Mean>Median\text{Mean} > \text{Median}).

    • Left-Skewed Curve: Tail extends left; the mean is pulled leftward by extreme values (Mean<Median\text{Mean} < \text{Median}).

  • Normal Distribution Parameters:

    • Notation: N(μ,σ)N(\mu, \sigma), where μ\mu is the mean and σ\sigma is the standard deviation.

    • Determined entirely by μ\mu (location of peak/center) and σ\sigma (width/spread).

    • Modifying μ\mu shifts the curve horizontally without altering shape.

    • Modifying σ\sigma alters spread: smaller σ\sigma produces a tall, narrow peak; larger σ\sigma produces a broad, flat curve.

The 68-95-99.7 Empirical Rule

  • Rule Definition: For any normally distributed variable N(μ,σ)N(\mu, \sigma):

    • Approximately 68%68\% of observations lie within 1σ1\sigma of the mean: [μ−σ,μ+σ][\mu - \sigma, \mu + \sigma].

    • Approximately 95%95\% of observations lie within 2σ2\sigma of the mean: [μ−2σ,μ+2σ][\mu - 2\sigma, \mu + 2\sigma].

    • Approximately 99.7%99.7\% of observations lie within 3σ3\sigma of the mean: [μ−3σ,μ+3σ][\mu - 3\sigma, \mu + 3\sigma].

  • Example 1: Women's Heights (N(64.5,2.5)N(64.5, 2.5)):

    • μ=64.5 inches\mu = 64.5\,\text{inches}, σ=2.5 inches\sigma = 2.5\,\text{inches}.

    • 68%68\% Interval: 64.5±2.5  ⟹  [62.0 inches,67.0 inches]64.5 \pm 2.5 \implies [62.0\,\text{inches}, 67.0\,\text{inches}].

    • 95%95\% Interval: 64.5±2(2.5)  ⟹  [59.5 inches,69.5 inches]64.5 \pm 2(2.5) \implies [59.5\,\text{inches}, 69.5\,\text{inches}].

      • The outer 5%5\% is split into 2.5%2.5\% below 59.5 inches59.5\,\text{inches} and 2.5%2.5\% above 69.5 inches69.5\,\text{inches}.

    • 99.7%99.7\% Interval: 64.5 \pm 3(2.5) \n\implies [57.0\,\text{inches}, 72.0\,\text{inches}].

  • Example 2: NAEP Eighth-Grade Mathematics Test (N(282,40)N(282, 40)):

    • μ=282 points\mu = 282\,\text{points}, σ=40 points\sigma = 40\,\text{points}.

    • 95%95\% Range (2σ2\sigma): 282±2(40)  ⟹  [202 points,362 points]282 \pm 2(40) \implies [202\,\text{points}, 362\,\text{points}].

    • 99.7%99.7\% Range (3σ3\sigma): 282±3(40)  ⟹  [162 points,402 points]282 \pm 3(40) \implies [162\,\text{points}, 402\,\text{points}].

    • Evaluating Score of 322 points322\,\text{points}:

      • Sitting at 282+40=322 points282 + 40 = 322\,\text{points} puts the score at exactly +1σ+1\sigma above the mean.

      • Proportion of students scoring below 322 points322\,\text{points} = 50%+34%=84%50\% + 34\% = 84\%.

Standardizing Data and Z-Scores

  • Definition: Standardizing converts observations from their original units into standardized units of standard deviations away from the mean.

  • Standard Normal Distribution:

    • Notation: N(0,1)N(0, 1), where μ=0\mu = 0 and σ=1\sigma = 1

  • Z-Score Formula:     z=x−μσz = \frac{x - \mu}{\sigma}

  • Interpretation of Z-Scores:

    • z=0z = 0: Observation equals the mean.

    • z>0z > 0: Observation lies above the mean.

    • z<0z < 0: Observation lies below the mean.

    • ∣z∣≥2.0|z| \ge 2.0 or ∣z∣≥3.0|z| \ge 3.0: Indicates an unusual or rare observation.

  • Step-by-Step Z-Score Calculation (Women's Heights N(64.5,2.5)N(64.5, 2.5) ):

    • Step 1: Recenter (Subtract mean μ\mu).

    • Step 2: Rescale (Divide by standard deviation σ\sigma).

    • Case 1: Height x=68 inchesx = 68\,\text{inches}:         z=68−64.52.5=3.52.5=1.40z = \frac{68 - 64.5}{2.5} = \frac{3.5}{2.5} = 1.40         (Height is 1.40 standard deviations1.40\,\text{standard deviations} above the mean).

    • Case 2: Height x=60 inchesx = 60\,\text{inches} (5 feet5\,\text{feet}):         z=60−64.52.5=−4.52.5=−1.80z = \frac{60 - 64.5}{2.5} = \frac{-4.5}{2.5} = -1.80         (Height is 1.80 standard deviations1.80\,\text{standard deviations} below the mean).

    • Case 3: Height x=66 inchesx = 66\,\text{inches}:         z=66−64.52.5=1.52.5=0.60z = \frac{66 - 64.5}{2.5} = \frac{1.5}{2.5} = 0.60         (Height is 0.60 standard deviations0.60\,\text{standard deviations} above the mean).

Cumulative Proportions and Utilizing Table A

  • Table A Definition: A standardized table listing cumulative proportions P(Z<z)P(Z < z), representing the area under the standard normal curve to the LEFT of a specific z-score.

  • Upper Tail Rule: To calculate the proportion greater than a given z-score (area to the right), subtract the Table A cumulative value from 1:     P(Z>z)=1−P(Z<z)P(Z > z) = 1 - P(Z < z)

  • How to Read Table A:

    • The left column gives the z-score to the tenths place (e.g., 1.41.4).

    • The top row gives the hundredths place (e.g., 0.070.07).

    • The intersection provides the cumulative left-tail area.

    • Example: For z=1.47z = 1.47, locate row 1.41.4 and column 0.070.07, yielding cumulative area 0.92920.9292 (92.92%92.92\% of observations lie to the left).

  • Worked Example: SAT Scores & NCAA Division I Eligibility:

    • Distribution: SAT scores for 1,400,000 students1{,}400{,}000\,\text{students} modeled as N(1026,209)N(1026, 209).

    • NCAA Eligibility Threshold: Minimum score of x=820x = 820.

    • Step 1: Compute Z-Score:         z=820−1026209=−206209≈−0.99z = \frac{820 - 1026}{209} = \frac{-206}{209} \approx -0.99

    • Step 2: Find Cumulative Area in Table A:         Lookup for z=−0.99z = -0.99 gives P(Z<−0.99)=0.1611P(Z < -0.99) = 0.1611 (16.11%16.11\% score below 820820).

    • Step 3: Calculate Qualifying Proportion (≥820\ge 820):         P(Z≥−0.99)=1−0.1611=0.8389  ⟹  83.89%P(Z \ge -0.99) = 1 - 0.1611 = 0.8389 \implies 83.89\%         Approximately 84%84\% of test-takers qualify.

    • Step 4: Range Calculation (720 to 820 points720\,\text{to}\,820\,\text{points}):         Subtract cumulative area at x=720x = 720 from cumulative area at x=820x = 820, yielding approximately 9%9\% of test-takers in this range.

Inverse Lookup Problems: Determining Threshold Scores

  • Concept: Finding an unknown raw value xx given a target percentile or percentage.

  • Problem: Determine the minimum SAT score required to reach the top 10%10\% of all test takers (90th90\text{th} percentile, cumulative area 0.90000.9000).

  • Step 1: Locate Target Cumulative Area in Table A:

    • Search the interior of Table A for the value closest to 0.90000.9000

    • z=1.28z = 1.28 gives cumulative area 0.89970.8997

    • z=1.29z = 1.29 gives cumulative area 0.90150.9015

    • Select z=1.28z = 1.28 as the closest approximation to 0.90000.9000

  • Step 2: Un-Standardize to Solve for Raw Score xx:     x=μ+z×σx = \mu + z \times \sigma     x=1026+(1.28×209)=1026+267.52=1293.52x = 1026 + (1.28 \times 209) = 1026 + 267.52 = 1293.52     A student must score approximately 1294 points1294\,\text{points} to place in the top 10%10\%.

Summary Quick Reference

  • Deviation: x−μx - \mu (Distance from mean; positive or negative).

  • Squared Deviation: (x−μ)2(x - \mu)^2 (Eliminates negative signs; weights extreme outliers heavily).

  • Variance: σ2=∑(x−μ)2N\sigma^2 = \frac{\sum (x - \mu)^2}{N} (Average squared deviation; measured in squared units).

  • Standard Deviation: σ=σ2\sigma = \sqrt{\sigma^2} (Typical distance from mean; in original units).

  • Z-Score: z=x−μσz = \frac{x - \mu}{\sigma} (Standardized distance from mean in standard deviation units).