Comprehensive Study Notes on Descriptive Statistics, Density Curves, Normal Distributions, and Z-Scores
Recap of Descriptive Measures: Center and Spread
Measures of Center:
Mode: The most frequently occurring value in a dataset.
Median: The exact middle value of an ordered distribution when values are sorted from smallest to largest ( percentile).
Mean: The arithmetic average of all data points in the distribution.
Measures of Spread: Resistant vs. Non-Resistant Categories:
Resistant Spread Measures (Robust to extreme outliers and extreme peripheral values):
Five-Number Summary: Minimum, First Quartile (), Median (), Third Quartile (), and Maximum.
Interquartile Range (): The distance between the percentile () and the percentile (). Adding an extreme outlier at the fringe of the distribution does not substantially alter the .
Non-Resistant Spread Measures (Highly sensitive to outliers):
Standard Deviation ( or ) and Variance ( or ).
Calculated by taking the deviation of every single value from the mean and squaring it. Squaring deviations weights larger peripheral distances disproportionately.
Activity Inequality and Public Health Case Study
Study Design & Dataset:
Analyzed anonymized smartphone step-count data from participants across to evaluate worldwide physical activity patterns.
Research Problem: Determining why obesity rates vary significantly across nations despite similar national mean daily step counts.
Findings on Mean vs. Spread:
The mean daily step count is not an effective predictor of national obesity rates.
Obesity rates are driven by activity inequality (the spread/distribution of physical activity across a population) rather than the center of the distribution.
Obesity concentrates primarily among the extreme lower tail of the physical activity distribution (highly inactive populations taking minimal steps).
Statistical Explanatory Power ():
Definition: A statistical metric representing the proportion of variance in an outcome variable predictable from an independent variable.
Activity Inequality : (approximately of variance in national obesity rates is explained by activity inequality).
Mean Daily Steps : (only of variance in national obesity rates is explained by mean daily steps).
Comparing Two Hypothetical Countries:
Country A: High clustering around the mean step count (). Low standard deviation, tight distribution, very small inactive tail, resulting in low obesity rates.
Country B: Identical mean step count (), but high standard deviation and spread. Features a large population segment in the highly inactive tail alongside a highly active segment, resulting in higher obesity rates.
Demographic Stratification & Sampling Biases:
Gender Stratification: While overall population distribution is approximately male and female, activity data is often unstratified by gender. Failing to break down data by gender hides gender gaps, risking public health interventions that amplify inequality.
Smartphone Sampling Biases:
Individual-Level Bias: Smartphone ownership favors younger and wealthier demographics.
Population Demographics: Older populations walk less on average due to age patterns. A healthy country with an aging population (e.g., Japan) may exhibit lower mean steps due to demographic composition while maintaining low obesity rates.
Core Methodological Principles:
Always report measures of spread alongside the mean.
Extreme values (tails), rather than central values, often drive real-world systemic outcomes.
Stratify descriptive statistics by sub-groups (e.g., gender, age, region) to uncover hidden patterns of inequality.
Evaluate sample characteristics against overall population demographics to spot sampling bias.
Building Blocks of Spread: Deviation to Standard Deviation
1. Deviation:
Formula: (sample) or (population).
Represents the distance and direction of a single data point from the mean. Can be positive or negative.
2. Squared Deviation:
Formula:
Squares the distance of an observation from the mean. Ensures all values are positive and rescales units into squared terms.
3. Variance:
Formula (Population):
Formula (Sample):
The average squared deviation. Measured in squared units (e.g., ), making direct intuitive interpretation difficult.
4. Standard Deviation:
Formula: or
The square root of the variance. Returns the metric to the original measurement units, representing the typical distance of data points from the mean.
Linear Transformations of Data
Core Rule: Linear transformations change the center and/or spread of a distribution, but NEVER change the underlying shape of the distribution.
Types of Linear Transformations:
Shift Transformation (Adding/Subtracting a Constant ):
Formula:
Shifts measures of center (mean, median, mode) by .
Measures of spread (standard deviation, , variance) remain unchanged.
Example: Converting temperature offsets (e.g., shifting Fahrenheit to Celsius scales).
Rescaling Transformation (Multiplying/Dividing by a Constant ):
Formula:
Rescales measures of center by and measures of spread (standard deviation, ) by .
Example: Converting miles to kilometers, cents to Canadian dollars, or rescaling an exam score out of .
Combined Transformation:
Formula:
Rescales spread by and shifts center by .
Density Curves and Normal Distributions
Density Curve Definition: A smooth mathematical curve that models and approximates the shape of an empirical data histogram (e.g., standard test scores of on the Iowa Test of Basic Skills).
Two Mandatory Rules for Density Curves:
The curve must lie entirely on or above the horizontal axis ( for all ). Negative concentration of observations is impossible.
The total area underneath the density curve must equal exactly ( of all observations).
Mean and Median on Density Curves:
Symmetric Unimodal Curve: Mean = Median = Mode, located at the central peak.
Right-Skewed Curve: Tail extends right; the mean is pulled rightward by extreme values ().
Left-Skewed Curve: Tail extends left; the mean is pulled leftward by extreme values ().
Normal Distribution Parameters:
Notation: , where is the mean and is the standard deviation.
Determined entirely by (location of peak/center) and (width/spread).
Modifying shifts the curve horizontally without altering shape.
Modifying alters spread: smaller produces a tall, narrow peak; larger produces a broad, flat curve.
The 68-95-99.7 Empirical Rule
Rule Definition: For any normally distributed variable :
Approximately of observations lie within of the mean: .
Approximately of observations lie within of the mean: .
Approximately of observations lie within of the mean: .
Example 1: Women's Heights ():
, .
Interval: .
Interval: .
The outer is split into below and above .
Interval: 64.5 \pm 3(2.5) \n\implies [57.0\,\text{inches}, 72.0\,\text{inches}].
Example 2: NAEP Eighth-Grade Mathematics Test ():
, .
Range (): .
Range (): .
Evaluating Score of :
Sitting at puts the score at exactly above the mean.
Proportion of students scoring below = .
Standardizing Data and Z-Scores
Definition: Standardizing converts observations from their original units into standardized units of standard deviations away from the mean.
Standard Normal Distribution:
Notation: , where and
Z-Score Formula:
Interpretation of Z-Scores:
: Observation equals the mean.
: Observation lies above the mean.
: Observation lies below the mean.
or : Indicates an unusual or rare observation.
Step-by-Step Z-Score Calculation (Women's Heights ):
Step 1: Recenter (Subtract mean ).
Step 2: Rescale (Divide by standard deviation ).
Case 1: Height : (Height is above the mean).
Case 2: Height (): (Height is below the mean).
Case 3: Height : (Height is above the mean).
Cumulative Proportions and Utilizing Table A
Table A Definition: A standardized table listing cumulative proportions , representing the area under the standard normal curve to the LEFT of a specific z-score.
Upper Tail Rule: To calculate the proportion greater than a given z-score (area to the right), subtract the Table A cumulative value from 1:
How to Read Table A:
The left column gives the z-score to the tenths place (e.g., ).
The top row gives the hundredths place (e.g., ).
The intersection provides the cumulative left-tail area.
Example: For , locate row and column , yielding cumulative area ( of observations lie to the left).
Worked Example: SAT Scores & NCAA Division I Eligibility:
Distribution: SAT scores for modeled as .
NCAA Eligibility Threshold: Minimum score of .
Step 1: Compute Z-Score:
Step 2: Find Cumulative Area in Table A: Lookup for gives ( score below ).
Step 3: Calculate Qualifying Proportion (): Approximately of test-takers qualify.
Step 4: Range Calculation (): Subtract cumulative area at from cumulative area at , yielding approximately of test-takers in this range.
Inverse Lookup Problems: Determining Threshold Scores
Concept: Finding an unknown raw value given a target percentile or percentage.
Problem: Determine the minimum SAT score required to reach the top of all test takers ( percentile, cumulative area ).
Step 1: Locate Target Cumulative Area in Table A:
Search the interior of Table A for the value closest to
gives cumulative area
gives cumulative area
Select as the closest approximation to
Step 2: Un-Standardize to Solve for Raw Score : A student must score approximately to place in the top .
Summary Quick Reference
Deviation: (Distance from mean; positive or negative).
Squared Deviation: (Eliminates negative signs; weights extreme outliers heavily).
Variance: (Average squared deviation; measured in squared units).
Standard Deviation: (Typical distance from mean; in original units).
Z-Score: (Standardized distance from mean in standard deviation units).