1/27
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
The Normal Distribution and Z-scores
The Normal Distribution and Z-scores
Distributions
Distributions
Different Types of Distributions
Unimodal Bimodal Uniform
Skew
Positive Skew: Mode, Median, Mean Symmetrical Distribution: Mean, Median, Mode Negative Skew: Mean, Median, Mode
Kurtosis
Kurtosis is a statistical measure that tells us whether a distribution is more or less peaked than the normal distribution. > 3: Leptokurtic (more peaked) = 3: Mesokurtic (normal distribution) < 3: Platykurtic (Less peaked)
Transforming our Data
● Many analytical methods rely on an assumption of distribution symmetry ● When a distribution is skewed, analyses yield invalid results ● Common transformation methods to handle skewed data ○ Log transformations: x = log(x) ○ Square root transformation: x = sqrt(x) ○ Cube root transformation: x = cbrt(x) ○ Box-cox transformation: uses a chosen parameter λ to optimally approximate a normal distribution ● Researchers should take care when transforming and interpreting data
Transforming a Variable in R
We’ve been looking at examining pre-existent variables… let’s “brew” some of our own!
Normal Distributions
Normal Distributions
Normal Distribution
● A probability distribution that is symmetric around the mean ● Also called a Gaussian distribution or a bell-shaped curve ● Holds unique characteristics ● Many distributions can be normal, even though they might have different means and standard deviations
Properties of the Normal Distribution
● Normal distributions are symmetric around their mean ● The mean, median, and mode are equal ● The area under the normal curve is equal to 1.0 ● Denser in the center and less dense in the tails ● Defined by two parameters ○ Mean (μ) determines the center ○ Standard Deviation (σ) determines the spread ● 68-95-99.7 Rule ○ ~68% within ± 1σ ○ ~95% within ± 2σ ○ ~99.7% within ± 3σ
The Standard Normal Distribution
The standard normal distribution is a specific type of normal distribution where μ is always 0 and σ is always 1. This distribution specifically is used to compare data from different normal distributions by converting values into standardized z-scores.
Transforming Raw Data
Raw Test Score Standardized Test Score
Z-scores
Z-scores
Z-scores
● Z-scores convert raw scores to standardized metrics ● Allows us to understand the relative location of a score within its distribution ○ Enables us to compare across data sets, even when original scale isn’t equal
Z-scores
Z-scores from a population z = (x-μ)/σ Z-scores from a sample z = (x-M)/s Z-scores combine information about characteristics of the distribution that help interpret raw scores
Interpreting Z-scores
Sign Positive: the score is above the mean (right tail) Negative: the score is below the mean (left tail) Magnitude How far away (in units of σ) the score is from μ Why is this usually between -3 & 3?
What Does This Mean?
● A z-score of 1.5 ● A z-score of –1.5 ● A z-score of –0.5 ● A z-score of 2.9 What if these values represented the age of dogs in a sample?
Converting Raw Scores
Imagine you have the following heights (in centimeters) across a population of 10 adults: 160, 165, 170, 175, 180, 182, 158, 165, 190, 175 Where does a height of 160 cm fall in this distribution? Where does a height of 190 cm fall in this distribution? z = (x-μ)/σ What’s μ? 172 What’s σ? 9.74
Finding Probabilities Using Z-Scores
● Standard Normal Table ○ Provides cumulative probabilities associated with each z score ● Example: ○ For z = 1.00, cumulative probability is 0.8413 ○ 84% of values fall below the value associated with a z-score of 1.00
Finding Probabilities Using Z-Scores
● Example: ○ For z = 1.27, cumulative probability is 0.8980 ○ 89% of values fall below the value associated with a z-score of 1.27 ● To find the percent of values that lie above a z-score, we can use the cumulative probability as well ○ Area under the curve = 1
Application
Imagine the heights of adults are normally distributed with μ = 172 and σ = 9.74. Find the percentage of adults with heights between 160 cm and 190 cm? What’s the cumulative probability for 160 cm? 0.1093 What’s the cumulative probability for 190 cm? 0.9678 0.9678 - 0.1093 = 0.8585 = 85.85%
Step-by-Step
Mean & Sample Size μ Deviation & Sums of Squares SS Variance σ² Standard Variation σ Z-Score z Probability %
Population vs. Sample
Population vs. Sample
Why Are They Different?
● Parameter vs. Statistic ○ Parameter: numerical value that describes a population ○ Statistic: numerical value that describes a sample ● A sample is a portion of the population selected for study ○ Used to make inferences about a larger population ● Sampling methods can do their best to minimize bias ○ Adjusting sample statistics (s²) and degrees of freedom (df) reduces bias as well ○ These differences ensure sample statistics accurately reflect the population
Population vs. Sample
Population Variance σ² = SS/N SD σ = √σ² Z-scores z = (x-μ)/σ Sample Variance s² = SS/(N-1) SD s = √s² Z-scores z = (x-M)/s
Think-Pair-Share
In a study of the migration patterns of the Monarch butterfly, researchers collected data on the number of butterflies observed in a specific region during their migration. The researchers found the following information: ● Population Standard Deviation (σ): 10 butterflies ● Population Size (N): 100 butterflies Now, suppose we treat this population as a sample of butterflies across multiple regions. What is the sample variance? What is the sample SD?
Think-Pair-Share Solution
Notes
● The SS for the population and the sample is the same ● The variance will always be larger for a sample than for if scores were the full population ○ Think about the denominators in both calculations ■ e.g., 99/11 vs. 99/10 ○ As sample size increases, the effect of subtracting becomes smaller ■ e.g., 99/33 vs. 99/32 ○ Larger N brings estimates of sample variance closer to that of the population variance Larger sample sizes better reflect the population