1/79
Wk3
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Statistical significance
How strong the evidence is against the null hypothesis.
Null hypothesis (H₀)
Default/status quo claim assumed true unless evidence suggests otherwise.
Alternative hypothesis (Hₐ)
The competing claim being tested.
Why assume H₀ true?
Hypothesis testing asks how unusual the sample result would be if H₀ were true.
Criminal trial analogy
Defendant innocent = H₀ true; evidence = sample statistic; guilty verdict = reject H₀.
p-value
Probability of observing the sample result (or more extreme) if H₀ is true.
Small p-value
Strong evidence against H₀.
Large p-value
Weak evidence against H₀.
If p-value = 0.03 and α = 0.05
Reject H₀.
If p-value = 0.21 and α = 0.05
Fail to reject H₀.
Simulation-based inference
Uses repeated random samples/randomization to approximate the null distribution.
Why simulation p-values differ
Random chance from repeated simulations.
Theory-based inference
Uses mathematical probability models (usually normal distribution).
Advantage of theory-based tests
Fast, easy, no simulation, same result for everyone.
Disadvantage of theory-based tests
Requires validity conditions.
Normal distribution
Bell-shaped, symmetric probability distribution common in nature.
Central Limit Theorem (CLT)
For large enough n, sample statistics become approximately normal.
CLT for sample proportions
Distribution of p̂ is approximately normal for large n.
Mean of sampling distribution of p̂
π
SD of sampling distribution of p̂
√[π(1−π)/n]
As sample size increases for p̂
Standard deviation decreases.
Centered at π
Means average sample proportion equals true population proportion.
One-proportion z-test
z = (p̂ − π₀) / √[π₀(1−π₀)/n]
What z-score measures
How many standard deviations the sample statistic is from the null value.
If z = 0
Sample statistic equals null value.
If z = 2.5
Sample is 2.5 SD above null; unusual under H₀.
If z = -1.8
Sample is 1.8 SD below null.
Rule of thumb for unusual z
|z| > 2
Validity condition for normal approximation
At least 10 successes and 10 failures.
If validity condition fails
Use simulation/randomization methods.
Why small n is a problem
Distribution may be too discrete or skewed for normal model.
Type I Error
Rejecting a true null hypothesis (false positive).
Type II Error
Failing to reject a false null hypothesis (false negative).
Example Type I Error
Concluding a treatment works when it does not.
Example Type II Error
Concluding a treatment does not work when it does.
Significance level α
Maximum tolerated probability of Type I error.
Common α value
0.05
Population
Entire group of interest.
Sample
Subset selected from the population.
Parameter
Numerical summary of a population (usually unknown).
Statistic
Numerical summary of a sample (calculated from data).
Symbol for population proportion
π
Symbol for sample proportion
p̂
Symbol for population mean
μ
Symbol for sample mean
x̄
Symbol for population SD
σ
Symbol for sample SD
s
Census
Data collected from every member of the population.
Why not always use census
Too expensive, slow, or difficult.
Sampling bias
Systematic tendency to overestimate or underestimate the truth.
Example of biased sample
Surveying only students on campus early morning about housing.
Bias is a property of
Method, not the individual sample.
Voluntary response bias
People choose themselves to respond.
Nonresponse bias
Selected individuals fail or refuse to respond.
Systematic exclusion
Some groups are left out entirely.
Simple Random Sample (SRS)
Every individual and every sample of size n has equal chance of selection.
Why SRS is valuable
Reduces bias and supports valid inference.
Can SRS still be unrepresentative?
Yes, due to random chance.
Small random sample vs huge biased sample
Small random sample is often better.
Sampling variability
Different random samples produce different statistics.
Is sampling variability normal?
Yes, it is expected randomness.
What reduces sampling variability
Larger sample size.
Sampling distribution
Distribution of a statistic over many random samples.
Mean of sampling distribution of x̄
μ
SD of sampling distribution of x̄
σ/√n
As n increases for x̄
Variability decreases.
Finite population condition
Population size should be more than 20 times sample size.
In 200 voters, 118 support candidate A
p̂ = 118/200 = 0.59
Population = all students; sample = 80 selected students
80 selected students is the sample.
If p-value = 0.60
Do not reject H₀.
If p-value = 0.002 and α = 0.05
Reject H₀.
If z = 3.1
Strong evidence against H₀.
If sample size quadruples
Standard deviation is cut in half.
Does bigger sample fix bias?
No, a bigger biased sample is still biased.
Randomness helps with
Reducing systematic bias.
Three pillars of inference
Representative sample, sampling variability, probability model.
Most important formula: z-test
z = (p̂ − π₀) / √[π₀(1−π₀)/n]
Most important formula: SD of p̂
√[π(1−π)/n]
Most important formula: SD of x̄
σ/√n