1/73
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress

unequal variances does what to mean tests?
unequal variances breaks mean tests

equal population means (μ1 = μ2)
unequal variances (σ²1 ≠ σ²2)
sample means (x̄) are separate
false difference (Type 1 Error)

different means (μ1 ≠ μ2)
unequal variances (σ²1 ≠ σ²2)
sample means (x̄) pulled together
missed difference (Type 2 Error)
When subtracting two independent sample means, V(x̄₁ − x̄₂), do you add or subtract their variances?
You add them: V(x̄₁ − x̄₂) = V(x̄₁) + V( x̄₂)
Subtraction adds uncertainty.
when does pooling make sense?
only when there is one variance to estimate
equal variances
otherwise, pooling estimates a parameter that doesn’t exist
what is the F test’s role in comparing 2 means?
it is the gatekeeper deciding pooled vs unpooled
what is the Behrens-Fisher problem?
The statistical challenge of testing the difference between two population means (μ1 - μ2) when:
The data comes from independent, normal populations.
The population variances σ1 and σ2 are unknown and unequal.
Common Solution: Welch’s t-test, which approximates the degrees of freedom to handle the unequal variances.
Welch-Satterthwaite fixes what?
The distribution (approximate t with v*df)
Cochran-Cox fixes what?
The critical value (weighted average of t critical values)
why do pairwise tests across g groups inflate error?
you inflate the Type I error rate (the probability of finding a false positive, or falsely claiming a significant difference when none exists) because of the cumulative probability of making at least one mistake.
the probability of making at least one false positive, known as the Family-Wise Error Rate (FWER), is calculated as 1 - (1-alpha)^c
what does 5% significance become at g = 5?
originally, there is a 5% chance of committing a Type I error (a false positive).
but at g = 5,
number of comparisons = (g(g-1))/2
= (5(5-1))/2 = (5 × 4)/2 = 10
number of pairwise tests © = 10
so, 1 − 0.95¹⁰ ≈ 0.4013
the family wise error rate is ~40%, which is the true significance level
How does αFW compare to the largest individual αj, and why?
αFW ≥ max αj, because a family error includes any single test's error
the family wide error rate is at least as big as the biggest error rate of any single test in it
When should you do pairwise tests (comparing groups two at a time)?
Only after a main test shows an overall difference, to find out which specific groups are different.
how does Hₐ say for multiple means?
at least ONE pair differs, not every pair
When comparing the means of 3 or more groups, how many hypotheses and p-values does ANOVA use?
ANOVA uses one overall (omnibus) null hypothesis: H₀: μ₁ = μ₂ = ... = μₖ
The alternative hypothesis is that at least one group mean differs from the others
A single F-test produces one p-value for that one hypothesis
The alternative would be running a separate t-test for every pair of groups, which gives many hypotheses and many p-values
why does ANOVA’s method of using one overall null hypotheses matter when comparing the means of 3 or more groups?
With 4 groups, that means 6 pairwise tests [ (4 × 3) / 2 = 12 / 2 = 6 ]
Each extra test adds a chance of a false positive, so the overall Type I error rate rises above α
ANOVA avoids this by testing all means at once and keeping the error rate at α
A significant F-test only says that some difference exists, not which groups differ
Post-hoc tests (such as Tukey's HSD) are used afterward to find which pairs differ
does the product formula always hold?
no, only if the tests are independent
pairwise tests on overlapping groups share data, so its an approximation
which family-wise error rate bounds always hold?
max αj ≤ αfw ≤ min{1, Σαj}
what does SST = SSE + SSM mean?
total variation (SST) equals within group variation plus between-group variation
SSM = variation explained by the model
SSE = variation NOT explained by the model
which family-wise error rate bounds always hold?
max αj ≤ αfw ≤ min{1, Σαj}
αj = significance level of one individual test j (for example 0.05 for each pairwise t-test)
αfw = family-wise error rate, the probability of making at least one false positive across all the tests in the family
Lower bound: max αj ≤ αfw
what does it mean?
If any single test has a false positive, then "at least one false positive" has happened
So the family-wise rate can never be smaller than the biggest individual rate
Running more tests can only keep αfw the same or push it up
Upper bound: αfw ≤ Σαj
what does it mean?
This is the union bound (Boole's inequality): P(A or B or C...) ≤ P(A) + P(B) + P(C)...
The sum can overcount because the "false positive" events may overlap
So adding up the individual rates gives a ceiling
in max αj ≤ αfw ≤ min{1, Σαj}
why is there the "1" in min{1, Σαj}
A probability can't exceed 1
With many tests, the sum can pass 1, so the bound becomes 1
Example with 6 pairwise tests at α = 0.05 each
Test out the fWER bounds: max αj ≤ αfw ≤ min{1, Σαj}
max αj = 0.05
Σαj = 6 × 0.05 = 0.30
So 0.05 ≤ αfw ≤ 0.30
If the tests were independent, αfw = 1 − (0.95)⁶ ≈ 0.265, which sits inside that range
In the ANOVA decomposition SST = SSW + SSB, why does the cross term 2(xij − x̄j)(x̄j − x̄) vanish?
Rewrite each deviation using the group mean: (xij − x̄) = (xij − x̄j) + (x̄j − x̄)
Squaring gives a within term, a between term, and a cross term: 2(xij − x̄j)(x̄j − x̄)
(x̄j − x̄) is constant within a group, so it factors out of the sum
Σ(xij − x̄j) = 0, since deviations from a group's own mean always cancel
Cross term = 0, so SST = SSW + SSB
what do we see if we DO NOT reject H0 in ANOVA?
SSM ≈ 0 and F ≈ 0
SST (Total Sum of Squares)
The total amount of variation or scatter in the dependent variable (Y) around its overall mean (ȳ). It treats the data as if no regression model has been applied yet.
SSR/SSM (Regression / Model Sum of Squares)
The amount of total variation in Y that is explained by your regression model (the distance of predicted values from the mean). A higher value means your model captures a lot of the pattern.
SSE (Error / Residual Sum of Squares)
The amount of total variation in Y that is not explained by your model (the leftover differences between actual observed values and what the model predicted). Lower values mean a tighter, more accurate model fit.
what do we see if we REJECT H0 in ANOVA?
F ≫ 0, and the rejection region is always the right tail.
why can heteroscedascity break regular ANOVA?
pooling into SSE distorts MSE, so F moves for reasons unrelated to the means
what test do you run before ANOVA?
run Bartlett first
fail to reject → regular ANOVA
reject → heteroscedastic ANOVA
what does Welch ANOVA repair?
the numerator (variance-weighted means)
what does Brown-Forsythe repair?
the denominator
which test has higher power on average (the probability that a statistical test correctly detects a true effect or relationship when it actually exists/rejecting a false null hypothesis): Welch or Brown-Forsythe?
Welch
which test is more robust to slight non-normality?
Brown-Forsythe
what numerical check should you always run on decomposition?
SST - (SSE + SSM) should always be 0
which tails reject for the F test on 2 variances?
both tails, with α/2 in each
what is the trap with the “larger variance on top” shortcut?
It still needs α/2 in the right tail.
Using α makes it a size 2α test.
which tails rejects for Bartlett?
right tail of χ²(g−1)
which tail rejects for regular and heteorsceastic ANOVA?
right tail
which tails rejects for the ⋆ sphercity statistic?
left tail
since ⋆ is in [0,1] and small ⋆ is evidence against
what does Wilks’ theorem say about a likelihood ratio statistic L?
−c·ln(L) is approximately chi-square for some positive c, so large -ln(L) rejects.
When the null hypothesis is true and the sample is large, the quantity -2 ln(L) follows approximately a chi-square distribution. The card writes this as -c·ln(L) with "some positive c", and the standard value is c = 2.
You rarely know the exact distribution of L, but this theorem gives you a usable one. You compute -2 ln(L), then compare it to a chi-square critical value.
How should you interpret R output?
Stat versus crit, p versus α, then one plain-English conclusion.
what is the follow-up if Bartlett rejects?
C(g,2) F tests, each with αk/2 per tail
"g choose 2," the number of possible pairs of groups. With g = 4 groups, that is 6 pairs
what the 3 assumptions across repeat across the 2 sample tests?
populations normal
populations independent of each other
observations independent within each sample
how do Bartlett and ANOVA extend the assumptions made in 2 sample tests?
they apply to all g groups
what is the 4th assumption of ANOVA?
Homoscedasticity / equal variance
The variance or spread of data within each comparison group should be roughly equal across all groups
Tested using Bartlett
when does regular ANOVA also assume equal variance?
always, unless you use Welch or Brown-Forsythe
what assumptions do paired tests drop and which do they keep?
paired tests don't require independence between the two populations
BUT
paired tests require independence across different pairs
homoscedastic
equal variances
equicovarianced
all pairwise covariances are equal
compound symmetric
Homoscedastic and equicovarianced
spherical
all pairwise differences have equal variance
which direction does the implication run between compound symmetry and sphericity?
compound symmetry → sphericity, but NOT the reverse
pairwise differences
calculate and store the numerical difference between each possible pair of values from two data columns or groups
why does compound symmetry imply sphericity?
What is Σ when variables are homoscedastic and independent?
Σ = σ²I
a set of random variables or dimensions has equal variance (σ²) and is completely uncorrelated with one another
what is the ⋆ sphericity statistic really checking?
closeness to a scaled identity matrix, which is stronger than sphericity alone
why can’t you use the independent formula for paired data? how do you fix this problem?
dependence makes Cov(X̄₁, X̄₂) unknown
to fix, take differences dk, giving one sample, and run a one-sample T with n − 1 df (n = number of pairs)
What is the covariance structure within a pair?
Cov(X_k⁽ⁱ⁾, X_k⁽ʲ⁾) is not necessarily 0 within the same pair k
What is the covariance structure across pairs?
Cov(X_k⁽ⁱ⁾, X_l⁽ⁱ⁾) = 0 for k ≠ l, which is why differencing works
which normality tests are CDF-based?
KS
CvM
AD
which normality tests are projection-based?
Shapiro-Francia and Shapiro-Wilk
what does Glivenko-Cantelli imply for power?
CDF-based tests have power → 1 almost surely
what do SF and SW need instead?
quantile convergence
what is monte carlo simulation used for?
estimating the distribution of a custom test statistics
which normality checks exist?
qualitative
CDF-based
projection-based
what factors do you check first before doing a test?
normality and independence
what do you check first if data is paired?
take differences and run a one-sample T
what do you check if you have 2 independent groups?
F test on variances
NOT rejected → pooled T
REJECTED → welch or cochran-cox
what do you check first if you have 3+ groups?
Bartlett
NOT rejected → ANOVA
rejected → welch or brown-forsythe
ANOVA rejects. Now what?
Pairwise follow ups
(Be mindful of family wise error rate)
Bartlett’s test for homoscedasticity
statistical method used to check if multiple independent samples have equal variances
follows an approximate Chi-Square distribution with k - 1 degrees of freedom
if the resulting p-value is less than your chosen significance level (commonly α = 0.05), you reject the null hypothesis and conclude that heteroscedasticity (unequal variance) is present