ANOVA and Correlation Notes Chapter 12: ANOVA (Analysis of Variance) Used to evaluate mean differences between two or more treatment groups. Uses sample data to draw conclusions about populations. Advantage over T-test: Can compare more than two groups simultaneously. Provides more flexibility in study design and interpretation. Goal: Determine if observed mean differences among samples are significant enough to conclude there are mean differences among populations. Hypotheses Null Hypothesis ( (H0) ) : No differences in populations; observed sample mean differences are due to chance or error. Represented as H < / e m > 0 : μ < e m > 1 = μ < / e m > 2 = μ 3 H</em>0 : \mu<em>1 = \mu</em>2 = \mu_3 H < / e m > 0 : μ < e m > 1 = μ < / e m > 2 = μ 3 .Alternative Hypothesis ( (H1) ) : Populations have different means, causing differences in sample means. Represented as H < / e m > 1 : μ < e m > 1 ≠ μ < / e m > 2 ≠ μ < e m > 3 H</em>1 : \mu<em>1 \neq \mu</em>2 \neq \mu<em>3 H < / e m > 1 : μ < e m > 1 = μ < / e m > 2 = μ < e m > 3 or H < / e m > 1 : μ < e m > 1 = μ < / e m > 2 H</em>1 : \mu<em>1 = \mu</em>2 H < / e m > 1 : μ < e m > 1 = μ < / e m > 2 , but (\mu_3) is different.Factors and Levels Factor : The independent variable designating the groups being compared.Levels of the Factor : The conditions that make up the factor.Factorial Design : A study that combines two or more factors.Type I Error and ANOVA ANOVA limits Type I error by performing all comparisons at once. Testwise Alpha Level : The risk of Type I error for a single hypothesis test.Experimentwise Alpha Level : The total probability of a Type I error accumulated from all individual tests. ANOVA has a lower experimentwise alpha error.ANOVA Test Statistic F-ratio: F = V a r i a n c e b e t w e e n s a m p l e m e a n s V a r i a n c e e x p e c t e d w i t h n o t r e a t m e n t e f f e c t F = \frac{Variance \ between \ sample \ means}{Variance \ expected \ with \ no \ treatment \ effect} F = V a r ian ce e x p ec t e d w i t h n o t r e a t m e n t e f f ec t V a r ian ce b e tw ee n s am pl e m e an s Conducting ANOVA Test Determine total variability.Between-Treatment Variance : Measures differences between sample means to get an overall variance measure.Within-Treatment Variance : Measures variability/differences between values within each treatment. Between-Treatment Variance Can occur for two reasons:Naturally occurring differences (sampling error). Differences caused by treatment effects. Measurement: Compare variances between treatments against sampling error without treatment effect. Within-Treatment Variance Provides a measure of how big the differences are when (H_0) is true. F-Ratio for Independent Measures Compares between- and within-treatment components. F = V a r i a n c e b e t w e e n t r e a t m e n t s ( d i f f e r e n c e s i n c l u d i n g a n y t r e a t m e n t e f f e c t s ) V a r i a n c e w i t h i n t r e a t m e n t s ( d i f f e r e n c e s w i t h n o t r e a t m e n t e f f e c t s ) F = \frac{Variance \ between \ treatments \ (differences \ including \ any \ treatment \ effects)}{Variance \ within \ treatments \ (differences \ with \ no \ treatment \ effects)} F = V a r ian ce w i t hin t r e a t m e n t s ( d i f f er e n ces w i t h n o t r e a t m e n t e f f ec t s ) V a r ian ce b e tw ee n t r e a t m e n t s ( d i f f er e n ces in c l u d in g an y t r e a t m e n t e f f ec t s ) If F-ratio is near 1.00, treatments are random and unsystematic, suggesting no treatment effect. If F-ratio is larger than 1.00, the numerator is significantly larger than the denominator, indicating significant differences between treatments. Error Term : The denominator of the F-ratio, measuring only random variability.ANOVA Notation (k): Number of treatment conditions/separate samples. (n): Number of scores in each treatment. (N): Total number of scores in the entire study. (T): (\sum X) sum of scores for each treatment condition. (G): Sum of all scores in the study. Sample Variance: (s^2) Sum of Squares Total: S S t o t a l = ∑ X 2 − G 2 N SS_{total} = \sum X^2 - \frac{G^2}{N} S S t o t a l = ∑ X 2 − N G 2 Degrees of Freedom Total: d f t o t a l = N − 1 df_{total} = N - 1 d f t o t a l = N − 1 Sum of Squares Within-Treatment: S S w i t h i n t r e a t m e n t = ∑ S S o f e a c h t r e a t m e n t g r o u p SS_{within\ treatment} = \sum SS \ of \ each \ treatment \ group S S w i t hin t r e a t m e n t = ∑ S S o f e a c h t r e a t m e n t g r o u p Degrees of Freedom Within: d f w i t h i n = ∑ ( n − 1 ) = ∑ d f i n e a c h t r e a t m e n t df_{within} = \sum (n - 1) = \sum df \ in \ each \ treatment d f w i t hin = ∑ ( n − 1 ) = ∑ df in e a c h t r e a t m e n t Sum of Squares Between-Treatment: S S < e m > b e t w e e n = S S < / e m > t o t a l − S S w i t h i n SS<em>{between} = SS</em>{total} - SS_{within} S S < e m > b e tw ee n = S S < / e m > t o t a l − S S w i t hin Degrees of Freedom Between: d f b e t w e e n = k − 1 df_{between} = k - 1 d f b e tw ee n = k − 1 F-Ratio Calculation: F = M S < e m > b e t w e e n M S < / e m > w i t h i n F = \frac{MS<em>{between}}{MS</em>{within}} F = M S < / e m > w i t hin M S < e m > b e tw ee n , where (MS) represents Mean Square. ANOVA and Hypothesis Testing Distribution of F: Expected to be around 1.00.F-ratios are always positive. When (H_0) is true, the ratio should be near 1.00. With smaller df, F values are more spread out. Steps for Hypothesis Testing State Hypotheses:(H0 : \mu 1 = \mu2 = \mu 3) (H_1): At least one of the treatment means is different. Locate the Critical Region for F:Find df total, within, and between. Use df within and df between to find the critical F-value. Compute the Observed F Test Statistic:Find all 3 SS (total, within, between). Calculate Mean Squares. Calculate F. Make a Statistical Decision About the Null Hypothesis:If F value is within the critical region, reject (H_0) and accept the alternative. If F value is not within the critical region, fail to reject (H_0). Chapter 14: Two-Factor Analysis Two-Factor ANOVA Allows examination of three types of mean differences within one analysis. Goal: Evaluate mean differences produced by factors acting independently or together. Main Effects Main Effect: Mean differences among levels of one factor. Example: (\delta - 4 = 4) is the main effect for factor gender. Evaluation of main effects involves 2/3 hypothesis tests for 2-factor ANOVA. Hypotheses:For Factor A: (H0 : M {a1} = M{a2}) and (H 1 : M{a1} \neq M {a2}) For Factor B: (H0 : M {b1} = M{b2}) and (H 1 : M{b1} \neq M {b2}) Interactions Interaction: Any extra differences not caused by the main effects. If the difference between factors is not constant, there is an interaction. Calculated by: F = V a r i a n c e ( m e a n d i f f e r e n c e s n o t e x p l a i n e d b y m a i n e f f e c t ) V a r i a n c e ( d i f f e r e n c e s e x p e c t e d i f n o t r e a t m e n t ) F = \frac{Variance \ (mean \ differences \ not \ explained \ by \ main \ effect)}{Variance \ (differences \ expected \ if \ no \ treatment)} F = V a r ian ce ( d i f f er e n ces e x p ec t e d i f n o t r e a t m e n t ) V a r ian ce ( m e an d i f f er e n ces n o t e x pl ain e d b y main e f f ec t ) Hypotheses:(H_0): There is no interaction between A and B. (H_1): There is an interaction between A and B. If levels of one factor depend on the levels of the other, there is an interaction. Non-parallel lines on a graph indicate an interaction. Two-Factor ANOVA Hypothesis Test A effect. B effect. (A \times B) Interaction. Examples of Main Effects and Interactions Main effect of 10, no interaction: +10, +10, +20, +20 Main effect, interaction: +10, +10, -10, -10 Analysis of Two-Factor ANOVA Total variance is split into between-treatment and within-treatment variance. Between-treatment variance is further divided into factor A variance, factor B variance, and interaction variance. Stages of Analysis Total Variability:S S t o t a l = ∑ X 2 − G 2 N SS_{total} = \sum X^2 - \frac{G^2}{N} S S t o t a l = ∑ X 2 − N G 2 d f t o t a l = N − 1 df_{total} = N - 1 d f t o t a l = N − 1 Within-Treatment:S S w i t h i n = ∑ S S e a c h t r e a t m e n t SS_{within} = \sum SS \ each \ treatment S S w i t hin = ∑ S S e a c h t r e a t m e n t d f w i t h i n = ∑ d f e a c h t r e a t m e n t df_{within} = \sum df \ each \ treatment d f w i t hin = ∑ df e a c h t r e a t m e n t Between-Treatment:S S < e m > b e t w e e n = S S < / e m > t o t a l − S S w i t h i n SS<em>{between} = SS</em>{total} - SS_{within} S S < e m > b e tw ee n = S S < / e m > t o t a l − S S w i t hin df_{between} = # \ of \ cells - 1 Factor A:S S A = … SS_A = … S S A = … df_A = # \ of \ Rows - 1 Factor B:S S B = … SS_B = … S S B = … df_B = # \ of \ Columns - 1 (A \times B) Interaction:S S < e m > A × B = S S < / e m > b e t w e e n − S S < e m > A − S S < / e m > B SS<em>{A \times B} = SS</em>{between} - SS<em>A - SS</em>B S S < e m > A × B = S S < / e m > b e tw ee n − S S < e m > A − S S < / e m > B d f < e m > A × B = d f < / e m > b e t w e e n − d f < e m > A − d f < / e m > B df<em>{A \times B} = df</em>{between} - df<em>A - df</em>B df < e m > A × B = df < / e m > b e tw ee n − df < e m > A − df < / e m > B F-Ratios and Critical Values 3 F-ratios for the 3 variances. Finding Critical F Value:Numerator: between groups df. Denominator: within groups df. d f < e m > w i t h i n = d f < / e m > t o t a l − d f < e m > A − d f < / e m > B − d f A × B df<em>{within} = df</em>{total} - df<em>A - df</em>B - df_{A \times B} df < e m > w i t hin = df < / e m > t o t a l − df < e m > A − df < / e m > B − d f A × B (\alpha = 0.05) Critical F = (F(df{between}, df {within})) Chapter 15: Correlation Correlation Measures and describes the relationship between two variables by measuring different variables for each individual. Characteristics of a Relationship Direction (negative or positive). Form (linear - Pearson correlation). Pearson Correlation Measures the degree and direction of a linear relationship. r = C o v a r i a b i l i t y o f X a n d Y V a r i a b i l i t y o f X a n d Y s e p a r a t e l y r = \frac{Covariability \ of \ X \ and \ Y}{Variability \ of \ X \ and \ Y \ separately} r = V a r iabi l i t y o f X an d Y se p a r a t e l y C o v a r iabi l i t y o f X an d Y Sum of Products of Deviations Similar to SS, measures the amount of covariability between two variables. Definitional Formula: S P = ∑ ( X − M < e m > X ) ( Y − M < / e m > Y ) SP = \sum (X - M<em>X)(Y - M</em>Y) S P = ∑ ( X − M < e m > X ) ( Y − M < / e m > Y ) r = S P S S < e m > X S S < / e m > Y r = \frac{SP}{\sqrt{SS<em>X SS</em>Y}} r = S S < e m > X S S < / e m > Y S P Hypothesis Testing Asks whether a correlation exists in the population. Hypotheses:(H_0 : \rho = 0) ((\rho) = population correlation) (H_1 : \rho \neq 0) Test Statistic: df = n - 2 Effect Size: