Multivariate + Regression
1. Basics of Multivariate Statistics
Multivariate statistics evaluate relationships among multiple variables to better reflect real-world complexity.
Goal: Understand complex structures and relationships across multiple Independent Variables (IVs) and Dependent Variables (DVs).
Pros: Offers richer realistic designs, multi-faceted analysis, and controls for Type I and Type II errors.
Cons: Requires larger sample sizes (N), harder to interpret, and highly sensitive to underlying assumption violations.
Outliers: Scores ≥2 standard deviations from the mean. They distort data, inflate/deflate means, create error noise, and reduce statistical power.
2. General Linear Model (GLM) & Popular Analyses
General Linear Model (GLM): Umbrella framework for ANOVA, linear regression, and logistic regression.
Multiple Regression (MR): more than one predictor variable; outcome is meausered on a continous scale
Logistic Regression: can be any number of predictor variables
ANCOVA: 11 continuous DVDV, ≥1≥1 categorical IVsIVs, controlling for ≥1≥1 continuous covariate.
MANOVA / MANCOVA: ≥2≥2 continuous DVsDVs, ≥1≥1 categorical IVsIVs (plus covariates for MANCOVA).
3. Role of "Third" Variables
Covariates/Controls: Secondary variables that influence the outcome and are controlled for.
Moderation: Answers when or for whom a relationship holds true.
Mediation: Explains how or why a relationship exists between two variables.
4. Regression Fundamentals
Regression models the relationship between predictor variables (XX-axis) and continuous outcome/criterion variables (YY-axis) by fitting a line to observed data.
Types:
Simple Regression: 1 predictor variable; outcome is measured on a continuous scale.
Multiple Regression: ≥2≥2 continuous predictors (XX), 11 continuous outcome (YY).
Logistic Regression: Categorical outcome (YY; binary or multinomial), any number of predictors (XX).
Regression Equations:
Line equation: Y′=bX+aY′=bX+a
Ordinary Least Squares (OLS) model with error: Y′=a+bX+eY′=a+bX+e
Where Y= predicted outcome, b = slope, a = Y-intercept, x=predictor variable
ee = residual/error.
Variance Explained: Quantified by the Coefficient of Determination (R2R2), ranging from 00 (no variance explained) to 11 (perfect prediction).
5. Regression Assumptions & SPSS Diagnostics
Linearity & Homoscedasticity: Checked via scatterplot of standardized residuals (ZRESIDZRESID) vs. standardized predicted values (ZPREDZPRED). Should appear as a random cloud centered at 00. A fan or cone shape indicates a violation of homoscedasticity.
Normality: Residuals should be normally distributed. Tested via a Normal P-P plot (points close to diagonal line) or bell-shaped residual histogram.
Independence of Errors: Evaluated using the Durbin-Watson statistic. Values between 1.51.5 and 2.52.5 are acceptable (ideal value near 2.02.0).
No Multicollinearity: Predictors must not be highly correlated (r<±0.8r<±0.8, Variance Inflation Factor VIF<5VIF<5 to 1010).
6. Sample Size & SPSS Execution
Sample Size Guidelines (per IVIV):
Assumptions met: 55–10+10+ subjects
Assumptions violated: 2020–30+30+ subjects
Stepwise MR: 4040–50+50+ subjects
Model Entry Methods:
Standard (Enter): All predictors entered simultaneously.
Hierarchical: Predictors entered in sequential blocks based on theory.
Key SPSS Output Tables:
ANOVA Table: Reports overall model significance (FF and pp).
Model Summary: Reports R2R2 (multiply by 100100 for percentage of outcome variance explained).
Coefficients Table: Reports standardized (ββ) and unstandardized (bb) regression weights.