Multivariate + Regression

1. Basics of Multivariate Statistics

Multivariate statistics evaluate relationships among multiple variables to better reflect real-world complexity.

  • Goal: Understand complex structures and relationships across multiple Independent Variables (IVs) and Dependent Variables (DVs).

  • Pros: Offers richer realistic designs, multi-faceted analysis, and controls for Type I and Type II errors.

  • Cons: Requires larger sample sizes (N), harder to interpret, and highly sensitive to underlying assumption violations.

  • Outliers: Scores ≥2 standard deviations from the mean. They distort data, inflate/deflate means, create error noise, and reduce statistical power.

2. General Linear Model (GLM) & Popular Analyses
  • General Linear Model (GLM): Umbrella framework for ANOVA, linear regression, and logistic regression.

  • Multiple Regression (MR): more than one predictor variable; outcome is meausered on a continous scale

  • Logistic Regression: can be any number of predictor variables

  • ANCOVA: 11 continuous DVDV, ≥1≥1 categorical IVsIVs, controlling for ≥1≥1 continuous covariate.

  • MANOVA / MANCOVA: ≥2≥2 continuous DVsDVs, ≥1≥1 categorical IVsIVs (plus covariates for MANCOVA).

3. Role of "Third" Variables
  • Covariates/Controls: Secondary variables that influence the outcome and are controlled for.

  • Moderation: Answers when or for whom a relationship holds true.

  • Mediation: Explains how or why a relationship exists between two variables.

4. Regression Fundamentals

Regression models the relationship between predictor variables (XX-axis) and continuous outcome/criterion variables (YY-axis) by fitting a line to observed data.

  • Types:

    • Simple Regression: 1 predictor variable; outcome is measured on a continuous scale.

    • Multiple Regression: ≥2≥2 continuous predictors (XX), 11 continuous outcome (YY).

    • Logistic Regression: Categorical outcome (YY; binary or multinomial), any number of predictors (XX).

  • Regression Equations:

    • Line equation: Y′=bX+aY′=bX+a

    • Ordinary Least Squares (OLS) model with error: Y′=a+bX+eY′=a+bX+e

    • Where Y= predicted outcome, b = slope, a = Y-intercept, x=predictor variable

    • ee = residual/error.

  • Variance Explained: Quantified by the Coefficient of Determination (R2R2), ranging from 00 (no variance explained) to 11 (perfect prediction).

5. Regression Assumptions & SPSS Diagnostics
  1. Linearity & Homoscedasticity: Checked via scatterplot of standardized residuals (ZRESIDZRESID) vs. standardized predicted values (ZPREDZPRED). Should appear as a random cloud centered at 00. A fan or cone shape indicates a violation of homoscedasticity.

  2. Normality: Residuals should be normally distributed. Tested via a Normal P-P plot (points close to diagonal line) or bell-shaped residual histogram.

  3. Independence of Errors: Evaluated using the Durbin-Watson statistic. Values between 1.51.5 and 2.52.5 are acceptable (ideal value near 2.02.0).

  4. No Multicollinearity: Predictors must not be highly correlated (r<±0.8r<±0.8, Variance Inflation Factor VIF<5VIF<5 to 1010).

6. Sample Size & SPSS Execution
  • Sample Size Guidelines (per IVIV):

    • Assumptions met: 55–10+10+ subjects

    • Assumptions violated: 2020–30+30+ subjects

    • Stepwise MR: 4040–50+50+ subjects

  • Model Entry Methods:

    • Standard (Enter): All predictors entered simultaneously.

    • Hierarchical: Predictors entered in sequential blocks based on theory.

  • Key SPSS Output Tables:

    • ANOVA Table: Reports overall model significance (FF and pp).

    • Model Summary: Reports R2R2 (multiply by 100100 for percentage of outcome variance explained).

    • Coefficients Table: Reports standardized (ββ) and unstandardized (bb) regression weights.