Describing Bivariate Numerical Data: Correlation and Linear Regression

Correlation (rr)

  • Bivariate Data: Observations on two numerical variables (x,y)(x, y), where xx is the predictor variable and yy is the response variable.
  • Pearson's Sample Correlation Coefficient (rr): Measures the strength and direction of a linear relationship between two numerical variables.
  • Standardized Formula:
    zx=x−xˉsxz_x = \frac{x - \bar{x}}{s_x}

  zy=y−yˉsyz_y = \frac{y - \bar{y}}{s_y}

  r=∑zxzyn−1r = \frac{\sum z_x z_y}{n - 1}

  • Properties of rr:
    • Range is bounded: −1≤r≤1-1 \le r \le 1.
    • Sign indicates direction: positive rr indicates an upward slope; negative rr indicates a downward slope.
    • Extreme values r=1r = 1 or r=−1r = -1 occur only when all points lie exactly on a straight line.
    • Unit-free: Independent of units of measurement for xx and yy
    • Measures linear relationships only; r≈0r \approx 0 means no linear relationship, but a strong nonlinear relationship may still exist.
    • Correlation does not imply causation.

Strength of a linear relationship based on value of r

Least Squares Regression Line

  • Regression Line Equation: Represented as y^=a+bx\hat{y} = a + bx, where y^\hat{y} is the predicted value of yy, aa is the y-intercept, and bb is the slope.
  • Slope (bb): Amount by which yy changes when xx increases by 1 unit1\, \text{unit}:   b=∑(x−xˉ)(y−yˉ)∑(x−xˉ)2b = \frac{\sum (x - \bar{x})(y - \bar{y})}{\sum (x - \bar{x})^2}
  • y-Intercept (aa): Predicted value of yy when x=0x = 0:   a=yˉ−bxˉa = \bar{y} - b\bar{x}
  • Principle of Least Squares: Finds the unique line minimizing the sum of squared vertical deviations (residuals):   SSResid=∑(y−y^)2\text{SSResid} = \sum (y - \hat{y})^2
  • Residual: Difference between observed value yy and predicted value y^\hat{y}:   residual=y−y^\text{residual} = y - \hat{y}
  • Extrapolation Danger: Making predictions for xx values outside the range of observed data is unreliable because the linear pattern may not extend beyond observed limits.

Assessing Model Fit

  • Residual Plot: A scatterplot of (x,residual)(x, \text{residual}) pairs.
    • Curvature or distinct patterns indicate a linear model is inappropriate.
    • Outliers show large vertical distances from zero.
    • Potentially influential observations have extreme xx values separated from main data that significantly alter slope or intercept.
  • Standard Deviation About the Regression Line (ses_e): Measures typical prediction error magnitude:   se=SSResidn−2=∑(y−y^)2n−2s_e = \sqrt{\frac{\text{SSResid}}{n - 2}} = \sqrt{\frac{\sum (y - \hat{y})^2}{n - 2}}
  • Coefficient of Determination (r2r^2): Proportion of total variation in yy explained by the linear relationship with xx:   r2=1−SSResidSSTotalr^2 = 1 - \frac{\text{SSResid}}{\text{SSTotal}}SSTotal=∑(y−yˉ)2\text{SSTotal} = \sum (y - \bar{y})^2
  • Model Quality Criterion: A useful linear model combines a small standard deviation about the line (ses_e) with a high coefficient of determination (r2r^2).

Steps in Linear Regression Analysis

  • Step 1: Construct a scatterplot of bivariate data.
  • Step 2: Visually verify whether the relationship appears approximately linear.
  • Step 3: Calculate slope bb and intercept aa for the least squares regression line y^=a+bx\hat{y} = a + bx
  • Step 4: Plot residuals against xx to confirm line appropriateness and check for outliers or influential points.
  • Step 5: Calculate and interpret ses_e and r2r^2
  • Step 6: Determine overall model suitability for predictions.
  • Step 7: Calculate predictions for xx values within the range of observed data.

Common Pitfalls

  • Concluding causation from statistical correlation.
  • Assuming r≈0r \approx 0 implies no relationship without inspecting scatterplots for non-linear patterns.
  • Swapping response (yy) and predictor (xx) variables; regression of yy on xx differs from xx on yy
  • Extrapolating beyond observed xx data ranges.
  • Interpreting intercept aa when x=0x = 0 lies outside the observed data range.
  • Evaluating a model using only r2r^2 or ses_e without examining residual plots.