Describing Bivariate Numerical Data: Correlation and Linear Regression
Correlation (r)
Bivariate Data: Observations on two numerical variables (x,y), where x is the predictor variable and y is the response variable.
Pearson's Sample Correlation Coefficient (r): Measures the strength and direction of a linear relationship between two numerical variables.
Standardized Formula: zx=sxx−xˉ
zy=syy−yˉ
r=n−1∑zxzy
Properties of r:
Range is bounded: −1≤r≤1.
Sign indicates direction: positive r indicates an upward slope; negative r indicates a downward slope.
Extreme values r=1 or r=−1 occur only when all points lie exactly on a straight line.
Unit-free: Independent of units of measurement for x and y
Measures linear relationships only; r≈0 means no linear relationship, but a strong nonlinear relationship may still exist.
Correlation does not imply causation.
Least Squares Regression Line
Regression Line Equation: Represented as y^=a+bx, where y^ is the predicted value of y, a is the y-intercept, and b is the slope.
Slope (b): Amount by which y changes when x increases by 1unit:
b=∑(x−xˉ)2∑(x−xˉ)(y−yˉ)
y-Intercept (a): Predicted value of y when x=0:
a=yˉ−bxˉ
Principle of Least Squares: Finds the unique line minimizing the sum of squared vertical deviations (residuals):
SSResid=∑(y−y^)2
Residual: Difference between observed value y and predicted value y^:
residual=y−y^
Extrapolation Danger: Making predictions for x values outside the range of observed data is unreliable because the linear pattern may not extend beyond observed limits.
Assessing Model Fit
Residual Plot: A scatterplot of (x,residual) pairs.
Curvature or distinct patterns indicate a linear model is inappropriate.
Outliers show large vertical distances from zero.
Potentially influential observations have extreme x values separated from main data that significantly alter slope or intercept.
Standard Deviation About the Regression Line (se): Measures typical prediction error magnitude:
se=n−2SSResid=n−2∑(y−y^)2
Coefficient of Determination (r2): Proportion of total variation in y explained by the linear relationship with x:
r2=1−SSTotalSSResidSSTotal=∑(y−yˉ)2
Model Quality Criterion: A useful linear model combines a small standard deviation about the line (se) with a high coefficient of determination (r2).
Steps in Linear Regression Analysis
Step 1: Construct a scatterplot of bivariate data.
Step 2: Visually verify whether the relationship appears approximately linear.
Step 3: Calculate slope b and intercept a for the least squares regression line y^=a+bx
Step 4: Plot residuals against x to confirm line appropriateness and check for outliers or influential points.
Step 5: Calculate and interpret se and r2
Step 6: Determine overall model suitability for predictions.
Step 7: Calculate predictions for x values within the range of observed data.
Common Pitfalls
Concluding causation from statistical correlation.
Assuming r≈0 implies no relationship without inspecting scatterplots for non-linear patterns.
Swapping response (y) and predictor (x) variables; regression of y on x differs from x on y
Extrapolating beyond observed x data ranges.
Interpreting intercept a when x=0 lies outside the observed data range.
Evaluating a model using only r2 or se without examining residual plots.