BIOSCI220 Lecture Notes: Regression and Multivariate Data Analysis (Lectures 9–10)
Linear Regression and Model Inference in R (Lecture 9)
- Context: Statistical inference for linear regression models using the penguin bill depth data (penguins_nafree). Topics include fitting different models, interpreting coefficients, making predictions, model selection (ANOVA and AIC), and confidence intervals.
7.1 Fitting regression models in R and interpreting coefficients
General approach: use lm() to fit linear models of the form Yi = α + β1X1i + β2X2i + … + εi, where Yi is the response and Xi are predictors (continuous or categorical).
Example data: billdepthmm as response; billlengthmm as a single continuous predictor (plus optional categorical predictors via dummy variables).
Example model with a single continuous predictor:
- slm <- lm(billdepthmm ~ billlengthmm, data = penguins_nafree)
- output highlights: intercept and slope with standard errors, t-values, and p-values.
Example coefficients (from the slide):
- Intercept ≈ , Std. Error ≈ , t value ≈ , Pr(>|t|) ≈
- billlengthmm coefficient ≈ , Std. Error ≈ , t value ≈ , Pr(>|t|) ≈
The plot typically shows bill depth (mm) vs bill length (mm) with the regression line.
Null (intercept-only) model (no predictors):
- Model: Yi = α + εi
- Purpose: Estimating the average value of the response in the population (the intercept serves as the estimated mean when predictors are absent).
- Example: slmnull
- Intercept estimate shown in the slides: α̂ ≈ (with additional statistics in the summary if shown).
- Null model corresponds to H0: average bill depth (mm) = 0 (as a baseline test of mean).
A note on interpretation:
- The intercept represents the estimated mean of billdepthmm when the predictor is at its reference level (or when the predictor is 0, depending on centering and coding).
- For a single continuous predictor, the slope indicates howbilldepthmm changes with a one-unit change in billlength_mm.
Regression with a single categorical explanatory variable using dummy coding (three penguin species):
- Categorical predictor: species with levels Adelie, Chinstrap, Gentoo.
- Dummy variables: SpeciesChinstrap (1 if Chinstrap, 0 otherwise), SpeciesGentoo (1 if Gentoo, 0 otherwise). Adelie is the reference level.
- Model form (with Adelie as reference):
ext{bill extdepthmm} = eta0 + eta1 ext{SpeciesChinstrap} + eta2 ext{SpeciesGentoo} + eta3 ( ext{billlengthmm}) +
floor - From the slides (example coefficients):
- Intercept (Adelie baseline) ≈
- Coefficient for billlengthmm ≈
- Coefficient for SpeciesChinstrap ≈ (note: in the final form the Chinstrap term adds to the intercept; the slide shows intercept + 0.07 for Chinstrap)
- Coefficient for SpeciesGentoo ≈
- Estimated mean bill depth by species (baseline Adelie):
- Adelie:
- Chinstrap:
- Gentoo:
One continuous variable plus a categorical factor (dummy variables):
- Model form (continuous + categorical):
- Example coefficients (Penguins data):
- Intercept ≈
- billlengthmm ≈
- SpeciesChinstrap ≈
- SpeciesGentoo ≈
- Model form (continuous + categorical):
Continuous predictor with a categorical factor and interaction term:
- Interaction allows the effect of a continuous predictor to differ by category.
- Model form with interaction:
- In R, this is fit as: slm.int <- lm(billdepthmm ~ billlengthmm * species, data = penguins_nafree)
- Alternatively: slm.int <- lm(billdepthmm ~ billlengthmm + species + billlengthmm:species, data = penguins_nafree)
Operators in R formulas: : vs *
- : denotes interaction only (no main effects): billdepthmm ~ billlengthmm : species
- * denotes main effects plus interaction (expanded): billdepthmm ~ billlengthmm * species is equivalent to billdepthmm ~ billlengthmm + species + billlengthmm:species
- Practical implication: use * when you want both main effects and their interaction; use : when you want only the interaction term
7.4 Point predictions from linear regression models
- Prediction with a single continuous predictor:
- For a given billlengthmm x0, predict billdepthmm using the fitted model:
- Prediction with a continuous predictor and a categorical factor:
- Use the fitted coefficients for the appropriate baseline or dummy category
- Example: Prediction for Chinstrap penguins with billlengthmm = 50 using the value-set: intercept 10.57, slope 0.20, Chinstrap dummy -1.93, Gentoo dummy -5.10
- For Chinstrap (Chinstrap = 1, Gentoo = 0):
- Note: The calculation shown in the slides demonstrates how to plug values into the coefficients for a prediction
- For Chinstrap (Chinstrap = 1, Gentoo = 0):
7.5 Model selection using anova() and AIC()
- ANOVA tests: use anova() to compare two (nested) models to see if adding predictors significantly improves fit
- AIC: use AIC() to compare multiple models; smaller AIC indicates a better trade-off between model fit and complexity
- Example: AIC(slmnull, slm, slmsp, slm_int) to compare null, single-predictor, single-predictor + species, and interaction models
- Important caveat: adding more predictors does not always improve predictive performance; always check model assumptions with diagnostics
7.6 Confidence intervals for parameter estimates (CI)
- CI concept: ranges around a population parameter that are believed to contain the true value with a specified probability (e.g., 95%)
- CI interpretation: for a 95% CI, we are 95% confident that the true parameter lies within the interval
- In R, confint() yields confidence intervals for model parameters:
- Example: cis <- confint(slm_sp) # level by default 0.95
- Interpreting CIs in context:
- We can interpret the intercept as the mean bill depth for the baseline category when other predictors are at their reference values
- We can interpret each coefficient as the difference in the mean response relative to the reference level, given other variables in the model
- Example interpretation (from slides):
- We are 95% confident that the average bill depth of an Adelie penguin is between 9.2 and 11.9 mm, given the other variables in the model.
- We are 95% confident that the average bill depth of a Chinstrap penguin is between 1.5 and 2.4 mm shallower than the Adelie penguin, given the other variables in the model.
Summary of regression notes
- There can be many different models for the same data; use them to estimate population parameters and test hypotheses
- Model selection methods (ANOVA and AIC) help choose a “best” model, but adding complexity is not always beneficial
- Confidence intervals provide a quantitative assessment of uncertainty in parameter estimates
Multivariate data and nMDS (Lecture 10)
8. Learning outcomes (8.1–8.6)
- 8.1 Define multivariate data: one observation has many variables; data form a multivariate cloud in a high-dimensional space
- 8.2 Dimension reduction: reduce to fewer dimensions while preserving as much relevant information as possible
- 8.3 Non-metric multidimensional scaling (nMDS): a dimension-reduction method for multivariate data that preserves rank-order relationships between samples using a distance matrix
- 8.4 Distance measures: Euclidean vs Bray-Curtis; guidance on when to use each
- 8.5 Implement nMDS in R, plot 2-D solutions, and diagnostic plots
- 8.6 Interpret and communicate nMDS outputs
8. What is multivariate data?
- An observation has multiple variables; you place observations in a multivariate data cloud in high-dimensional space
- Challenge: interpretation and inference can be difficult; ordination provides a practical approximation of true relationships
8. Dimension reduction
- Goal: represent high-dimensional data in fewer dimensions while retaining as much information as possible
- Common methods: PCA, PCoA, cluster analysis, nMDS, etc.
- Rationale: improves interpretability and visualization for complex datasets
8. Why dimension reduction?
- Improves interpretability and visualization
- Aids communication of complex relationships
8.3 Non-metric multidimensional scaling (nMDS)
- Purpose: create a low-dimensional (usually 2-D or 3-D) representation of samples that preserves the rank order of pairwise distances
- Data input: any multivariate data; uses a distance matrix and rank-order distances
- Process: iterative optimization to place samples in a reduced-dimension space so that the rank-order of distances in the ordination matches the rank-order in the original distance matrix
- Key output concept: ordination coordinates (e.g., NMDS1, NMDS2) and a goodness-of-fit measure (stress)
- Diagnostic: Kruskal’s stress statistic indicates how well the ordination preserves the rank-order distances
8. The concept of distance in multivariate space
- Closer samples are more similar; farther samples are more different
- Distances can be computed in the original high-dimensional space; nMDS uses these distances to produce a low-dimensional plot
8.4 Distance matrices
- Definition: a square matrix of pairwise distances between samples
- Example visualization: a “distance matrix” showing distances between N samples
8.4 Distance measures: Euclidean vs Bray-Curtis
- Euclidean distance (L2):
d_E(i,j) = igg( rac{1}{p} igg)^{1/2} imes rac{1}{?} ext{(basic form)}
Let me present the standard form:
d_E(i,j) = igg(igg)- The standard definition is:
- Bray-Curtis distance (BC):
- In ecological and multivariate contexts, BC is commonly used and emphasizes contributions of zeros and small values
- Practical note (from slides):
- BC distance weighting matters; double zeros can be treated as unimportant; zeros and low values weight differently than high values
- Euclidean distance (L2):
8.4 Your Turn: distance calculations
- Hands-on exercises compare Euclidean and Bray-Curtis distances between samples using given data
- Observations: double values and double zeros influence the calculated distance differently depending on the metric; BC is sensitive to weights from the denominator (sum of values)
8.3–8.4 How NMDS works (operational outline)
- Start with an initial random configuration of samples in k dimensions
- Compute original distance matrix (rank-ordered)
- Place samples in k-D space to minimize Kruskal’s stress (fit between original distances and ordination distances)
- Use iterative improvements: progressively better configurations (ordinal alignment improves) until convergence
- Kruskal’s stress: a measure of goodness-of-fit; lower values indicate a better representation
- Interpretive thresholds (from slides):
- Stress < 0.05 = excellent
- Stress < 0.10 = good
- Important note: adding more samples or variables often increases stress unless the model improves; stress tends to increase with data complexity but decreases with more dimensions (careful interpretation required)
8.5 Visual interpretation and diagnostics
- 2-D NMDS plots show samples in a plane, with closeness indicating similarity
- Diagnostic plots and stress values help assess the reliability of the ordination
8.6 Communicating NMDS results
- Report NMDS coordinates (e.g., NMDS1, NMDS2) for samples or groups
- Report stress value and interpretation of fit
- Discuss clustering or separation of sample groups and potential ecological or biological interpretations
9. Practical notes and tips
- Distances and ordinations depend on the data scale and transformation; sometimes standardize or normalize data before computing distances
- The choice of distance measure matters; Euclidean treats all variables similarly, while Bray-Curtis emphasizes abundance and relative differences, with particular handling for zeros and low counts
- In environmental and ecological studies, BC is commonly used due to its interpretation in terms of community composition
- Example visuals and exercises shown in slides
- Example plots show billlengthmm vs billdepthmm with species labeling and ordinal axes for NMDS (e.g., NMDS1, NMDS2) to illustrate group separation
- Example distance matrices and NMDS ordination are used to demonstrate how different distance measures produce different representations of similarity
- Administrative notes
- The module includes practical exercises and tutorials (e.g., H5P NMDS tutorial, quizzes)
- Office hours and lecturers listed for questions about Module 1 and the test
- Key takeaways
- You can create multiple regression models and compare them using ANOVA and AIC to determine the best balance of fit and complexity
- Confidence intervals quantify parameter uncertainty and can be interpreted in context to make probabilistic statements about population parameters
- Multivariate data analysis with NMDS provides a way to visualize and interpret high-dimensional data via distance-based ordination, with stress as a primary diagnostic metric
Quick reference formulas (LaTeX)
- Linear model (single continuous predictor):
- Null model:
- Categorical predictors with dummy coding (Adelie as reference):
- Interaction model: full form
- Model comparison: AIC
- Confidence intervals for parameters (example):
- Euclidean distance between samples i and j:
- Bray-Curtis distance between samples i and j:
- Kruskal’s stress (NMDS goodness-of-fit):
- Linear model (single continuous predictor):
End of notes