Comprehensive Notes on Linear Regression, Multiple Testing, and Publication Bias
Course Logistics and Administrative Information
Course Operations & Communication Protocols:
- Holiday Note: Labor Day was acknowledged in the lecture chat.
- Question Submission: All questions during lecture must be directed to "everyone" in the chat window.
- Course Website: Serves as the primary hub for syllabus materials, homework assignments, and exam instruction sheets.
- Homework 4: Official due date is set for this Friday.
- Final Exam Schedule: Scheduled for Wednesday, exactly one week from the lecture date, running from to .
Final Exam Logistics & Protocols:
- Instruction Release Time: The exam instruction document (
finalinstructions.text) will be posted on the course website at on Wednesday, September 9. - Zoom Login Requirement: Students must log into the designated Zoom room at and remain present on Zoom for the duration of the exam.
- Answer Submission: Exam answers must be emailed directly to the instructor following the precise instructions outlined in the official released document.
- Exam Length and Question Format: Consists of approximately questions total, primarily multiple-choice with potentially one short-answer question.
- Scope and Content Weighting: Cumulative covering all material across the term, but weighted more heavily toward concepts introduced during the final week of class.
- Permitted Testing Resources: The exam is open-note, allowing students to reference printed or digital notes during the exam.
- Instruction Release Time: The exam instruction document (
Multiple Testing and Publication Bias
Fundamentals of -Values and Significance Cutoffs:
- Meaning of -Value: A -value of () indicates that if the null hypothesis () is true, there is a probability of obtaining sample data as extreme as or more extreme than the observed data.
- Standard Significance Threshold (): The conventional significance level cutoff is set at ().
- False Positive Rate under : When evaluating unrelated variables where the null hypothesis is true, a statistically significant result will be generated of the time by pure chance, yielding a Type I error.
- Necessity of Confirmation Studies: Because any single study carries an inherent false-positive rate under , initial significant findings (such as drug efficacy) require independent follow-up confirmation studies before conclusions can be validated.
The Multiple Testing Problem:
- Context and Definition: Occurs when a researcher tests many explanatory variables or performs multiple hypothesis tests simultaneously within a single dataset.
- Survey Example: Investigating potential causes of a rare condition by surveying participants on dozens of behavioral, dietary, and exercise variables increases the likelihood of finding at least one false positive.
- Mathematical Demonstration ( Independent Tests under True Null Hypotheses at ):
- Expected Number of Type I Errors:
- Probability of No Type I Error on a Single Test:
- Probability of Zero Type I Errors across Independent Tests:
- Probability of Making AT LEAST ONE Type I Error across Independent Tests:
Bonferroni's Correction:
- Purpose: A method used to control the overall family-wise error rate when conducting hypothesis tests or evaluating explanatory variables.
- Adjusted Significance Level Formula:
- Theoretical Guarantee: Guarantees that the overall probability of making at least one Type I error across all tests combined remains less than or equal to (
- Application Examples:
- For tests, corrected significance threshold is:
- For tests, corrected significance threshold is:
Publication Bias:
- Definition and Process: Occurs when journal editors and reviewers publish studies that demonstrate statistically significant findings () while rejecting or ignoring studies that show non-significant findings.
- Illustrative Scenario: Assume independent research teams test a drug that actually has zero true effect ( is true).
- Approximately teams will find no significant effect ().
- Approximately teams will find a statistically significant effect () due to random sampling variation.
- If only the significant studies are published while the non-significant studies are suppressed, the published record falsely indicates that the drug works.
- Scientific Solution: Journals should publish methodologically sound research regardless of whether the results achieve statistical significance.
Case Study on ADHD Stimulants (Staley et al., 2024):
- Background: Numerous published studies claimed that stimulant medications like methylphenidate are effective treatments for ADHD. Out of published articles in the literature, the majority concluded that methylphenidate was effective.
- Meta-Analysis Methodology: Staley et al. (2024) conducted a comprehensive audit analyzing all available experimental data, comprising published articles and unpublished articles ( total studies).
- Statistical Analysis: Pooled data comparing methylphenidate treatment groups to placebo groups using a two-sample -test (due to large pooled sample size).
- Pooled Test Results:
- Test Statistic:
- Calculated -value:
- Standard Cutoff:
- Conclusion: Incorporating unpublished studies demonstrates that methylphenidate performs no better than a placebo (). The null hypothesis is consistent with the complete dataset, highlighting significant publication bias within the published literature.
Testing Correlation via Simulation Methods
Statistical Framework for Correlation Testing:
- Objective: Evaluate whether an observed sample correlation between two quantitative variables indicates a true population correlation () or is attributable to random sampling variability under
Case Study on Emotional Loneliness and Prolonged Grief (Vetter et al., 2025):
- Study Sample: Danish individuals whose spouses died within the past year.
- Variables Measured:
- Explanatory Variable (): Emotional Loneliness (EL) score.
- Response Variable (): Prolonged Grief Syndrome (PGS) score.
- Formal Hypotheses:
- Null Hypothesis (): (no linear relationship between EL and PGS in the population of bereaved spouses).
- Alternative Hypothesis (): (a linear relationship exists between EL and PGS in the population).
- Sample Raw Data Excerpt ( paired observations):
- Subject 1: ,
- Subject 2: ,
- Subject 3: ,
- Observed Sample Correlation: (Represents a very weak positive linear relationship).
Simulation / Permutation Procedure for Two Quantitative Variables:
- Step 1: Record the original sample correlation ().
- Step 2: Decouple the paired observations under . Fix the explanatory variable values (, EL scores) in place.
- Step 3: Randomly shuffle (permute) the response variable values (, PGS scores) across subjects.
- Step 4: Re-pair the shuffled response values with the fixed explanatory values and compute the simulated correlation ().
- Step 5: Repeat the permutation process many times to generate a null distribution of simulated correlations.
- Step 6: Compute the -value as the proportion of simulated correlations at least as extreme as in either direction.
Simulation Results in Vetter et al. (2025):
- Single Simulation Realization Example:
- Small Trial Run ( simulations):
- simulations yielded correlations between and
- simulations yielded correlations as extreme as or more extreme than
- -simulation -value estimate:
- Extended Simulation Run ( simulations):
- Final -value estimate:
- Statistical Conclusion: Because , the data fails to reject . The observed sample correlation () is not statistically significant and is consistent with chance alone.
Linear Regression and Least Squares Foundations
Core Concepts of Linear Regression:
- Definition: A statistical technique used to fit a straight line to quantitative data to model relationships and predict a response variable () using an explanatory variable ().
Algebraic Equation Conventions:
- Standard College/Statistics Notation:
- Parameter Notation for Population Model:
- Variable Definitions:
- : Explanatory variable value.
- : Predicted response variable value.
- : Actual observed response variable value.
- : Sample -intercept (predicted response when ).
- : Sample slope (predicted change in per 1-unit increase in ).
- High School vs. College Notation: In college statistics, represents the intercept and represents the slope coefficient attached to (unlearn ).
Residuals and the Least Squares Criterion:
- Definition of Residual: The vertical distance between an observed data point and the predicted point on the fitted line:
- Least Squares Regression Criterion: The unique line that minimizes the Sum of Squared Residuals ():
- Fundamental Properties of Least Squares Residuals:
- Sum of residuals is always zero:
- Mean of residuals is always zero:
Mathematical Formulas for Regression Parameters:
- Regression Slope (): where is sample correlation, is standard deviation of response variable , and is standard deviation of explanatory variable
- Regression Intercept (): where is the mean of the explanatory variable and is the mean of the response variable. The regression line always passes through the centroid point .
- Standard Deviation of Residuals (): Measures the typical size of prediction errors (typical distance of actual values from predicted values).
- Sign Consistency: The slope and correlation always share the exact same sign (both positive, both negative, or both zero).
Coefficient of Determination () and Error Metrics
Definition and Calculation of :
- Defined as the square of the sample correlation coefficient :
Contextual Interpretation of :
- Represents the proportion (or percentage) of the total variation in the response variable that is explained by the linear relationship with the explanatory variable (or explained by the regression line).
Relative vs. Absolute Prediction Performance:
- provides a relative measure of variance explained ( or
- Standard deviation of residuals () provides an absolute measure of typical prediction error expressed in the original units of
Case Study: Dinner Plate Sizes and Delboeuf Illusion
Delboeuf Illusion Context:
- Visual Optical Illusion: Surrounding outer concentric rings alter the human brain's perception of inner circle sizes.
- Experiment Findings: Comparing two black circles inside rings shows the right circle is actually larger than the left circle, despite observers perceiving them as equal.
- Behavioral Hypothesis: Serving food on larger dinner plates makes food portions appear smaller, prompting individuals to serve and consume more food.
Cornell Food and Brand Lab Study (2011):
- Research Objective: Investigate historical growth in dinner plate sizes using eBay listings of antique and modern dinner plates.
- Explanatory Variable (): Year of plate manufacture.
- Response Variable (): Plate diameter in inches.
- Dataset Summary Statistics & Regression Output:
- Sample Correlation:
- Coefficient of Determination:
- Least Squares Regression Equation:
Interpretations and Calculations:
- Predicted Diameter for Year 2000:
- Predicted Diameter for Year 2001:
- Interpretation of Slope (): Plate diameters increased on average by per year in the sample dataset.
- Interpretation of Intercept (): Predicting plate diameter for Year yields , illustrating extreme and physically impossible extrapolation.
- Interpretation of : of the variation in dinner plate diameters is accounted for by the linear relationship with the year of manufacture.
Critical Pitfalls and Limitations in Linear Regression
Pitfall 1: Inferring Causation from Correlation (Observational Data & Confounding):
- Case Study on Dietary Fat and Breast Cancer Mortality:
- Ecological aggregate data shows a strong positive correlation between national per-capita fat consumption () and per-capita breast cancer mortality rates per () across nations (e.g., Italy, UK, US).
- Led to the "Dietary Fat Hypothesis" advocating fat reduction to prevent breast cancer.
- Willett (2004) Systematic Review: Reviewed all prospective cohort studies with at least breast cancer cases; found NO statistically significant positive association between total, saturated, monounsaturated, or polyunsaturated fat intake and breast cancer risk.
- Underlying Confounding Variables: Economic development level, lower parity, later age at first birth, overall body fat percentage, and lower physical activity levels in Western nations confound the ecological correlation.
Pitfall 2: Extrapolation:
- Definition: Predicting response values () for explanatory values () far outside the range of observed sample data.
- Unreliability: Linear relationships observed locally rarely hold across extreme ranges.
- Real-World Examples:
- Brookings Institute Demographics Prediction: Extrapolated a South Korean birth rate of children per woman (based on data points over years) to claim South Korea faces natural extinction by year
- High-to-Low Dose Toxicological Animal Studies: Giving high chemical doses to small cohorts of laboratory rodents to achieve statistical test power, then extrapolating linearly to low human environmental exposure levels.
Pitfall 3: Non-Linearity and Curvature:
- Kowalski et al. (2005) Re-analysis of Framingham Heart Study:
- Re-examined blood glucose mortality data originally reported as linear.
- Showed data fits an exponential curve followed by a flat line, proving blood sugar reduction alters mortality risk only at elevated levels.
- Kowalski et al. (2005) on BMI and Mortality Risk:
- Demonstrated relationship between Body Mass Index (BMI) and mortality in males aged 40 and 50 is U-shaped (elevated mortality below BMI and above BMI or
- Fitting a linear model incorrectly advises individuals with low baseline BMI to lower their BMI, increasing mortality risk.
Pitfall 4: Heteroscedasticity:
- Definition: Non-constant residual variance across values of explanatory variable (e.g., fan-shaped residual pattern).
- Contrast with Homoscedasticity: Constant vertical variance/spread of points across all slices of
- Practical Impact: Invalidates global standard error estimates (), causing prediction margin-of-error bounds to be severely inaccurate at different regions of
Pitfall 5: High-Leverage Influential Outliers:
- A single extreme outlier can drastically alter calculated values of correlation , intercept , and slope
Theory-Based Inference for Regression and Correlation (-Test)
Mathematical Equivalence of Tests:
- Testing Population Correlation :
- Testing Population Regression Slope :
- Equivalence: Both tests produce identical -statistics, degrees of freedom, and -values.
Test Statistic Formula (-statistic): Degrees of Freedom:
Standard Error Formulas:
- Standard Error of Correlation :
- Standard Error of Slope :
Confidence Intervals:
- Confidence Interval for Correlation :
- Confidence Interval for Slope :
- For large sample sizes ( large), for confidence.
Validity Conditions for Theory-Based -Test:
- Linearity: Scatter plot shows linear pattern without curvature.
- Independence: Observations are Independent and Identically Distributed (IID).
- Normality / Symmetry: Residuals distributed symmetrically above and below regression line.
- Homoscedasticity: Equal variance of vertical slices across explanatory variable
Student Discussion, Audience Questions, and Course Review
Dialogue & Question Exchanges:
- Student Question (Sophia): Clarification on homework item showing positive scatter plot but negative regression equation (). Answer: Disregard plot graphic flaw; perform calculations using stated algebraic equation.
- Student Question (Sophia): Standardized statistic subtraction parameter for null distribution mean. Answer: Null mean is theoretically ; subtract in test statistic formula ().
- Student Question (Jasmine): Cumulative nature of final exam. Answer: Final is cumulative, but heavily weighted toward final week material.
- Student Question (Izzy): Multiplier for confidence intervals. Answer: Large sample sizes use for confidence; small sample sizes use critical -multiplier from software with
Comprehensive Topic Checklist:
- Summary statistics: Mean, standard deviation, parameters vs. statistics, z-scores.
- Simulation inference: Permutation tests, null distributions, -value calculation.
- Central Limit Theorem and validity conditions for 1-sample and 2-sample tests.
- Error types: Type I error, Type II error, statistical power, significance level
- Proportions: Two-proportion z-tests, pooled vs. unpooled standard errors, confidence intervals.
- Quantitative data: Two-sample t-tests, paired data analysis, five-number summary, IQR.
- Observational biases: Placebo effect, adherer bias, non-response bias, publication bias.
- Regression & Correlation: Slope , intercept , residuals, , , extrapolation, curvature, heteroscedasticity.
Named Studies Reference Guide (Open-Note Exam Guide):
- Woodward (1970) Study: Real study on specific research conclusions.
- Cantrell (1958) Study: Evaluated specific statistical methodology.
- Feldman Study: Evaluated randomization procedures.
- Marine Cancer-Sniffing Canine Study: Evaluation of cancer detection sensitivity.
Worked Final Exam Example Problems and Solutions
Problems 1–6: Adult Height and Weight Regression Dataset
Dataset Specifications ( adults):
Explanatory Variable (, Height): ,
Response Variable (, Weight): ,
Sample Correlation:
Problem 1: What does this correlation of imply?
Options: A. of variation in weight is explained by height | B. Typical variation of heights is of weights | C. There is strong association between height and weight in the sample | D. For every inch increase in height, predict weight increase | E. Person weighing expected to be tall.
Solution: Option C (There is strong association between height and weight in the sample).
Problem 2: Calculate estimated slope in lbs per inch.
Calculation:
Solution: Option A ().
Problem 3: How much would a prediction using this regression line typically be off by?
Calculation:
Solution: Option E ().
Problem 4: Typical deviation of a random adult's height from mean height ().
Calculation: Equal to standard deviation of height .
Solution: Option F ().
Problem 5: Why shouldn't one trust this line to predict weight of someone tall?
Solution: Option C (The value of is too far outside the range of most observations — Extrapolation).
Problem 6: How should one interpret the estimated slope of ?
Solution: Option E (For each extra inch taller you are, your predicted weight increases by ).
Problems 7–11: Two-Group Acne Rate Analysis
Sample Details:
Group A (Aged 21–30): , with acne (
Group B (Aged 31–40): , with acne (
Problem 7: Calculate pooled sample percentage with acne.
Calculation:
Problem 8: Calculate pooled standard error () under
Calculation:
Problem 9: Calculate -statistic for testing difference between percentages.
Calculation:
Problem 10: Find confidence interval for difference () using unpooled standard error.
Unpooled Standard Error Calculation:
Margin of Error ():
Confidence Interval:
Solution: Option E (
Problem 11: What can we conclude from the confidence interval ()?
Solution: Option D (A statistically significantly higher percentage of people aged 21 to 30 have acne than those aged 31 to 40, because zero is not contained within the confidence interval).