Comprehensive Notes on Linear Regression, Multiple Testing, and Publication Bias

Course Logistics and Administrative Information

  • Course Operations & Communication Protocols:

    • Holiday Note: Labor Day was acknowledged in the lecture chat.
    • Question Submission: All questions during lecture must be directed to "everyone" in the chat window.
    • Course Website: Serves as the primary hub for syllabus materials, homework assignments, and exam instruction sheets.
    • Homework 4: Official due date is set for this Friday.
    • Final Exam Schedule: Scheduled for Wednesday, exactly one week from the lecture date, running from 10:00 AM10:00\,\text{AM} to 11:50 AM11:50\,\text{AM}.
  • Final Exam Logistics & Protocols:

    • Instruction Release Time: The exam instruction document (finalinstructions.text) will be posted on the course website at 09:30 AM09:30\,\text{AM} on Wednesday, September 9.
    • Zoom Login Requirement: Students must log into the designated Zoom room at 10:00 AM10:00\,\text{AM} and remain present on Zoom for the duration of the exam.
    • Answer Submission: Exam answers must be emailed directly to the instructor following the precise instructions outlined in the official released document.
    • Exam Length and Question Format: Consists of approximately 2020 questions total, primarily multiple-choice with potentially one short-answer question.
    • Scope and Content Weighting: Cumulative covering all material across the term, but weighted more heavily toward concepts introduced during the final week of class.
    • Permitted Testing Resources: The exam is open-note, allowing students to reference printed or digital notes during the exam.

Multiple Testing and Publication Bias

  • Fundamentals of pp-Values and Significance Cutoffs:

    • Meaning of pp-Value: A pp-value of 3%3\% (0.030.03) indicates that if the null hypothesis (H0H_0) is true, there is a 3%3\% probability of obtaining sample data as extreme as or more extreme than the observed data.
    • Standard Significance Threshold (α\alpha): The conventional significance level cutoff is set at 5%5\% (0.050.05).
    • False Positive Rate under H0H_0: When evaluating unrelated variables where the null hypothesis is true, a statistically significant result will be generated 5%5\% of the time by pure chance, yielding a Type I error.
    • Necessity of Confirmation Studies: Because any single study carries an inherent 5%5\% false-positive rate under H0H_0, initial significant findings (such as drug efficacy) require independent follow-up confirmation studies before conclusions can be validated.
  • The Multiple Testing Problem:

    • Context and Definition: Occurs when a researcher tests many explanatory variables or performs multiple hypothesis tests simultaneously within a single dataset.
    • Survey Example: Investigating potential causes of a rare condition by surveying participants on dozens of behavioral, dietary, and exercise variables increases the likelihood of finding at least one false positive.
    • Mathematical Demonstration (m=100m = 100 Independent Tests under True Null Hypotheses at α=0.05\alpha = 0.05):
    • Expected Number of Type I Errors:       Expected Errors=m×α=100×0.05=5\text{Expected Errors} = m \times \alpha = 100 \times 0.05 = 5
    • Probability of No Type I Error on a Single Test:       1−α=1−0.05=0.95=95%1 - \alpha = 1 - 0.05 = 0.95 = 95\%
    • Probability of Zero Type I Errors across 100100 Independent Tests:       P(Zero Errors)=(0.95)100≈0.006=0.6%P(\text{Zero Errors}) = (0.95)^{100} \approx 0.006 = 0.6\%
    • Probability of Making AT LEAST ONE Type I Error across 100100 Independent Tests:       P(At Least One Error)=1−P(Zero Errors)=1−0.006=0.994=99.4%P(\text{At Least One Error}) = 1 - P(\text{Zero Errors}) = 1 - 0.006 = 0.994 = 99.4\%
  • Bonferroni's Correction:

    • Purpose: A method used to control the overall family-wise error rate when conducting mm hypothesis tests or evaluating mm explanatory variables.
    • Adjusted Significance Level Formula:     αadj=αm=0.05m\alpha_{adj} = \frac{\alpha}{m} = \frac{0.05}{m}
    • Theoretical Guarantee: Guarantees that the overall probability of making at least one Type I error across all mm tests combined remains less than or equal to 5%5\% (≤5%\le 5\%
    • Application Examples:
    • For m=10m = 10 tests, corrected significance threshold is:       αadj=5%10=0.5%=0.005\alpha_{adj} = \frac{5\%}{10} = 0.5\% = 0.005
    • For m=100m = 100 tests, corrected significance threshold is:       αadj=5%100=0.05%=0.0005\alpha_{adj} = \frac{5\%}{100} = 0.05\% = 0.0005
  • Publication Bias:

    • Definition and Process: Occurs when journal editors and reviewers publish studies that demonstrate statistically significant findings (p≤0.05p \le 0.05) while rejecting or ignoring studies that show non-significant findings.
    • Illustrative Scenario: Assume 100100 independent research teams test a drug that actually has zero true effect (H0H_0 is true).
    • Approximately 9595 teams will find no significant effect (p>0.05p > 0.05).
    • Approximately 55 teams will find a statistically significant effect (p≤0.05p \le 0.05) due to random sampling variation.
    • If only the 55 significant studies are published while the 9595 non-significant studies are suppressed, the published record falsely indicates that the drug works.
    • Scientific Solution: Journals should publish methodologically sound research regardless of whether the results achieve statistical significance.
  • Case Study on ADHD Stimulants (Staley et al., 2024):

    • Background: Numerous published studies claimed that stimulant medications like methylphenidate are effective treatments for ADHD. Out of 1414 published articles in the literature, the majority concluded that methylphenidate was effective.
    • Meta-Analysis Methodology: Staley et al. (2024) conducted a comprehensive audit analyzing all available experimental data, comprising 1414 published articles and 1717 unpublished articles (3131 total studies).
    • Statistical Analysis: Pooled data comparing methylphenidate treatment groups to placebo groups using a two-sample zz-test (due to large pooled sample size).
    • Pooled Test Results:
    • Test Statistic:       z=0.62z = 0.62
    • Calculated pp-value:       p=53.5%=0.535p = 53.5\% = 0.535
    • Standard Cutoff:       α=5%=0.05\alpha = 5\% = 0.05
    • Conclusion: Incorporating unpublished studies demonstrates that methylphenidate performs no better than a placebo (p=53.5%≫5%p = 53.5\% \gg 5\%). The null hypothesis is consistent with the complete dataset, highlighting significant publication bias within the published literature.

Testing Correlation via Simulation Methods

  • Statistical Framework for Correlation Testing:

    • Objective: Evaluate whether an observed sample correlation rr between two quantitative variables indicates a true population correlation (ρ≠0\rho \neq 0) or is attributable to random sampling variability under ρ=0\rho = 0
  • Case Study on Emotional Loneliness and Prolonged Grief (Vetter et al., 2025):

    • Study Sample: n=20n = 20 Danish individuals whose spouses died within the past year.
    • Variables Measured:
    • Explanatory Variable (xx): Emotional Loneliness (EL) score.
    • Response Variable (yy): Prolonged Grief Syndrome (PGS) score.
    • Formal Hypotheses:
    • Null Hypothesis (H0H_0): ρ=0\rho = 0 (no linear relationship between EL and PGS in the population of bereaved spouses).
    • Alternative Hypothesis (HaH_a): ρ≠0\rho \neq 0 (a linear relationship exists between EL and PGS in the population).
    • Sample Raw Data Excerpt (2020 paired observations):
    • Subject 1: EL=9.3EL = 9.3, PGS=17PGS = 17
    • Subject 2: EL=8.2EL = 8.2, PGS=9PGS = 9
    • Subject 3: EL=8.2EL = 8.2, PGS=11PGS = 11
    • Observed Sample Correlation:     r=0.069r = 0.069     (Represents a very weak positive linear relationship).
  • Simulation / Permutation Procedure for Two Quantitative Variables:

    • Step 1: Record the original sample correlation (robs=0.069r_{obs} = 0.069).
    • Step 2: Decouple the paired observations under H0H_0. Fix the explanatory variable values (xx, EL scores) in place.
    • Step 3: Randomly shuffle (permute) the response variable values (yy, PGS scores) across subjects.
    • Step 4: Re-pair the shuffled response values with the fixed explanatory values and compute the simulated correlation (rsimr_{sim}).
    • Step 5: Repeat the permutation process many times to generate a null distribution of simulated correlations.
    • Step 6: Compute the pp-value as the proportion of simulated correlations at least as extreme as ∣robs∣=0.069|r_{obs}| = 0.069 in either direction.
  • Simulation Results in Vetter et al. (2025):

    • Single Simulation Realization Example: rsim=0.073r_{sim} = 0.073
    • Small Trial Run (3030 simulations):
    • 77 simulations yielded correlations between −0.069-0.069 and +0.069+0.069
    • 2323 simulations yielded correlations as extreme as or more extreme than ∣0.069∣|0.069|
    • 3030-simulation pp-value estimate:       p=2330≈77%=0.767p = \frac{23}{30} \approx 77\% = 0.767
    • Extended Simulation Run (1,0001,000 simulations):
    • Final pp-value estimate:       p=76.4%=0.764p = 76.4\% = 0.764
    • Statistical Conclusion: Because p=76.4%≫5%p = 76.4\% \gg 5\%, the data fails to reject H0H_0. The observed sample correlation (r=0.069r = 0.069) is not statistically significant and is consistent with chance alone.

Linear Regression and Least Squares Foundations

  • Core Concepts of Linear Regression:

    • Definition: A statistical technique used to fit a straight line to quantitative data to model relationships and predict a response variable (yy) using an explanatory variable (xx).
  • Algebraic Equation Conventions:

    • Standard College/Statistics Notation:     y^=a+b⋅x\hat{y} = a + b \cdot x
    • Parameter Notation for Population Model:     μy=α+β⋅x\mu_y = \alpha + \beta \cdot x
    • Variable Definitions:
    • xx: Explanatory variable value.
    • y^\hat{y}: Predicted response variable value.
    • yy: Actual observed response variable value.
    • aa: Sample yy-intercept (predicted response when x=0x = 0).
    • bb: Sample slope (predicted change in yy per 1-unit increase in xx).
    • High School vs. College Notation: In college statistics, aa represents the intercept and bb represents the slope coefficient attached to xx (unlearn y=m⋅x+by = m \cdot x + b).
  • Residuals and the Least Squares Criterion:

    • Definition of Residual: The vertical distance between an observed data point (x,y)(x, y) and the predicted point (x,y^)(x, \hat{y}) on the fitted line:     Residual=y−y^=Actual y−Predicted y^\text{Residual} = y - \hat{y} = \text{Actual } y - \text{Predicted } \hat{y}
    • Least Squares Regression Criterion: The unique line that minimizes the Sum of Squared Residuals (SSRSSR):     Minimize ∑(y−y^)2=∑(y−(a+b⋅x))2\text{Minimize } \sum (y - \hat{y})^2 = \sum (y - (a + b \cdot x))^2
    • Fundamental Properties of Least Squares Residuals:
    • Sum of residuals is always zero:       ∑(y−y^)=0\sum (y - \hat{y}) = 0
    • Mean of residuals is always zero:       Mean of Residuals=0\text{Mean of Residuals} = 0
  • Mathematical Formulas for Regression Parameters:

    • Regression Slope (bb):     b=r×sysxb = r \times \frac{s_y}{s_x}     where rr is sample correlation, sys_y is standard deviation of response variable yy, and sxs_x is standard deviation of explanatory variable xx
    • Regression Intercept (aa):     a=yˉ−b⋅xˉa = \bar{y} - b \cdot \bar{x}     where xˉ\bar{x} is the mean of the explanatory variable and yˉ\bar{y} is the mean of the response variable. The regression line always passes through the centroid point (xˉ,yˉ)(\bar{x}, \bar{y}).
    • Standard Deviation of Residuals (sress_{res}):     sres=1−r2×sys_{res} = \sqrt{1 - r^2} \times s_y     Measures the typical size of prediction errors (typical distance of actual values from predicted values).
    • Sign Consistency: The slope bb and correlation rr always share the exact same sign (both positive, both negative, or both zero).

Coefficient of Determination (R2R^2) and Error Metrics

  • Definition and Calculation of R2R^2:

    • Defined as the square of the sample correlation coefficient rr:     R2=(r)2R^2 = (r)^2
  • Contextual Interpretation of R2R^2:

    • Represents the proportion (or percentage) of the total variation in the response variable yy that is explained by the linear relationship with the explanatory variable xx (or explained by the regression line).
  • Relative vs. Absolute Prediction Performance:

    • R2R^2 provides a relative measure of variance explained (0≤R2≤10 \le R^2 \le 1 or 0%≤R2≤100%0\% \le R^2 \le 100\%
    • Standard deviation of residuals (sres=1−r2×sys_{res} = \sqrt{1 - r^2} \times s_y) provides an absolute measure of typical prediction error expressed in the original units of yy

Case Study: Dinner Plate Sizes and Delboeuf Illusion

  • Delboeuf Illusion Context:

    • Visual Optical Illusion: Surrounding outer concentric rings alter the human brain's perception of inner circle sizes.
    • Experiment Findings: Comparing two black circles inside rings shows the right circle is actually 20%20\% larger than the left circle, despite observers perceiving them as equal.
    • Behavioral Hypothesis: Serving food on larger dinner plates makes food portions appear smaller, prompting individuals to serve and consume more food.
  • Cornell Food and Brand Lab Study (2011):

    • Research Objective: Investigate historical growth in dinner plate sizes using eBay listings of antique and modern dinner plates.
    • Explanatory Variable (xx): Year of plate manufacture.
    • Response Variable (yy): Plate diameter in inches.
    • Dataset Summary Statistics & Regression Output:
    • Sample Correlation:       r=0.604r = 0.604
    • Coefficient of Determination:       R2=(0.604)2=0.365=36.5%R^2 = (0.604)^2 = 0.365 = 36.5\%
    • Least Squares Regression Equation:       y^=−14.8+0.0128⋅x\hat{y} = -14.8 + 0.0128 \cdot x
  • Interpretations and Calculations:

    • Predicted Diameter for Year 2000:     y^=−14.8+0.0128×2000=10.8 inches\hat{y} = -14.8 + 0.0128 \times 2000 = 10.8\,\text{inches}
    • Predicted Diameter for Year 2001:     y^=−14.8+0.0128×2001=10.8128 inches\hat{y} = -14.8 + 0.0128 \times 2001 = 10.8128\,\text{inches}
    • Interpretation of Slope (b=0.0128b = 0.0128): Plate diameters increased on average by 0.0128 inches0.0128\,\text{inches} per year in the sample dataset.
    • Interpretation of Intercept (a=−14.8a = -14.8): Predicting plate diameter for Year x=0x = 0 yields −14.8 inches-14.8\,\text{inches}, illustrating extreme and physically impossible extrapolation.
    • Interpretation of R2=36.5%R^2 = 36.5\%: 36.5%36.5\% of the variation in dinner plate diameters is accounted for by the linear relationship with the year of manufacture.

Critical Pitfalls and Limitations in Linear Regression

  • Pitfall 1: Inferring Causation from Correlation (Observational Data & Confounding):

    • Case Study on Dietary Fat and Breast Cancer Mortality:
    • Ecological aggregate data shows a strong positive correlation between national per-capita fat consumption (xx) and per-capita breast cancer mortality rates per 100,000100,000 (yy) across nations (e.g., Italy, UK, US).
    • Led to the "Dietary Fat Hypothesis" advocating fat reduction to prevent breast cancer.
    • Willett (2004) Systematic Review: Reviewed all prospective cohort studies with at least 200200 breast cancer cases; found NO statistically significant positive association between total, saturated, monounsaturated, or polyunsaturated fat intake and breast cancer risk.
    • Underlying Confounding Variables: Economic development level, lower parity, later age at first birth, overall body fat percentage, and lower physical activity levels in Western nations confound the ecological correlation.
  • Pitfall 2: Extrapolation:

    • Definition: Predicting response values (y^\hat{y}) for explanatory values (xx) far outside the range of observed sample data.
    • Unreliability: Linear relationships observed locally rarely hold across extreme ranges.
    • Real-World Examples:
    • Brookings Institute Demographics Prediction: Extrapolated a South Korean birth rate of 1.191.19 children per woman (based on 55 data points over 2020 years) to claim South Korea faces natural extinction by year 27502750
    • High-to-Low Dose Toxicological Animal Studies: Giving high chemical doses to small cohorts of laboratory rodents to achieve statistical test power, then extrapolating linearly to low human environmental exposure levels.
  • Pitfall 3: Non-Linearity and Curvature:

    • Kowalski et al. (2005) Re-analysis of Framingham Heart Study:
    • Re-examined blood glucose mortality data originally reported as linear.
    • Showed data fits an exponential curve followed by a flat line, proving blood sugar reduction alters mortality risk only at elevated levels.
    • Kowalski et al. (2005) on BMI and Mortality Risk:
    • Demonstrated relationship between Body Mass Index (BMI) and mortality in males aged 40 and 50 is U-shaped (elevated mortality below BMI 2020 and above BMI 3535 or 4040
    • Fitting a linear model incorrectly advises individuals with low baseline BMI to lower their BMI, increasing mortality risk.
  • Pitfall 4: Heteroscedasticity:

    • Definition: Non-constant residual variance across values of explanatory variable xx (e.g., fan-shaped residual pattern).
    • Contrast with Homoscedasticity: Constant vertical variance/spread of points across all slices of xx
    • Practical Impact: Invalidates global standard error estimates (sres=1−r2×sys_{res} = \sqrt{1 - r^2} \times s_y), causing prediction margin-of-error bounds to be severely inaccurate at different regions of xx
  • Pitfall 5: High-Leverage Influential Outliers:

    • A single extreme outlier can drastically alter calculated values of correlation rr, intercept aa, and slope bb

Theory-Based Inference for Regression and Correlation (tt-Test)

  • Mathematical Equivalence of Tests:

    • Testing Population Correlation ρ\rho:     H0:ρ=0vs.Ha:ρ≠0H_0: \rho = 0 \quad \text{vs.} \quad H_a: \rho \neq 0
    • Testing Population Regression Slope β\beta:     H0:β=0vs.Ha:β≠0H_0: \beta = 0 \quad \text{vs.} \quad H_a: \beta \neq 0
    • Equivalence: Both tests produce identical tt-statistics, degrees of freedom, and pp-values.
  • Test Statistic Formula (tt-statistic):   t=rSEr=r1−r2n−2t = \frac{r}{SE_r} = \frac{r}{\sqrt{\frac{1 - r^2}{n - 2}}}   Degrees of Freedom:   df=n−2df = n - 2

  • Standard Error Formulas:

    • Standard Error of Correlation rr:     SEr=1−r2n−2SE_r = \sqrt{\frac{1 - r^2}{n - 2}}
    • Standard Error of Slope bb:     SEb=SEr×sysx=1−r2n−2×sysxSE_b = SE_r \times \frac{s_y}{s_x} = \sqrt{\frac{1 - r^2}{n - 2}} \times \frac{s_y}{s_x}
  • Confidence Intervals:

    • Confidence Interval for Correlation ρ\rho:     r±t∗×SErr \pm t^* \times SE_r
    • Confidence Interval for Slope β\beta:     b±t∗×SEbb \pm t^* \times SE_b
    • For large sample sizes (nn large), t∗≈1.96t^* \approx 1.96 for 95%95\% confidence.
  • Validity Conditions for Theory-Based tt-Test:

    • Linearity: Scatter plot shows linear pattern without curvature.
    • Independence: Observations are Independent and Identically Distributed (IID).
    • Normality / Symmetry: Residuals distributed symmetrically above and below regression line.
    • Homoscedasticity: Equal variance of vertical slices across explanatory variable xx

Student Discussion, Audience Questions, and Course Review

  • Dialogue & Question Exchanges:

    • Student Question (Sophia): Clarification on homework item showing positive scatter plot but negative regression equation (y^=162.56−0.9658⋅BMI\hat{y} = 162.56 - 0.9658 \cdot \text{BMI}). Answer: Disregard plot graphic flaw; perform calculations using stated algebraic equation.
    • Student Question (Sophia): Standardized statistic subtraction parameter for null distribution mean. Answer: Null mean is theoretically 00; subtract 00 in test statistic formula (statistic−0SE\frac{\text{statistic} - 0}{SE}).
    • Student Question (Jasmine): Cumulative nature of final exam. Answer: Final is cumulative, but heavily weighted toward final week material.
    • Student Question (Izzy): Multiplier for confidence intervals. Answer: Large sample sizes use 1.961.96 for 95%95\% confidence; small sample sizes use critical t∗t^*-multiplier from software with df=n−2df = n - 2
  • Comprehensive Topic Checklist:

    • Summary statistics: Mean, standard deviation, parameters vs. statistics, z-scores.
    • Simulation inference: Permutation tests, null distributions, pp-value calculation.
    • Central Limit Theorem and validity conditions for 1-sample and 2-sample tests.
    • Error types: Type I error, Type II error, statistical power, significance level α\alpha
    • Proportions: Two-proportion z-tests, pooled vs. unpooled standard errors, confidence intervals.
    • Quantitative data: Two-sample t-tests, paired data analysis, five-number summary, IQR.
    • Observational biases: Placebo effect, adherer bias, non-response bias, publication bias.
    • Regression & Correlation: Slope bb, intercept aa, residuals, R2R^2, sress_{res}, extrapolation, curvature, heteroscedasticity.
  • Named Studies Reference Guide (Open-Note Exam Guide):

    • Woodward (1970) Study: Real study on specific research conclusions.
    • Cantrell (1958) Study: Evaluated specific statistical methodology.
    • Feldman Study: Evaluated randomization procedures.
    • Marine Cancer-Sniffing Canine Study: Evaluation of cancer detection sensitivity.

Worked Final Exam Example Problems and Solutions

  • Problems 1–6: Adult Height and Weight Regression Dataset

    • Dataset Specifications (n=100n = 100 adults):

    • Explanatory Variable (xx, Height): xˉ=65 inches\bar{x} = 65\,\text{inches}, sx=5 inchess_x = 5\,\text{inches}

    • Response Variable (yy, Weight): yˉ=160 lbs\bar{y} = 160\,\text{lbs}, sy=40 lbss_y = 40\,\text{lbs}

    • Sample Correlation: r=0.82r = 0.82

    • Problem 1: What does this correlation of 0.820.82 imply?

    • Options: A. 82%82\% of variation in weight is explained by height | B. Typical variation of heights is 82%82\% of weights | C. There is strong association between height and weight in the sample | D. For every inch increase in height, predict 0.82 lbs0.82\,\text{lbs} weight increase | E. Person weighing 100 lbs100\,\text{lbs} expected to be 82 inches82\,\text{inches} tall.

    • Solution: Option C (There is strong association between height and weight in the sample).

    • Problem 2: Calculate estimated slope bb in lbs per inch.

    • Calculation:       b=r×sysx=0.82×405=0.82×8=6.56 lbs/inchb = r \times \frac{s_y}{s_x} = 0.82 \times \frac{40}{5} = 0.82 \times 8 = 6.56\,\text{lbs/inch}

    • Solution: Option A (6.566.56).

    • Problem 3: How much would a prediction using this regression line typically be off by?

    • Calculation:       sres=1−r2×sy=1−(0.82)2×40=1−0.6724×40=0.3276×40≈0.57236×40=22.89 lbs≈22.9 lbss_{res} = \sqrt{1 - r^2} \times s_y = \sqrt{1 - (0.82)^2} \times 40 = \sqrt{1 - 0.6724} \times 40 = \sqrt{0.3276} \times 40 \approx 0.57236 \times 40 = 22.89\,\text{lbs} \approx 22.9\,\text{lbs}

    • Solution: Option E (22.9 lbs22.9\,\text{lbs}).

    • Problem 4: Typical deviation of a random adult's height from mean height (65 inches65\,\text{inches}).

    • Calculation: Equal to standard deviation of height sx=5.0 inchess_x = 5.0\,\text{inches}.

    • Solution: Option F (5.0 inches5.0\,\text{inches}).

    • Problem 5: Why shouldn't one trust this line to predict weight of someone 25 inches25\,\text{inches} tall?

    • Solution: Option C (The value of 25 inches25\,\text{inches} is too far outside the range of most observations — Extrapolation).

    • Problem 6: How should one interpret the estimated slope of 6.566.56?

    • Solution: Option E (For each extra inch taller you are, your predicted weight increases by 6.56 pounds6.56\,\text{pounds}).

  • Problems 7–11: Two-Group Acne Rate Analysis

    • Sample Details:

    • Group A (Aged 21–30): n1=400n_1 = 400, x1=88x_1 = 88 with acne (p^1=88400=0.22=22%\hat{p}_1 = \frac{88}{400} = 0.22 = 22\%

    • Group B (Aged 31–40): n2=200n_2 = 200, x2=30x_2 = 30 with acne (p^2=30200=0.15=15%\hat{p}_2 = \frac{30}{200} = 0.15 = 15\%

    • Problem 7: Calculate pooled sample percentage with acne.

    • Calculation:       p^pooled=88+30400+200=118600≈0.1967=19.67%\hat{p}_{pooled} = \frac{88 + 30}{400 + 200} = \frac{118}{600} \approx 0.1967 = 19.67\%

    • Problem 8: Calculate pooled standard error (SEpooledSE_{pooled}) under H0:p1=p2H_0: p_1 = p_2

    • Calculation:       SEpooled=p^pooled(1−p^pooled)(1n1+1n2)=0.1967×0.8033(1400+1200)=0.1580×0.0075≈0.0344=3.44%SE_{pooled} = \sqrt{\hat{p}_{pooled}(1 - \hat{p}_{pooled}) \left( \frac{1}{n_1} + \frac{1}{n_2} \right)} = \sqrt{0.1967 \times 0.8033 \left( \frac{1}{400} + \frac{1}{200} \right)} = \sqrt{0.1580 \times 0.0075} \approx 0.0344 = 3.44\%

    • Problem 9: Calculate zz-statistic for testing difference between percentages.

    • Calculation:       Observed Difference=p^1−p^2=0.22−0.15=0.07=7%\text{Observed Difference} = \hat{p}_1 - \hat{p}_2 = 0.22 - 0.15 = 0.07 = 7\%z=p^1−p^2SEpooled=0.070.0344≈2.03z = \frac{\hat{p}_1 - \hat{p}_2}{SE_{pooled}} = \frac{0.07}{0.0344} \approx 2.03

    • Problem 10: Find 95%95\% confidence interval for difference (p1−p2p_1 - p_2) using unpooled standard error.

    • Unpooled Standard Error Calculation:       SEunpooled=0.22×0.78400+0.15×0.85200=0.1716400+0.1275200=0.000429+0.0006375=0.00010665≈0.03266=3.266%SE_{unpooled} = \sqrt{\frac{0.22 \times 0.78}{400} + \frac{0.15 \times 0.85}{200}} = \sqrt{\frac{0.1716}{400} + \frac{0.1275}{200}} = \sqrt{0.000429 + 0.0006375} = \sqrt{0.00010665} \approx 0.03266 = 3.266\%

    • Margin of Error (95%95\%):       ME=1.96×3.266%≈6.41%ME = 1.96 \times 3.266\% \approx 6.41\%

    • Confidence Interval:       7%±6.41%=[0.59%,13.41%]7\% \pm 6.41\% = [0.59\%, 13.41\%]

    • Solution: Option E (7%±6.41%7\% \pm 6.41\%

    • Problem 11: What can we conclude from the 95%95\% confidence interval (7%±6.41%=[0.59%,13.41%]7\% \pm 6.41\% = [0.59\%, 13.41\%])?

    • Solution: Option D (A statistically significantly higher percentage of people aged 21 to 30 have acne than those aged 31 to 40, because zero 0%0\% is not contained within the confidence interval).