Logistic Regression: Predictions, Transformations, Diagnostics, and Model Validation

Foundations of Logistic Regression and Binary Outcome Modeling

  • Binary Outcome Structure:

    • Binary outcomes represent variable measurements that can only take on two mutually exclusive values.
    • Standard numerical encoding assigns 00 to represent the negative state (absence, failure, or "no") and 11 to represent the positive state (presence, success, or "yes").
  • Probability Distribution Mechanics:

    • Binary data points arise from a Binomial or Bernoulli distribution (a Binomial distribution with sample size trial count n=1n = 1).
    • For an outcome Y∼B(n,p)Y \thicksim B(n, p), the probability mass function defining the occurrence of yy successes in nn trials is given by:         P(Y=y)=(ny)py(1−p)n−yP(Y = y) = \binom{n}{y} p^y (1-p)^{n-y}
    • The expected mean of a binomial distribution is μ=n×p\mu = n \times p.
    • The variance of a binomial distribution is dependent on the underlying probability and is calculated as:         σ2=n×p×(1−p)\sigma^2 = n \times p \times (1-p)
  • The Logit Link Function:

    • Standard linear regression cannot be directly fitted to binary 0/10/1 raw outcomes because linear combinations of predictors can produce predicted values outside the probability bounds of [0,1][0, 1].
    • Logistic regression models the expected proportion or probability of the positive outcome (pp) rather than raw binary occurrences.
    • To allow linear modeling against continuous or categorical predictors across an unbounded outcome space from (−∞,∞)(-\infty, \infty), the outcome probability is transformed using the logit link function:         logit(p)=ln⁡(p1−p)=β0+β1X1+β2X2+⋯+βkXk\text{logit}(p) = \ln\left(\frac{p}{1-p}\right) = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \dots + \beta_k X_k
    • The term p1−p\frac{p}{1-p} represents the odds of the outcome occurring. The natural logarithm of the odds transforms bounded probabilities into an unbounded continuous scale (−∞-\infty to +∞+\infty).
  • Parameter Estimation via Maximum Likelihood:

    • Unlike ordinary least squares (OLS) regression, parameters (β\beta) in logistic regression are estimated using Maximum Likelihood Estimation (MLE).
    • The maximum likelihood method selects parameter estimates that maximize the probability of observing the sample data.
    • The likelihood function L(p∣data)L(p | \text{data}) constructed across nn independent observations is written as:         L(p∣data)=∏i=1nP(yi∣pi)=∏i=1n(niyi)piyi(1−pi)ni−yiL(p | \text{data}) = \prod_{i=1}^n P(y_i | p_i) = \prod_{i=1}^n \binom{n_i}{y_i} p_i^{y_i} (1-p_i)^{n_i - y_i}

Parameter Interpretation Across Measurement Scales

  • Interpretation of the Model Intercept (β0\beta_0):

    • On the log-odds scale, β0\beta_0 represents the expected log-odds of the outcome when all continuous predictor variables X1,…,XkX_1, \dots, X_k equal 00, or when all categorical predictors reside in their respective baseline/reference categories.
    • On the odds scale, exp⁡(β0)\exp(\beta_0) represents the expected baseline odds of the outcome when all predictors equal 00 or sit at baseline.
  • Interpretation of Continuous Predictor Coefficients (β1\beta_1):

    • On the log-odds scale, β1\beta_1 is a log odds ratio representing the average change in the expected log-odds of the outcome per single unit increase in X1X_1.
    • On the odds scale, exp⁡(β1)\exp(\beta_1) is an Odds Ratio (OR). It estimates the multiplicative fold, factor, or percentage change in the odds of the outcome per single unit increase in X1X_1
  • Interpretation of Categorical Predictor Coefficients (β1\beta_1):

    • On the log-odds scale, β1\beta_1 measures the step-change in expected log-odds when transitioning from the baseline/reference category to a specified comparison category.
    • On the odds scale, exp⁡(β1)\exp(\beta_1) represents the expected odds ratio comparing the odds of the outcome in the specified category relative to the baseline/reference category.
  • Interpretation in Multivariable Models:

    • In a multiple logistic regression model containing kk predictors, βk\beta_k reflects the adjusted log odds ratio for variable XkX_k, isolating its effect while holding all other variables constant in the model.
  • Scale Invariance and Non-Linearity Properties:

    • Log-odds change linearly and additively with respect to unit changes in predictor variables XX
    • Odds change multiplicatively by a factor of exp⁡(βk)\exp(\beta_k) for unit changes in predictor variables XX
    • Probability changes non-linearly with respect to changes in XX. The change in predicted probability per unit increase in XX varies depending on the baseline value of XX

Probability, Odds, and Logit across continuous Age

Prediction Mechanics and Mathematical Conversions

  • Scale Predictions:

    • Log-odds Scale:ln⁡(p^(Y=1)1−p^(Y=1))=β^0+β^1X1+β^2X2+⋯+β^kXk\ln\left(\frac{\hat{p}(Y = 1)}{1 - \hat{p}(Y = 1)}\right) = \hat{\beta}_0 + \hat{\beta}_1 X_1 + \hat{\beta}_2 X_2 + \dots + \hat{\beta}_k X_k
    • Odds Scale:p^(Y=1)1−p^(Y=1)=exp⁡(β^0+β^1X1+β^2X2+⋯+β^kXk)\frac{\hat{p}(Y = 1)}{1 - \hat{p}(Y = 1)} = \exp\left(\hat{\beta}_0 + \hat{\beta}_1 X_1 + \hat{\beta}_2 X_2 + \dots + \hat{\beta}_k X_k\right)
    • Probability Scale:p^(Y=1)=exp⁡(β^0+β^1X1+⋯+β^kXk)1+exp⁡(β^0+β^1X1+⋯+β^kXk)\hat{p}(Y = 1) = \frac{\exp\left(\hat{\beta}_0 + \hat{\beta}_1 X_1 + \dots + \hat{\beta}_k X_k\right)}{1 + \exp\left(\hat{\beta}_0 + \hat{\beta}_1 X_1 + \dots + \hat{\beta}_k X_k\right)}
  • Algebraic Derivation of the Inverse Logit Function:

    • Beginning with the odds equation:         p^1−p^=exp⁡(β^0+⋯+β^kXk)\frac{\hat{p}}{1 - \hat{p}} = \exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right)
    • Multiplying both sides by (1−p^)(1 - \hat{p}):         p^=exp⁡(β^0+⋯+β^kXk)−p^×exp⁡(β^0+⋯+β^kXk)\hat{p} = \exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right) - \hat{p} \times \exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right)
    • Grouping terms containing p^\hat{p} on the left-hand side:         p^+p^×exp⁡(β^0+⋯+β^kXk)=exp⁡(β^0+⋯+β^kXk)\hat{p} + \hat{p} \times \exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right) = \exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right)
    • Factoring out p^\hat{p}:         p^[1+exp⁡(β^0+⋯+β^kXk)]=exp⁡(β^0+⋯+β^kXk)\hat{p} \left[1 + \exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right)\right] = \exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right)
    • Solving explicitly for p^\hat{p} yields the standard inverse logit equation:         p^=exp⁡(β^0+⋯+β^kXk)1+exp⁡(β^0+⋯+β^kXk)\hat{p} = \frac{\exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right)}{1 + \exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right)}
    • Dividing numerator and denominator by exp⁡(β^0+⋯+β^kXk)\exp\left(\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right) yields the alternative sigmoid formulation:         p^=11+exp⁡(−[β^0+⋯+β^kXk])\hat{p} = \frac{1}{1 + \exp\left(-\left[\hat{\beta}_0 + \dots + \hat{\beta}_k X_k\right]\right)}
    • The inverse logit function constrains all predicted values strictly between 00 and 11, forming a characteristic S-shaped (sigmoid) curve.

Inverse Logit Function Curve

  • Statistical Inference and Wald Confidence Intervals:
    • Testing population parameter hypotheses involves evaluating H0:βk=0H_0: \beta_k = 0 against HA:βk≠0H_A: \beta_k \neq 0.
    • In large sample sizes, parameter estimates follow an asymptotic normal distribution.
    • The 95%95\% Wald Confidence Interval on the log-odds scale is calculated as:         β^k±1.96×SE(β^k)\hat{\beta}_k \pm 1.96 \times \text{SE}(\hat{\beta}_k)
    • The 95%95\% Confidence Interval on the odds scale (Odds Ratio scale) is obtained by exponentiating the log-odds confidence boundaries:         [exp⁡(β^k−1.96×SE(β^k)),exp⁡(β^k+1.96×SE(β^k))]\left[\exp\left(\hat{\beta}_k - 1.96 \times \text{SE}(\hat{\beta}_k)\right), \exp\left(\hat{\beta}_k + 1.96 \times \text{SE}(\hat{\beta}_k)\right)\right]
    • An odds ratio confidence interval overlapping 1.01.0 indicates that the association is not statistically significant at the α=0.05\alpha = 0.05 level (corresponding to a log-scale confidence interval overlapping 0.00.0).

Model Diagnostics, Assumptions, and Data Considerations

  • Core Assumptions of Logistic Regression:

    • Binary Outcome: The dependent variable must be strictly binary and follow a Bernoulli or Binomial distribution.
    • Linearity in the Logit: Continuous independent variables XX must exhibit a strictly linear relationship with the log-odds (logit) of the outcome YY.
    • Independence of Observations: Sample observations must be independent of one another (satisfied via appropriate random sampling without clustering or repeated measures over time).
  • Residual Metrics in Generalized Linear Models:

    • Pearson Residuals: Measures the standardized discrepancy between observed outcomes and fitted probabilities:         Pearsoni=yi−p^ip^i(1−p^i)\text{Pearson}_i = \frac{y_i - \hat{p}_i}{\sqrt{\hat{p}_i (1 - \hat{p}_i)}}
    • Deviance Residuals: The default residual type in Generalized Linear Models (GLMs). Represents the individual contribution of each point to the overall model deviance:         Deviance Residuali=sign(yi−p^i)−2[yiln⁡(p^i)+(1−yi)ln⁡(1−p^i)]\text{Deviance Residual}_i = \text{sign}(y_i - \hat{p}_i) \sqrt{-2 \left[y_i \ln(\hat{p}_i) + (1 - y_i) \ln(1 - \hat{p}_i)\right]}
    • Ineffectiveness of Traditional OLS Diagnostic Plots: Standard OLS diagnostic plots generated by routine software commands (e.g., Residuals vs Fitted, Normal Q-Q, Scale-Location) are uninformative for binary outcomes because residuals naturally cluster into two distinct parallel lines bounded by the discrete 00 and 11 outcomes.
  • Diagnostic Tools for Assumption Verification:

    • Linearity Verification: Assessed using visual regression methods such as partial residual plots (visreg). Nonlinear patterns require variable transformations (e.g., polynomial terms, splines, or categorization).
    • Independence Verification: Evaluated via Autocorrelation Function plots (acf()) on residuals to ensure no systematic lag correlations exist.

Linearly checking log-odds using visreg plot

  • Identification and Treatment of Anomalous Data:
    • Outliers: Observations with extreme, unpredicted outcome values that produce large deviance residuals.
    • High Leverage Points: Observations containing extreme exposure/predictor values (XX), characterized by large hat values.
    • Influential Points: Observations possessing a combination of extreme exposure (XX) and extreme outcome (YY) values that substantially alter the slope of the regression line. Influential points are identified using Cook's Distance values:
      • Cut-off thresholds for Cook's Distance are typically set at >0.5> 0.5 or >1.0> 1.0
      • Remediation Procedure: Sequentially evaluate and remove the most influential observation, refit the regression model, and repeat until the estimated regression coefficients (β\beta) stabilize.
    • Multicollinearity: Strong linear inter-correlations between two or more independent variables. It inflates coefficient variance and leads to unstable parameter estimates.
      • Quantified using the Variance Inflation Factor (VIF).
      • Actionable multicollinearity threshold: VIF>4\text{VIF} > 4 or VIF>5\text{VIF} > 5

Cook's Distance and Leverage Diagnostics

  • Variable Transformations:

    • Algebraic Scaling: Rescaling continuous variables (e.g., dividing age by 10 using I(age/10) or converting income from single currency units to thousands) alters the magnitude of the log-odds coefficient and standard error by the identical scale factor.
    • Effect of Scaling: The resulting odds ratio reflects the outcome change per multi-unit step (e.g., per 10-year age increase). Test statistics (zz-values), p-values, overall deviance, and model fit metrics remain mathematically identical.
  • Missing Data Mechanics:

    • Generalized Linear Model estimation routines default to complete-case analysis (listwise deletion), discarding any observation containing missing values in either the outcome or included predictors.
    • Sample Size Impact: Dropping incomplete records reduces statistical power and introduces selection bias if data are not missing completely at random (MCAR).
  • Linear Separability (Complete and Quasi-Complete Separation):

    • Occurs when predictor values perfectly (or near-perfectly) separate outcome states 00 and 11 (e.g., a 2×22 \times 2 contingency table with a zero cell count).
    • Mathematical Failure: Maximum likelihood estimation breaks down because the optimal coefficient β^\hat{\beta} approaches ±∞\pm \infty.
    • Software Manifestation: Model summaries output extremely large parameter estimates (β^≈18.6\hat{\beta} \approx 18.6), massive standard errors (SE≈1586.96\text{SE} \approx 1586.96), uninformative p-values near 1.01.0 (p≈0.9906p \approx 0.9906), and non-convergent profiling warnings.

Goodness of Fit and Model Comparison

  • Goodness-of-Fit Assessment Frameworks:

    • Calibration Plots: Visualizes agreement between observed and predicted probabilities by grouping data into quantile bins (e.g., deciles) and plotting mean predicted probability against observed event frequency per group. Ideal fit aligns directly along the 45-degree diagonal.
    • Hosmer-Lemeshow Test: Computes a summary χ2\chi^2 statistic evaluating discrepancies between observed and predicted outcome frequencies across risk deciles. Caution: The test lacks statistical power in small samples and can be overly sensitive in large samples.
  • Deviance Concepts and Formulas:

    • Saturated Model: A theoretical model containing as many estimated parameters as sample observations (k=nk = n), resulting in perfect fit where predicted probabilities match observed outcome values exactly (LL^saturated\widehat{\text{LL}}_{\text{saturated}}).
    • Null Model Deviance (D0D_0 or DnullD_{\text{null}}): Deviance derived from an intercept-only model containing no predictor variables (k=0k = 0). Represents the baseline unexplainable variation.
    • Residual Model Deviance (DcurrentD_{\text{current}}): Deviance remaining after fitting the specified model with predictor variables.
    • Mathematical Definition of Deviance (DD):D=−2(LL^current−LL^saturated)D = -2 \left(\widehat{\text{LL}}_{\text{current}} - \widehat{\text{LL}}_{\text{saturated}}\right)
    • Smaller deviance values indicate superior model fit as the current model log-likelihood approaches the saturated log-likelihood.
  • Pseudo-R2R^2 (Percentage of Deviance Explained):

    • Calculated as:         Pseudo R2=(Dnull−DcurrentDnull)×100\text{Pseudo } R^2 = \left(\frac{D_{\text{null}} - D_{\text{current}}}{D_{\text{null}}}\right) \times 100
    • Limitation: Does not penalize for model complexity or additional parameters. Adding any predictor automatically reduces residual deviance, making pseudo-R2R^2 prone to overfitting.
  • Comparison of Nested Models via Likelihood Ratio Test (LRT):

    • Nested Model Definition: Model A is nested within Model B if all predictor terms present in Model A are simultaneously contained within Model B.
    • Likelihood Ratio Test Statistic: Evaluates whether adding parameter terms yields a statistically significant reduction in deviance:         ΔD=DModel 1−DModel 2=−2(ln⁡LModel 1−ln⁡LModel 2)=−2ln⁡(LModel 1LModel 2)\Delta D = D_{\text{Model 1}} - D_{\text{Model 2}} = -2 \left(\ln L_{\text{Model 1}} - \ln L_{\text{Model 2}}\right) = -2 \ln\left(\frac{L_{\text{Model 1}}}{L_{\text{Model 2}}}\right)
    • The test statistic ΔD\Delta D follows a Chi-square (χ2\chi^2) distribution with degrees of freedom equal to the difference in estimated parameter counts (ΔDF=DF<em>1−DF</em>2\Delta \text{DF} = \text{DF}<em>{1} - \text{DF}</em>{2}).
  • Comparison of Non-Nested Models via Akaike Information Criterion (AIC):

    • Used to compare nested or non-nested candidate models fitted to the identical dataset.
    • Mathematical Formula:AIC=−2ln⁡(L)+2(k+1)\text{AIC} = -2 \ln(L) + 2(k + 1)         where kk is the number of predictor variables in the model, and k+1k + 1 accounts for total parameters including the intercept.
    • AIC explicitly penalizes model complexity by adding 2(k+1)2(k + 1) to the deviance. Smaller (or more negative) AIC values designate the superior trade-off between model fit and parsimony.

Classification, Decision Boundaries, and Discrimination

  • Explanatory vs Predictive Modeling Frameworks:

    • Explanatory Modeling (Inference): Purpose is testing causal hypotheses and isolating parameter estimates (β\beta). Requires adjusting for known confounders regardless of p-values, verifying model assumptions, and interpreting adjusted odds ratios.
    • Predictive Modeling: Purpose is accurately assigning outcome status (Y^\hat{Y}) to unobserved instances. Focuses on classification accuracy, choice of decision boundaries, and discrimination metrics.
  • Decision Boundaries and Classification Matrices:

    • Models convert fitted probabilities p^\hat{p} into binary classifications (00 or 11) by setting a decision boundary threshold (τ\tau):
      • If p^≥τ\hat{p} \ge \tau, the instance is classified as Positive (11).
      • If p^<τ\hat{p} < \tau, the instance is classified as Negative (00).
    • Default decision boundary threshold is typically τ=0.50\tau = 0.50, treating false positive and false negative classification errors as equally costly.
    • Classifications evaluated against true observed outcomes generate a Confusion Matrix:
Observed OutcomeModel Prediction: Positive (11)Model Prediction: Negative (00)
Positive (11)True Positives (aa)False Negatives (cc)
Negative (00)False Positives (bb)True Negatives (dd)
  • Performance Metrics Formulations:

    • Overall Classification Accuracy:Accuracy=a+da+b+c+d\text{Accuracy} = \frac{a + d}{a + b + c + d}
    • Sensitivity (True Positive Rate / Recall):Sensitivity=aa+c\text{Sensitivity} = \frac{a}{a + c}
    • Specificity (True Negative Rate):Specificity=db+d\text{Specificity} = \frac{d}{b + d}
    • False Positive Rate (FPR):FPR=1−Specificity=bb+d\text{FPR} = 1 - \text{Specificity} = \frac{b}{b + d}
    • Positive Predictive Value (PPV / Precision):PPV=aa+b\text{PPV} = \frac{a}{a + b}
    • Negative Predictive Value (NPV):NPV=dc+d\text{NPV} = \frac{d}{c + d}
  • Impact of Modifying Decision Boundaries:

    • Decreasing the threshold τ\tau (e.g., from 0.500.50 to 0.100.10) lowers the barrier for positive classification, increasing Sensitivity (capturing more true positive events), but increases False Positives (lowering Specificity and PPV).
    • Increasing the threshold τ\tau increases Specificity and PPV, but risks missing true cases (lowering Sensitivity).
    • Accuracy Paradox: In heavily imbalanced datasets (e.g., low outcome prevalence), a naive classifier predicting all instances as negative can achieve high accuracy (e.g., 99%99\% accuracy when 99%99\% of cases are class 0) while possessing zero clinical utility.
  • Receiver Operating Characteristic (ROC) Curve and AUC:

    • The ROC curve plots Sensitivity (True Positive Rate) on the Y-axis against 1−Specificity1 - \text{Specificity} (False Positive Rate) on the X-axis across all decision boundary thresholds (τ∈[0,1]\tau \in [0, 1]).
    • Area Under the Curve (AUC / ROC AUC): Quantifies the discrimination capacity of the model:
      • AUC=0.50\text{AUC} = 0.50 indicates performance no better than random guessing.
      • AUC=1.00\text{AUC} = 1.00 indicates perfect classification discrimination.
      • Higher AUC values reflect superior classifier performance across all threshold settings.
  • Precision-Recall (PR) Curves and AUCPR:

    • Plots Precision (PPV) on the Y-axis against Recall (Sensitivity) on the X-axis.
    • Recommended over standard ROC curves when evaluating models trained on highly imbalanced datasets (e.g., rare outcome states).
    • The baseline Area Under the Precision-Recall Curve (AUCPR) equals the baseline prevalence of the positive outcome class in the dataset.

Precision-Recall Curves

Comprehensive Case Studies and Calculations

  • Case Study 1: Coronary Heart Disease (CHD) Risk Study (N=3479N = 3479)

    • Dataset Overview: Sample size N=3479N = 3479 individuals; 25152515 without CHD, 964964 with CHD. Sample characteristics: Age 55±9 years55 \pm 9\,\text{years}, 57%57\% female, 43%43\% current smokers, Diastolic Blood Pressure 83±12 mmHg83 \pm 12\,\text{mmHg}, Body Mass Index 26±4 kg/m226 \pm 4\,\text{kg/m}^2
    • Model 1: Intercept-Only Modellogit(p^CHD)=−0.96(SE=0.04,p<0.001)\text{logit}(\hat{p}_{\text{CHD}}) = -0.96 \quad (\text{SE} = 0.04, p < 0.001)
      • Null Deviance = 41074107 on 34783478 DF; Residual Deviance = 41074107 on 34783478 DF.
      • Expected log-odds of CHD = −0.96-0.96
      • Expected odds of CHD = exp⁡(−0.96)=0.38\exp(-0.96) = 0.38
      • Expected probability of CHD = 0.381+0.38=0.28\frac{0.38}{1 + 0.38} = 0.28
      • Classification at Threshold τ=0.30\tau = 0.30: Outcome assigned as "No CHD" (0.28<0.300.28 < 0.30).
      • Classification at Threshold τ=0.20\tau = 0.20: Outcome assigned as "CHD" (0.28≥0.200.28 \ge 0.20).
    • Model 2: Single Predictor Model (Age)logit(p^CHD)=−2.85+0.03×Age(βAge SE=0.004,p<0.001)\text{logit}(\hat{p}_{\text{CHD}}) = -2.85 + 0.03 \times \text{Age} \quad (\beta_{\text{Age}} \text{ SE} = 0.004, p < 0.001)
      • Residual Deviance = 40364036 on 34773477 DF.
      • Baseline log-odds at Age 0 = −2.85-2.85
      • Odds Ratio per 1-year age increase = exp⁡(0.03)=1.03\exp(0.03) = 1.03
      • Expected log-odds for age 55 = −2.85+0.03(55)=−1.20-2.85 + 0.03(55) = -1.20
      • Expected odds for age 55 = exp⁡(−1.20)=0.30\exp(-1.20) = 0.30
      • Expected probability for age 55 = 0.301+0.30=0.23\frac{0.30}{1 + 0.30} = 0.23
    • Model 3: Single Predictor Model (Sex, Reference = Male)logit(p^CHD)=−0.59−0.68×SexFemale(βFemale SE=0.08,p<0.001)\text{logit}(\hat{p}_{\text{CHD}}) = -0.59 - 0.68 \times \text{Sex}_{\text{Female}} \quad (\beta_{\text{Female}} \text{ SE} = 0.08, p < 0.001)
      • Residual Deviance = 40264026 on 34773477 DF.
      • Males: Log-odds = −0.59-0.59, Odds = exp⁡(−0.59)=0.55\exp(-0.59) = 0.55, Probability = 0.551.55=0.355\frac{0.55}{1.55} = 0.355
      • Females: Log-odds = −0.59−0.68=−1.27-0.59 - 0.68 = -1.27, Odds = exp⁡(−1.27)=0.28\exp(-1.27) = 0.28, Probability = 0.281.28=0.219\frac{0.28}{1.28} = 0.219
      • Odds Ratio Females vs Males = exp⁡(−0.68)=0.51\exp(-0.68) = 0.51
    • Model 4: Multivariable Model (Age + Sex)logit(p^CHD)=−2.58+0.04×Age−0.72×SexFemale\text{logit}(\hat{p}_{\text{CHD}}) = -2.58 + 0.04 \times \text{Age} - 0.72 \times \text{Sex}_{\text{Female}}
      • Residual Deviance = 39503950 on 34763476 DF.
      • Predictions for Male Aged 55: Log-odds = −2.58+0.04(55)=−0.38-2.58 + 0.04(55) = -0.38; Odds = exp⁡(−0.38)=0.68\exp(-0.38) = 0.68; Probability = 0.681.68=0.40\frac{0.68}{1.68} = 0.40
      • Predictions for Male Aged 56: Log-odds = −2.58+0.04(56)=−0.34-2.58 + 0.04(56) = -0.34; Odds = exp⁡(−0.34)=0.71\exp(-0.34) = 0.71; Probability = 0.711.71=0.415\frac{0.71}{1.71} = 0.415
      • Predictions for Female Aged 55: Log-odds = −2.58+0.04(55)−0.72=−1.10-2.58 + 0.04(55) - 0.72 = -1.10; Odds = exp⁡(−1.10)=0.33\exp(-1.10) = 0.33; Probability = 0.331.33=0.25\frac{0.33}{1.33} = 0.25
      • Predictions for Female Aged 56: Log-odds = −2.58+0.04(56)−0.72=−1.06-2.58 + 0.04(56) - 0.72 = -1.06; Odds = exp⁡(−1.06)=0.35\exp(-1.06) = 0.35; Probability = 0.351.35=0.26\frac{0.35}{1.35} = 0.26
    • Model Fit Comparisons:
      • Comparing Age Model vs Multivariable Model: ΔD=4036−3950=86\Delta D = 4036 - 3950 = 86 on ΔDF=1\Delta \text{DF} = 1 (p<0.001p < 0.001). Adding sex significantly improves model fit.
      • Comparing Sex Model vs Multivariable Model: ΔD=4026−3950=76\Delta D = 4026 - 3950 = 76 on ΔDF=1\Delta \text{DF} = 1 (p<0.001p < 0.001). Adding age significantly improves model fit.
      • AUC ROC Comparisons: Intercept + Age (AUC=0.59\text{AUC} = 0.59); Intercept + Sex (AUC=0.58\text{AUC} = 0.58); Intercept + Age + Sex (AUC=0.64\text{AUC} = 0.64).
  • Case Study 2: Laryngeal Cancer Survival Analysis

    • Outcome: Mortality from laryngeal cancer (11 = Dead, 00 = Alive).
    • Model Estimates Output: Intercept = 24.5273524.52735 (SE 9.231279.23127); Stage 2 = 0.300730.30073 (SE 0.692120.69212); Stage 3 = 1.043591.04359 (SE 0.593140.59314); Stage 4 = 2.678712.67871 (SE 0.960260.96026); Age = 0.044090.04409 (SE 0.022520.02252); Year of Diagnosis = −0.37471-0.37471 (SE 0.125600.12560).
    • Fitted Prediction Equation:ln⁡(p^1−p^)=24.53+0.30(Stage2)+1.04(Stage3)+2.68(Stage4)+0.04(Age)−0.37(Year)\ln\left(\frac{\hat{p}}{1 - \hat{p}}\right) = 24.53 + 0.30(\text{Stage}_2) + 1.04(\text{Stage}_3) + 2.68(\text{Stage}_4) + 0.04(\text{Age}) - 0.37(\text{Year})
    • Stage 2 Effect Confidence Interval:
      • Log scale: 0.30±1.96(0.69)=[−1.05,1.65]0.30 \pm 1.96(0.69) = [-1.05, 1.65]
      • Odds scale: [exp⁡(−1.05),exp⁡(1.65)]=[0.35,5.21][\exp(-1.05), \exp(1.65)] = [0.35, 5.21]
      • Interpretation: Interval overlaps 1.01.0, indicating no statistically significant difference in odds of death between Stage 2 and Stage 1 (p=0.664p = 0.664).
    • Interpretation of Stages 3 & 4:
      • Stage 3 Odds Ratio = exp⁡(1.04)=2.83\exp(1.04) = 2.83. The odds of death are 2.832.83 times higher for Stage 3 patients compared to Stage 1, adjusting for age and year (p=0.0785p = 0.0785).
      • Stage 4 Odds Ratio = exp⁡(2.68)=14.59\exp(2.68) = 14.59. The odds of death are 14.5914.59 times higher for Stage 4 patients compared to Stage 1, adjusting for covariates (p=0.00528p = 0.00528).
    • Year Effect: Odds Ratio = exp⁡(−0.37)=0.69\exp(-0.37) = 0.69. Each calendar year progression decreases the adjusted odds of mortality by 31%31\% (p=0.00285p = 0.00285).
    • Probability Calculation: Patient with Stage 3 cancer, aged 72, diagnosed in year '74:
      • Log-odds = 24.53+1.04(1)+0.04(72)−0.37(74)=1.0724.53 + 1.04(1) + 0.04(72) - 0.37(74) = 1.07
      • Odds = exp⁡(1.07)=2.91\exp(1.07) = 2.91
      • Probability = 2.911+2.91=0.75\frac{2.91}{1 + 2.91} = 0.75
  • Case Study 3: Meteorological Prediction Model

    • Fitted Equation:logit(p^sunny)=1.19−0.33(Rainyes)−0.02(Windkph)\text{logit}(\hat{p}_{\text{sunny}}) = 1.19 - 0.33(\text{Rain}_{\text{yes}}) - 0.02(\text{Wind}_{\text{kph}})
    • Rain Effect: Odds Ratio = exp⁡(−0.33)=0.72\exp(-0.33) = 0.72. Rain presence reduces the odds of sunny weather by 28%28\%
    • Wind Effect: Odds Ratio = exp⁡(−0.02)=0.98\exp(-0.02) = 0.98. Each 1 kph wind increase reduces sunny weather odds by 2%2\%
    • No Rain, No Wind (Rain=0,Wind=0\text{Rain} = 0, \text{Wind} = 0):
      • Expected Log-odds = 1.191.19
      • Expected Odds = exp⁡(1.19)=3.29\exp(1.19) = 3.29
      • Expected Probability = 3.294.29=0.77\frac{3.29}{4.29} = 0.77
    • No Rain, Wind = 38 kph (Rain=0,Wind=38\text{Rain} = 0, \text{Wind} = 38):
      • Log-odds = 1.19−0.02(38)=0.431.19 - 0.02(38) = 0.43
      • Expected Odds = exp⁡(0.43)=1.54\exp(0.43) = 1.54
      • Expected Probability = 1.542.54=0.61\frac{1.54}{2.54} = 0.61
    • Rain Present, Wind = 38 kph (Rain=1,Wind=38\text{Rain} = 1, \text{Wind} = 38):
      • Log-odds = 1.19−0.33(1)−0.02(38)=0.101.19 - 0.33(1) - 0.02(38) = 0.10
      • Expected Odds = exp⁡(0.10)=1.11\exp(0.10) = 1.11
      • Expected Probability = 1.112.11=0.526\frac{1.11}{2.11} = 0.526
  • Case Study 4: Caregiver Distress Multivariable Analysis

    • Model Summary Table: Intercept OR = 0.750.75 (95%95\% CI: 0.60,0.930.60, 0.93); Hours of care/week OR = 1.021.02 (95%95\% CI: 1.00,1.041.00, 1.04); Cognitive Impairment Yes OR = 2.192.19 (95%95\% CI: 1.42,3.371.42, 3.37), Ref = No.
    • Fitted Logit Equation:logit(p^)=ln⁡(0.75)+ln⁡(1.02)×Hours+ln⁡(2.19)×CognitiveImpairment\text{logit}(\hat{p}) = \ln(0.75) + \ln(1.02) \times \text{Hours} + \ln(2.19) \times \text{CognitiveImpairment}
    • 10-Hour Weekly Care Increase Effect:OR10hours=1.0210=1.219\text{OR}_{10\text{hours}} = 1.02^{10} = 1.219Interpretation: A 10-hour weekly increase in care hours increases distress odds by 21.9%21.9\%
    • Distress Odds/Probability for Caregiver Working 40 Hours/Week for Cognitively Impaired Individual:
      • Log-odds = ln⁡(0.75)+40×ln⁡(1.02)+ln⁡(2.19)=−0.2877+0.3921+0.7839=0.8883\ln(0.75) + 40 \times \ln(1.02) + \ln(2.19) = -0.2877 + 0.3921 + 0.7839 = 0.8883
      • Odds = exp⁡(0.8883)=2.43\exp(0.8883) = 2.43
      • Probability = 2.431+2.43=0.708  (70.8%\frac{2.43}{1 + 2.43} = 0.708\; (70.8\%
  • Case Study 5: Confusion Matrix Diagnostics at Varying Decision Boundaries

    • Model: logit(p^hypertension)=−3.22+0.05(Age)−1.05(SexFemale)\text{logit}(\hat{p}_{\text{hypertension}}) = -3.22 + 0.05(\text{Age}) - 1.05(\text{Sex}_{\text{Female}})
    • Matrix Scenario A (Decision Boundary Threshold τ=0.30\tau = 0.30):
      • True Positives (aa) = 1717, False Positives (bb) = 1919
      • False Negatives (cc) = 1414, True Negatives (dd) = 102102
      • Total observations = 152152
      • Accuracy=17+102152=119152=0.783  (78.3%)\text{Accuracy} = \frac{17 + 102}{152} = \frac{119}{152} = 0.783\; (78.3\%)
      • Sensitivity=1717+14=1731=0.548  (54.8%)\text{Sensitivity} = \frac{17}{17 + 14} = \frac{17}{31} = 0.548\; (54.8\%)
      • Specificity=10219+102=102121=0.843  (84.3%)\text{Specificity} = \frac{102}{19 + 102} = \frac{102}{121} = 0.843\; (84.3\%)
      • PPV=1717+19=1736=0.472  (47.2%)\text{PPV} = \frac{17}{17 + 19} = \frac{17}{36} = 0.472\; (47.2\%)
      • NPV=102102+14=102116=0.879  (87.9%)\text{NPV} = \frac{102}{102 + 14} = \frac{102}{116} = 0.879\; (87.9\%)
    • Matrix Scenario B (Decision Boundary Threshold τ=0.10\tau = 0.10):
      • True Positives (aa) = 2929, False Positives (bb) = 6262
      • False Negatives (cc) = 22, True Negatives (dd) = 5959
      • Total observations = 152152
      • Accuracy=29+59152=88152=0.579  (57.9%)\text{Accuracy} = \frac{29 + 59}{152} = \frac{88}{152} = 0.579\; (57.9\%)
      • Sensitivity=2929+2=2931=0.935  (93.5%)\text{Sensitivity} = \frac{29}{29 + 2} = \frac{29}{31} = 0.935\; (93.5\%)
      • Specificity=5962+59=59121=0.488  (48.8%)\text{Specificity} = \frac{59}{62 + 59} = \frac{59}{121} = 0.488\; (48.8\%)
      • PPV=2929+62=2991=0.319  (31.9%)\text{PPV} = \frac{29}{29 + 62} = \frac{29}{91} = 0.319\; (31.9\%)
      • NPV=5959+2=5961=0.967  (96.7%)\text{NPV} = \frac{59}{59 + 2} = \frac{59}{61} = 0.967\; (96.7\%)
    • Comparison: Lowering threshold τ\tau from 0.300.30 to 0.100.10 increased Sensitivity from 54.8%54.8\% to 93.5%93.5\% (reducing missed cases), but reduced Specificity from 84.3%84.3\% to 48.8%48.8\% and dropped overall Accuracy from 78.3%78.3\% to 57.9%57.9\%