Lesson 6 Study Guide: Relationships Between Categorical Variables

Introduction to Analyzing Categorical Variables

  • Lesson 6 focuses on the relationship between categorical variables, specifically how to analyze them visually and numerically to determine if a relationship exists.

  • The primary tools discussed include organizing data into tables, calculating probabilities (risk and odds), and understanding the cautions involved in interpreting these statistics.

  • This lesson moves away from measurement variables to explore how data categories interact.

Analyzing Categorical Data: The Two-Way Table

  • Definition of a Two-Way Table: A summary table where each cell contains the count (frequency) of cases that fall into both a specific row category and a specific column category.

  • Example: Binge-Watching Survey:

    • Organization: YouGov.

    • Data Source: A 2022 survey of 1,0001,000 U.S. adults regarding TV viewing habits.

    • Research Question: Is there a relationship between binge-watching frequency and geographical census regions in the U.S.?

    • Variables: There are exactly two variables shown in this analysis:

    1. Binge-watching frequency (e.g., Always, Never).

    2. Geographical region (Northeast, Midwest, South, West).

    • Sample Size: While 1,0001,000 adults were surveyed, only 932932 provided responses to this specific question, making 932932 the grand total for the table.

Reading and Testing Knowledge of Two-Way Tables

  • Summarized vs. Raw Data: Two-way tables contain summarized data (totals/counts), not raw individual data points.

  • Navigating Rows and Columns:

    • To find the total who reported "Never" binge-watching, look at the end of that specific row (Total = 128128).

    • To compare specific regions for the "Always" category, look across the row: there were more "Always" responses from the South than from the West.

    • Total respondents from the Northeast reached 167167 (found at the bottom of the column).

    • The "Always" category totaled 6868 out of 932932, making the statement that it was the "most selected" false.

Calculating Proportions and Conditional Proportions

  • Proportion of Total Sample:

    • Proportion who always binge-watch: 689320.073\frac{68}{932} \approx 0.073 or approximately 7.3%7.3\%.

  • Combined Proportions:

    • Proportion from the South (364364) or the West (209209): 364+209932=5739320.615\frac{364 + 209}{932} = \frac{573}{932} \approx 0.615 or 61.5%61.5\%.

  • Conditional Proportions:

    • These limit the denominator to a specific subgroup (a "double conditional").

    • Proportion of Northeast residents who always binge-watch:

    • Number who always watch in the Northeast = 1717.

    • Total in the Northeast = 167167.

    • Calculation: 171670.10\frac{17}{167} \approx 0.10 or 10%10\%.

Visualizing Categorical Relationships

  • Side-by-Side Bar Charts: Used to compare categories across different groups (e.g., region vs. binge frequency).

  • Segmented Bar Charts: A single bar represents a whole group (like a region), and different colors within that bar represent the frequency levels of binge-watching, stacking them to show total proportions.

Quantitative Methods: Risk and Odds

  • In research studies, particularly randomized comparative experiments, groups are often divided into treatment and placebo to measure differences.

  • Risk:

    • Formula: Risk=Number with the traitTotal number in the group\text{Risk} = \frac{\text{Number with the trait}}{\text{Total number in the group}}.

    • Example: In a set of 5 items where 2 are positive, the risk is 25=0.4\frac{2}{5} = 0.4.

  • Odds:

    • Formula: Odds=Number with the traitNumber without the trait\text{Odds} = \frac{\text{Number with the trait}}{\text{Number without the trait}}.

    • Example: In a set of 5 items where 2 are positive and 3 are negative, the odds are 22 to 33, or 23\frac{2}{3}.

    • Note: Odds can result in improper fractions (values greater than 1), unlike risk.

Relative Risk (RR) and Odds Ratios (OR)

  • Relative Risk (RR):

    • Formula: Relative Risk=Risk in Group 1Risk in Group 2\text{Relative Risk} = \frac{\text{Risk in Group 1}}{\text{Risk in Group 2}}.

    • Interpretation:

    • RR=1RR = 1: The risk is identical for both groups.

    • RR=3RR = 3: The first group is 3 times more likely to have the outcome than the second.

    • Increased Risk: If RR=1.25RR = 1.25, the increased risk is 25%25\% (calculated as RR1RR - 1).

  • Odds Ratio (OR):

    • Formula: Odds Ratio=Odds in Group 1Odds in Group 2\text{Odds Ratio} = \frac{\text{Odds in Group 1}}{\text{Odds in Group 2}}.

    • Interpretation:

    • OR=1OR = 1: The odds are the same for both groups.

    • OR=3OR = 3: The odds for the first group are 3 times the odds of the second group.

Case Study: Rosiglitazone for Type 2 Diabetes

  • Study Goal: Evaluating the success of glycemic control in youth with Type 2 diabetes using three treatments (N=699N = 699 participants randomly assigned).

  • Explanatory Variables (3 Groups):

    1. Metformin: Known drug alone (n=232n = 232 total).

    2. ROSI: Metformin plus Rosiglitazone (n=233n = 233 total; 143143 successes, 9090 failures).

    3. Lifestyle: Metformin plus weight loss/exercise program (n=234n = 234 total; 125125 successes, 109109 failures).

  • Response Variable: Glycemic control (Success vs. Failure).

  • Calculations for the ROSI Group:

    • Risk of Success: 143233\frac{143}{233}.

    • Odds of Success: 14390\frac{143}{90}.

  • Relative Risk Calculation (ROSI vs. Lifestyle):

    • RR=143/233125/2341.15\text{RR} = \frac{143/233}{125/234} \approx 1.15.

    • Interpretation: Patients using rosiglitazone had a 1.151.15 times greater risk of success, or a 15%15\% increased risk of success compared to the lifestyle group.

  • Odds Ratio Calculation (ROSI vs. Metformin Alone):

    • Odds for ROSI = 143901.59\frac{143}{90} \approx 1.59.

    • Odds for Metformin = 1121200.93\frac{112}{120} \approx 0.93.

    • OR=1.590.931.71\text{OR} = \frac{1.59}{0.93} \approx 1.71.

    • Interpretation: The odds of success for ROSI patients are approximately 1.71.7 times the odds for those on metformin alone.

Cautions and Interpretations of Risk

  • Baseline Risk: It is crucial to know the starting risk level. A reported "5 times higher risk" could mean an increase from 22 per 100,000100,000 to 1010 per 100,000100,000. While significant, the absolute risk remains very low.

  • Protective Effect: A relative risk less than 1 (e.g., RR < 1.0) indicates that the factor being studied provides a protective effect, lowering the risk below the baseline.

  • Specificity: Risk is rarely universal. It is usually conditional on factors like age, sex, health history, and behavior.

    • Example: Men over 50 have 55 times the risk of sudden cardiac events during marathons compared to women under 40. This risk does not apply equally to a 22-year-old woman.

  • Relative Risk vs. Baseline Risk:

    • Relative risk answers: "Compared to what?"

    • Baseline risk answers: "How big is the problem in the first place?"

Identifying Statistical Statements

  • Relative Risk: "Participants taking ROSI were 1.51.5 times as likely to maintain control."

  • Odds: "17 reported always binge-watching, while 159 did not."

  • Odds Ratio: "The odds of reporting 'always' among Northeast respondents were 2.12.1 times the odds among respondents from the Midwest."

  • Risk: "Across all 932932 respondents, 7.3%7.3\% reported always binge-watching."

Simpson’s Paradox

  • Definition: An observed association between two variables that changes or reverses direction when a third confounding variable is introduced that interacts strongly with both.

  • Speed Limit Example:

    • In a scatterplot of speed vs. accident rates, the overall trend might appear negative (as speed limits increase, injuries go down).

    • However, when looking within individual speed limit groups (e.g., 30,50,80,100km/h30, 50, 80, 100\,\text{km/h}), the trend in each group is positive (higher average speed within that limit leads to more injuries).

  • Restaurant Recommendation Example:

    • Carlos' restaurant may have a higher recommendation percentage among males and a higher percentage among females individually.

    • Yet, when the data is merged, Sophia's restaurant might have a higher overall percentage due to differences in sample sizes within the gender subgroups (e.g., 200200 vs. 3636).

  • Prevention: To avoid Simpson's Paradox, researchers must identify potential confounding variables during the study design phase and analyze data within subgroups rather than just in aggregate.