3032 Week 7
Univariate Count Data
On the Usefulness of Nominal and Ordinal Data
Independent Variables: Commonly categorical.
Control versus Treatment: A or B.
Dependent Variables: Can also be categorical.
Examples include:
Success of treatment (versus failure)
Adherence to policy (versus non-adherence)
Count Data
Key Aspects:
Mutual Exclusivity: Observations cannot belong to more than one category.
Exhaustiveness: Categories must represent all options present in the dataset; no observations should be left uncategorized.
Binomial Experiments
Definitions:
Dichotomous Data: Data categorized into two groups only.
Characteristics of Binomial Experiments:
Must consist of a set number of identical trials.
Each trial has only two possible outcomes.
Outcomes are mutually exclusive and exhaustive.
Each trial must have equal probabilities.
Probability Notation:
Let be the probability of the "outcome of interest."
Relationship: , hence .
Bernoulli Trials
A specific type of binomial experiment involving a set number of identical trials.
Example: Consider a 10-item test where (probability of guessing) = 0.2.
Calculation: What is the probability a student guesses 5 or more questions correctly?
Bernoulli Trial:
For , what is the probability a student guesses a particular question correctly?
Distribution Parameters:
For Bernoulli distributions: Mean = and Standard Deviation (SD) = .
Distributions of Proportions
Notation:
Sample proportions are denoted with a "hat": .
Distribution Center: The center of the distribution of sample proportions is equal to .
SD of Sample Proportions:
Formula: .
Maximal Variability: Occurs when .
Flashback!
Z-Score formula:
Where:
= Sample mean
= Population mean
= Population standard deviation
= Sample size
Second Z-Score formula relates to proportions:
Confidence Interval for Proportions
Formula:
Where:
Point estimate = Sample proportion ()
Confidence level =
Standard Error =
Example #1 - Hypothesis Testing
Step 1: Hypothesis Statements
Null Hypothesis ():
Alternative Hypothesis ():
Step 2: Choose Test
Test Statistic:
Step 3: Rejection Region / Critical Value
Assumed Distribution: Reject if observed z-value () < -1.645
Set significance level ( = 0.05) leads to:
Critical value ()
Step 4: Calculate Obtained Value
Using given values:
Step 5a: Comparison, Decision, Conclusion
Comparing values:
() < ()
Conclusion: Reject . Evidence suggests fewer Canadians suffer from allergic rhinitis than Europeans.
Step 5b: Effect Size and Confidence Interval
Effect Size Calculation:
Confidence Interval:
Based on earlier calculations:
Final confidence interval: (i.e., between 28.9% and 33.1%).
Example #2
Results indicate we are 99% confident that between 9.4% and 10.4% of university students meet all four components of the guidelines:
Calculation:
Given values:
Example #3 - Estimating Sample Size for Confidence Interval
Required values:
Critical Value:
Margin of Error:
Formula for Sample Size:
Final calculation for required sample size:
Count Data
Presentation of Proportional Statistics:
Proportional data can also be presented as frequency counts.
Advantages include:
Offers researchers a reminder to state the total sample size (essential for interpretation).
Allows for easier comparison of more than two proportions.
Proportions to Count Data (Example #1 Revisited)
Observations:
419 Canadians report suffering from hay fever.
931 Canadians report not suffering from hay fever.
Expectations:
European prevalence of hay fever is assumed to be 0.35.
Count Breakdown:
Hay Fever () and No Hay Fever ()
Total Observations:
Observed: 419 (hay fever), 931 (no hay fever) out of 1350 total.
Expected Counts:
Hay Fever:
No Hay Fever:
Chi-square Calculations
For each cell in the count data matrix:
Compare observed score () with expected value ().
Chi-square is calculated as:
Using previous scores:
Observed: 419 and 931
Expected: 472.5 and 877.5
Chi-square Result:
From the calculations:
Test of Model Fit
Chi-square is a model fit test applicable for any number of categories.
Example #4
Setup for Chi-square Test
Data distribution:
Categories: 1, 2, 3, 4, 5, 6 with respective counts: 30, 150, 340, 330, 130, 20.
Number of pairwise comparisons calculated as .
Step 1a: Hypothesis Statements
Question #1: Is the distribution uniform?
: , , etc.
: At least one observed proportion differs from expectation.
Question #2: Is the distribution normal?
: Expected proportions based on normal distribution.
: At least one observed proportion differs from expectation.
Step 2: Choosing Test Statistic
Test Statistic calculation:
Step 3: Rejection Region / Critical Value
Degrees of Freedom Calculation: df = where = number of categories.
For , df = 5.
Critical Value at is .
Reject if .
Step 4: Calculate Obtained Value (Uniform Distribution)
Using expected count: Each cell equals . Calculate:
Step 4: Calculate Obtained Value (Normal Distribution)
Expected counts based on normal distribution proportions:
Calculate:
Step 5: Conclusion from Chi-square Results
Comparison: (For uniform distribution) versus .
Conclusion: Fail to reject (uniform distribution).
For normal distribution:
Reject .
Step 5b: Evaluation of Standardized Residuals
Chi-square is an omnibus test, indicating at least one probability differs.
Use standardized residuals to identify specific contributions:
.
Values calculated:
.
Residuals greater than 2 indicate meaningful contributions to chi-square values.
Next Week…
Bivariate Count Data!
The confidence interval in a z-test of proportions is centered around: The sample proportion.
When comparing two dichotomous variables, one could use either a chi-square or a z-test of proportions.
The standard error for the confidence interval surrounding the sample proportion is calculated using sample proportions.
The center of the distribution of sample proportions is p̂.
Variability is maximized for a binomial variable when p̂ = 0.50.
A significant chi-square in a simple comparison of observed frequencies suggests that all of the above are correct (individuals are not assorted into categories randomly, at least one category's values are different from expectation, the null hypothesis is not true).
The denominator for the z-test of proportions is the standard error.
The expected values used in simple comparisons of cell size are determined by multiplying the total number of participants by the expected proportion in each category.
When the results of a chi-square test are significant, standardized residuals that are greater than or equal to +2; -2 are considered to have contributed significantly to the overall difference between categories.
We use the sample proportion in our confidence interval calculations in a z-test of proportions because we are assuming that the sample proportion is our best estimate of the population proportion.
Which of the following is false regarding the distribution of a binomial variable? p̂ + q̂ = ± 1.0 is false (it should equal 1.0).
Variance in a Bernoulli distribution is maximized when p = 0.50.
Which of the following is inaccurate regarding binomial experiments? The distribution of proportions must be normal across all trials is inaccurate.
An example of a Bernoulli trial is: The probability that a child will guess the correct colour of the ball you're holding behind your back in one try.
A z-test of proportions is appropriate when comparing an observed proportion to a hypothetical proportion.
When the categories represent all possible options for a variable, the categories are said to be exhaustive.
In a simple comparison of group sizes, with 6 categories and 101 participants, the degrees of freedom would be 5 (df = k - 1 = 6 - 1).
In a z-test where we’re interested in determining if one's sample is less likely to suffer from a disease than a population, we will reject the null hypothesis when z-obtained is less than -2.575 (for alpha of 0.01).
For the Nursing students' activity guidelines study, reject the null hypothesis since the z-obtained was 2.23 and the z-critical was 1.645, we conclude that the population from which we drew our sample has a significantly different proportion than 0.653.
To find the chi-square for the study on migraines, the obtained value could be 5.0239 (applying the chi-square formula based on observed vs expected values).