Advanced Hypothesis Testing: Type I & II Errors, Effect Size, and Statistical Power
The Nature of Sample-Based Research and Uncertainty
Fundamental Limitation of Research: All research is based on a sample rather than an entire population. Even modern censuses are often conducted using statistical sampling to generalize to whole zip codes or populations.
Lack of Certainty: Looking at a sample can never provide absolute certainty about a population. Because research involves making an inference about a population based on a subset, the process is not airtight.
The Problem of Unusual Samples: A researcher can always inadvertently obtain an unusual sample by chance.
- Ginkgo Metaphor: If testing Ginkgo Biloba on memory using a sample of people, it is possible to get a group that naturally has extra-great memory. The researcher might conclude the Ginkgo worked, when in reality, it was just the luck of the draw.
- Depression Therapy Metaphor: In a study with a sample of people, a researcher might find lower depression levels after therapy. However, the sample might have consisted of individuals with naturally milder cases of depression, skewing the result.
- Athletic Performance Metaphor: A study on weight training might conclude it improves performance simply because the handful of athletes in the sample were naturally high performers.
Inherent Risk of Error: Because conclusions about a population (whether to reject or fail to reject ) are based on samples, they are always open to being mistaken.
Type I Errors (False Positives)
Definition: A Type I error occurs when a researcher rejects the null hypothesis () when is actually true.
Logical Meaning: It is the error of rejecting something that is true. For example, if the Earth is round () and you reject that to conclude it is flat, you have made an error.
Treatment Context:
- states that the treatment causes no effect in the population.
- Rejecting means concluding that the treatment works.
- Therefore, a Type I error is concluding a treatment works when it actually does not.
False Positive Terminology: This is often called a "false positive." In a research context, a "positive" result is seeing the treatment as working (which researchers often hope for). A false positive means you had that positive conclusion, but it was wrong.
Causes: Typically caused by obtaining an unusually high-performing sample by chance, which makes a useless treatment (like a sugar pill) look effective.
Type II Errors (False Negatives)
Definition: A Type II error occurs when a researcher fails to reject the null hypothesis () when is actually false.
Logical Meaning: Failing to reject means retaining or keeping the starting assumption (). It is an error to retain a belief that is actually false.
Treatment Context:
- states the treatment doesn't work.
- Failing to reject means concluding the treatment doesn't work.
- Therefore, a Type II error is concluding a treatment does not work when it actually does work.
False Negative Terminology: This is called a "false negative." It represents a missed opportunity to identify a real effect.
Hypothetical Consequence: If Ginkgo Biloba actually is a real memory-enhancing "pill," but a study concludes it doesn't work, people might stop studying it or using it, which is a significant error.
Causes: This can happen if a researcher gets a sample by chance that has unusually poor baseline memory. Even if the treatment raises their memory scores, they only reach "average" levels, leading the researcher to think the treatment did nothing.
Researcher Bias and Error Control
Researcher Bias: Researchers are naturally biased toward making Type I errors. This is because most researchers want their treatment, therapy, or training program to work. They "root" for rejecting .
Statistical Constraints: Because researchers are human and prone to wanting "action" (results), statistical rules are built to tightly cap Type I errors.
Comparing Errors: Neither error is objectively "worse"; both lead to false information. However, historical statistical practice focuses heavily on minimizing Type I errors because researchers might otherwise "cut corners" to find an effect where none exists.
Probability Levels: Alpha () and Beta ()
Alpha (): This is the probability of making a Type I error.
- It is the significance level chosen by the researcher (e.g., or ).
- If , there is a chance of rejecting even if the treatment is useless.
- Tic Tac Example: If a researcher tests if eating a Tic Tac makes you grow taller (a useless treatment) and repeats the study times using , in of those studies, they will likely make a Type I error and conclude Tic Tacs work.
- Using a stricter alpha (like ) makes it harder to reject , which decreases the probability of a false rejection.
Beta (): This is the probability of making a Type II error. Unlike alpha, beta is not as easily "chosen" or precisely stated in introductory stats; it is messier and often requires advanced estimation.
Effect Size and Cohen's
The Question of "How Much": Rejecting only tells you that it is statistically likely the treatment has some effect. It does not communicate the magnitude of that effect.
Contextual Examples:
- Income: If a college degree has an effect on income at age , a student wants to know if that means a or a .
- Diet Pill: If a pill is clinically proven to make you lose weight, you want to know if that is or over months before paying .
Necessity: Effect size is only necessary when you reject . If you fail to reject , you are concluding there is no effect, so the size of the effect is moot.
Standardized Measures (Cohen's ): Often, research uses specialized units (e.g., the "Stanford Wechler memory test"). Most people don't know if a on that test is significant. Cohen's standardizes the effect by measuring how many standard deviations the treatment moves the scores.
Interpreting Cohen's (Ballpark Figures):
- : Small effect.
- : Medium effect.
- : Large effect.
- Example: A Cohen's of is in the small-to-medium range.
Absolute Value: When judging the size of Cohen's , researchers look at the absolute value. A and a are both medium effects; one is a medium decrease, and one is a medium increase.
Statistical Power
Definition: Power is the probability that a study will detect a treatment effect that truly exists in the population.
Relationship to Beta: Power is defined as .
- If a study has power, it has an chance of making a Type II error (falsely concluding the treatment doesn't work).
Importance: A study with low power (e.g., ) is generally not worth doing. Researchers aim for power levels around or .
Factors Increasing Power:
- Larger Effect Size: If a treatment has a huge impact on the population, it is easier to detect. This is usually outside the researcher's control.
- Larger Sample Size (): This is fully within the researcher's control. Increasing the number of participants directly increases the power of the study.
Sample Size Guidelines:
- To ensure a valid normal distribution of means, a sample of at least to is usually required unless the parent population is already normal.
- A researcher should not just stop at the minimum ( to ); bigger is always better (e.g., , , or ) to maximize the chance of detecting a real effect.