Hypothesis Testing Notes
Introduction to Hypothesis Testing
Hypothesis testing involves materials, procedures, forming a hypothesis, and drawing a conclusion.
What is a Hypothesis?
A hypothesis is a statement or claim that requires supporting evidence.
The evidence should either support or refute the theoretical proposition.
It is a statement about a population.
A researcher collects data to test the claim.
A hypothesis test evaluates two mutually exclusive statements about a population to determine which statement is best supported by the sample data.
Studying the entire population might not be feasible, so researchers collect sample data.
The sample data serves as evidence to check the reasonableness of the claim.
A hypothesis test is a statistical inferential method to determine if there is enough evidence in a sample to prove a hypothesis is true for the entire population.
Example: Academic Dean claims recent graduates make an average of $35,000.
Data is collected to prove or refute the claim by randomly selecting 50 graduates to obtain their annual income.
Suppose the sample group makes an average income of $34,200. Question: Is $34,200 close enough to the claimed $35,000? Is the difference of $800 significant?
If the difference of is statistically significant, we reject the claim.
If the difference is not significant, we accept the claim.
Types of Hypothesis Statements
Alternative (or Research) Hypothesis (H1)
States there is a significant difference between the hypothetical value and the observed value.
Symbolized by .
The characteristics of the two values are not equal (the difference is not equal to zero).
Assumes means of two groups are not equal, statistically not the same, and have dissimilar characteristics.
The difference is so large it could not have occurred by random chance or sampling error.
Mathematically stated as:
Null Hypothesis (H0)
States there is no significant difference between the hypothetical value and the observed value.
Symbolized by .
The characteristics of the two values are equal, hence the difference is equal to zero.
Assumes that the means of the two groups are equal, statistically the same, and have the same characteristic.
The difference is so small that it could have occurred by random chance or sampling error.
Mathematically stated as:
Examples of Null and Alternative Hypotheses
Null Hypothesis: There is no significant difference in the perception of public safety between males and females.
Alternative Hypothesis: There is a significant difference in the perception of public safety between males and females.
Null Hypothesis: The average cost of tuition fee is equal to $6,500 per year.
Alternative Hypothesis: The average cost of tuition fee is more than $6,500 per year.
Null Hypothesis: The day shift workers have the same productivity as the night shift workers.
Alternative Hypothesis: The day shift workers are more productive than the night shift workers.
Null Hypothesis: An average social worker earns $60,000 a year.
Alternative Hypothesis: An average social worker earns less than $60,000 a year.
Null Hypothesis: The defendant is not guilty.
Alternative Hypothesis: The defendant is guilty.
Default Assumption and Testing
The default assumption is that the null hypothesis is true (no significant difference) until evidence is shown to the contrary.
In a criminal trial, the defendant is assumed not guilty until proven in court.
The researcher will directly test for the validity of the null hypothesis with the hope that the null hypothesis can be rejected in order to prove that his/her research hypothesis is valid.
Since the null hypothesis is assumed to be true, data is collected to seek evidence against it.
The researcher can only prove his hypothesis is true only when he/she has sufficient evidence to reject the null hypothesis.
We do not directly test for the alternative/research hypothesis.
It is only when the null hypothesis is rejected that we consider the alternative/research hypothesis to be true.
The decision to “reject” or “accept” is made only in reference to the null hypothesis. Statistically, we say “not to reject” as opposed to “accept” because we can never truly confirm a hypothesis given that we may later find evidence to the contrary.
In a criminal trial, the prosecutor (researcher) brings a case to court to prove their hypothesis.
The null hypothesis is the defendant is not guilty, and the research hypothesis is the defendant is guilty.
The prosecutor can only prove that the defendant is guilty (research hypothesis) when he/she is able to provide enough evidence (beyond reasonable doubt) to reject the presumption of innocent (not guilty).
The onus is on the prosecutor to prove that the null hypothesis is false by seeking sufficient evidence against the null hypothesis in order to reject it.
Police collect data through investigation to seek evidence against the null hypothesis with the goal of rejecting the “not guilty” (null hypothesis) assumption and proving that the research hypothesis is valid.
If sufficient evidence (beyond reasonable doubt) is established, the jury will reject the null hypothesis and return a guilty verdict against the defendant (that is research hypothesis is valid).
If the prosecutor could not provide sufficient evidence against the null hypothesis then the jury will accept the null hypothesis and return a verdict that the defendant is not guilty. We can not say the defendant is innocent because we can later find evidence to the contrary.
Types of Tail Test
Two-Tail Test
The research or alternative hypothesis states that a parameter is different from the hypothesized value.
Seeks difference in either direction (less than or more than).
Example: Daily calorie intake of a Canadian adult male is different from 2500 kcal.
One-Tail Test
Left-Tail or Lower Tail: When the research or alternative hypothesis states that a parameter is less than the hypothesized value. Example: An average female weighs less than 170 pounds.
Right-Tail or Upper Tail: When the research or alternative hypothesis states that a parameter is more than the hypothesized value. Example: An average male is taller than 70 inches.
Types of Error
Decision is based on data evidence collected from the sample rather than the entire population, so there is a possibility of making a wrong decision.
The only way to be 100% certain is to test the entire population (highly impossible).
It is proper and expected to allow some room for error when testing a hypothesis.
Rather than prove a hypothesis, the researcher estimates the likelihood of the assumption being correct given a certain probability of error.
Type 1 Error
Occurs when a null hypothesis is rejected when in fact it is true.
This refers to the probability of rejecting the null hypothesis when in fact it is true.
Also called the significance level and is represented by the value of alpha ().
This is the false rejection of the null hypothesis. The test falsely provides evidence in support of the alternative hypothesis. A false positive.
A researcher concludes that there is a significant different between the two values whereas in fact the difference is not significant.
Example: a researcher concludes that there is a significant difference in the heights between boys and girls whereas in fact the difference is not significant.
Type 2 Error
Occurs when a null hypothesis is accepted when in fact it is false.
This refers to the probability of accepting the null hypothesis when in fact it is false.
The probability of committing this type of error is represented by the value of beta ().
This is the false acceptance of the null hypothesis. The test falsely provides evidence against the alternative hypothesis. A false negative.
A researcher concludes that there is no significant difference between the two values whereas in fact the difference is significant.
Example: a researcher concludes that there is no significant difference in the weights between boys and girls whereas in fact the difference is significant.
Examples illustrating Type I and Type II Errors
Justice System
Null Hypothesis: A defendant is not guilty.
Alternative Hypothesis: A defendant is guilty.
Type 1 Error: A defendant did not commit the crime, but the jury found him guilty.
Type 2 Error: A defendant committed the crime, but the jury find him not guilty.
Medical System
Null Hypothesis: A person does not have the disease.
Alternative Hypothesis: A person does have the disease.
Type 1 Error: A person tested positive for the disease that he/she does not have.
Type 2 Error: A person tested negative for the disease that he/she does have.
Drug Testing
Null Hypothesis: A drug has no effect on the disease.
Alternative Hypothesis: A drug has effect on the disease.
Type 1 Error: A drug is falsely claimed to have an effect on the disease.
Type 2 Error: A drug is falsely claimed to have no effect on the disease.
Consequences of Each Error Type:
Type 1 error (false positive) in medical testing: Patient tests positive but doesn't have the disease, may cause anxiety, but further tests will reveal the error.
Type 2 error (false negative) in medical testing: Patient tests negative but has the disease, leads to a false sense of assurance, and the disease goes untreated.
Type 1 error in criminal trial: The defendant is innocent but is found guilty.
Type 2 error in criminal trial: The defendant is guilty but is found not guilty.
Type 1 error will result in a defendant found guilty of a murder that he did not commit and could be sentenced to death.
A Type 2 error occurs when the defendant committed a murder, but he was found not guilty, is bad outcome for society as a whole. However, should more evidence resurface later there is an opportunity to correctly convict the defendant
Cost of Errors and Significance Level
If the cost of committing Type 1 error is expected to be higher, then we should demand a stringent significance level – certainly 1%. The case of medical sciences (eg. surgery).
If the cost of committing Type 2 error is expected to be higher, then we should demand a relax significance level – somewhere 10%.
Type 1 and Type 2 error probabilities are inversely related.
Reducing one increases the other.
Accepting weak evidence of guilt risks punishing innocent people.
Ignoring strong evidence allows guilty people to go unpunished.
Critical Value
Separates the area of rejection (shaded area) from the area of acceptance (non-shaded area).
The area of rejection is also known as the critical region.
Provides the point on the scale of the test statistic beyond which we reject the null hypothesis.
Derived from the significance level (alpha).
For left-tail test, the critical region lies on the left (negative values).
For right-tail test, the critical region lies on the right (positive values).
For two-tail test, the critical region lies on both sides (positive and negative values).
Test Value
A test value is a standardized value that is calculated from a sample data. It involves the use of statistical formula to obtain the test value. This is also known as a test statistic.
Provides a numerical summary of a dataset, reducing it to one value that can be used to perform the hypothesis test.
Quantifies the behavior of the observed data to help distinguish the null hypothesis from the alternative or research hypothesis.
The test value is the criterion upon which we base our decision about the hypothesis. In a criminal justice analogy, the test value refers to the evidence presented in a court case.
If the test value is close enough to the null hypothesis (consistent), then we accept the null hypothesis and conclude that there is no significant difference (i.e., null hypothesis is true).
If the test value is far different from the null hypothesis (inconsistent), then we reject the null hypothesis and conclude that there is a significant difference (i.e., the alternative hypothesis is true).
Significance Level
Type one error is also known as significance level or alpha.
In a hypothesis, the null hypothesis is tested against the probability of committing type one error.
A significance level is a theoretical limit of error that one can allow in the hypothesis test. This limit must be established prior to the hypothesis being tested.
The significance level refers to a pre-established maximum probability of error allowed in your decision to reject the null hypothesis.
The complement of significance level is called the confidence level.
The confidence level is the probability that the null hypothesis is true.
Examples:
Significance Level = 5%, Confidence Level = 95%
Significance Level = 1%, Confidence Level = 99%
Significance Level = 10%, Confidence Level = 90%
P-Value
The primary interest of a researcher is to reject the null hypothesis.
The decision to reject the null hypothesis is usually the desired outcome for the researcher.
We want to find support for our research hypothesis by showing that there is no support for the null hypothesis.
The goal of the p-value is to answer the following question: What is the likelihood of committing an error should the null hypothesis be rejected?
The p-value quantifies the strength of the evidence against the null hypothesis in favor of the alternative or research hypothesis.
The p-value, or calculated probability, is the probability of finding the observed, or more extreme, results when the null hypothesis () of a study question is true. For instance, a p-value of 2% means we have 2% chance of being wrong when we reject the null hypothesis.
Making Decision
In hypothesis testing, there are two ways to determine whether there is enough evidence from the sample to reject or accept the null hypothesis:
critical value approach: compare the critical value and the test value.
p-value approach: compare the p-value and the significance level.
Critical Value Approach
Involves determining whether an observed test value is more extreme than would be expected if the null hypothesis were true.
A critical value is a point on the test distribution that is compared to the test value to determine whether to reject or accept the null hypothesis.
This approach entails comparing the observed test value to the cut-off value, called the critical value.
If the test value is more extreme than the critical value, then the null hypothesis is rejected in favor of the alternative hypothesis.
If the test value is not as extreme as the critical value, then the null hypothesis is accepted.
In simple terms, if the absolute test value is greater than the absolute critical value, you can declare statistical significance and reject the null hypothesis. However, if the absolute test value is less than the absolute critical value, the difference is not significant, and we accept the null hypothesis.
Analogy - Speed Limit
Posted speed limit is 100 km/h.
Normal driving is expected to conform to this speed limit.
Allowing for some degree of flexibility (standard error), a police officer can tolerate speeds up to 120 km/h.
120 km/h is the critical point that will guide the officer’s decision whether or not to prosecute the motorist for over speeding.
A speed of 102 km/h will be considered not significantly different from the posted speed limit of 100 km/h.
The officer will accept the speed of 102 km as this will present a weak evidence in court and the motorist will likely be found not guilty of over speeding (the null hypothesis is accepted).
However, a speed of 145 km/h is significantly more than posted speed limit of 100 km/h.
This presents a sufficient evidence in court to found the driver the guilty of over- speeding (the null hypothesis is rejected).
Null Hypothesis: The speed is not significantly different from the posted limit of 100 km/hr (driver is not guilty).
Alternative Hypothesis: The speed is significantly more than the posted speeding limit of 100km/hr (driver is guilty).
A speed of 102 km/h will be accepted and a speed of 145 km/h will be rejected.
Comparing Test Value and Critical Value
Left-Tail Test: If the test value (-2.80) is less than the critical value (-1.65) then reject the null hypothesis. However, if the test value (-1.24) is greater than the critical value (-1.65) then accept the null hypothesis.
Right-Tail Test: If the test value (2.80) is greater than the critical value (1.65) then reject the null hypothesis. However, if the test value (1.24) is less than the critical value (1.65) then accept the null hypothesis.
Two-Tail Test: If the test value (- 1.40) lies within the lower critical value (-1.65) and upper critical value (1.65) then accept the null hypothesis. If the test value (-2.80) is less than the lower critical value (-1.40) then reject the null hypothesis. If the test value (2.80) is greater than the upper critical value (1.65) then reject the null hypothesis.
Comparing Significance Level and P-Value
Significance level is determined before the test, while the p-value is calculated after the test.
The significance level is the allowable maximum error you are willing to commit prior to the test, while the p-value is the actual calculated error you have committed after the test is conducted.
Scenario 1: The significance level is 5%. After the test, you calculated the actual error (p-value) to be 8%. Since the p-value (8%) is more than the maximum allowable limit of 5% (significance level), then accept the null hypothesis.
Scenario 2: The significance level is 5%. After the test, you calculated the actual error (p-value) to be 2%. Since the p-value (2%) is less than (or falls within) the maximum allowable limit of 5% (significance level), then reject the null hypothesis.
If the p-value is less than the significance level, then reject the null hypothesis.
If the p-value is greater than the significance level, then accept the null hypothesis.
When You Have Only The P-Value
The smaller the p-value, the stronger the evidence against the null hypothesis and the higher chance of rejecting the null hypothesis.
Without pre-established threshold (significance level), the p-value can be interpreted in the following way when making inferential decision:
If the p-value is between to , then reject the null hypothesis.
If the p-value is , , or , then accept the null hypothesis.
Seven Steps in Hypothesis Testing
State the Hypotheses
Select the Method
Determine the Tail Test
Establish the Critical Value
Calculate the Test Value
Make the Decision
Interpret the Error (P-Value)
Details of Each Step
State the null and research hypotheses
State the null hypothesis of “no significant difference” and the alternative hypothesis of “significance difference”.
Choose the appropriate statistical method
Based on the sample size and knowledge of population standard deviation, select either z-test or t-test
Determine the tail test
Depending on the direction of the alternative hypothesis, determine the type of tail test – left tail, right tail or two tail test.
Establish the critical value
Use the significance level and tail test to find the critical value from the statistical table.
Calculate the test value
Use a formula to calculate a single, standardized value for the test. This requires the knowledge of the mean, standard deviation and sample size.
Make the decision
Compare the test value with the critical value. Reject the null hypothesis if the test value falls in the “reject” area. Accept the null hypothesis if the test value falls in the “accept” area.
Find p-value and interpret the error
Find the p-value and interpret the error for rejecting the null hypothesis. Reject the null hypothesis if the p-value is less than the significance level. Accept the null hypothesis if the p-value is greater than the significance level. 1 2 3 4 5 6 7