1/64
Sampling and Data- week 1: September 9th, 2026
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Why do we use a sample instead of studying the entire population?
too time-consuming
expensive
impractical
What is a parameter?
A numerical characteristic of the entire population.
Calculated the average GPA of every student in the population:
Their average GPA = parameter
Parameter → Population
What is qualitative data?
Data consisting of labels or names rather than numerical measurements.
What is quantitative data?
Data consisting of numerical values.
Age, height, income, or number of children.
What is discrete data?
Quantitative data that can be counted and usually take separate, distinct values.
Ex: Number of students in a class.
What is continuous data?
Quantitative data that can be measured and can take any value within a range.
Ex: A person's height or weight.
What is stratified sampling?
Divide the population into groups called strata.
Take a proportion from each group.
Use random sampling within each group.
What is cluster sampling?
Divide the population into clusters
randomly select some entier clusters
Include all members of the selected clusters.
What is systematic sampling?
Start at a random point and select every Kth individual.
What is sampling error?
Errors caused by the sampling process itself.
Example:
The sample isn't large enough.
What is nonsampling error?
Errors caused by factors not related to the sampling process.
Examples:
Defective counting device
Data-entry errors
Flawed/poorly worded survey questions
Nonresponse/refusal
What is sampling bias?
Some members of the population are more likely to be selected than others.
What is cumulative relative frequency?
Total of the relative frequencies up to a particular value.
What is a proportion/How do you calculate it?
Proportion = number of successes ÷ total number of observations.
What is variation in data?
The fact that individual observations can differ from one another.
Why can data vary?
Different:
individuals
measurements
amounts
methods
accuracy
Why can two random samples give different results?
Random sampling produces natural variation from sample to sample.
What is nominal measurement?
Categories or labels with no meaningful order.
Ex: Eye colour or type of car.
What is ordinal measurement?
Categories that CAN be ordered/ranked.
But the differences between rankings cannot be measured.
Ex: Class ranking or satisfaction level.
What is interval measurement?
Numerical data with meaningful differences between values
no true zero
Ex: Temperature measured in Celsius.
What is ratio measurement?
Numerical data with meaningful differences and a true zero.
Ex: Height, weight, or income.
What is an explanatory variable?
The variable researchers control or manipulate that might cause or explain a change in another variable.
Example: Type of fertilizer → might cause a change in plant growth.
What is a response variable?
Variable being measured to determine whether it changes in response to the explanatory variable.
Example: Whether or not a person has a heart attack when testing aspirin.
What are treatments?
The different values of the explanatory variable
Ex: Aspirin vs. placebo
What is random assignment?
Researchers randomly assign experimental units to different treatment groups.
Why is random assignment important?
Spread lurking variables among the treatment groups
allows researchers to better determine cause-and-effect
What is a lurking variable?
An additional variable that can interfere with the relationship being studied.
Ex: Researchers notice that people who regularly take vitamin E have better health.
Does that prove vitamin E causes better health?
No.
People who take vitamin E may also:
Exercise
Eat healthier
Take other supplements
Not smoke
What is a control group?
Group in a randomized experiment that receives an inactive treatment
is otherwise managed in the same way as the treatment group
What is a placebo?
Inactive treatment that cannot directly affect the response
used to account for the power of suggestion
What is causality?
One variable causes a change in another variable
What is confounding?
Effects of multiple factors cannot be separated
difficult to determine which variable is actually causing the observed result
Ex: Studying more is related to higher grades, but sleep or tutoring could also affect grades, so we can't tell if studying itself caused the higher grades.
What does N represent in sampling formulas?
The population size — the total number of individuals in the population.
What is the formula for k in systematic sampling?
k = N ÷ n
What does k represent in systematic sampling?
The sampling interval — how many individuals you skip/select between selections.
Ex: every 5th person
What are the five main sampling methods?
Simple random sampling
Stratified sampling
Cluster sampling
Systematic sampling
Convenience sampling
What is the formula for relative frequency?
RF = f ÷ n
f = frequency
n = total number of observations