1/63
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Overall Steps of the Scientific Method
ID the research Questions
Conduct background research
Form a hypothesis
Experiment/collect data
Explore, summarize and analyze data
Make conclusions
Population
The entire set of individuals in which we are interested (another way to frame it - what are you interested in studying)
Denoted N
Census
Recording information about all individuals in a population
What are the main issues with using a census
Populations change and it takes a lot of time and $ - usually not feasible
Sample
A subset of the population from which we collect information
Sample Size
-denoted n
-the number of people in the sample
Variable
The characteristic of the individuals that we want to study
Parameter
A summary of a variable for the entire population
Statistics
A summary of a variable for a sample
Statistical inference
Process of using sample information to make conclusions about the population.
Not necessary for a full census because then we’d have the parameter directly
Biased Samples
A sample that is more likely to produce some outcomes more than others.
This is because the sample is consistently too low or too high - causing innacurate parameter
Convenience sample
Samples that are easy to take
Volunteer response sample
Self selected sample of people who respond to a general appeal
Why could a convenience sample be bad?
Don’t represent the population well
Often biased
Why could a volunteer response sample be bad?
Those who volunteer may be different from the general population
Tend to get people who feel strongly about topic
Or could get mismatch between target population and who actually sees/answers the survey
How do we avoid biased samples and why?
We use a probability sample (also known as a random sample)
This avoids bias in the process of selecting participants because observations are independent and identically distributed
What does it mean for a value to be independent?
Independent means the value isn’t affected by other variables
What does it mean for a value to be identically distributed?
Means all values follow the same pattern. Also means assuming parameter is same for everyone
Simple Random Sample (SRS)
Sample taken in such a way that every set of n units has equal chance of being chosen
Sampling Frame
List of every single unit in the population
How does the SRS work?
First compile a list of every single unit in the population (sampling frame)
Then select units from the sampling frame using a random process
The sample is made of the units that were randomly selected
What is stratified sampling?
When population is divided into groups (strata) based on a characteristic (gender, age, etc) and random samples are taken from those groups.
How does stratified sampling work?
Compile sampling frame for each stratum in population
Take an SRS of units from each sampling frame/strata
Example: Can sample uni students by selected 5 grad students and 5 undergrad students
What is cluster sampling?
Population divided into groups (clusters) and a random sample of clusters is taken. The sample is composed of every subject in each selected cluster
How does cluster sampling work?
Take a sampling frame of all clusters in the population (for example a cluster could be a district)
Take an SRS of clusters
Sample is everyone in each of the selected clusters
Example: Randomly selecting 10 high schools and talking to all students within each school
Difference between Strata and Clusters
Strata: Units are similar within group and different between groups
“some from all”
Clusters: Units are different within group and similar between groups
“all from some”
Selection bias
Only a particular subset of people is selected in the sample
How to we limit selection bias?
Use a probability sampling method
Undercoverage
Type of selection bias where sampling frame doesn’t include all of the population
How do we limit under coverage?
By using the most current and complete sampling frame possible
What are errors due to the sampling process?
Selection Bias
Undercoverage
What are errors not due to the sampling process (but can still cause problems?)
Data entry/processing errors
Non response bias
Response bias
Non response bias
Some part of population may not respond or refuses to participate
Why might non response bias be a problem?
The people who don’t respond could differ in an important way from the people who do respond - messing up any inferences from the study
What are some ways to limit non response bias?
Contact multiple times in multiple ways
Incentive for participation like a gift card
Assure anonymity or confidentiality
Response bias
Responses given are not an accurate reflection of the truth
How could we limit response bias?
Randomize order of questions
Test wording of questions (do participants understand what is being asked)
Assure anonymity/confidentiality
Specialized survey techniques for sensitive topics
Observational Study
A study that doesn’t formally assign people to variables. Instead it involves a researcher observing differences and making conclusions from that.
Subject chooses for themselves the group they want to be in
Ex: A researcher following people around the grocery store to see if people with baskets buy more or people with grocery carts
Lurking variables
Variables that may influence the response but are often not studied explicitly
Observational studies are vulnerable to them
Response Variable
Also known as dependent variable
Denoted ‘y’
Measures outcome of interest in the study
Explanatory Variable
Also known as independent variable
Denoted ‘x’
Variable that may cause changes in y
Treatment
Specific regimen or procedure assigned to subjects: different levels of the explanatory variable y
Units
Also known as subjects or participants
The units whose data we record
Experiment
When we impose a difference in the explanatory variables to see if there is a difference in the outcome
What is Random Assignment
When subjects are assigned to different levels of the explanatory variable by a random mechanism.
In simpler terms, when subjects randomly get their treatment
How does Random Assignment help avoid bias
Equally distributes lurking variables
What is replication
Having more than 1 subject in each experimental group
How does replication help against bias?
Helps determine whether treatment effect random or not
What is a control?
Absence of treatment or a baseline/standard of care used for comparison
How does a control protect against bias?
It determines if the treatment actually does anything by comparing it to the current standard of care
Placebo
Something given to units in the control group
Anything that seems like a treatment but lacks the active ingredient (example: sugar pill)
Placebo Effect
A person’s tendency to react to a treatment regardless of what the treatment actually does
Looks at effect of getting any treatment at all
Occurs in both treatment and control group
Difference between random sampling and random assignment
Random sampling produces sample to represent the population
Random assignment balances lurking variables so any differences between groups can be attributed to the treatment
What is the Hawthorne Effect?
People act differently when they know they are in an experiment
What is the best way to limit Hawthorne effect?
Through blinding
It helps avoids any bias due to subject and/or researcher beliefs or preferences
Single Blinded Study
One party involved in the experiment (either subjects or experimenter) don’t know which group a subject is assigned to
Double Blinded Study
Neither the subjects nor the experimenters know which group the subjects are assigned to
Completely Randomized Design (CRD)
Each unit randomly assigned to receive exactly one level of explanatory variable (without taking other variables into consideration)
Simplest to design
Can be hard to arrange
Example: Get 40 cars and randomly assign 20 to have new tires (other 20 keep old tires)
Matched Pairs Design
Each “unit” gets each level of the treatment
Means either each person measured multiple times (subject serves as their control)
or
Multiple units are matched together (one receives treatment the other gets control)
Example: A car gets two of its tires randomly assigned to be new
Block Design
Units divided into similar groups called blocks
Each level of explanatory variable applied in each block
Example: Take multiple vehicle types. 5 cars from each type are randomly assigned new tires.
What are some ways to reduce response variation
Control
Blocked Design
Matched Pairs Design
What are some ways to reduce potential for bias
Random Assignment
Blinding
Placebos
Why would we want to reduce response variation?
Allows us to increase ability to detect what the treatment effect was (which effects were actually from the treatment)
Why would we want to reduce bias potential?
Increases our ability to say any effect is due to the treatment and not other variables