1/73
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
scientific claims
Data drive/objective
Verifiable
Public
nonscientific claims
Often anecdotal
Lack evidence
Not falsifiable/replicable
empirical research article structure
Introduction
Method
Participants
Materials
Procedure
Results
Discussion
Philosophy of empiricism
Causal chain of events
Things do not just happen randomly out of nowhere
I.e. when you drop something, you know it will fall to the floor and not just teleport randomly
Causes of events are knowable
Empirical method
Investigating via direct observation
Testing causes with experiments
scientific method
observation
background research
formulate hypothesis
experiment
analyze results
report conclusions
observation
Most research beings with observation
Cats sometimes use only 1 paw when playing
Observation used to construct research question
Do cats have handedness preferences? Do they prefer to use one paw over the other for different tasks
background research
Literature search for relevant research
Has anyone researched this before?
Fun fact: male cats are fairly ambidextrous & female cats prefer their right paw
From the literature review, you can decide to conduct your own experiment
formulate hypothesis
Female cats will show a preference for using their right paw over their left paw
Our prediction
experiment
Create study to test hypothesis
Convince cats to hit stick
Record which paw used
analyze results
Conduct statistical tests to determine whether evidence supports hypothesis
This course: understand the results of statistical tests
report
Interpret analyses
Write up results for dissemination
This course: summarize findings of empirical research
hypothesis
falsifiable prediction of results of experiment
You cannot make ALL or NOTHING statements because they aren’t really falsifiable → you cannot measure all people that’s literally impossible
Derived from theory/observation
Outlines that provide a basic framework of what we should expect people to do
Provides us pieces that we can break apart and test
null hypothesis
predicts no change, no difference, no relationship in population
The “no” hypothesis
Notation: H0
Default assumption
The boring hypothesis
Eating sugar is unrelated to hyperactivity in children
alternative hypothesis
predicts change, difference, or relationship in population
Experimental hypothesis
Notation: HA
Stated hypothesis in research articles
Eating sugar is related to hyperactivity in children
correct reject
your sample data provides strong enough statistical evidence to conclude that the default assumption (the null hypothesis) is false
The effect genuinely exists in reality, and the study accurately identified and proved it.
correct retention
your statistical test did not find enough strong evidence to reject the default assumption that there is no effect, difference, or relationship between variables
The study found no effect on hyperactivity and sugar because no real effect exists.
type i error
rejecting null when null is true
Conclusion: change/difference/relationship exists
Reality: no change/difference/relationship exists
False positive
The researchers rejected null hypothesis and claimed an effect exists when it doesn't—perhaps because birthday party excitement or confounding variables skewed their data
type ii error
failing to reject H0 when H0 is false
Conclusion: no change/difference/relationship exists
Reality: change/difference/relationship exists
False negative
The researchers retained null hypothesis when it was actually false—perhaps because their sample size was too small or the sugar dose tested wasn't high enough to detect the real effect.
variable
takes on range of values to be measured
some can be measured directly
construct
internal characteristic, cannot be directly measured
operational definition
how variable is defined in order to be measured
independent variable (sugar intake)
a child drinks an 8-ounce beverage containing 35 grams of cane sugar on an empty window within a 5-minute window
control group drinks an 8-ounce beverage with artificial sweetener
dependent variable (hyperactivity)
hyperactivity is measured through physical movement, the amount of times the child rotates in their desk more than 45 degrees
hyperactivity can also be measured through fidgeting (leg bouncing, finger fidgeting/tapping, object manipulating behaviors)
imagine you are testing whether sugar causes hyperactivity in kids. what would your operational definition be?
measurement scales
Nominal scales
Ordinal scales
Interval scales
Ratio scales
ordinal scales
Set of categories in ordered sequence
Used for ranking
Cannot quantify size of difference
If all we have is the place numbers, we don’t know how far apart the people in the race crossed the finish line
Examples:
Places in a race
Shirt sizes
interval scales
Ordered categories, fixed distance between scale points
Quantifies difference between observations
No/arbitrary zero
Most common in psychological research
Examples:
Temperature scales
Likert-type scales
ratio scales
Interval scale + meaningful zero
Quantifies differences
0 = absence of construct
Examples:
Time
Weight
ordinal
Which type of measurement scale would be most appropriate for each of the following constructs?
Ranking of highest-grossing movies
ratio
Which type of measurement scale would be most appropriate for each of the following constructs?
height
interval
Which type of measurement scale would be most appropriate for each of the following constructs?
numerical product ratings
nominal
Which type of measurement scale would be most appropriate for each of the following constructs?
favorite animal
construct validity
how well a measurement captures concept
Reliable, valid measured variables
Manipulation designed well
The researcher counts how many times a child raises their hand or speaks out during a 30-minute lesson.
High academic engagement and curiosity.
Extroversion.
Confusion about the instructions.
tool is measuring verbal participation or curiosity
what is an example of low construct validity on childhood hyperactivity
The researcher uses a wrist sensor that logs physical activity & heart rate
Captures the actual physical and behavioral components of hyperactivity
Children who score high on it should also score high on established clinical ADHD assessments and physiological arousal markers
what is an example of high construct validity on childhood hyperactivity
Reliability
Measure’s ability to detect differences (or lack thereof)
Common types:
Internal consistency
Interrater
Test-retest
Belmont Report guidelines
Respect for persons
Right to decide without coercion
Beneficence
Maximize benefits of research
Minimize risk to participants
Justice
Risks and benefits equally distributed
APA Ethical principles
Informed consent
Freedom from coercion
Protection from harm
Weight risks vs. benefits
Use of deception
Debriefing
Confidentiality
population
all people of interest
sample
subset of population measured
representativeness
how much sample reflects target population
Samples used to make inferences about unknown population
Representative samples more accurate
sampling strategies
Ways to collect research data
Independent random sampling
Stratified random sampling
Convenience sampling
independent random sampling
Ideal strategy
Random sampling: each population member has equal chance of selection
Independent: probability of selection stays constant
Ex: population is 1,000. sample only needs 50. each person is assigned a random number and a number generator picks 50 random numbers
stratified random sampling
Sampling from within predetermined subsets
Targets small demographic groups
Example: income stratification
Ex: population is high school students. seperate them by freshman, sophomore, junior senior. choose 50 of each demographic at random
convenience sampling
Sampling from readily available subset of population
Often used in psychological research
Snowball sampling: each participant asked to add another
Particularly bad for representativeness
A professor hands out a survey to students in their own 8:00 a.m. class to study college stress
sample size
Bigger is better
Minimum 30 participants per group for accuracy
Sufficient sample size varies by:
Study design
Research question
Subfield conventions
Research articles often include sample size justifications
validity
Appropriateness of claim/conclusion
Types:
Internal validity
External validity
Statistical validity
Construct validity
internal validity
How well a study rules out alternative explanations
Isolates process
Keep extraneous variables constant
Control within lab setting
external validity
How well study captures real world process
Generalizability: whether results apply to other contexts
Often tradeoffs with internal validity
statistical validity
How well statistics support conclusions
How accurate is the estimate?
How precise is it?
Is the effect meaningful?
Claims
Data used to draw different types of conclusions
Frequency claims
Association claims
Causal claims
frequency claims
Claims about how often something happens
Opinion polls
Less common in research
frequency claims: validity
External validity: was the population accurately represented?
Statistical validity: how precise is the estimate? Does it replicate?
Construct validity: how well was the construct measured?
association claims
Claims about whether variables are related
Describe whether/how variables change together
association claims: validity
External validity: do findings generalize?
Statistical validity: how strong is the association? how precise is the estimate?
Construct validity: how well were the variables operationalized?
causal claims
Claims that one variable directly influences another
Tested with experiments
causal claims: validity
Internal validity: have alternative explanations been ruled out?
External validity: do results apply to other situations?
Statistical validity: how large is the effect size? Does it replicate?
Construct validity: how well was the manipulation designed? How well were variables measured?
describing data
strategies:
graphs
measures of central tendenc
bar g
bar graph vs. histogram
bar graph
categoriacal (discrete groups)
x-axis: gaps between bars
y-axis, any measure: count, total sales, averages, etc.
flexible order of bars: can be sorted alphebetically, high-to-low, etc.
bar width arbitrary: width has no mathematical meaning
histogram
continuous/quantitative
no gaps: bars touch to show continuous data flow
x-axis: numerical intervals
y-axis: frequency (count or percentage)
order of bars: fixed (must strictly follow numeric order)
bar width: re
good graph construction
Simple, concise
Informative labels
Represent full range of data
Or include axis break
Consistent scaling
No 3D
descriptive statistics
Summarize data numerically
Condense large datasets to few numbers
Quickly see trends
frequencies
counts of category membership
Notation: n or N
percentages
proportions of dataset in specific category
Most often used for demographics
measures of central tendency
Center of distribution of scores
Typical/”average” value
Mean
Median
Mode
mean
Arithmetic average of dataset
Most common measure of central tendency
Notion: M or μ
median
Midpoint of distribution ordered smallest to largest
50% of distribution below median
No standardized notation
Not typically used in research to describe a sample
mode
Most frequently occurring value
No standardized notation
Least often used in research
measures of variability
Describe spread of scores in distribution
Generally smaller = better
Standard deviation
Variance
Standard error
standard deviation
Average distance between given score and mean
Describes how clustered scores are
Notation: SD, s, or σ
variance
standard deviation squared
Notation: s2, σ2
Not often reported by itself
Used in calculations for statistical tests
standard error
Standard deviation of distribution of sample means accounting for sample size
Describes how precise estimate of mean is
Notation: SE
Reported with some statistical tests