1/59
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
steps to conducting a statistical analysis
establish the research question
formulate a hypothesis
select an appropriate statistical test
sample correctly
collect data
perform statistical test(s)
make decisions
strong research is done
intentionally
what is a priori?
concept that means your research question was pre-determined
what is a fishing expedition?
skipping steps 1-4 of conducting a statistical analysis, essentially playing around with data to see what you get
secondary data analysis
steps 4 and 5 of the statistical analysis are conducted first and the data is collected for a non-research purpose (steps 1-3 must be specified before 6)
discrete data
set values in a particular range (whole number, category that a measurement fully belongs into)
nominal data
type of discrete data that has no meaningful order (differences in categories are not incremental)
ordinal data
type of discrete data that has a natural order, but the difference between categories is not necessarily incremental or equal
continuous data
can be any value in a particular range (including decimals)
interval data
type of continuous data that has a natural order and the difference between each unit is the same, has a zero value, but it does not mean the absence of that factor
ratio data
type of continuous data that has a natural order and the difference between each unit is the same, but a zero means the absence of that value
eye colour is an example of ____ data
nominal
comparing BSc vs MSc is a type of ____ data
ordinal
a measurement in degrees Celsius is a type of ____ data
interval
a measurement of a patient’s weight is a type of ____ data
ratio
hierarchy of data
continuous>ordinal>nominal
can you convert continuous data to categorical? vice-versa?
yes, no
measure of ‘spread’ for mean
standard deviation
measure of ‘spread’ for median
interquartile range (difference between 75th and 25th percentile values)
measure of ‘spread’ for mode
range
what is most commonly used: mean, median, or mode?
mean
what is preferred if measurements are skewed?
median
point estimate
result calculated from dataset
confidence interval
describes uncertainty around the point estimate
is a narrower or wider CI range more certain?
narrower
what factors influence the width of a CI?
size of CI calculated. larger % = wider range of values that fall within
variability around point estimate observed. CI based on point estimate and SD so a greater variability = higher SD = wider range
sample size. larger = more certainty
null hypothesis
independent variable has no effect on the dependent variable (or if you are trying to see if two options are equivalent, it is that there is a different between two options)
alternative hypothesis
independent variable has a significant effect on the dependent variable
p value
probability that the null hypothesis is correct (observed effect is due to chance alone)
factors that influence the p-value
effect size: larger difference = less likely it is due to chance
variability around point estimate observed in sample: more variability = less certainty = higher p-value
sample size: more subjects = more certainty = lower p-value
how do you determine p-value if you are given a point estimate with a CI?
determine null value
measuring a difference between groups = 0 (number subtracted from itself)
ratio or proportion = 1 (number divided by itself is 1)
if 95% CI crosses null value = p-value is greater than 0.5 (not significant)
if 95% CI does not cross null value = p-value is less than 0.5 (significant)
type I statistical error
occurs when you reject H0 and it is true (false positive)
type II statistical error
occurs when you accept H0 and it is false (false-negative)
probability of a type I error
denoted by alpha and set to 5% or 0.05 (p-value that is set)
probability of a type II error
denoted by beta and typically set to 20%
statistical power
likelihood that a study will detect an effect if there is one (also the probability of rejecting H0 when it is false, ie probability of being right?)
how do you calculate statistical power?
1-beta (beta is usually 20%, therefore power is 80%)
factors that influence statistical power
sample size: larger = higher power
standard deviation: smaller = higher power
effect size: greater effect size (difference between groups) = higher power
significance level (alpha): lower alpha = lower power (more evidence needed to reject null hypothesis
one or two-tailed test: one-tailed = higher power, two-tailed = lower power (need a more extreme p-value to fall into shaded area)
two-tailed test
unsure of which side HA falls on so you use 2.5% on each side of the distribution
one-tailed test
you know HA falls on one side so you use 5% on each side of the distribution (be cautious of tests that use this)
how is sample size of subjects calculated?
based on change expected to see in primary outcome
primary outcome
answers the primary/most important question
secondary outcome
other relevant outcomes from a study, but less important, usually hypothesis-generating only
composite outcome
pooling outcomes together
what type of test would you use if you had a categorical independent variable and a categorial dependent variable?
Chi-square (x2)
what type of test would you use if you had a categorical independent variable and a continuous dependent variable?
t-test, ANOVA
what type of test would you use if you had a continuous independent variable and a continuous dependent variable?
regression
paired t-test
both sets of measurements come from the same subject
unpaired t-test
outcome measures across groups come from different subjects
when to use ANOVA over t-test?
if you have more than 2 independent variables
why is doing multiple t-tests problematic?
increases chance of type I error
efficacy
capacity of a treatment to produce the desired effect in a controlled environment
effectiveness
actual effect of treatment in real life
accuracy
how close the measured value is to the true value
precision/reliability
variability across measurements
validity
how well the variable assesses outcome of interest
sensitivity
can a test correctly identify those with a disease (sensitive = low rate of false negative - type II error)
specificity
can a test identify those without a disease (specific = low rate false positive - type I error)
statistical significance
difference between 2 interventions resulting in p<0.05
clinical significance
difference between two interventions that is meaningful to a patient and their health outcomes