1/18
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
reliability
"shows how much noise is in a psychological measurement; theoretical feature of the test scores, not of the test
two ways to think about reliability:
Classical Test Theory
"most widely used theory in Psychological tests.
four models
"can't calculate reliability since true scores are unknown
All methods to determine test reliability are based on the idea of multiple tests:
Parallel test models
"identical true scores, identical measurement error variance
model for true score: Xt2=Xt1
model for the observed score of test 1: Xo1 = Xt1+Xe1 and
model for the observed score of test 1: Xo2 = Xt1+Xe2
Implications:
Tau-equivalent test model
"model for the true score: Xt2=Xt1
model for the observed score of test 1: Xo1=Xt1+Xe1
model for the observed score of test 2: Xo2=Xt1+Xe2
model for the observed scores: Xo1=Xo2
Implications:
Essentially tau-equivalent test model
"
Implications:
Congeneric test model
"Least restrictive model.
Implications:
three methods of reliability estimation
a. split half
Depends highly on split used (undesirable), still frequently used
COTAN guidelines about reliability
item discrimination
Item total correlation = correlation between item scores and sum scores
factors affecting reliability
"
estimating true scores
"1. True score estimate = summed item score
attenuation
"since we use observed scores rather than true scores
-> effect sizes/correlations will be smaller than effect sizes of the true scores.
We should always interpret effect size/significance in light of test reliability!
Cohen's d (number of st. dev. that the groups differ) is smaller for less reliable tests.
corrections for attenuation
some people use it, but it's better to just focus on using reliable tests.
index of reliability
= unsquared correlation between observed scores and true scores (rot)
standard error of measurement (sem)
"= st. dev. of error scores; average size of error scores
shared assumptions across all four models