1/57
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is reliability?
The consistency of measurement
What does reliability ask?
Can i trust the score?
Higher reliability =
less random error, lower SEM
What is the reliability coefficient range?
0-1. closer to 1 = better reliability
Reliability vs validity
Reliability = consistency
Validity = accuracy
A test can be reliable without being valid
What is random error?
Error that occurs unpredictably
fatigue
illness
distraction
phone ringing
Random error lowers reliability
What is systematic error
Error that occurs consistently
dyslexia
ESL
vision problems
Systematic error may not reduce reliability
What does test-retest reliabilty ask?
Will i get the same score later?
What is test-retest reliability?
Same test + Same people + Different times
What error source does test-retest assess?
Time sampling error
Why is 2 weeks the gold standard?
Short enough to reduce true change. Long enough to reduce memory effects.
Why not test 1 hour later?
practice effects & memory effects
Why not test 2 years later?
Real change may occur
What question does parallel forms reliability ask?
Do form A and form B produce similar scores
Major assumption of parallel forms
Both forms should produce the same true score
Why use parallel forms?
to reduce practice effects
What coefficient is used for parallel forms?
Coefficient of equivalence
What error source is assessed when forms are adminstered immediately (parallel forms)
Content sampling error
What error source is assessed when forms are adminsitered 2 weeks apart (parallel forms)
Content sampling + time sampling error
High same-day correlation but low 2 week correlation suggests…
the test may be measuring a state instead of a trait
What is a state?
temporary condition
mood, stress, fatigue
What is a trait?
Stable characteristic
ex: trait anxiety, introversion
What is internal consistency
Whether items “hang together”
Main question of internal consistency
Do the items measure the same construct?
What happens if a depression scale includes a pizza item?
Internal consistency decreases. The item does not represent the construct
What psychometric property is being assessed by Cronbach Alpha?
Internal consistency reliability
How is split half reliability calculated ?
Split test into 2 halves and correlate scores
Main weakness of split half reliability?
Reliability depends on how the test is split
Why can speeded test inflate split-hald reliabilty
Both halves contain unanswered items when people run out of time. The clock creates the correlation
Problem with speeded test and corelation
correlation may reflect the time limit rather than the construct.
What does KR-20 measure
internal consistency reliability
When is KR-20 used
Dichotomous scoring
Right/wrong
Correct/incorrect
Examples of KR-20 tests
Spelling tests, math tests, achievement tests
Can KR-20 be used on a Likert scale
No. only dichotomous measures
Cronbach’s alppha measures….
internal consistency
Can an alpha be used on likert scales?
yes
Can alpha be used on dichotomous tests?
yes
High alpha means ….
items are highly related and hang together
Does high alpha = unidimensionality
No
AP history test example demonstrates
High alpha can occur even when multiple constructs are measured
Why can Cronbach’s alpha be misleading in a homogeneous sample?
Everyone responds similarly
Artifically inflating alpha
What type of sample gives a better estimate of reliability?
heterogeneous sample
Homogeneous sample =
people are very similar
Heterogeneous sample =
Wide range of scores
What is the spearman brown prochecy formula used for?
Predict reliability after changing test length
What happens to reliability when quality items are added?
Reliability generally increases
Why does adding items increase reliability
Random error has less influence on the total score
Why cant reliability increase forever?
Diminishing returns
Items become redundant
What does interrater reliability assess?
Agreement among judges
Main question of interrater reliability
Do the judges agree?
Main error source of interrater reliabilty?
Interrater differences
When would you use Kendall’s coefficient of concordance?
When judges rank people (beauty pageant)
When would you use Cohen’s Kappa?
When judges classify people into categories
Teacher and parent disagree on a CBCL. What reliability issue is involved
Interrater reliability
Split half assesses what?
content sampling error
KR-20 / Cronbach’s alpha assess what
Content heterogeneity and internal consistency
What is SEM
Standard error of measurement
Estimate of error around an observed score
What does SEM ask?
How much wiggle room exists around the score