1/29
Core EPPP review cards covering reliability, validity, standard scores, measurement error, screening accuracy, base rates, norms, and test bias. Includes direct recall and applied discrimination items.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is reliability?
The consistency or dependability of scores produced by a measurement instrument.
What does test–retest reliability assess?
The stability of test scores across two administrations over time.
What does interrater reliability assess?
The degree to which two or more raters or observers agree in their ratings.
What does internal consistency assess?
How consistently items on the same test measure the same construct.
What statistic is commonly used to estimate internal consistency?
Cronbach’s alpha; higher values generally indicate greater consistency among the test items.
What is validity?
The degree to which evidence and theory support the intended interpretation and use of test scores.
What is content validity?
The extent to which a test adequately samples the full domain of content or behaviour it is intended to measure.
What is construct validity?
The extent to which a test actually measures the theoretical construct it claims to measure.
What is convergent validity?
Evidence that a measure correlates strongly with other measures of the same or closely related constructs.
What is discriminant validity?
Evidence that a measure has weak relationships with measures of distinct or unrelated constructs.
What is criterion-related validity?
The extent to which test scores relate to a relevant external criterion, standard, or outcome.
What is concurrent validity?
Criterion validity established when the test and criterion are measured at approximately the same time.
What is predictive validity?
Criterion validity established when test scores accurately predict a future criterion or outcome.
Can a test be reliable without being valid?
Yes. A test can consistently produce the same result while measuring the wrong construct. Reliability is necessary but not sufficient for validity.
What is the standard error of measurement (SEM)?
An estimate of how much an observed score is expected to fluctuate because of measurement error.
How does reliability affect the SEM?
As reliability increases, SEM decreases; more reliable tests have less measurement error.
How is a confidence interval around an observed test score interpreted?
It gives a range within which the person’s true score is likely to fall, based on the SEM and the selected confidence level.
What does a z-score indicate?
How many standard deviations a score lies above or below the mean; z scores have a mean of 0 and standard deviation of 1.
What does a T-score indicate?
A standardized score with a mean of 50 and standard deviation of 10; T = 50 + 10z.
What is sensitivity?
The proportion of people who truly have a condition who test positive: the true-positive rate.
What is specificity?
The proportion of people who truly do not have a condition who test negative: the true-negative rate.
A highly sensitive test minimizes which kind of error?
False negatives. High sensitivity means few true cases are missed.
A highly specific test minimizes which kind of error?
False positives. High specificity means few non-cases are incorrectly identified.
How do low base rates affect positive test results?
When a condition is rare, false positives can comprise a large proportion of all positive results, lowering positive predictive value.
What is norm-referenced interpretation?
Interpreting a person’s score by comparing it with the scores of an appropriate reference or normative group.
What is test bias?
Systematic error that causes test scores to have different meanings, validity, or predictive accuracy across groups.
Applied: A depression screener identifies nearly everyone who actually has depression but also flags many people without it. What is its likely profile?
High sensitivity but low specificity.
Applied: An employment test is administered now and correlated with job performance one year later. Which validity evidence is being assessed?
Predictive criterion-related validity.
Applied: A new anxiety scale correlates strongly with established anxiety scales and weakly with a measure of physical strength. What two forms of validity does this support?
Convergent validity and discriminant validity.
Applied: Two clinicians independently rate the same interviews. Which reliability estimate is most relevant?
Interrater reliability.