1/13
test theory
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
construct validity
the extent to which a test, survey, or measurement tool accurately measures a theoretical, unobservable concept—known as a construct—rather than an unrelated or extraneous variable
convergent validity: Shows that your measure strongly correlates with other tests that measure the same or a conceptually similar concept.
divergent validity: Shows that your measure does not correlate with unrelated or distinct concepts, proving it stays focused only on the target trait
content validity
the degree to which a test or measurement tool fully represents all important parts of the concept or domain it aims to measure
face validity
he subjective, surface-level judgment of whether a test, survey, or research tool appears to measure what it claims to measure upon first inspection
weakest form of validity
criterion validity
How well predicts test score X criterion Y?
We need: test scores and criterion data in a representative sample
Determine relation between test scores and criterion
Correlation test score and criterion score = validity coefficient r(X,Y)
Multiple predictors: R (multiple correlation) = validity coefficient
predictive validity
to what extent are the predictions confirmed by the criterian data obtained in the future?
e.g. how well does the SAT predict later study behavior in college
concurrent validity
to what extent is the agreement between the test results and criterion data obtained at the same time
How strongly are the scores on a new depression inventory related to the scores on the Beck Depression Inventory, administered at the same time
criterion related validity in practice
often not larger than r = .60 (r = correlation)
explained varianace R2 = .36
does not seem much, but we can explain sometimes a reasonable amount of the explained variance with on or a small number of test scores
rules of thumb r
.10 = small
.30 = moderate
.50 = large
reasons for low validity coefficients
low reliability criterion
underestimation predictive validity
to assume a linear relation
may not be linear: underestimation predictive validity
consider the relation, more is not always better
range restriction
there are only criterion scores from the selected group: underestimation validity
often encountered problem when determining the predictive validity in selection
selection: only persons with a high score on the predictors are selected
effect: there are only criterion scores Y available for the highest scoring candidates
reduced spread in X
underestimation of the predictive validity
we can correct this with stats
incremental validity
what is the additional value of a test on top of existing information?
Test with relative low correlation with criterion can sometimes add important information on top of other predictors: when the relation with existing predictors is low
If they are highly correlated, they don’t add much new information
Correlation should be no higher than .4 – .5
utility of a test
Practical importance of a test depends on the quality of the decision made on the basis of the test. To determine what the test adds to the decision: Compare the decisions made with and without the test
E.g. How do students do when selecting them randomly instead of using SAT
Contribution of a test is not equal to predictive validity
base rate
natural occurring frequency of a condition
proportion of applicants that is suited for the job within the group of applicants
proportions of persons with a depression in the population
selection ratio
proportion applicants that is hired
proportion of persons that get a diagnosis of depression
success ratio
the aim is to optimize the success ratio
proportion of persons that is succesful
proportion of diagnosed persons that really have a depression