PTPR 2: Classical Test Theory

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/18

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 3:08 PM on 9/14/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

19 Terms

1
New cards

reliability

"shows how much noise is in a psychological measurement; theoretical feature of the test scores, not of the test

two ways to think about reliability:

  • proportion of variance: proportion of true score variance from observed variance
so2=st2+se2
  • shared variance: true scored variance shared with observed score variance
squaring a correlation gives variance shared by the two variables
"

2
New cards

Classical Test Theory

"most widely used theory in Psychological tests.

  • central statistic: summed item score = sum of item scores on each subject (sum score/test score/score on the test)
  • central idea: every test taker has a true score on a test = score you'd get using a perfect measurement instrument Xt
    • due to measurement error (=other influences that cause random noise in observed score), observed score is not equal to the true score
  • two core assumptions:
  1. observed scores = true scores + measurement error: Xo = Xt + Xe
  2. measurement error is random.
  • error cancels itself out across participants
  • I<3 roeland
-> 
  • Xe = mean of the measurement error is equal to 0 (a nonzero mean would make the measurement error systematic)
  • rte = 0 ;correlation between true scores and error is equal to 0; because Xe is 0 for all true values
  • so2=st2+se2; observed scores variance = true scores variance + error variance
    • observed scores variance > true scores variance
"

3
New cards

four models

"can't calculate reliability since true scores are unknown
All methods to determine test reliability are based on the idea of multiple tests:

  • split-halves: two tests
  • test-retest: same test twice
  • Cronbach's alpha and Omega: each item is its own test
-> four models that give assumption to make calculating reliability possible (from most to least restrictive):

1. Parallel test model: true scores are equal, measurement error variance is equal.
2. Tau-equivalent test: true scores are equal, measurement error variance is different.
3. Essentially tau-equivalent test: assume that the true scores on the 2nd test are equal to true scores on the 1st + some number.
4. Congeneric test model: true scores on 2nd test = true scores on 1st x slope (b) + some number - allows for differences in variance between the two scores

"

4
New cards

Parallel test models

"identical true scores, identical measurement error variance
model for true score: Xt2=Xt1 
model for the observed score of test 1: Xo1 = Xt1+Xe1 and 
model for the observed score of test 1: Xo2 = Xt1+Xe2

Implications:

  • mean of true scores on both tests are equal
    • variance of true scores on both tests are equal
    • correlation between true scores on both tests is 1
  • mean of observed scores on both tests are equal
  • reliability of both tests is equal (correlation between true scores and observed scores on test 1 and 2)
  • variance of observed scores of both tests is equal 
  • reliability = correlation between two tests
test-retest and split-halves reliability based on this model

"

5
New cards

Tau-equivalent test model

"model for the true score: Xt2=Xt1 
model for the observed score of test 1: Xo1=Xt1+Xe1
model for the observed score of test 2: Xo2=Xt1+Xe2
model for the observed scores: Xo1=Xo2

Implications:

  • mean of true scores of both tests are equal
    • variance of true scores of both tests are equal
  • mean of observed scores on both tests are equal
    • observed variances don't have to be equal 
    • (it's not assumed that variances of error are equal)
  • correlation between true scores on both tests is 1model for the observed scores: Xo1=Xo2
  • reliability doesn't have to be equal
-> reliability is not correlation
"

6
New cards

Essentially tau-equivalent test model

"
Implications:

  • mean of true scores are different 
  • variance of true scores are equal  
  • correlation between true scores = 1
Cronbach's alpha is based on this
"

7
New cards

Congeneric test model

"Least restrictive model.

Implications:

  • mean of true scores are different
  • variance of true scores are different
  • correlation between true scores is 1
Omega is based on this model.
"

8
New cards

three methods of reliability estimation

  1. alternate forms
    Hardly feasible in practice, only in specific situations.
    • assumes parallel test model
    • two alternative versions of the same test
    • correlation between tests = reliability
    • hard to construct an almost identical test that is still different
    • carry over effects, practice effects

    2. test-retest
    Important to establish reliability for a new test (COTAN requirement), only sometimes used in research
    • assumes parallel test model
    • apply the exact same test twice
    • correlation = reliability
    • challenges: carry over effects, change in true score (e.g. mood questionnaire)
    3. internal consistency
    • assumes parallel or essential tau-equivalent test model
    • consider blocks of items as seperate test
    • formula will give reliability
    • challenge: carry over effects
9
New cards
  1. internal consistency

a. split half
Depends highly on split used (undesirable), still frequently used

  • assumes parallel test model
  • split test in two parts
  • formula gives reliability
2. cronbach's alpha (KR20 for binary items)
Very popular in research due to ease, assumption (essential-tau equivalence) is hardl met -> lower bound to the reliability - Cronbach's alpha underestimates reliability
  • assumes an essential tau-equivalent test model
  • each item is considered a separate part
  • formula gives reliability
3. omega (best reliability index)
  • assumes a congeneric test model (or stricter)
  • estimate true score variance using unidimensional factor analysis
  • reliability = true score variance/observed score variance

10
New cards

COTAN guidelines about reliability

COTAN reviews all published tests for quality (reliability, variability...)
  • test used for high-impact inferences at individual level: mistake might affect someone's life gravely
    • good: >0.9
    • sufficient: 0.8-0.9
    • insufficient: <0.8
  • test used for less impact inference at individual level: consequences for mistakes, but smaller
    • good: >0.8
    • sufficient: 0.7-0.8
    • insufficient: <0.7
  • test used at group level
    • good: >0.7
    • sufficient: 0.6-0.7
    • insufficient: <0.6
11
New cards

item discrimination

Item total correlation = correlation between item scores and sum scores

  • whether an item score can predict the total sum score
  • but you correlate item with itself; biased upwards, will never correlate 0 because it's in there
Corrected item total correlation = correlation betweem item scores and rest scores
  • whether an item score can predict sum score - itself
  • rest score = corrected total score, sum score without the item to be correlated

12
New cards

factors affecting reliability

"

  • test length: longer tests are more reliable
  • sample heterogeneity: in homogenous samples, reliability will be smaller than heterogenous samples
    • in homogenous samples, st2 is smaller, people are relatively similar
    • in heterogenous samples, st2 is larger, people are relatively dissimilar
Homogeneity is undesirable, reliability should be a property of the test, not of the sample.
  • correlation between pretest and posttest scores 
Popular in psychology, test an improvement after an intervention.
    • If correlation between pretest and posttest is large -> reliability is small
    • difference reliability depends on reliability of the pretest and posttest
    • Rd sensitive to difference in variance between Xi and Yi (not a big issue in pretest-posttest)
"

13
New cards

estimating true scores

"1. True score estimate = summed item score

Mostly used in practice
2. True score estimate = Xest = Xo+Rxx(Xo-X-o)
Corrects for regression to the mean (due to unreliability, high scoring people will likely score lower on a new test)
-> the lower the reliability, the more we pull the true score estimate toward the mean.

Standard error (""of measurement"")
-> higher reliability, smaller se 
-> lower reliability, larger se 
Can be used to construct 95% CI around true score estimate.
"

14
New cards

attenuation

"since we use observed scores rather than true scores
-> effect sizes/correlations will be smaller than effect sizes of the true scores.
We should always interpret effect size/significance in light of test reliability!


Cohen's d (number of st. dev. that the groups differ) is smaller for less reliable tests.

  • Group difference is less likely to be significant for less reliable tests.
-> correlation and reliability also affect results
  • correlation is smaller for less reliable tests
  • correlation is less likely to be significant for less reliable tests

"

15
New cards

corrections for attenuation

some people use it, but it's better to just focus on using reliable tests.

16
New cards

index of reliability

= unsquared correlation between observed scores and true scores (rot)

  • if squared, = coefficient of reliability (Rxx)

17
New cards

standard error of measurement (sem)

"= st. dev. of error scores; average size of error scores

  • the larger semthe larger the average difference between observed scores and true scores -> the less reliable the test
  • closely related to reliability:
  • sem can never be larger than so 
"

18
New cards

shared assumptions across all four models

  • error is random: test error score 1 and 2 are uncorrelated
  • two tests reflect same construct
  • true scores are linearly related
19
New cards