Ch 4

1. History and Theory of Reliability

  • Complexity of Measurement in Psychology:

    • Measuring psychological traits is more complex than physical measurements like height or weight due to the nature of psychological constructs.

    • Psychological measurements often involve traits that are not directly observable.

    • This complexity can lead to overestimated or underestimated scores.


2. Basics of Test Score Theory

  • Concept of Error:

    • Error is defined as the discrepancy between a true score (the actual value) and the observed score (the value that is measured).

    • The relationship can be expressed mathematically as:
      X=T+EX = T + E

    • Where:

    • X= Observed score

    • T = True score

    • E = Error

    • Classical Test Theory Assumptions:

    • Errors of measurement are random.


3. Standard Error of Measurement

  • Understanding Standard Error:

    • Classical test theory posits that the true score remains stable over repeated test administrations.

    • In practice, repeated testing can yield different scores due to random measurement errors.

    • The distribution of errors aids in calculating the standard deviation of errors, known as the standard error of measurement (SEM).


4. The Domain Sampling Model

  • Measurement Limitations:

    • Utilizing a limited number of observations to represent a larger construct poses challenges because it generates error due to under-sampling.

    • The reliability estimate arises from comparing sample variance to population variance.

Formula for Reliability:
  • Reliability (r{1t}) is given by: r{1t} = rac{ ext{scores on test 1}}{ ext{true score}} = ext{average correlation between test 1 and other random parallel tests}


5. Item Response Theory

  • Overview of Item Response Theory (IRT):

    • IRT diverges from classical test theory by using adaptive assessments that change according to the test taker’s performance.

    • Requires a significant pool of items and advanced computer programs for adjustment.

    • IRT emphasizes evaluating questions based on their difficulty to accurately assess an individual's ability level.


6. Models of Reliability

  • Determining Reliability Sources:

    • Reliability must be established before making high-stakes decisions (e.g., employment, educational assessments).

    • Reliability is represented through correlations and can be mathematically viewed as ratios.

Sources of Error:
  1. Situational factors affecting test performance.

  2. Test-retest reliability: Evaluate consistency across time.

  3. Parallel forms reliability: Compare variations of the same test.

  4. Internal consistency reliability: Measures agreed consistency among items in a test.


7. Time Sampling: The Test–Retest Method

  • Test-Retest Methodology:

    • Involves administering the same test at two different times to the same participant; effective for stable traits (e.g., IQ).

    • Correlates scores from different time points, accounting for carryover effects (e.g., practice effects).

    • Consideration of time intervals between tests is crucial for accuracy.

Visual Representation:
  • Time Points:

    • Time 1 → Test Administered

    • Time 2 → Test Administered Again


8. Item Sampling: Parallel Forms Method

  • Alternate Forms Reliability:

    • Involves comparing scores from two different tests that measure the same quality (also known as equivalent forms method).

    • Conducted on the same day; variability in scores should only reflect error or test differences.


9. Split-Half Method

  • Definition and Application:

    • A method where a single test is divided into two halves, with each half compared for consistency.

    • Can be segmented randomly, by odd/even, or by first/second halves.

Reliability Formula: Spearman–Brown Correction
  • Given by: r{SB} = rac{2r{1}}{1 + r_{1}}

    • Dedicated for cases with unequal variance, where Cronbach’s coefficient alpha may be applied.


10. KR20 and Coefficient Alpha

  • KR20 Formula:

    • An estimate of reliability for tests with dichotomous criteria (yes/no, true/false).

    • Shows reliability comparable to the split-half method but is particularly useful for certain scenarios.

Coefficient Alpha:
  • Evaluates reliability when there isn’t a single correct answer per item.

    • Similar assessment properties to KR20 but handles more complex testing situations effectively.

  • Low internal consistency indicates deviation among test items fueling the need for further factor analysis.


11. Reliability of a Difference Score

  • Understanding Difference Scores:

    • Occur when one measurement is subtracted from another, where error rates are often higher than observed or true scores.

    • Recommended to convert to Z scores before analysis for accuracy in interpretation.


12. Knowledge Check Examples

  • To assess reliability of a questionnaire:

    • Test–Retest Method: Appropriate when measuring stable traits over different time points.

Assessing Research Assistant Accuracy:
  • Employ Interrater Reliability to measure how students correlate in their observations of content in media.


13. Reliability in Behavioral Observation Studies

  • Complications in Behavioral Observation:

    • Observability constraints result in sampling errors that can distort true behavior representation.

    • Interrater reliability assesses the correlation among different observers' evaluations.

Kappa Statistic:
  • A reliable measure ranging from -1 to 1; interpreted as follows:

    • > 0.75 = Excellent

    • 0.40 - 0.75 = Satisfactory

    • < 0.40 = Poor


14. Sources of Error Connection with Reliability Assessment Methods

Source of Error

Example

Method

How Assessed



Time Sampling

Same test at two points

Test-Retest

Correlation of scores from both occasions



Item Sampling

Different items for same attribute

Alternate or Parallel Forms

Correlate equivalent test forms with different items



Internal Consistency

Item reliability within a test

Split-half, KR20, Alpha

Consistency evaluation of items in a test structuring



Observer Differences

Different observers scoring

Kappa Statistic

Reveals inter-observer correlation












15. Addressing Low Reliability

  • Methods to enhance reliability when findings are insufficient:

    • Increase Items:

    • Particularly beneficial in domain sampling; use Spearman-Brown to evaluate the necessary item count.

    • Reliability upgrade can be gauged through formulas to predict item augmentation requirements:
      N = rac{(1 - ro)}{(1 - rd)}

    • Additional approaches include factor/item analysis or corrections for attenuation for improved correlation assessments.