Psychological Testing and Assessment: Chapter 5 - Reliability
Fundamental Principles of Reliability
Conceptual Overview: Reliability refers to the consistency of measurement. While in colloquial language it implies dependability or high quality, in psychometrics it strictly denotes the extent to which a measurement process produces similar results across repeated applications. It is not an all-or-none property; a test may be reliable in one setting and unreliable in another.
Reliability Coefficient: This is a statistic used to quantify the degree of reliability. It ranges from a value of (indicating a total lack of reliability) to (indicating perfect reliability).
Measurement Error and Score Theory
Measurement Error: In a scientific context, this refers to the inherent uncertainty and imprecision associated with any measurement. It encompasses both preventable mistakes and inevitable aspects of measurement fluctuation. Even when procedures are followed perfectly, measurement error is present.
Theoretical Scores:
True Score (): A theoretical value representing a measurement if no error were present. True scores cannot be observed directly; they are approximated by averaging multiple measurements while holding time and conditions constant.
Observed Score (): The actual score obtained from a measurement process.
Construct Score: A value representing a person’s standing on a theoretical variable (e.g., Depression or Reading Ability) independent of a specific instrument. While reliable tests approximate true scores, valid tests approximate construct scores.
Measurement Interference:
Elapse of Time: Psychological variables like mood or alertness flux constantly, causing true scores to change over time.
Carryover Effects: The act of measurement can change the variable being measured. This includes Practice Effects (increased skill through taking the test) and Fatigue Effects (decreased motivation or energy).
The Mathematical Foundations of Reliability
The Classical Test Theory (CTT) Equation: The relationship between observed, true, and error scores is defined as:
Where is the observed score, is the true score, and is the measurement error.
Variance Components: In a distribution of test scores, total observed variance () is the sum of true variance () and error variance ():
Reliability is defined as the proportion of total variance attributed to true variance. Higher proportions of true variance yield higher reliability.
Types of Error Variance
Random Error: Unpredictable fluctuations with no discernible pattern. Often called "noise," these events may raise or lower scores inconsistently (e.g., a sudden increase in a testtaker's blood pressure or an accidental distraction). Long-term, random errors tend to cancel each other out.
Systematic Error: Consistent errors that influence scores in a fixed direction. This is referred to as Bias. For example, a ruler that is actually inches long will consistently overestimate measurements by inches. Because systematic error is consistent, it does not necessarily affect the consistency (reliability) of the scores, though it compromises validity.
Major Sources of Error Variance
Test Construction:
Item Sampling or Content Sampling: Variation in item wording or the specific domain content selected for a test can cause scores to fluctuate across different versions of a test.
Test Administration:
Environmental Factors: Room temperature, lighting, ventilation, noise, or even the quality of the writing surface.
Testtaker Variables: Emotional state, physiological discomfort, lack of sleep, medication, or fasting glucose levels (which can affect cognitive performance).
Examiner Variables: Physical appearance, nonverbal cues (head nodding, eye movements), or deviations from prescribed testing procedures.
Test Scoring and Interpretation:
Scorer Subjectivity: In non-objective assessments (essays, behavioral ratings, projective tests), different scorers may interpret the same response differently.
Technical Glitches: Glitches in computer scoring systems can contaminate data.
Other Sources:
Sampling Error: The extent to which a sample (e.g., in a public opinion poll) does not represent the actual population.
Methodological Error: Ambiguous wording in questionnaires or poorly trained interviewers.
Specific Methods for Estimating Reliability
Test-Retest Reliability: Pairs of scores are correlated from the same individuals on two different administrations of the same test. This measures stability over time. When the interval exceeds six months, it is called a Coefficient of Stability.
Parallel-Forms and Alternate-Forms Reliability:
Parallel Forms: Different versions of a test where means and variances of observed scores are equal.
Alternate Forms: Different versions designed to be equivalent in content and difficulty but not necessarily meeting the strict mathematical criteria of parallel forms.
Coefficient of Equivalence: The correlation between the different forms.
Inter-Scorer Reliability: The degree of agreement between two or more raters. It is often measured using Cohen’s Kappa, Pearson r, or Spearman rho. High inter-scorer reliability ensures that scores are derived systematically rather than through rater idiosyncrasies.
Psychology's Replicability Crisis
The Crisis: Findings published in leading journals were found to be non-replicable in of cases.
Causal Factors: Lack of published replication attempts, editorial preference for positive findings (rejecting the null hypothesis), and Questionable Research Practices (QRPs) such as "peeking" at data to decide if more should be collected.
Remedies: Preregistration (publicly committing to procedures and analysis plans before data collection) and open science initiatives.
Internal Consistency Estimates
Internal Consistency: Evaluates the extent to which items on a single scale correlate with one another (inter-item consistency).
Split-Half Reliability: A single test is divided into halves, and the scores are correlated. Common methods include Odd-Even Reliability (correlating odd-numbered items with even-numbered items).
The Spearman-Brown Formula: Used to adjust split-half reliability. Since reliability increases with test length, reducing a test to halves underestimates the reliability of the whole test:
For a whole test based on half-test correlation ():
Coefficient Alpha (Cronbach's Alpha): The mean of all possible split-half correlations, corrected by Spearman-Brown. It is the most common measure of internal consistency:
Where is the number of items, is the sum of item variances, and is the total test variance.
McDonald’s Omega: An alternative to Alpha that accurately estimates internal consistency even when items have unequal loadings (strengths of relationship to the true score).
Diagnostic Reliability and Methodology
Comparison of Methods: Inter-rater reliability in diagnostic settings varies significantly based on the estimation method.
Audio-Recording Method: One clinician interviews, a second listens to the recording. This often inflates reliability because the second rater is constrained by the same information and follow-up questions provided by the first.
Test-Retest Method: Two independent interviews are conducted. This is more ecologically valid as it accounts for patient information variance.
DSM-5 Data: The field trials for DSM-5 used the test-retest method and showed a mean kappa of only ("fair"), whereas older versions using audio-recording showed "excellent" reliability.
Factors Influencing Reliability Estimates
Homogeneity vs. Heterogeneity: Homogeneous tests (measuring one factor) should have high internal consistency. Heterogeneous tests may have low internal consistency but high test-retest reliability.
Dynamic vs. Static Characteristics: Static traits (e.g., Intelligence) are stable; dynamic states (e.g., Anxiety) flux. Test-retest is inappropriate for dynamic states.
Restriction/Inflation of Range: Correlational coefficients decrease if the score range is restricted (e.g., testing only hired high-performers) and increase if the range is inflated.
Speed vs. Power Tests:
Power Test: Time is sufficient, but items are difficult.
Speed Test: Items are easy but time is limited. Split-half reliability calculated from a single administration of a speed test will be spuriously high (near ).
Criterion-Referenced Tests: Often used for mastery (pass/fail). Traditional reliability measures requiring variability are often inappropriate for these tests because the goal is not to differentiate between individuals but to ensure a criterion is met.
Alternatives to Classical Test Theory
Generalizability Theory: Instead of a single true score, this theory proposes a Universe Score. It examines Facets (sources of variation like time of day or scorer training). It includes Generalizability Studies (testing the impact of facets) and Decision Studies (assessing the utility of scores for making specific decisions).
Item Response Theory (IRT): Often called Latent-Trait Theory. It models the probability of a specific response based on the testtaker's level of a underlying trait.
Item Difficulty: How hard an item is to solve or perform.
Item Discrimination: The degree to which an item identifies people with higher vs. lower levels of the trait.
Rasch Model: A specific IRT model where all items are assumed to have an equivalent relationship with the construct.
Standard Error of Measurement (SEM)
Definition: The standard deviation of a theoretically normal distribution of test scores obtained by one person on equivalent tests. It represents the precision of an individual observed score. There is an inverse relationship between SEM and reliability:
Confidence Intervals: A range of scores likely to contain the true score based on the SEM.
confidence interval: Observed Score
confidence interval: Observed Score (often rounded to )
confidence interval: Observed Score (often rounded to )
Standard Error of the Difference Between Two Scores
Standard Error of the Difference (SED): A tool to determine if the difference between two scores (either from two people or two different tests for one person) is statistically significant. The SED is larger than the SEM of either individual score because it accounts for error in both measurements:
Or, using reliability coefficients on the same scale:
To be confident that two scores represent a true difference, the scores must be separated by approximately .