Chapter 5
Reliability in Assessment
Overview
Course: PSYC 385 Psychological Test
Semester: Fall 2025
Outline
Concept of Reliability
Sources of Error Variance
Types of Reliability Estimates
True Score and Alternative Models
Learning Objectives
Understand the concept of reliability and identify different reliability estimates.
Be able to interpret the usefulness of different estimates of reliability depending on the context and intent of the psychological test.
Interpret standard error terms and understand their importance in creating confidence intervals.
The Concept of Reliability
Reliability:
Defined as the or _ in measurement.
Reliability coefficient is an index of reliability, defined as a proportion that indicates the ratio between the true score variance on a test and the total variance.
The observed score can be represented as follows:
Where:
represents the observed score.
represents the true score.
signifies the error associated with measurement.
Error: Refers to the component of the observed score that does not pertain to the test taker’s true ability or the trait being measured.
Reliability Estimates
Variance:
Variance is defined as the standard deviation squared.
The total variance can be expressed as:
Reliability:
Defined as the proportion of the total variance attributed to true variance.
Measurement Error:
Refers to all factors associated with the measurement process that do not relate to the variable being measured.
The Concept of Measurement Error
Random Error:
A source of error in measuring a targeted variable caused by and inconsistencies of other variables in the measurement process (i.e., noise).
Systematic Error:
A source of error that is typically constant or proportionate to the presumed true value of the variable being measured.
Sources of Error Variance
Potential Sources of Error Variance:
Example: Duck or Rabbit?
Test Construction:
Variation may exist within items on a test or between different tests (also known as item sampling or content sampling).
Test Administration:
Error may arise from the testing environment and situational factors, such as pressing emotional issues, physical discomfort, lack of sleep, or drug effects.
Examiner-related variables—including physical appearance, training, and demeanor—can also contribute.
Test Scoring and Interpretation:
Computer testing can reduce scoring errors, but many tests still require expert interpretation (e.g., projective tests).
Subjectivity in scoring may influence behavioral assessments.
Other Sources of Error:
Sampling Error:
Represents the extent to which the sample in a study is truly representative of the population.
Methodological Error:
Includes factors like ambiguous wording in questionnaires and biased question framing.
Reliability Estimates
Test-Retest Reliability:
An estimate obtained by correlating pairs of scores from the same individual on two different administrations of the same test.
Most applicable for variables expected to be stable over time (e.g., personality) and not suitable for those expected to fluctuate (e.g., mood).
Reliability estimates tend to decrease as intervals broaden, known as the coefficient of stability when intervals exceed six months.
Test-Retest Reliability Details
Execution:
Same person, same test at two different time points illustrated as:
Time 1: 3
Time 2: 8
Reliability Estimates via Alternate Forms
Parallel-Forms or Alternate-Forms:
Coefficient of equivalence measures the relationship degree between various test forms.
Parallel Forms:
Mean and variance of observed test scores are equal across forms.
Alternate Forms:
Different test versions that do not meet parallel forms' strict requirements but maintain similar content and difficulty.
Reliability is verified by administering two forms to the same group, acknowledging potential error factors related to test-takers' states (practice, fatigue, etc.) or item sampling.
Split-Half Reliability
Split-half reliability is obtained through correlating scores from two halves of a single administered test.
Process Steps:
Step 1: Divide the test into two halves (e.g., odd vs. even items).
Step 2: Calculate Pearson r between scores on both halves.
Step 3: Adjust reliability using the Spearman-Brown formula, estimating internal consistency reliability from the correlation of both halves.
Other Methods of Estimating Internal Consistency
Inter-Item Consistency:
Assesses the relatedness of items within a test to gauge its _.
Kuder-Richardson Formula 20:
A preferred statistic for determining inter-item consistency of .
Coefficient Alpha:
Represents the mean of all possible split-half correlations, adjusted via the Spearman-Brown formula.
Popular for estimating internal consistency with values ranging from 0 to 1.
Coefficient Alpha Ratings
Coefficient Alpha (α) | Interpretation |
|---|---|
≥ 0.9 | Excellent |
0.9 ≥ α ≥ 0.8 | Good |
0.8 > α ≥ 0.7 | Acceptable |
0.7 > α ≥ 0.6 | Questionable |
0.6 > α ≥ 0.5 | Poor |
0.5 > α | Unacceptable |
Measures of Inter-Scorer or Inter-Rater Reliability
Inter-Rater Reliability:
Assesses the degree of agreement among raters with respect to a particular measure.
Commonly applied in behavioral measures to minimize biases or idiosyncrasies in scoring.
Coefficient of Inter-Score Reliability:
Correlates scores from different raters.
Practical Example
Table 1. Scale Characteristics for the Deployment Communication Inventory.
Scale
Mean
SD
Range
α
r
Frequency
25.12
5.11
13-42
--
1
Assurance
19.09
4.21
5-25
.89
.64
Conflict
8.41
3.43
5-22
.86
.56
PS/Disclosure
16.46
3.92
6-25
.83
.50
Perceived benefits
32.87
5.49
8-40
.91
.56
Perceived costs
9.50
3.85
4-20
.83
.55
Reliability Estimates Overview
Reliability estimates may vary based on the nature of the variables under study.
It is crucial, yet challenging, to decompose error variance into its individual components.
Determining Reliability Metrics
The test's nature influences the reliability metric utilized. - Considerations:
Homogeneity or heterogeneity of test items.
Whether the measured characteristic, ability, or trait is dynamic or static.
Range restrictions of test scores.
Whether the test is a speed or power test.
Whether the test is criterion-referenced.
True-Score Model vs. Alternatives
True-Score Model:
Often known as Classical Test Theory (CTT), it is the most commonly used model due to its simplicity.
True Score:
A value that genuinely reflects a person's ability or trait level as per classical test theory.
Contention:
While CTT assumptions are generally acceptable, a significant controversial assumption is the equivalence of test items, generally yielding longer tests.
Alternative Models
Generalizability Theory:
Suggests that test scores may vary due to variables present in the testing scenario and sampling.
Encourages test developers and researchers to elucidate the specifics of the test situation leading to a certain score.
Item-Response Theory (IRT):
Models the probability of individuals of ability X to score at level Y on a test.
Key Elements:
A family of methods encompassing item difficulty and discrimination.
Discrimination (a):
Refers to an item's effectiveness in distinguishing individuals with different levels of the trait being measured.
Difficulty (b):
Relates to the complexity of an item.
Example of IRT Parameters
Table 1: Item discrimination (a) and difficulty (b) parameters for 10 items of the Marital Satisfaction Inventory-Brief Form.
Item
a
b
Endorsement threshold
1.
1.17
0.30
L
2.
3.25
0.89
H
3.
2.32
1.97
H
4.
3.25
0.77
L
5.
1.36
-0.03
L
…
…
…
…
The Standard Error of Measurement
Standard Error of Measurement (SEM):
Represents a measure of the of an ____.
Estimates the inherent error in an observed score or measurement.
Generally, greater test reliability leads to a reduction in standard error.
SEM can help estimate how much an observed score deviates from a true score.
Confidence Interval:
A range of test scores likely to contain the true score.
Can be expressed using the formula:
Example of Confidence Consideration
Example: If one scores 74 on an IQ test with a cut-off of 75 for suggesting intellectual disability, how confident can we be that the individual falls below the threshold if the SEM is 3?
Standard Error of the Difference
Standard Error of Difference:
A measure assisting test users to discern how large a score difference should be deemed _ ___ .
Can address three types of queries:
Comparison of an individual's performance across different tests (test 1 vs. test 2).
Comparison of an individual's performance against another individual's performance on the same test (test 1).
Comparison of an individual's performance on one test against another individual's performance on a different test (test 1 vs. test 2).