Ch 4
1. History and Theory of Reliability
Complexity of Measurement in Psychology:
Measuring psychological traits is more complex than physical measurements like height or weight due to the nature of psychological constructs.
Psychological measurements often involve traits that are not directly observable.
This complexity can lead to overestimated or underestimated scores.
2. Basics of Test Score Theory
Concept of Error:
Error is defined as the discrepancy between a true score (the actual value) and the observed score (the value that is measured).
The relationship can be expressed mathematically as:
Where:
X= Observed score
T = True score
E = Error
Classical Test Theory Assumptions:
Errors of measurement are random.
3. Standard Error of Measurement
Understanding Standard Error:
Classical test theory posits that the true score remains stable over repeated test administrations.
In practice, repeated testing can yield different scores due to random measurement errors.
The distribution of errors aids in calculating the standard deviation of errors, known as the standard error of measurement (SEM).
4. The Domain Sampling Model
Measurement Limitations:
Utilizing a limited number of observations to represent a larger construct poses challenges because it generates error due to under-sampling.
The reliability estimate arises from comparing sample variance to population variance.
Formula for Reliability:
Reliability (r{1t}) is given by: r{1t} = rac{ ext{scores on test 1}}{ ext{true score}} = ext{average correlation between test 1 and other random parallel tests}
5. Item Response Theory
Overview of Item Response Theory (IRT):
IRT diverges from classical test theory by using adaptive assessments that change according to the test taker’s performance.
Requires a significant pool of items and advanced computer programs for adjustment.
IRT emphasizes evaluating questions based on their difficulty to accurately assess an individual's ability level.
6. Models of Reliability
Determining Reliability Sources:
Reliability must be established before making high-stakes decisions (e.g., employment, educational assessments).
Reliability is represented through correlations and can be mathematically viewed as ratios.
Sources of Error:
Situational factors affecting test performance.
Test-retest reliability: Evaluate consistency across time.
Parallel forms reliability: Compare variations of the same test.
Internal consistency reliability: Measures agreed consistency among items in a test.
7. Time Sampling: The Test–Retest Method
Test-Retest Methodology:
Involves administering the same test at two different times to the same participant; effective for stable traits (e.g., IQ).
Correlates scores from different time points, accounting for carryover effects (e.g., practice effects).
Consideration of time intervals between tests is crucial for accuracy.
Visual Representation:
Time Points:
Time 1 → Test Administered
Time 2 → Test Administered Again
8. Item Sampling: Parallel Forms Method
Alternate Forms Reliability:
Involves comparing scores from two different tests that measure the same quality (also known as equivalent forms method).
Conducted on the same day; variability in scores should only reflect error or test differences.
9. Split-Half Method
Definition and Application:
A method where a single test is divided into two halves, with each half compared for consistency.
Can be segmented randomly, by odd/even, or by first/second halves.
Reliability Formula: Spearman–Brown Correction
Given by: r{SB} = rac{2r{1}}{1 + r_{1}}
Dedicated for cases with unequal variance, where Cronbach’s coefficient alpha may be applied.
10. KR20 and Coefficient Alpha
KR20 Formula:
An estimate of reliability for tests with dichotomous criteria (yes/no, true/false).
Shows reliability comparable to the split-half method but is particularly useful for certain scenarios.
Coefficient Alpha:
Evaluates reliability when there isn’t a single correct answer per item.
Similar assessment properties to KR20 but handles more complex testing situations effectively.
Low internal consistency indicates deviation among test items fueling the need for further factor analysis.
11. Reliability of a Difference Score
Understanding Difference Scores:
Occur when one measurement is subtracted from another, where error rates are often higher than observed or true scores.
Recommended to convert to Z scores before analysis for accuracy in interpretation.
12. Knowledge Check Examples
To assess reliability of a questionnaire:
Test–Retest Method: Appropriate when measuring stable traits over different time points.
Assessing Research Assistant Accuracy:
Employ Interrater Reliability to measure how students correlate in their observations of content in media.
13. Reliability in Behavioral Observation Studies
Complications in Behavioral Observation:
Observability constraints result in sampling errors that can distort true behavior representation.
Interrater reliability assesses the correlation among different observers' evaluations.
Kappa Statistic:
A reliable measure ranging from -1 to 1; interpreted as follows:
> 0.75 = Excellent
0.40 - 0.75 = Satisfactory
< 0.40 = Poor
14. Sources of Error Connection with Reliability Assessment Methods
Source of Error | Example | Method | How Assessed | |
|---|---|---|---|---|
Time Sampling | Same test at two points | Test-Retest | Correlation of scores from both occasions | |
Item Sampling | Different items for same attribute | Alternate or Parallel Forms | Correlate equivalent test forms with different items | |
Internal Consistency | Item reliability within a test | Split-half, KR20, Alpha | Consistency evaluation of items in a test structuring | |
Observer Differences | Different observers scoring | Kappa Statistic | Reveals inter-observer correlation | |
15. Addressing Low Reliability
Methods to enhance reliability when findings are insufficient:
Increase Items:
Particularly beneficial in domain sampling; use Spearman-Brown to evaluate the necessary item count.
Reliability upgrade can be gauged through formulas to predict item augmentation requirements:
N = rac{(1 - ro)}{(1 - rd)}Additional approaches include factor/item analysis or corrections for attenuation for improved correlation assessments.