Notes on Pearson Correlation and Measurement Theory
Reliability
Core idea: a test is good if it is reliable; reliability initially means stability over time.
Underlying assumption: measurement X = true score T + random error e (classical test theory).
Definitions:
Test-retest reliability (stability over time): measure X by giving it at Time 1 and again at Time 2 to the same people; correlate the two sets of scores. The resulting correlation is denoted and is interpreted as the test’s reliability over time.
Alternate forms reliability: administer two equivalent forms of a test (X1 and X2) to the same people at different occasions; the correlation between the two forms is the alternate forms reliability, labeled similarly as a correlation between two measures of the same construct.
Split-half reliability: a practical, time-saving approach to estimate reliability. The test is split into two halves (e.g., odd items vs. even items); compute the Pearson correlation between the two halves. This is often averaged over many possible splits and reported as the split-half reliability. It assesses internal consistency (how well the items on a test measure the same construct) rather than stability over time.
Internal consistency and item-total relationships: to ensure a test is unidimensional, assess how each item relates to the total score (item-total correlations, sometimes denoted as r_it or rir). High item-total correlations support keeping the item; low ones suggest removing it.
Important caveats:
Split-half reliability measures internal consistency, not stability over time; thus it is not a substitute for test-retest or alternate forms reliability.
High split-half reliability does not guarantee homogeneity (unidimensionality); it indicates items relate to a common attribute but may reflect multiple subfactors.
Reliability reporting in applied research often emphasizes split-half/Cronbach’s alpha; true test-retest reliability is rarer in published studies.
Practical notes from the chapter:
Long-term reliability (IQ-type measures) can be very high (often in the .90s across years). Short-term reliability is easier to obtain but may need balance with practical utility.
Learning and practice effects can influence retest results; when using time-separated designs, the choice of time interval matters.
For cognitive ability tests, the system can yield high reliability; for attitudinal measures, reliability can vary more.
Reliability coefficients often reported in literature (illustrative):
Typical split-half reliabilities for some measures around .70–.92; some studies report as low as .44 or even lower for problematic measures.
Meta-analytic findings: median Cronbach’s alpha around $.79$, supervisory ratings around $.86$, etc.
Key takeaway: measuring reliability requires choosing the appropriate method for the context (time-based stability vs. internal consistency) and understanding that reliability and validity are distinct properties of a measure.
Validity
Reliability is necessary but not sufficient for validity: a test can be highly reliable (stable) but fail to measure the intended construct.
Four broad validity notions (brief overview):
Face validity: subjective judgment that the test looks like the right kind of measure for the intended construct.
Example: an exam item that clearly tests statistics should look valid; a Rorschach inkblot example illustrates potential face validity concerns.
Content validity: the content of the measure systematically reflects the intended construct; involves defining the universe of content and sampling from it to reflect the construct.
Evaluation criteria: (1) how well the universe is defined, (2) how systematically items sample from this universe, (3) how accurately items reflect the sampled domains.
Applications: end-of-term exams aligned with course content; thorough job analyses preceding performance appraisals or selection test development.
Criterion-related validity: the degree to which a test predicts an outcome (the criterion Y). The correlation between X (predictor) and Y (criterion) is the criterion-related validity (often labeled as simply validity in practice):
Predictive validity: administer X, wait, then collect Y; compute r_{xy} between X and later Y.
Concurrent validity: collect X and Y at the same time; compute r_{xy} to assess validity.
Important caveat: the value of r_{xy} is context-specific and depends on the chosen performance measure and the context; validity can vary across settings and over time (validity generalization is discussed later).
Construct validity: evidence that a measure of a construct behaves as theory would predict with respect to other variables (the nomological net).
Approach: measure Y and a battery of other relevant variables; examine a pattern of correlations with Y that aligns with theory (positive, negative, or zero relations as expected).
Key point: construct validity is supported by a network of correlations rather than a single criterion; it is not provable but strengthened by converging evidence.
Nomological net example (construct validity): measure of job insecurity (Y) expected to relate in specific ways to other variables (e.g., positive relation with organizational changes, role ambiguity, somatic complaints; negative relation with organizational commitment and trust). Observed correlations example: .46 with organizational changes, .09 with somatic complaints, .46 with role ambiguity, -.47 with organizational commitment, -.51 with trust; correlations with preexisting insecurity measures: .48 and .35. These supporting correlations increase confidence in the measure’s validity without proving it.
Important caveat about validity claims:
There is no universal “validity” of a measure; validity is the extent to which a measure is appropriate for a particular use in a specific context. The same measure can be valid in one context and not in another.
Validity generalization attempts to address portability of validity evidence across settings by correcting for unreliability and other artifacts across studies.
Classical Test Theory and Core Equations
Core model: for a given measurement, observed score X is the sum of true score T and error e:
Assumptions: errors are random, independent of true scores and of each other:
Variance decomposition (population):
Let
Hence:
Reliability can also be interpreted as the squared correlation between the observed score X and its true score T:
The classic result:
The relationship with two testing occasions (the same true score across occasions):
If you measure X at two occasions X and X', with the same true T and independent errors, then:
(under the assumption that errors are uncorrelated across occasions)
The reliability of X is also the correlation between X and X' (the same test measured twice):
Consequences:
Reliability equals the proportion of observed score variance that is due to true score variance, i.e., a measure of consistency.
Higher reliability implies smaller standard error of measurement (SEM).
Standard Error of Measurement (SEM)
SEM is the standard deviation of an individual’s observed scores across repeated measurements, holding true score constant:
Equivalently, with notation from the chapter: (the key idea: SEM decreases as reliability increases).
Numerical example (IQ testing): assume ; then
A 95% CI for an observed score of 120 would be: 120 ± 1.96 × SEM, i.e., approximately 111.69 to 128.31.
Another example (performance appraisal Y on a 5-point scale with reliability r_{YY} = 0.70, SD of Y ≈ 1.0):
95% CI around a score of 3 would be 3.0 ± 1.96 × 0.548 ≈ [1.93, 4.07], which could cross a raise cutoff point.
Corrections for Unreliability (Attenuation) and Hypothesis Testing
If X and Y each have unreliability, the observed correlation is attenuated relative to the true correlation between the underlying constructs Tx and Ty:
For population correlations: (the chapter derives the relationship)
Corrected correlation (attenuation correction):
In sample notation:
Example: A study found r{YX} = -0.17 for negative affect (Y) and prosocial behavior (X). With reliabilities r{YY} = 0.87 and r_{XX} = 0.88, the corrected correlation is
Hypothesis testing using r:
When testing the null hypothesis that the true correlation is zero, use the uncorrected sample r and the test statistic
The corrected r should not be used in the t-test directly because the standard error of r changes when correcting for unreliability (the denominator is not simply the standard error of r).
Practical implications for validity studies in selection: often correct r_{XY} for unreliability in Y but not for unreliability in X when assessing the validity of a test for its intended use; a simplified form used in practice is
Bounds on validity (test length relationships):
The maximum possible observed validity given reliabilities is
In particular, if the criterion Y is perfectly reliable (r{YY} = 1), then the maximum possible observed validity is
This demonstrates that reliability is necessary but not sufficient for validity; even with perfect X reliability, the observed validity can be limited by the reliability of Y.
The Spearman-Brown Prophecy Formula
Purpose: predict how reliability changes when test length changes (e.g., doubling the number of items) under the assumption that added items are equivalent in quality and difficulty.
General form for length change by factor k:
This formula is valid for noninteger k as well and can model increases or decreases in length.
Example from the chapter: original r_{XX} = 0.784; if length is increased by a factor k = 2 (doubling items):
Intuition: increasing test length generally increases reliability, but with diminishing returns as reliability approaches 1.0.
Practical use: the Spearman-Brown formula helps decide how long a test should be to achieve a desired reliability without conducting a new empirical study.
Validity Generalization, Meta-analysis, and Related Topics
Validity generalization: the idea that a set of correlations (validities) across studies, populations, and measurement methods may be transportable in sign and often in magnitude after correcting for unreliability and other artifacts.
Hypothesis: if a predictor X is empirically valid in one organization, it may be valid in similar organizations after correcting for differences in unreliability.
Caveat: differences in reliability across studies affect observed validities; corrected correlations are compared to assess whether true relationships are consistent.
Meta-analysis: a broader analytic approach that aggregates corrected effect sizes across studies, adjusts for unreliability and other artifacts, and examines remaining variation to identify moderators and general patterns.
Limitations and criticisms: meta-analytic corrections are debated; reliability corrections are part of broader validity-generalization discussions and have been reconsidered in some contexts.
Generalizability Theory
Extension of classical test theory to account for multiple sources of error (facets), not just time or alternate forms.
In practice, reliability can be decomposed into components due to different facets such as raters, items, contexts, and occasions.
Decision (D) studies: use generalizability theory to decide which facets to increase (or adjust) to maximize reliability most efficiently, possibly by combining facets (e.g., two types of items like observation and interview) rather than simply increasing the number of raters.
Example: two types of items (simulation and interview) may yield larger reliability gains than adding more raters.
Benefits: provides a principled way to optimize measurement procedures across multiple sources of error; can be integrated with test-length considerations.
Item Length, Response Scales, and Validity Considerations
Reliability tends to plateau after a moderate number of response categories; 5–7 categories often suffice for reliability, with diminishing returns beyond that range.
However, validity may depend on the number of scale points: more scale points can improve detection of interactions and the ability to model relationships in regression analyses.
Practical takeaway: use a scale with enough points to capture nuances (near-continuous measures are preferable when feasible); reliability does not improve indefinitely with more categories, but validity of certain models can benefit from finer scaling.
Software and tools have been developed to allow near-continuous scoring in practice (e.g., Aguinis, Bommer, and Pierce, 1996).
Practical and Conceptual Takeaways
Reliability and validity are distinct but interrelated properties of measurement:
Reliability concerns consistency and freedom from measurement error.
Validity concerns whether a test measures what it is intended to measure and whether it serves its predicted purpose.
When evaluating a measure for selection or assessment purposes, it is common to report:
Reliability (e.g., rxx, ryy, alpha) and its method (test-retest, split-half, alpha),
Validity evidence (criterion-related validity rxy, predictive vs concurrent validity, and construct validity evidence via nomological networks).
In practice, researchers correct observed correlations for unreliability to estimate the true relationships, but caution is needed in hypothesis testing and interpretation, especially regarding which reliabilities to correct and how to interpret corrected values.
Theoretical and empirical work in measurement theory continues to address generalizability, range restriction, and complex sources of error via generalizability theory and meta-analytic methods.
Worked Examples and Key Formulas (recap)
Classical Test Theory model:
Reliability as variance proportion:
Standard Error of Measurement:
Correction for unreliability (attenuation):
Maximum possible observed validity given reliabilities:
Spearman-Brown prophecy formula (change test length by k):
Relationship between true and observed correlations in two measures:
If X and Y have reliabilities r{XX} and r{YY}, then the true correlation between Tx and Ty is:
Predictive vs concurrent validity distinctions and cautions on interpretation and stability across contexts.
Problems and Practice (themes)
Problem-type practice prompts include: selecting appropriate reliability type, correcting observed validities for unreliability, adjusting test length to achieve a target reliability, and correcting correlation matrices for unreliability in multiple measures.
These problems reinforce the mathematical relationships among reliability, validity, and test length, as well as the practical implications for measurement in selection and assessment contexts.
Connections to broader measurement theory
This chapter links Pearson r to fundamental measurement concepts in psychometrics (reliability, validity) and shows how classical test theory underpins practical decisions in test construction, validation, and throughput in organizational settings.
It also situates reliability and validity in a broader research program, including validity generalization, meta-analysis, and generalizability theory, showing how correlations and their corrections are used across studies, populations, and measurement methods.
References to foundational ideas mentioned in the chapter
Classic sources and related readings: Cronbach (alpha), Spearman-Brown, Campbell & Fiske (multitrait-multimethod), reliability and validity debates, and measurement theory texts by Ghiselli, Campbell, and Zedeck; Nunnally; and contemporary applications in organizational psychology and personnel selection.