Validity & Test Construction
Definition (#f7aeae)
Important (#edcae9)
Extra (#fffe9d)
Validity:
How well a test measures what it is supposed to measure.
A test is only valid if its scores are meaningful and appropriate for a specific purpose.
Key point: No test is valid for all people, all uses, or forever—validity depends on context, population, and intended application.
Why ‘valid test’ is misleading:
Tests are not universally valid.
Only the inferences we draw from test scores can be considered valid or invalid.
A test may be valid for diagnosing depression but not for predicting suicide risk.
Norms that were valid in 20026 may be outdated by 2036.
Validity is always context-dependent, population-specific and time limited.
Validation: the process
Test developers provide validity evidence in test manuals.
Test users may conduct local validation studies when needed for their specific context or population.
Local validation is essential when:
The test is translated to another language, items or instructions are modified, or the test is used with a new population not represented in original norms.
3 types of validity: (Trinitarian model)
Content Validity: Ensures the test covers the entire domain of knowledge, skills, or behaviors it claims to measure.
Criterion-Related Validity: Examines how well scores correlate with external criteria like academic success.
Construct Validity: Evaluates whether the test captures the underlying psychological concept it intends to assess.
Validity diagram:
The "umbrella model" views construct validity as the overarching framework that encompasses all validity evidence.
Content validity and criterion-related validity are not separate types, but rather different sources of evidence that support construct validity.
This unified approach recognizes that all validation efforts ultimately aim to demonstrate that test scores meaningfully measure the intended psychological construct.

Face validity:
Refers to the perceived relevance of test items from the test-taker's perspective, not actual statistical validity.
Content validity:
Ensures comprehensive coverage of all relevant areas.
Determined by the degree to which the questions, tasks, or items on a test are representative of the universe of behavior the test was designed to sample.
Quantification of content validity.
Content validity = D/(A + B + C + D).
Test blueprint:
Ensures representative content sampling across all relevant domains.
For an Assertiveness Test, items are distributed systematically:
Home: 10 items (25%) - family interactions, domestic situations.
Work: 15 items (40%) - professional settings, colleague relations.
Social context: 10 items (25%) - friendships, public situations.
Miscellaneous: 5 items (10%) - other relevant contexts.
This structured approach guides item development and ensures comprehensive domain coverage.
Lawshe’s content validity ratio: (CVR)
Experts rate each item as Essential.
Useful but not essential, or not needed.
Formula: CVR = (ne – N/2) / (N/2)
Where ne = number rating "Essential" and N = total experts.
Ex:
If 9 of 10 experts rate an item Essential → CVR = 0.80 → Accept item.
If only 4 of 10 rate Essential → CVR = -0.20 → Reject item from the test.
Criterion-related validity:
Concurrent validity:
Test and criterion measured at the same time.
Ex: A new depression screening tool is validated against an established clinical diagnosis made on the same day.
Predictive validity:
Test scores predict future performance or outcomes.
Ex: SAT scores administered in high school predict college GPA years later.
Validity coefficient:
The correlation between test scores and criterion measures.
Typical values range from 0.00 to 0.50, and rarely exceed 0.60.
A coefficient must be high enough to meaningfully aid decision-making in practical contexts.
Context matters when interpreting these values—a coefficient of 0.30 may be excellent for personality assessment but inadequate for high-stakes selection decisions.
Expectancy table:
Scores 90–100 → 80% likelihood of achieving GPA > 3.0.
Scores 70–89 → 60% likelihood.
Scores 50–69 → 30% likelihood.
Scores 30–49 → only 10% likelihood of success.
Expectancy tables translate test scores into probability statements, showing the likelihood of achieving a specific criterion (ex: GPA above 3.0) based on score ranges.
These tables have practical utility in education (predicting academic success), hiring decisions (forecasting job performance), and establishing meaningful cutoff scores for selection.
Ratings range from Poor to Excellent, with each higher rating corresponding to a greater probability of job success.
This visual tool helps employers make informed hiring decisions.
Bar or line charts illustrate the trend clearly:
Poor ratings may show 20% success likelihood.
Excellent ratings demonstrate 80% or higher probability of successful job performance.

Taylor-Russell tables:
Demonstrate how 3 factors combine to determine hiring success:
The validity coefficient: How well the test predicts job performance.
The base rate: Percentage who would succeed without testing.
The selection ratio: Proportion of applicants hired.
Higher predictive validity significantly increases the accuracy of hiring decisions.
Key insight: Tests become most useful when the selection ratio is low— when hiring few applicants from a large pool.
In competitive hiring scenarios, modest validity coefficients can substantially improve the proportion of successful hires compared to random selection, making validation studies essential for organizational decision-making.
Incremental validity:
Refers to the additional predictive power a new measure provides beyond existing predictors.
This concept is crucial when selecting which tests or measures to include in assessment batteries.
For ex:
When predicting exam performance, consider 3 predictors: time spent studying, time in the library, and sleep quality before exams.
If sleep quality explains variance not captured by study time alone, it demonstrates high incremental validity.
Construct validity:
Addresses whether a test truly measures the underlying psychological concept it claims to assess.
Constructs are unobservable theoretical entities—such as intelligence, anxiety—that cannot be directly seen or touched.
Establishing construct validity requires a solid theoretical foundation, testable hypotheses, and multiple sources of converging evidence.
Includes examining internal consistency, correlations with related measures, group differences, and factor structure.
Serves as the foundation for meaningful test interpretation and is considered by many to be the overarching framework encompassing all other validity types.
Construct validity map:
Internal Evidence:
Homogeneity and internal consistency demonstrate items measure the same construct.
Age-related changes show expected developmental patterns; pre/post treatment changes confirm sensitivity to interventions.
External Evidence:
Known group differences validate the test distinguishes between relevant populations; convergent validity shows correlation with similar measures; discriminant validity confirms low correlation with unrelated constructs.
Statistical Evidence:
Factor analysis identifies underlying dimensions and confirms the test measures intended constructs.
A comprehensive approach integrating all evidence sources strengthens construct validity claims.
Convergent validity:
High correlations with other tests measuring the same construct demonstrate convergent validity.
Ex: Test A ↔ Test B (r = .78) shows strong agreement between similar measures.
Ex: Test A ↔ Test C (r = .63) further supports convergent validity.
Multiple high correlations with theoretically related measures indicate the test truly captures the intended construct.
Discriminant validity:
A test should show low or near-zero correlations with measures of unrelated constructs.
Low correlations with unrelated constructs confirm the test measures its intended construct specifically, not general response tendencies or irrelevant factors.
Factor analysis:
Identifies the underlying dimensions measured by a test, confirming it measures intended constructs.
Ex: A mood assessment might reveal:
Factor 1: Anxiety (Items 1–4 load highly)
Factor 2: Depression (Items 5–7 load highly)
This statistical technique helps researchers refine and validate test structure by showing which items cluster together.
Validity, bias & fairness:
Validity ≠ Fairness, Validity ≠ Bias: These concepts are related but distinct.
A test may be:
Valid but unfair: Accurately measures construct but used inappropriately.
Fair but invalid: Equal treatment but poor measurement.
Biased but predictive: Systematic error yet still useful.
Unbiased but misused: Technically sound but applied incorrectly).
Ethical Testing Requirements: Responsible psychological assessment requires 3 integrated components:
Valid scores that accurately measure intended constructs.
Fair use that treats all test-takers equitably.
Cultural awareness that recognizes diverse backgrounds and contexts.
All 3 must work together for ethical practice.
Test bias:
Slope bias: Different regression slopes between groups mean the test predicts outcomes differently for each group. A test may predict well for one group but poorly for another.
Intercept bias: Same slope but different intercepts means the test systematically over or under-predicts for certain groups.
Both types impact fairness and validity of test use.
Rating errors:
Central tendency error: Raters avoid using extreme scores, clustering all ratings around the middle of the scale. This reduces score variability and masks true differences between individuals being assessed.
Halo effect: When one positive characteristic influences the rater's perception of unrelated traits. A person seen as friendly may be rated higher on competence, even without evidence of actual skill.
Leniency/severity errors: Some raters consistently give overly generous scores (leniency) while others are habitually harsh (severity). Rater training and calibration sessions are essential to minimize these biases.
Test fairness:
Fairness is essential alongside validity.
A test may produce valid scores yet still be used unfairly.
Ethical testing requires impartial, just application that considers the test-taker's background and context.
Examples of unfair test uses include administering IQ tests in non-native languages, using punitive testing practices, relying on outdated norms, and ignoring cultural differences in test interpretation.
Evryday psychometric: affirmative action
Score adjustment techniques have been used historically to improve representation:
Adding points to minority scores
Differential cutoffs for different groups
Within-group norming (comparing individuals to their own group)
Banding (treating similar scores as equivalent).
These methods were developed to correct historical inequalities in testing. However, they remain highly controversial.
In the U.S., within-group norming and race-based score adjustments in employment testing were banned by the Civil Rights Act of 1991.