Validity & Test Construction

Definition (#f7aeae)

Important (#edcae9)

Extra (#fffe9d)


Validity:

  • How well a test measures what it is supposed to measure.

  • A test is only valid if its scores are meaningful and appropriate for a specific purpose.

  • Key point: No test is valid for all people, all uses, or forever—validity depends on context, population, and intended application.


Why ‘valid test’ is misleading:

  • Tests are not universally valid.

  • Only the inferences we draw from test scores can be considered valid or invalid.

  • A test may be valid for diagnosing depression but not for predicting suicide risk.

  • Norms that were valid in 20026 may be outdated by 2036.

  • Validity is always context-dependent, population-specific and time limited.


Validation: the process

  • Test developers provide validity evidence in test manuals.

    • Test users may conduct local validation studies when needed for their specific context or population.

  • Local validation is essential when:

    • The test is translated to another language, items or instructions are modified, or the test is used with a new population not represented in original norms.


3 types of validity: (Trinitarian model)

  1. Content Validity: Ensures the test covers the entire domain of knowledge, skills, or behaviors it claims to measure.

  2. Criterion-Related Validity: Examines how well scores correlate with external criteria like academic success.

  3. Construct Validity: Evaluates whether the test captures the underlying psychological concept it intends to assess.


Validity diagram:

  • The "umbrella model" views construct validity as the overarching framework that encompasses all validity evidence.

  • Content validity and criterion-related validity are not separate types, but rather different sources of evidence that support construct validity.

  • This unified approach recognizes that all validation efforts ultimately aim to demonstrate that test scores meaningfully measure the intended psychological construct.


Face validity:

  • Refers to the perceived relevance of test items from the test-taker's perspective, not actual statistical validity.


Content validity:

  • Ensures comprehensive coverage of all relevant areas.

  • Determined by the degree to which the questions, tasks, or items on a test are representative of the universe of behavior the test was designed to sample.

  • Quantification of content validity.

  • Content validity = D/(A + B + C + D).


Test blueprint:

  • Ensures representative content sampling across all relevant domains.

  • For an Assertiveness Test, items are distributed systematically:

  • Home: 10 items (25%) - family interactions, domestic situations.

  • Work: 15 items (40%) - professional settings, colleague relations.

  • Social context: 10 items (25%) - friendships, public situations.

  • Miscellaneous: 5 items (10%) - other relevant contexts.

  • This structured approach guides item development and ensures comprehensive domain coverage.


Lawshe’s content validity ratio: (CVR)

  • Experts rate each item as Essential.

  • Useful but not essential, or not needed.

  • Formula: CVR = (ne – N/2) / (N/2)

  • Where ne = number rating "Essential" and N = total experts.

  • Ex:

    • If 9 of 10 experts rate an item Essential → CVR = 0.80 → Accept item.

    • If only 4 of 10 rate Essential → CVR = -0.20 → Reject item from the test.


Criterion-related validity:

  • Concurrent validity:

    • Test and criterion measured at the same time.

    • Ex: A new depression screening tool is validated against an established clinical diagnosis made on the same day.

  • Predictive validity:

    • Test scores predict future performance or outcomes.

    • Ex: SAT scores administered in high school predict college GPA years later.


Validity coefficient:

  • The correlation between test scores and criterion measures.

  • Typical values range from 0.00 to 0.50, and rarely exceed 0.60.

  • A coefficient must be high enough to meaningfully aid decision-making in practical contexts.

  • Context matters when interpreting these values—a coefficient of 0.30 may be excellent for personality assessment but inadequate for high-stakes selection decisions.


Expectancy table:

  • Scores 90–100 → 80% likelihood of achieving GPA > 3.0.

  • Scores 70–89 → 60% likelihood.

  • Scores 50–69 → 30% likelihood.

  • Scores 30–49 → only 10% likelihood of success.

  • Expectancy tables translate test scores into probability statements, showing the likelihood of achieving a specific criterion (ex: GPA above 3.0) based on score ranges.

  • These tables have practical utility in education (predicting academic success), hiring decisions (forecasting job performance), and establishing meaningful cutoff scores for selection.

  • Ratings range from Poor to Excellent, with each higher rating corresponding to a greater probability of job success.

  • This visual tool helps employers make informed hiring decisions.

  • Bar or line charts illustrate the trend clearly:

    • Poor ratings may show 20% success likelihood.

    • Excellent ratings demonstrate 80% or higher probability of successful job performance.

Expectancy Table


Taylor-Russell tables:

  • Demonstrate how 3 factors combine to determine hiring success:

    • The validity coefficient: How well the test predicts job performance.

    • The base rate: Percentage who would succeed without testing.

    • The selection ratio: Proportion of applicants hired.

  • Higher predictive validity significantly increases the accuracy of hiring decisions.

  • Key insight: Tests become most useful when the selection ratio is low— when hiring few applicants from a large pool.

  • In competitive hiring scenarios, modest validity coefficients can substantially improve the proportion of successful hires compared to random selection, making validation studies essential for organizational decision-making.


Incremental validity:

  • Refers to the additional predictive power a new measure provides beyond existing predictors.

  • This concept is crucial when selecting which tests or measures to include in assessment batteries.

  • For ex:

    • When predicting exam performance, consider 3 predictors: time spent studying, time in the library, and sleep quality before exams.

    • If sleep quality explains variance not captured by study time alone, it demonstrates high incremental validity.


Construct validity:

  • Addresses whether a test truly measures the underlying psychological concept it claims to assess.

  • Constructs are unobservable theoretical entities—such as intelligence, anxiety—that cannot be directly seen or touched.

  • Establishing construct validity requires a solid theoretical foundation, testable hypotheses, and multiple sources of converging evidence.

  • Includes examining internal consistency, correlations with related measures, group differences, and factor structure.

  • Serves as the foundation for meaningful test interpretation and is considered by many to be the overarching framework encompassing all other validity types.


Construct validity map:

  • Internal Evidence:

    • Homogeneity and internal consistency demonstrate items measure the same construct.

    • Age-related changes show expected developmental patterns; pre/post treatment changes confirm sensitivity to interventions.

  • External Evidence:

    • Known group differences validate the test distinguishes between relevant populations; convergent validity shows correlation with similar measures; discriminant validity confirms low correlation with unrelated constructs.

  • Statistical Evidence:

    • Factor analysis identifies underlying dimensions and confirms the test measures intended constructs.

    • A comprehensive approach integrating all evidence sources strengthens construct validity claims.


Convergent validity:

  • High correlations with other tests measuring the same construct demonstrate convergent validity.

  • Ex: Test A ↔ Test B (r = .78) shows strong agreement between similar measures.

  • Ex: Test A ↔ Test C (r = .63) further supports convergent validity.

  • Multiple high correlations with theoretically related measures indicate the test truly captures the intended construct.


Discriminant validity:

  • A test should show low or near-zero correlations with measures of unrelated constructs.

  • Low correlations with unrelated constructs confirm the test measures its intended construct specifically, not general response tendencies or irrelevant factors.


Factor analysis:

  • Identifies the underlying dimensions measured by a test, confirming it measures intended constructs.

  • Ex: A mood assessment might reveal:

    • Factor 1: Anxiety (Items 1–4 load highly)

    • Factor 2: Depression (Items 5–7 load highly)

  • This statistical technique helps researchers refine and validate test structure by showing which items cluster together.


Validity, bias & fairness:

  • Validity ≠ Fairness, Validity ≠ Bias: These concepts are related but distinct.

  • A test may be:

    • Valid but unfair: Accurately measures construct but used inappropriately.

    • Fair but invalid: Equal treatment but poor measurement.

    • Biased but predictive: Systematic error yet still useful.

    • Unbiased but misused: Technically sound but applied incorrectly).

  • Ethical Testing Requirements: Responsible psychological assessment requires 3 integrated components:

    • Valid scores that accurately measure intended constructs.

    • Fair use that treats all test-takers equitably.

    • Cultural awareness that recognizes diverse backgrounds and contexts.

  • All 3 must work together for ethical practice.


Test bias:

  • Slope bias: Different regression slopes between groups mean the test predicts outcomes differently for each group. A test may predict well for one group but poorly for another.

  • Intercept bias: Same slope but different intercepts means the test systematically over or under-predicts for certain groups.

  • Both types impact fairness and validity of test use.


Rating errors:

  • Central tendency error: Raters avoid using extreme scores, clustering all ratings around the middle of the scale. This reduces score variability and masks true differences between individuals being assessed.

  • Halo effect: When one positive characteristic influences the rater's perception of unrelated traits. A person seen as friendly may be rated higher on competence, even without evidence of actual skill.

  • Leniency/severity errors: Some raters consistently give overly generous scores (leniency) while others are habitually harsh (severity). Rater training and calibration sessions are essential to minimize these biases.


Test fairness:

  • Fairness is essential alongside validity.

  • A test may produce valid scores yet still be used unfairly.

  • Ethical testing requires impartial, just application that considers the test-taker's background and context.

  • Examples of unfair test uses include administering IQ tests in non-native languages, using punitive testing practices, relying on outdated norms, and ignoring cultural differences in test interpretation.


Evryday psychometric: affirmative action

  • Score adjustment techniques have been used historically to improve representation:

    • Adding points to minority scores

    • Differential cutoffs for different groups

    • Within-group norming (comparing individuals to their own group)

    • Banding (treating similar scores as equivalent).

  • These methods were developed to correct historical inequalities in testing. However, they remain highly controversial.

  • In the U.S., within-group norming and race-based score adjustments in employment testing were banned by the Civil Rights Act of 1991.