Chapter 6
Validity and Test Bias
PSYC 385 Psychological Test - Fall 2025
Outline
Concept of Validity
Types of Validity
Face Validity
Content Validity
Criterion-Related Validity
Construct Validity
Validity and Test Bias
Learning Objectives
Three primary learning objectives:
Understand the steps involved in establishing validity.
Decipher the bounds and the strength a measure’s validity through a thorough understanding of the types of validity.
Identify potential sources of test bias.
The Concept of Validity
Validity: A judgment or estimate of how well a test measures what it is intended to measure within a particular context.
Validation: The process of gathering and evaluating evidence about validity.
Both test developers and test users may play a role in the validation of a test.
Test users may validate a test with their own group of test-takers – a process known as local validation.
Types of Validity
Validity is often conceptualized according to three categories:
Content Validity:
Evaluation of the subjects, topics, or content covered by the items in the test.
Criterion-Related Validity:
Evaluating the relationship of scores obtained on the test to scores on other tests or measures.
Construct Validity:
This is a measure of validity that is arrived at by executing a comprehensive analysis of:
a. How scores on the test relate to other test scores and measures.
b. How test scores can be interpreted within a theoretical framework that explains the construct the test was designed to measure.
Face Validity
Face Validity: A judgment concerning how relevant the test items appear to be in measuring what is intended.
If a test appears to measure what it is supposed to be measuring “on the face of it,” it is considered to be high in face validity.
A perceived lack of face validity may contribute to a lack of confidence in the test.
Content Validity
Content Validity: How well a test samples behaviors that are representative of the broader set of behaviors it was designed to measure.
Key questions:
Do the test items adequately represent the content that should be included in the test?
Test blueprint: A plan regarding the types of information to be covered by the items, the number of items tapping each area of coverage, and the organization of the items in the test, among others.
Content validity is typically established by recruiting a subject matter expert and obtaining insights about item importance and scrutiny regarding what is missing from the measure.
It is important to remember that content validity of a test can vary across cultures and over time.
Criterion-Related Validity
A criterion is the standard against which a test or a test score is evaluated.
Characteristics of a Criterion:
An adequate criterion is relevant for the matter at hand, suitable for the purpose for which it is being used, and independent, meaning it is not part of the predictor.
Criterion-Related Validity: A judgment of how adequately a test score can be used to infer an individual’s most probable standing on some measure of interest (the criterion).
Concurrent Validity: An index of the degree to which a test score is related to some criterion measure obtained at the same time (concurrently).
Predictive Validity: An index of the degree to which a test score predicts some criterion, or outcome, measure in the future. Tests are evaluated for their predictive validity.
Predictive Validity Considerations
Base Rate: The extent to which the phenomenon exists in the population.
Hit Rate: Accurate identification; includes True-positive and True-negative.
Miss Rate: Failure to identify accurately.
False-positive: Incorrectly identifying a presence of the characteristic.
False-negative: Incorrectly identifying the absence of the characteristic.
Validity Coefficient and Incremental Validity
The validity coefficient: A correlation coefficient between test scores and scores on the criterion measure.
Validity coefficients are affected by restriction or inflation of range.
Incremental Validity: The degree to which an additional predictor explains something about the criterion measure beyond what is already explained by other predictors.
To what extent does a test predict the criterion over and above other variables?
Construct Validity
Construct Validity: The ability of a test to measure a theorized construct (e.g., intelligence, aggression, personality, etc.) that it aims to measure.
If a test is a valid measure of a construct, high scorers and low scorers should behave as theorized.
Construct validity serves as an umbrella term; content and criterion-related validity provide evidence for construct validity and expand its insights.
Evidence of Construct Validity
Evidence of Homogeneity: How consistent a test is in measuring a single concept.
Evidence of Changes: Some constructs are expected to change over time (e.g., reading rate).
Evidence of Pretest/Posttest Changes: Test scores change as a result of some experience between a pretest and a posttest (e.g., after therapy).
Evidence from Distinct Groups: Scores on a test vary predictably based on membership in specific groups (e.g., scores on the Psychopathy Checklist for prisoners vs. civilians).
Convergent Evidence: Correlates highly in the predicted direction with scores on established tests designed to measure the same constructs.
Discriminant Evidence: Demonstrates little relationship between test scores and other variables with which scores on the test theoretically should not be correlated.
Factor Analysis: A new test should load on a common factor with other tests of the same construct.
CLOSER LOOK: Deployment Communication Inventory
Table 2: Correlation among Deployment Communication Inventory Scales.
Frequency
Assurance/Support
PS/Disclosure
Conflict
Perceived Benefits
Perceived Costs
*Note: Correlations are significant if *p* < .05. Soldier correlations are below the diagonal and partner correlations are above the diagonal.*
Correlations of Deployment Communication Inventory and Relationship, Family, and Individual Functioning Scales
Table 4 - Relationships:
Relationship Functioning: DAS (Dyadic Adjustment Scale)
Individual Mental Health: PTSD symptoms, Depressive symptoms, Alcohol use.
A Synthetic Multitrait-Multimethod Matrix
Matrix indicates various correlation coefficients between methods and traits relevant to construct measurement.
Validity and Test Bias
Bias: A factor inherent in a test that systematically prevents accurate, impartial measurement.
Bias implies systematic variation in test scores.
Prevention during test development is the best remedy for test bias.
Rating Error: A judgment resulting from the intentional or unintentional misuse of a rating scale. Raters may display biases such as:
Being too lenient
Being too severe
Reluctance to give extreme ratings (central tendency error)
Halo Effect: A tendency to give one individual a higher rating than they objectively deserve due to a positive impression on another dimension.
Fairness: The extent to which a test is used in an impartial, just, and equitable manner.