Construct Validity, Incremental Validity, and Factor Analysis Notes
Applied Measurement and Incremental Validity
Applied Nature of Psychological Testing: Psychological and educational measures are applied tools designed to solve explicit, real-world problems (e.g., selection, placement, and diagnostic issues) rather than existing purely as abstract theoretical exercises.
Practical Selection Case Study (Research Assistant Selection):
- Context: Faculty members managing research grants require research assistants for essential laboratory tasks, including data entry, data collection, data analysis, and software execution.
- Criteria & Predictors: Applicants' performance scores in an introductory statistics course serve as a predictive selection measure.
- Required Skill Alignment: High statistics scores indicate mastery of research methodology, proficiency in operating statistical software like SAS to crunch numeric data, and the ability to evaluate statistical significance from output files.
- Decision Logic: High scorers on statistics tests are systematically projected to perform research assistant duties far more effectively than low scorers.
Incremental Validity & Explanatory Power:
- Definition: Incremental validity measures the exact increase in decision accuracy or explanatory power provided by introducing a new predictor over an existing decision-making process.
- Baseline Selection Dynamics: Current university selection processes often rely on baseline criteria (e.g., minimum application requirements, GPA, or SAT scores). Under standard baseline selection, approximately (slightly under half) of admitted students graduate within a -year timeframe.
- Evaluation of New Predictors: To test incremental validity, an institution might introduce a supplemental assessment—such as a group problem-solving task scored on collaborative performance. Incremental validity is formally demonstrated if incorporating this task leads to a statistically measurable increase in predicting student success above the standard baseline.
The Construct Validity Framework
Definition of Construct Validity:
- Construct validity evaluates whether a test truly measures the theoretical construct it claims to measure.
- It cannot be established by a single study or metric; it relies on a cumulative, overarching body of empirical evidence.
Theoretical Constructs:
- Constructs are unobservable, latent theoretical attributes, states, or traits (e.g., intelligence, introversion, depression, shyness).
Integrated Sources of Construct Validity Evidence:
- Content Validity: Ensures the test items fully cover all theoretical components defining the construct.
- Criterion-Related Validity: Demonstrates that test scores accurately predict specific empirical behaviors or criteria.
- Convergent Evidence: Demonstrates high correlations with independent measures targeting the same construct.
- Discriminant Evidence: Demonstrates low or non-existent correlations with measures targeting conceptually distinct constructs.
- Factor Analysis: Evaluates internal structural relationships among test items to verify underlying latent dimensions.
Operationalizing and Assessing Shyness
Construct Definition of Shyness:
- Theoretical definition (Cheek & Buss): Characterized by experiencing affective anxiety and behavioral inhibition in social situations.
- Context Specificity: Primary triggers involve novel, unique social environments or interactions with unfamiliar individuals (strangers) and authority figures; social anxiety is markedly lower in familiar environments or around family members.
Item Content Analysis (Revised Cheek and Buss Shyness Scale):
- The Revised Cheek and Buss Shyness Scale uses concise item counts (, , , or items; the -item version is commonly selected for research efficiency).
- Affective/Anxiety Items:
- "I feel tense when I'm with people I don't know well."
- Having doubts regarding social confidence.
- Feeling nervous when speaking to individuals in authority.
- Feeling uncomfortable at parties.
- Behavioral Inhibition Items:
- Having trouble looking someone directly in the eye.
- Finding it difficult to ask other people for information.
- Experiencing social awkwardness and finding it hard to act naturally.
- Finding it difficult to talk to strangers.
Criterion-Related Evidence for Shyness:
- Observational Study Design: Administer the shyness scale to participants, then place them in a standardized behavioral scenario involving interaction with a stranger.
- Empirical Prediction: Individuals scoring high in shyness will exhibit measurable behavioral inhibition (e.g., speaking significantly less or refraining from initiating dialogue), whereas low-shyness individuals will speak more frequently.
Convergent and Discriminant Evidence
Convergent Evidence:
- Requires positive correlations between the target scale and alternative scales measuring the same or overlapping constructs.
- Comparative Shyness Instruments:
- Cheek and Buss Shyness Scale.
- McCroskey Shyness Scale.
- Social Avoidance and Distress Scale (SAD).
- Interaction Anxiousness Scale (IAS): Assesses pure subjective/affective anxiety while explicitly omitting external behavioral inhibition (capturing individuals who appear outwardly calm but experience severe internal nervousness).
Discriminant Evidence:
- Requires demonstrating that a test does not correlate with constructs from which it should theoretically diverge.
- Shyness vs. Intelligence: Theoretically unrelated; empirical correlations should approach .
- Shyness vs. Extraversion/Introversion:
- Introversion: Represents a structural preference for low-stimulation environments (e.g., preferring a quiet dinner with or close friends or staying home to read a book over attending a loud party with strangers). It does not stem from fear or anxiety.
- Shyness: Involves genuine distress, fear, and high anxiety in social settings.
- Shyness vs. Neuroticism (Adjustment):
- Neuroticism: Reflects general negative affectivity, chronic worry, and unspecific anxiety not bound to social contexts.
- Shyness: Specifically ties negative affect and behavioral inhibition to social contexts.
- Empirical Correlation Benchmarks:
- The Cheek and Buss Shyness Scale demonstrates moderate correlations ranging between and when evaluated against Introversion and Neuroticism measures.
- These moderate values confirm partial conceptual overlap while remaining low enough to prove construct distinctness. Correlations in the range of to would indicate a failure of discriminant validity, signifying that the scale merely measures general introversion or neuroticism.
Structural Evaluation: Internal Consistency and Factor Analysis
Internal Consistency Evaluation:
- Assesses whether all items on a test measure a single, unified construct.
- Split-Half Reliability: A subtype of internal consistency where test items are divided into two equal halves to calculate their correlation.
- Cronbach's Coefficient Alpha (): Calculates the mean of all possible split-half correlation combinations across a test.
Factor Analysis Principles:
- Goes beyond basic internal consistency by identifying specific sub-clusters of items within a multi-item instrument.
- Factors: The underlying latent concepts, domains, or constructs assessed by distinct groups of test items.
- Factor Loadings: Represent the correlation coefficients between individual test items and a specific underlying factor.
Subtypes of Factor Analysis:
- Exploratory Factor Analysis (EFA): An inductive approach used to uncover hidden factor structures without prior mathematical assumptions.
- Confirmatory Factor Analysis (CFA): A deductive hypothesis-testing approach used to verify whether data conforms to a pre-defined theoretical factor structure.
Empirical Demonstration: Correlation Matrices and Factor Structures
Correlation Matrix Properties:
- A symmetric mathematical table displaying pairwise correlation coefficients () across multiple continuous test variables.
- The identity diagonal displays values of (representing each variable correlated with itself).
- The upper and lower triangles mirror each other identically across the diagonal.
Six-Variable Academic Subject Matrix Example:
- Evaluates student performance across academic disciplines: Algebra, Biology, Calculus, Chemistry, Geology, and Statistics.
- Empirical Pairwise Correlations ():
- Algebra and Calculus: (Strong positive mathematical relationship).
- Algebra and Biology: (Weak relationship).
- Algebra and Chemistry: (Weak relationship).
- Algebra and Geology: (Weak relationship).
- Algebra and Statistics: (Moderate relationship; weaker than Calculus because statistics emphasizes theoretical sampling and probability rather than pure algebraic manipulations).
- Chemistry and Biology: (Strong positive natural science relationship).
- Chemistry and Geology: (Strong positive natural science relationship).
- Statistics and Calculus: (Moderate positive relationship).
- Factor Structure Discovery: Visual inspection reveals two distinct empirical latent factors: a Mathematics Factor (Algebra, Calculus, Statistics) and a Natural Science Factor (Biology, Chemistry, Geology).
Matrix Dimensionality and Scaling:
- A -variable test produces a correlation matrix.
- A -item scale (e.g., Cheek and Buss) produces a correlation matrix.
- A -item scale (e.g., NEO Five-Factor Inventory short form) produces a matrix, necessitating computer-driven factor analysis algorithms.
Applied CFA Selection Example:
- An employment interview battery contains items designed to measure job performance dimensions: Communication ( items), Decision Making ( items), Problem Solving ( items), and Customer Focus ( items).
- Confirmatory Factor Analysis evaluates whether the empirical item responses map directly onto these hypothesized latent factor loadings.
Questions and Discussion
Question: What is a specific, named measure of reliability?
Response: Split-half reliability.
Question: Split-half reliability is a subtype of what broader class of reliability?
Response: Internal consistency reliability.