Factor Analysis Validation, Reliability Analysis, Scale Construction, and Construct Measurement

Step 7: Validating Factor Analysis Results

Validating a factor analysis model requires rigorous evaluation across internal consistency, outlier impact, and replicability across independent samples.

Internal Consistency Assessment

Conclude the preliminary factor structure evaluation by evaluating internal consistency for each extracted factor or constructed scale.

/

Outlier Assessment

Evaluate the sensitivity of the factor structure to extreme values by comparing analyses conducted with outliers included against analyses conducted with outliers excluded.

Replication Strategies

Replicate the exploratory factor structure using Confirmatory Factor Analysis (CFA) across independent datasets. This is performed using either:

  • Holdout Sample: A randomly selected subset of the original sample that was deliberately withheld from initial exploratory analyses.

  • New Sample: An independently collected, fresh dataset gathered from the target population.

Validation Steps and Objectives

Scale Construction Principles

Scale construction guidelines depend on the dimensionality of the underlying assessment instrument:

  1. Unidimensional Instruments: If an assessment instrument measures a single underlying construct, conduct an internal consistency analysis across all items in the full instrument.

  2. Multidimensional Instruments: If an assessment instrument measures multiple distinct constructs, conduct separate internal consistency analyses for each individual construct or scale. For example, in an instrument yielding a two-factor structure, separate scale reliability analyses must be performed for each of the two extracted scales.

Scale Construction Guidelines

Internal Consistency Analysis: Cronbach's Alpha and McDonald's Omega

Internal consistency quantifies the degree to which items within a scale measure the same underlying psychological construct.

Fundamental Logic

If a set of items measures a common target construct, individuals who score high on one item should logically score high on the remaining items within that scale.

Internal Consistency Concepts

Metrics and Evaluation Standards

  • Cronbach's Alpha (α\alpha): Historically the most widely reported metric. Cronbach's alpha is influenced by two main factors:

    1. Inter-Item Correlations: Higher average correlations among items generate higher alpha values.

    2. Scale Length: Holding inter-item correlations constant, scales with a greater number of items yield higher alpha values.

Standard Alpha Thresholds
  • α≥.80\alpha \ge .80: High internal consistency.

  • α≥.70\alpha \ge .70: Acceptable internal consistency.

  • α<.70\alpha < .70: Potentially problematic internal consistency, though frequently reported in published literature.

Cronbach's Alpha Considerations
  • McDonald's Omega (ω\omega): An increasingly preferred metric among contemporary methodologists. McDonald's omega provides a more accurate and representative estimate of internal consistency because it relies on factor loadings rather than assuming essential tau-equivalence (equal item weighting).

Handling Reverse-Scored Items

Before computing Cronbach's α\alpha or McDonald's ω\omega, ensure all scale items are scored in a uniform directional orientation.

  • Detection: Inspect the component or pattern loading matrices. Items loading negatively on a factor indicate opposite polarity.

  • Remediation: Recode negative-loading variables (reverse scoring) prior to running reliability algorithms.

  • Impact of Omission: Failing to reverse-score negatively loaded items depresses the magnitude of α\alpha and can produce artificially negative reliability coefficients.

Reverse Scoring and Internal Consistency

Factor Naming and Pattern Matrix Loadings

Examine item content loading onto derived components to define appropriate construct names.

Factor 1: Network Stressful Life Events (NSLE)

Items loading strongly on Component 1 reflect adverse events occurring to significant individuals within the participant's social network.

Pattern Matrix Part 1
Pattern Matrix Loadings for Component 1
  • Spouse's death: .371.371

  • Child's death: .360.360

  • Parents death: .361.361

  • Twins death: .386.386

  • Siblings death: .375.375

  • Relatives death: .460.460

  • Others death: .480.480

  • Spouses illness: .339.339

  • Childs illness: .410.410

  • Parents illness: .369.369

  • Twins illness: .436.436

  • Siblings illness: .420.420

  • Relatives illness: .490.490

  • Others illness: .526.526

  • Childs crisis: .400.400

  • Parents crisis: .504.504

  • Twins crisis: .451.451

  • Siblings crisis: .475.475

  • Relatives crisis: .509.509

  • Others crisis: .554.554

Factor 2: Personal Stressful Life Events (PSLE)

Items loading strongly on Component 2 reflect adverse life events directly experienced by the respondent.

Pattern Matrix Part 2
Pattern Matrix Loadings for Component 2
  • Problems with spouse: .542.542

  • Problems with twin: .497.497

  • Housemate problems: .472.472

  • Problems with friend: .453.453

  • Neighbour problems: .390.390

  • Problems with coworker: .454.454

  • Marital separation: .407.407

  • Broken relationship: .448.448

  • Separation: .456.456

  • Illness or injury: .367.367

  • Serious accident: .461.461

  • Robbed or burgled: .469.469

  • Laid off: .388.388

  • Work difficulties: .469.469

  • Financial problems: .516.516

  • Legal troubles: .400.400

  • Poor living conditions: .546.546

Note: Three items were excluded from the final solution due to univariate/multivariate outlier status or severe cross-loadings.

SPSS Reliability Analysis Procedures and Output Interpretation

SPSS Navigation Protocol

To calculate Cronbach's alpha for scale factors:

  1. Select Analyze -> Scale -> Reliability Analysis.

  2. Select all items loading onto the target factor and transfer them to the Items list.

  3. Verify that Alpha is selected in the Model dropdown menu.

  4. Click Statistics, and under Descriptives for, check Scale if item deleted.

  5. Click Continue, then OK.

SPSS Reliability Analysis Menu Setup

NSLE Scale Reliability Results

  • Overall Scale Cronbach's Alpha: α=.780\alpha = .780 (based on standardized items α=.780\alpha = .780, N=20N = 20 items). This value represents acceptable internal consistency.

  • Item Deletion Analysis: Reviewing the "Cronbach's Alpha if Item Deleted" column confirms that removing any single item does not increase overall scale reliability above .780.780. If an "alpha if item deleted" value exceeded .780.780, that item would be deleted to optimize scale reliability.

Reliability Statistics and Item-Total Statistics Output
Detailed Item-Total Statistics for NSLE Scale
  • Spouses death: Scale Mean if Deleted = 38.29938.299, Scale Variance if Deleted = 72.35172.351, Corrected Item-Total Correlation = .322.322, Squared Multiple Correlation = .177.177, Cronbach's Alpha if Deleted = .772.772

  • Childs death: Scale Mean if Deleted = 38.23238.232, Scale Variance if Deleted = 72.99472.994, Corrected Item-Total Correlation = .305.305, Squared Multiple Correlation = .158.158, Cronbach's Alpha if Deleted = .773.773

  • Parents death: Scale Mean if Deleted = 38.28038.280, Scale Variance if Deleted = 72.79572.795, Corrected Item-Total Correlation = .305.305, Squared Multiple Correlation = .151.151, Cronbach's Alpha if Deleted = .773.773

  • Twins death: Scale Mean if Deleted = 38.24438.244, Scale Variance if Deleted = 72.95572.955, Corrected Item-Total Correlation = .306.306, Squared Multiple Correlation = .153.153, Cronbach's Alpha if Deleted = .773.773

  • Siblings death: Scale Mean if Deleted = 38.24038.240, Scale Variance if Deleted = 72.50272.502, Corrected Item-Total Correlation = .322.322, Squared Multiple Correlation = .163.163, Cronbach's Alpha if Deleted = .772.772

  • Relatives death: Scale Mean if Deleted = 38.26938.269, Scale Variance if Deleted = 72.30172.301, Corrected Item-Total Correlation = .356.356, Squared Multiple Correlation = .263.263, Cronbach's Alpha if Deleted = .770.770

  • Others death: Scale Mean if Deleted = 38.20338.203, Scale Variance if Deleted = 72.08172.081, Corrected Item-Total Correlation = .361.361, Squared Multiple Correlation = .307.307, Cronbach's Alpha if Deleted = .770.770

  • Spouses illness: Scale Mean if Deleted = 38.24438.244, Scale Variance if Deleted = 73.08173.081, Corrected Item-Total Correlation = .285.285, Squared Multiple Correlation = .121.121, Cronbach's Alpha if Deleted = .775.775

  • Childs illness: Scale Mean if Deleted = 38.23638.236, Scale Variance if Deleted = 72.21872.218, Corrected Item-Total Correlation = .339.339, Squared Multiple Correlation = .172.172, Cronbach's Alpha if Deleted = .771.771

  • Parents illness: Scale Mean if Deleted = 38.22538.225, Scale Variance if Deleted = 73.06473.064, Corrected Item-Total Correlation = .291.291, Squared Multiple Correlation = .142.142, Cronbach's Alpha if Deleted = .774.774

  • Twins illness: Scale Mean if Deleted = 38.21438.214, Scale Variance if Deleted = 71.99171.991, Corrected Item-Total Correlation = .340.340, Squared Multiple Correlation = .154.154, Cronbach's Alpha if Deleted = .771.771

  • Siblings illness: Scale Mean if Deleted = 38.20338.203, Scale Variance if Deleted = 72.48172.481, Corrected Item-Total Correlation = .329.329, Squared Multiple Correlation = .180.180, Cronbach's Alpha if Deleted = .772.772

  • Relatives illness: Scale Mean if Deleted = 38.18138.181, Scale Variance if Deleted = 71.74971.749, Corrected Item-Total Correlation = .371.371, Squared Multiple Correlation = .285.285, Cronbach's Alpha if Deleted = .769.769

  • Others illness: Scale Mean if Deleted = 38.19938.199, Scale Variance if Deleted = 71.21271.212, Corrected Item-Total Correlation = .402.402, Squared Multiple Correlation = .350.350, Cronbach's Alpha if Deleted = .767.767

  • Childs crisis: Scale Mean if Deleted = 38.20338.203, Scale Variance if Deleted = 73.05973.059, Corrected Item-Total Correlation = .312.312, Squared Multiple Correlation = .146.146, Cronbach's Alpha if Deleted = .773.773

  • Parents crisis: Scale Mean if Deleted = 38.25838.258, Scale Variance if Deleted = 71.34071.340, Corrected Item-Total Correlation = .399.399, Squared Multiple Correlation = .230.230, Cronbach's Alpha if Deleted = .767.767

  • Twins crisis: Scale Mean if Deleted = 38.21438.214, Scale Variance if Deleted = 72.33272.332, Corrected Item-Total Correlation = .336.336, Squared Multiple Correlation = .175.175, Cronbach's Alpha if Deleted = .771.771

  • Siblings crisis: Scale Mean if Deleted = 38.23638.236, Scale Variance if Deleted = 72.25572.255, Corrected Item-Total Correlation = .358.358, Squared Multiple Correlation = .170.170, Cronbach's Alpha if Deleted = .770.770

  • Relatives crisis: Scale Mean if Deleted = 38.20738.207, Scale Variance if Deleted = 72.19472.194, Corrected Item-Total Correlation = .367.367, Squared Multiple Correlation = .251.251, Cronbach's Alpha if Deleted = .769.769

  • Others crisis: Scale Mean if Deleted = 38.17038.170, Scale Variance if Deleted = 71.12771.127, Corrected Item-Total Correlation = .414.414, Squared Multiple Correlation = .275.275, Cronbach's Alpha if Deleted = .766.766

Factor Score Computation and Applications

A factor score summarizes a participant's standing on an extracted latent factor into a single composite variable.

Factor Scores Overview

Computation Protocol (Scale Scores Approach)

Compute composite factor scores by calculating the arithmetic mean across all items loading on that factor for each case:

  1. Select Transform -> Compute Variable.

  2. Enter the target factor variable name (e.g., NSLE or PSLE) in Target Variable.

  3. In Function group, select Statistical, then double-click Mean in Functions and special variables.

  4. Move all loading items into the Numeric Expression field inside MEAN(...), separated by commas.

  5. Click OK.

Creating Factor Scores in SPSS

Descriptive Statistics for Computed Factor Scores

Sample size: N=300N = 300 valid cases, 00 missing values.

Metric

NSLE

PSLE

Mean

2.00212.0021

2.01312.0131

Median

1.95001.9500

1.94121.9412

Mode

2.052.05

1.821.82

Std. Deviation

.43687.43687

.45265.45265

Skewness

.629.629 (Std. Error = .141.141)

.422.422 (Std. Error = .141.141)

Kurtosis

.176.176 (Std. Error = .281.281)

−.357-.357 (Std. Error = .281.281)

Minimum

1.201.20

1.121.12

Maximum

3.353.35

3.413.41

Descriptive Statistics Table and Histograms for NSLE and PSLE

Downstream Research Applications

Factor scores reduce dimensionality (e.g., condensing 40 individual items into 2 scale scores), allowing researchers to:

  • Test demographic differences (e.g., comparing NSLE/PSLE across gender or age groups).

  • Analyze relationships with mental health constructs (e.g., depression and anxiety).

  • Construct longitudinal prediction models (e.g., assessing whether baseline resilience or self-efficacy predicts future NSLE or PSLE scores).

  • Utilize factor scores as predictors or outcome variables in structural or regression analyses (e.g., Network Stressful Life Events and Personal Stressful Life Events predicting Resilience).

Factor Scores as Variables Model Diagram

Simplified Demonstration Structure

In a standard 14-variable dataset (B1B1 through B14B14):

  • Variable B5B5 was dropped due to univariate/multivariate outlier status.

  • No items required reverse scoring.

  • Two components emerged:

    • Factor 1: Items B1,B2,B3,B4,B6,B7,B8,B9,B13B1, B2, B3, B4, B6, B7, B8, B9, B13

    • Factor 2: Items B10,B11,B12,B14B10, B11, B12, B14

  • Scale consistency across factors is verified using Split-half reliability, Cronbach's alpha, and McDonald's omega.

Measurement Error and Reliability Theory

Sources of Measurement Error in Testing

  1. Test Construction: Variance resulting from item sampling or content sampling across test items or test versions.

  2. Test Administration: Error variance originating from physical or psychological environmental conditions during administration.

  3. Test-Taker Variables: Participant-level noise including active emotional distress, physical discomfort, sleep deprivation, or pharmacological/medication effects.

  4. Examiner Variables: Variations in administrator demeanor, physical appearance, attention to protocol detail, and professional conduct.

  5. Test Scoring and Interpretation: Subjectivity, coding errors, or inconsistent rating criteria.

Reliability Quantification

Reliability reflects measurement consistency. It is quantified via a reliability coefficient, ranging from 00 (completely unreliable) to 11 (perfect reliability).

Methods for Estimating Reliability

  1. Test-Retest Reliability: Calculated by correlating scores from two distinct administrations of the same test to the same group.

    • Best suited for stable, enduring traits (e.g., personality traits).

    • Unsuitable for dynamic constructs expected to fluctuate over time.

    • Confounded by practice effects, memory recall, temporal changes, and environment shifts.

    • Over intervals exceeding 6 months, test-retest reliability is defined as the coefficient of stability.

  2. Alternate-Forms and Parallel-Forms Reliability: Quantified via the coefficient of equivalence by administering two test forms to the same group.

    • Parallel Forms: Test forms share identical score means and variances.

    • Alternate Forms: Test forms share equivalent content coverage and difficulty levels.

    • Confounded by fatigue, practice effects, and item sampling variation.

  3. Split-Half Reliability: Derived by administering a single test once, splitting it into equivalent halves, and correlating half-scores.

    • Step 1: Divide test into two equivalent halves.

    • Step 2: Compute Pearson correlation (rr) between half-scores.

    • Step 3: Adjust the correlation using the Spearman-Brown formula to estimate total test internal consistency.

  4. Inter-Item Consistency: Measures correlation among all scale items simultaneously.

    • Cronbach's Alpha (α\alpha): Mean of all possible split-half correlations corrected by Spearman-Brown formula (0≤α≤10 \le \alpha \le 1).

    • McDonald's Omega (ω\omega): Calculates reliability from factor loadings, avoiding the restrictive assumption of equal item weighting.

  5. Inter-Scorer Reliability: Measures scoring consistency across two or more independent raters/judges.

    • Essential for observational and nonverbal behavior coding.

    • Prevents individual observer biases or idiosyncratic scoring.

    • Quantified using the coefficient of inter-scorer reliability.

Test Characteristics Influencing Reliability Evaluation

Reliability choice depends on:

  • Item homogeneity vs. heterogeneity.

  • Trait nature (dynamic vs. static).

  • Presence or absence of score range restriction.

  • Speed tests vs. power tests.

  • Criterion-referenced vs. norm-referenced testing.

Limitations of Time-Dependent Reliability Indices

  • Psychological traits fluctuate or experience constant flux over time.

  • Measurement interference produces carryover effects, practice/learning gains, and testing fatigue.

Validity Theory and Assessment Methods

Validity represents an overall evaluative judgment regarding how well a test measures its intended theoretical construct within a specific population, context, and time period. Validation is the formal process of gathering empirical and theoretical support. Local validation studies compare specific target populations against standardized norming samples.

Primary Categories of Validity

1. Content Validity

Evaluates whether test items adequately sample the full domain of behavior or knowledge the test was designed to measure.

  • Test Blueprint: A structured matrix defining topic coverage, item allocation per area, item organization, and cognitive target levels.

2. Face Validity

A superficial judgment regarding whether test items appear relevant on their face. Evaluated during initial test construction by soliciting feedback from panels of domain experts.

3. Criterion-Related Validity

Evaluates how effectively a test score predicts an individual's standing on an independent outcome benchmark (criterion).

  • Concurrent Validity: Test score and criterion measure are obtained simultaneously.

  • Predictive Validity: Test score predicts a criterion measured at a future point in time.

  • Properties of an Adequate Criterion:

    • Relevant: Directly applicable to the domain.

    • Valid: Validated for the specific predictive purpose.

    • Uncontaminated: Measured independently of predictor variables (e.g., preventing identical raters from scoring both predictor and criterion).

  • Validity Metrics:

    • Validity Coefficient: Correlation coefficient (rr) linking test score to criterion.

    • Incremental Validity: Additional variance explained in the criterion score when adding a new predictor to an existing set (evaluated via hierarchical regression).

4. Construct Validity

The overarching framework integrating all validity evidence to confirm that test scores reflect the theoretical construct. High and low scorers must behave strictly as theoretical models predict.

Sources of Evidence for Construct Validity

  • Homogeneity Evidence: Internal uniform measurement of a single concept.

  • Developmental/Age-Related Changes: Scores shift systematically across age brackets as predicted by developmental theory.

  • Pretest-Posttest Changes: Scores change systematically following targeted intervention or therapy.

  • Distinct Groups Evidence: Test scores discriminate predictably between distinct demographic or clinical populations.

  • Convergent Evidence: Test scores correlate highly with established measures of the same or similar constructs.

  • Discriminant Evidence: Validity coefficients demonstrate negligible or non-significant correlations with constructs theoretically unrelated to the target measure.

  • Factor Analysis: Mathematical dimensionality reduction demonstrating that items group into factor structures aligned with theory.

Factor Analysis Reporting Protocols and Sample Write-Up

8 Required Elements in a Factor Analysis Report

  1. Purpose Statement: Clear outline of the analysis objective.

  2. Variable Description and Factorability: Detailed description of input variables and tests of factorability (e.g., KMO, Bartlett's Test of Sphericity).

  3. Outlier Handling: Documentation of univariate and multivariate outlier evaluations, specifying whether cases/items were retained, deleted, or recoded.

  4. Normality and Transformations: Evaluation of distributional normality and details of transformations performed prior to factor analysis.

  5. Analytical Decisions: Explanation of factor retention rules, rotation techniques, and interpretive thresholds for factor naming.

  6. Summary Table: Comprehensive table presenting factors, percentage of variance explained, item loadings, and inter-factor correlations (for oblique rotations).

  7. Scale Score and Reliability Metrics: Procedure used to construct composite scale scores, reporting Cronbach's α\alpha, McDonald's ω\omega, and split-half reliability.

  8. Substantive Interpretation: Clear plain-English summary of what the statistical findings mean.

Sample Principal Components Analysis Results Section

To determine the factor structure of our new stressful events measure, we conducted an exploratory Principal Components Analysis on the 40 items from the scale. The Kaiser-Meyer-Olkin measure of sampling adequacy was .79.79, indicating that there were strong linear relationships within the item set. There were no outliers for any of the items, however two items had both communalities below .10.10 and did not load above .30.30 on either of the retained factors. Additionally one item loaded above .30.30 on both factors. These three items were therefore omitted from the final analysis.

Cattell's (1966) scree test indicated that two components should be retained. The solution was rotated using the direct oblimin approach, with Δ\Delta set to 00 to permit moderate correlations among the components. The solution exhibited good simple structure with pattern-matrix loadings greater than .40.40 for 25 of the 37 items and no cross-loadings above .30.30. Examination of the loadings indicated that the first component assessed stressful life events impacting on other people in the individual's network, and the second assessed stressful life events that had happened to them personally. The correlation between the two retained components was small and negative (r=−.10r = -.10), suggesting that they were assessing two quite distinct constructs. Together the two factors accounted for 20%20\% of the variance in the data.

Scale scores were computed by averaging the items that loaded above .30.30 on each component. Cronbach's α\alpha for both scales exceeded .75.75 and McDonald's ω\omega exceeded .70.70, reflecting adequate internal consistency. Split-half reliability was good, with both components exceeding .80.80.