Chapter 5

Reliability in Assessment

Overview

  • Course: PSYC 385 Psychological Test

  • Semester: Fall 2025

Outline

  • Concept of Reliability

  • Sources of Error Variance

  • Types of Reliability Estimates

  • True Score and Alternative Models

Learning Objectives

  • Understand the concept of reliability and identify different reliability estimates.

  • Be able to interpret the usefulness of different estimates of reliability depending on the context and intent of the psychological test.

  • Interpret standard error terms and understand their importance in creating confidence intervals.

The Concept of Reliability

  • Reliability:

    • Defined as the or _ in measurement.

    • Reliability coefficient is an index of reliability, defined as a proportion that indicates the ratio between the true score variance on a test and the total variance.

    • The observed score can be represented as follows: (X=T+E)(X = T + E)

      • Where:

      • XX represents the observed score.

      • TT represents the true score.

      • EE signifies the error associated with measurement.

    • Error: Refers to the component of the observed score that does not pertain to the test taker’s true ability or the trait being measured.

Reliability Estimates

  • Variance:

    • Variance is defined as the standard deviation squared.

    • The total variance can be expressed as:
      (extVariance=extTrueVariance+extErrorVariance)( ext{Variance} = ext{True Variance} + ext{Error Variance})

    • Reliability:

    • Defined as the proportion of the total variance attributed to true variance.

    • Measurement Error:

    • Refers to all factors associated with the measurement process that do not relate to the variable being measured.

The Concept of Measurement Error

  • Random Error:

    • A source of error in measuring a targeted variable caused by and inconsistencies of other variables in the measurement process (i.e., noise).

  • Systematic Error:

    • A source of error that is typically constant or proportionate to the presumed true value of the variable being measured.

Sources of Error Variance

  • Potential Sources of Error Variance:

    • Example: Duck or Rabbit?

  • Test Construction:

    • Variation may exist within items on a test or between different tests (also known as item sampling or content sampling).

  • Test Administration:

    • Error may arise from the testing environment and situational factors, such as pressing emotional issues, physical discomfort, lack of sleep, or drug effects.

    • Examiner-related variables—including physical appearance, training, and demeanor—can also contribute.

  • Test Scoring and Interpretation:

    • Computer testing can reduce scoring errors, but many tests still require expert interpretation (e.g., projective tests).

    • Subjectivity in scoring may influence behavioral assessments.

  • Other Sources of Error:

    • Sampling Error:

    • Represents the extent to which the sample in a study is truly representative of the population.

    • Methodological Error:

    • Includes factors like ambiguous wording in questionnaires and biased question framing.

Reliability Estimates

  • Test-Retest Reliability:

    • An estimate obtained by correlating pairs of scores from the same individual on two different administrations of the same test.

    • Most applicable for variables expected to be stable over time (e.g., personality) and not suitable for those expected to fluctuate (e.g., mood).

    • Reliability estimates tend to decrease as intervals broaden, known as the coefficient of stability when intervals exceed six months.

Test-Retest Reliability Details

  • Execution:

    • Same person, same test at two different time points illustrated as:

    • Time 1: 3

    • Time 2: 8

Reliability Estimates via Alternate Forms

  • Parallel-Forms or Alternate-Forms:

    • Coefficient of equivalence measures the relationship degree between various test forms.

    • Parallel Forms:

      • Mean and variance of observed test scores are equal across forms.

    • Alternate Forms:

      • Different test versions that do not meet parallel forms' strict requirements but maintain similar content and difficulty.

    • Reliability is verified by administering two forms to the same group, acknowledging potential error factors related to test-takers' states (practice, fatigue, etc.) or item sampling.

Split-Half Reliability

  • Split-half reliability is obtained through correlating scores from two halves of a single administered test.

  • Process Steps:

    • Step 1: Divide the test into two halves (e.g., odd vs. even items).

    • Step 2: Calculate Pearson r between scores on both halves.

    • Step 3: Adjust reliability using the Spearman-Brown formula, estimating internal consistency reliability from the correlation of both halves.

Other Methods of Estimating Internal Consistency

  • Inter-Item Consistency:

    • Assesses the relatedness of items within a test to gauge its _.

    • Kuder-Richardson Formula 20:

    • A preferred statistic for determining inter-item consistency of .

    • Coefficient Alpha:

    • Represents the mean of all possible split-half correlations, adjusted via the Spearman-Brown formula.

    • Popular for estimating internal consistency with values ranging from 0 to 1.

Coefficient Alpha Ratings

Coefficient Alpha (α)

Interpretation

≥ 0.9

Excellent

0.9 ≥ α ≥ 0.8

Good

0.8 > α ≥ 0.7

Acceptable

0.7 > α ≥ 0.6

Questionable

0.6 > α ≥ 0.5

Poor

0.5 > α

Unacceptable

Measures of Inter-Scorer or Inter-Rater Reliability

  • Inter-Rater Reliability:

    • Assesses the degree of agreement among raters with respect to a particular measure.

    • Commonly applied in behavioral measures to minimize biases or idiosyncrasies in scoring.

    • Coefficient of Inter-Score Reliability:

    • Correlates scores from different raters.

Practical Example




  • Table 1. Scale Characteristics for the Deployment Communication Inventory.


    Scale

    Mean

    SD

    Range

    α

    r



    Frequency

    25.12

    5.11

    13-42

    --

    1



    Assurance

    19.09

    4.21

    5-25

    .89

    .64



    Conflict

    8.41

    3.43

    5-22

    .86

    .56



    PS/Disclosure

    16.46

    3.92

    6-25

    .83

    .50



    Perceived benefits

    32.87

    5.49

    8-40

    .91

    .56



    Perceived costs

    9.50

    3.85

    4-20

    .83

    .55









    Reliability Estimates Overview







    • Reliability estimates may vary based on the nature of the variables under study.

    • It is crucial, yet challenging, to decompose error variance into its individual components.

    Determining Reliability Metrics

    • The test's nature influences the reliability metric utilized. - Considerations:

      • Homogeneity or heterogeneity of test items.

      • Whether the measured characteristic, ability, or trait is dynamic or static.

      • Range restrictions of test scores.

      • Whether the test is a speed or power test.

      • Whether the test is criterion-referenced.

    True-Score Model vs. Alternatives

    • True-Score Model:

      • Often known as Classical Test Theory (CTT), it is the most commonly used model due to its simplicity.

      • True Score:

      • A value that genuinely reflects a person's ability or trait level as per classical test theory.

      • Contention:

      • While CTT assumptions are generally acceptable, a significant controversial assumption is the equivalence of test items, generally yielding longer tests.

    Alternative Models

    • Generalizability Theory:

      • Suggests that test scores may vary due to variables present in the testing scenario and sampling.

      • Encourages test developers and researchers to elucidate the specifics of the test situation leading to a certain score.

    • Item-Response Theory (IRT):

      • Models the probability of individuals of ability X to score at level Y on a test.

      • Key Elements:

      • A family of methods encompassing item difficulty and discrimination.

      • Discrimination (a):

      • Refers to an item's effectiveness in distinguishing individuals with different levels of the trait being measured.

      • Difficulty (b):

      • Relates to the complexity of an item.

    Example of IRT Parameters




    • Table 1: Item discrimination (a) and difficulty (b) parameters for 10 items of the Marital Satisfaction Inventory-Brief Form.


      Item

      a

      b

      Endorsement threshold



      1.

      1.17

      0.30

      L



      2.

      3.25

      0.89

      H



      3.

      2.32

      1.97

      H



      4.

      3.25

      0.77

      L



      5.

      1.36

      -0.03

      L









      The Standard Error of Measurement





      • Standard Error of Measurement (SEM):

        • Represents a measure of the of an ____.

        • Estimates the inherent error in an observed score or measurement.

        • Generally, greater test reliability leads to a reduction in standard error.

        • SEM can help estimate how much an observed score deviates from a true score.

        • Confidence Interval:

        • A range of test scores likely to contain the true score.

        • Can be expressed using the formula:
          extObservedScore<br>ightarrowextConfidenceLevelzextscore(SEM)ext{Observed Score} <br>ightarrow ext{Confidence Level} z ext{-score (SEM)}

      Example of Confidence Consideration

      • Example: If one scores 74 on an IQ test with a cut-off of 75 for suggesting intellectual disability, how confident can we be that the individual falls below the threshold if the SEM is 3?

      Standard Error of the Difference

      • Standard Error of Difference:

        • A measure assisting test users to discern how large a score difference should be deemed _ ___ .

        • Can address three types of queries:

        1. Comparison of an individual's performance across different tests (test 1 vs. test 2).

        2. Comparison of an individual's performance against another individual's performance on the same test (test 1).

        3. Comparison of an individual's performance on one test against another individual's performance on a different test (test 1 vs. test 2).