Chapter 8

Test Development

Outline of Test Development Process

  • Test Conceptualization: Identifying the purpose and need for the test.

  • Test Construction: The actual building of the test, including item writing and scaling.

  • Scaling: Determining how attributes will be measured.

  • Writing: Development of test items.

  • Scoring: Methods of scoring and interpreting scores.

  • Item Analysis: Assessing the effectiveness of items in measuring constructs.

  • Quantitative: Statistical methods to evaluate item effectiveness.

  • Qualitative: Non-statistical methods for in-depth item assessment.

  • Test Revision: Updating and improving the test based on feedback and analysis.

Learning Objectives

  • Objective 1: Understand the primary considerations in test development, including:

    • Conceptualization

    • Construction

    • Analysis

    • Revision

  • Objective 2: Acquire knowledge of item analysis techniques:

    • Quantitative techniques: Numerical analysis of test items.

    • Qualitative techniques: Descriptive analysis of test items.

What is Test Development?

  • Definition: The systematic process of creating, evaluating, and revising psychological tests.

  • Five Major Stages of Test Development:

    1. Test Conceptualization

    2. Test Construction

    3. Test Tryout: Administering the test to gather data on its functionality.

    4. Analysis: Reviewing the performance of the items and the overall test.

    5. Revision: Making necessary changes to improve the test.

Test Conceptualization

  • Initial Motivation: Development often stems from the thought: "there ought to be a test for…"

  • Reason for Need:

    • To solve psychometric issues in existing tests.

    • Addressing new social phenomena or emerging occupations.

Preliminary Questions in Test Conceptualization
  • What is the test designed to measure?

  • What are the test's objectives?

  • Is there a demonstrated need for the test?

  • Who are the intended users and test-takers?

  • What content will the test assess?

  • How should the test be administered?

  • What is the ideal format?

  • Should multiple forms be developed?

  • What qualifications are necessary for test administrators?

  • What types of responses will be required from test-takers?

  • Who stands to benefit, and are there any risks?

  • How will scores be interpreted?

Test Construction

  • Pilot Work: Develop a test prototype and gather feedback from focus groups and expert panels.

  • Scaling: Calibration of responses on various types of scales:

    • Unidimensional: Measures a single trait.

    • Multidimensional: Measures multiple traits.

    • Categorical: Classifies responses into distinct categories.

    • Ordinal: Ranks responses.

  • Rating Scales: Provide respondents with words, statements, or symbols to indicate the strength of a particular attribute.

Likert Scale Example
  • Patient Health Questionnaire-9 (PHQ-9): Used to measure emotional concerns.

    • Respondents indicate frequency of various problems, rated from "Not at all" to "Nearly every day."

    • Items include feelings of hopelessness, sleep troubles, and self-worth among others.

Test Construction: Scaling Methods

  • Likert Scales: Widely used in psychology for measuring attitudes due to their reliability.

  • Comparative Scaling: Involves comparing items directly with one another.

  • Categorical Scaling: Classifies stimuli into predefined categories.

  • Guttman Scale: Sequential items where agreement on stronger statements guarantees agreement on milder ones.

Guttman Scale Example
  • Rank items with cumulative implications:

    1. Concerns about unlikely events.

    2. Difficulties in controlling worries.

    3. Physical symptoms related to worrying.

    4. Interference with daily functioning.

Item Writing in Test Construction

  • Item Pool: A comprehensive collection from which final test items are selected.

  • Item Format includes variables such as arrangement and structure:

    • Selected-response format: Respondents choose from given answers.

    • Constructed-response format: Require test-takers to generate their answers (e.g., essay responses).

Selected-Response Format Example
  • Multiple-choice questions: Include a stem and various answer choices, with distractors.

  • Other formats include matching and true-false items.

Computerized Adaptive Testing (CAT)
  • Definition: An interactive testing method where subsequent items depend on previous responses.

  • Advantages:

    • Increases testing efficiency and reduces instances of misclassification at both ends of the ability spectrum (floor and ceiling effects).

Scoring Methods in Test Construction

  • Cumulative Scoring: Higher scores reflect higher levels of the measured trait.

  • Class Scoring: Classifies responses into diagnostic categories.

  • Ipsative Scoring: Compares scores within the same test.

Test Tryout
  • Requirements: The test must be tried on the same population it targets.

  • Rule of Thumb: Aim for 5-10 respondents per item.

  • Qualities of a Good Item: Items must be reliable (consistent results) and valid (accurately measure the intended construct).

Item Analysis Techniques

  • Item-Difficulty Index: Proportion of respondents answering correctly.

  • Item-Endorsement Index: Percentage of affirmative responses.

  • Item Reliability Index: Internal consistency of the scale, often evaluated through factor analysis.

  • Item-Validity Index: Assesses the validity related to a criterion measure.

  • Item-Discrimination Index: Measures how well an item separates high-scoring from low-scoring test-takers, calculated as: d=P<em>highP</em>lowd = P<em>{high} - P</em>{low} where:

    • PhighP_{high}: Proportion of high scorers answering correctly

    • PlowP_{low}: Proportion of low scorers answering correctly.

Item Characteristic Curves (ICCs)
  • α Parameter: Indicates the slope relating the item to the latent construct.

  • b Parameter: Represents the difficulty level where the probability of endorsing the item is 50%.

Qualitative Item Analysis

  • Utilizes Qualitative Methods: Emphasizes descriptive over numerical techniques.

    • Thinkaloud Method: Respondents articulate their thought processes during testing.

    • Expert Panels: Leverage expert opinions to enhance item analysis.

    • Sensitivity Review: Ensure fairness and cultural sensitivity in item content.

Test Revision

  • Process: Evaluate items for strengths/weaknesses; remove or replace ineffective items.

  • Implementation: Administer revised tests standardly to a new sample.

  • Standardization: Norms are established to finalize the test.

Reasons for Revising Existing Tests
  • Address dated materials or terms, update norms, improve psychometric properties, or update theory behind testing.

Cross-validation and Co-validation

  • Cross-validation: Validating a test on a different sample to ensure the stability of predictive validity.

  • Co-validation: Conducting tests with similar populations using the same methods, enhancing cost-effectiveness and reliability.

Use of IRT in Test Development and Revision

  • Item Response Theory (IRT) Application:

    • Evaluate existing tests for revisions.

    • Assess measurement equivalence across diverse test-taker populations.

    • Create item banks for test construction.