Chapter 8
Test Development
Outline of Test Development Process
Test Conceptualization: Identifying the purpose and need for the test.
Test Construction: The actual building of the test, including item writing and scaling.
Scaling: Determining how attributes will be measured.
Writing: Development of test items.
Scoring: Methods of scoring and interpreting scores.
Item Analysis: Assessing the effectiveness of items in measuring constructs.
Quantitative: Statistical methods to evaluate item effectiveness.
Qualitative: Non-statistical methods for in-depth item assessment.
Test Revision: Updating and improving the test based on feedback and analysis.
Learning Objectives
Objective 1: Understand the primary considerations in test development, including:
Conceptualization
Construction
Analysis
Revision
Objective 2: Acquire knowledge of item analysis techniques:
Quantitative techniques: Numerical analysis of test items.
Qualitative techniques: Descriptive analysis of test items.
What is Test Development?
Definition: The systematic process of creating, evaluating, and revising psychological tests.
Five Major Stages of Test Development:
Test Conceptualization
Test Construction
Test Tryout: Administering the test to gather data on its functionality.
Analysis: Reviewing the performance of the items and the overall test.
Revision: Making necessary changes to improve the test.
Test Conceptualization
Initial Motivation: Development often stems from the thought: "there ought to be a test for…"
Reason for Need:
To solve psychometric issues in existing tests.
Addressing new social phenomena or emerging occupations.
Preliminary Questions in Test Conceptualization
What is the test designed to measure?
What are the test's objectives?
Is there a demonstrated need for the test?
Who are the intended users and test-takers?
What content will the test assess?
How should the test be administered?
What is the ideal format?
Should multiple forms be developed?
What qualifications are necessary for test administrators?
What types of responses will be required from test-takers?
Who stands to benefit, and are there any risks?
How will scores be interpreted?
Test Construction
Pilot Work: Develop a test prototype and gather feedback from focus groups and expert panels.
Scaling: Calibration of responses on various types of scales:
Unidimensional: Measures a single trait.
Multidimensional: Measures multiple traits.
Categorical: Classifies responses into distinct categories.
Ordinal: Ranks responses.
Rating Scales: Provide respondents with words, statements, or symbols to indicate the strength of a particular attribute.
Likert Scale Example
Patient Health Questionnaire-9 (PHQ-9): Used to measure emotional concerns.
Respondents indicate frequency of various problems, rated from "Not at all" to "Nearly every day."
Items include feelings of hopelessness, sleep troubles, and self-worth among others.
Test Construction: Scaling Methods
Likert Scales: Widely used in psychology for measuring attitudes due to their reliability.
Comparative Scaling: Involves comparing items directly with one another.
Categorical Scaling: Classifies stimuli into predefined categories.
Guttman Scale: Sequential items where agreement on stronger statements guarantees agreement on milder ones.
Guttman Scale Example
Rank items with cumulative implications:
Concerns about unlikely events.
Difficulties in controlling worries.
Physical symptoms related to worrying.
Interference with daily functioning.
Item Writing in Test Construction
Item Pool: A comprehensive collection from which final test items are selected.
Item Format includes variables such as arrangement and structure:
Selected-response format: Respondents choose from given answers.
Constructed-response format: Require test-takers to generate their answers (e.g., essay responses).
Selected-Response Format Example
Multiple-choice questions: Include a stem and various answer choices, with distractors.
Other formats include matching and true-false items.
Computerized Adaptive Testing (CAT)
Definition: An interactive testing method where subsequent items depend on previous responses.
Advantages:
Increases testing efficiency and reduces instances of misclassification at both ends of the ability spectrum (floor and ceiling effects).
Scoring Methods in Test Construction
Cumulative Scoring: Higher scores reflect higher levels of the measured trait.
Class Scoring: Classifies responses into diagnostic categories.
Ipsative Scoring: Compares scores within the same test.
Test Tryout
Requirements: The test must be tried on the same population it targets.
Rule of Thumb: Aim for 5-10 respondents per item.
Qualities of a Good Item: Items must be reliable (consistent results) and valid (accurately measure the intended construct).
Item Analysis Techniques
Item-Difficulty Index: Proportion of respondents answering correctly.
Item-Endorsement Index: Percentage of affirmative responses.
Item Reliability Index: Internal consistency of the scale, often evaluated through factor analysis.
Item-Validity Index: Assesses the validity related to a criterion measure.
Item-Discrimination Index: Measures how well an item separates high-scoring from low-scoring test-takers, calculated as: where:
: Proportion of high scorers answering correctly
: Proportion of low scorers answering correctly.
Item Characteristic Curves (ICCs)
α Parameter: Indicates the slope relating the item to the latent construct.
b Parameter: Represents the difficulty level where the probability of endorsing the item is 50%.
Qualitative Item Analysis
Utilizes Qualitative Methods: Emphasizes descriptive over numerical techniques.
Thinkaloud Method: Respondents articulate their thought processes during testing.
Expert Panels: Leverage expert opinions to enhance item analysis.
Sensitivity Review: Ensure fairness and cultural sensitivity in item content.
Test Revision
Process: Evaluate items for strengths/weaknesses; remove or replace ineffective items.
Implementation: Administer revised tests standardly to a new sample.
Standardization: Norms are established to finalize the test.
Reasons for Revising Existing Tests
Address dated materials or terms, update norms, improve psychometric properties, or update theory behind testing.
Cross-validation and Co-validation
Cross-validation: Validating a test on a different sample to ensure the stability of predictive validity.
Co-validation: Conducting tests with similar populations using the same methods, enhancing cost-effectiveness and reliability.
Use of IRT in Test Development and Revision
Item Response Theory (IRT) Application:
Evaluate existing tests for revisions.
Assess measurement equivalence across diverse test-taker populations.
Create item banks for test construction.