Week 3
Stages of Test Development
Stage 1: Test Construction
- Test format considerations
- Item selection and writing
- Scale construction
- Response sets
Stage 2: Standardization
- Development of normative data
- Establishment of test validity and reliability
Stage 3: Revision of Tests
- Rewriting, adding, or deleting items
- Development of new norms
- References: Murphy & Davidshofer, 2005
Test Construction: Format Considerations
Consider the following in construction:
- Purpose of the test
- Interpretation of individual scores
Example:
- A test designed to identify individuals needing remedial interventions in a subject
- Emphasizes a well-represented lower ability level rather than overly difficult items
- Reference: Dorfman & Hersen, 2001
Distinction between:
- Multiple choice or structured items
- Constructed response items (essays, open-ended items)
- Benefits and disadvantages of each type must be assessed
Item Wording
- Good item characteristics:
- Concise length
- Appropriate vocabulary for target group
- Absence of grammatical issues (e.g., double negatives)
- Avoidance of offensive language (e.g., sexist, racist)
Item Selection: Strategies
Rational Scales:
- Guided by theory
- Effectiveness based on theory soundness
Empirical/Criterion-Keyed Approach:
- Selection based on ability to differentiate between groups
- Potential lack of transparency for test-takers regarding item meanings
Analytic Approach:
- Start with guided theory for item selection
- Test administration to large samples followed by factor analysis
- Evaluate if item clusters align with theory
Response Set Measurement
- This test measures response sets including:
- Social desirability
- Faking good/bad
- “Yea” saying or acquiescence
- “Nay” saying or criticalness
- Inattention/random responding
- Validity scales assess these characteristics, distinct from overall test validity
Variables and Measurement
- Measurement: Assigning numbers or symbols to represent variables using logic rules
- Examples: Age, relationship satisfaction, relationship status
Types of Variables (Drummond, Sheperis, & Jones, 2016)
- Quantitative: Numerical value
- Examples: Age, income, test score
- Qualitative: Non-numeric categories
- Examples: Gender, relationship status
- Continuous Variables: Subdivisible; involve “how much”
- Examples: Height, weight, anxiety degree, time
- Discrete Variables: Basic unit measurement; non-subdivisible
- Example: Number of household members
- Observable Variables: Directly measured
- Examples: Hair color, height
- Latent Variables: Inferred from other measurable variables
- Example: Personality traits
Scales of Measurement
| Scale Type | Defining Features | Examples |
|---|---|---|
| Nominal | Named categories, no magnitude; cannot perform arithmetic operations | Relationship status, gender |
| Ordinal | Rank order classification; uneven intervals | Student rank by GPA |
| Interval | Ordered categories with equal intervals, no absolute zero point | Temperature |
| Ratio | Same as interval but with a true zero point (meaningful ratios can be calculated) | Height, weight |
Criterion-Referenced Testing
- Raw scores compared to a predetermined score
- Focus on mastery of specific skills or instructional objectives
- Scores represented as:
- Percentage correct (e.g., 90%)
- Performance categories (e.g., proficient, basic)
- Cutoff scores often define minimally acceptable performance
Norm-Referenced Testing
- Comparison of individual scores against a normative sample
- Normative sample is crucial for valid interpretation of scores
- Example Comparison:
- Jean's scores: Math = 50, English = 75
- Questions arise about individual skill levels in each area
Measures of Central Tendency
| Measure | Description |
|---|---|
| Mean | Arithmetic average; easily influenced by outliers |
| Median | Middle score; provides better representation when skewed |
| Mode | Most frequently occurring value |
Normative Sample
- The group against which individual scores are compared
- Should be representative of the intended test population
- Factors such as gender, age, ethnicity need consideration
Derived Scores
- Standard scores used to express individual performance relative to normative scores
- Examples of derived scores include:
- Z-scores:
- T-scores:
- Deviation IQ:
Z-Score Example Calculation
- Test taker’s score = 65, Mean = 50, SD = 10
- Calculation:
Converting Z-Scores to Percentile Ranks
- Percentile ranks indicate how a test performance compares with peers
- Example: 75th percentile signifies better than 75% of test-takers
- To convert Z-scores to percentile:
- Use a Z-score conversion table to find corresponding percentile rank
- Consider sign of Z-score for the correct column in the table