Chapter 4
Learning Objectives
Primary learning objectives for PSYC 385 Psychological Test, Fall 2024:
Evaluate whether a test fulfills the basic assumptions and scrutinize the quality of the psychological test.
Derive meaning from test scores by understanding characteristics of the standardization sample.
Assumptions about Psychological Testing
Key Assumptions
Psychological traits and states exist
Definition: A trait is “any distinguishable, relatively enduring way in which one individual varies from another.”
Traits are distinguished from states, which are relatively less enduring.
Psychological traits exist as constructs, which are informed, scientific concepts used to describe or explain behavior.
Constructs cannot be directly observed; their existence is inferred from overt behavior.
Psychological traits and states can be quantified and measured
Test developers may define and measure constructs differently.
After defining a construct, item content and weighting are created.
A scoring system and interpretation method are developed.
Test-related behavior predicts non-test-related behavior
Responses can predict both test-related and non-test-related behaviors.
Tests and measurement techniques have strengths and weaknesses
Familiarity with the test development process and appropriate application and interpretation of tests are crucial.
Cyber physical limitations of the test must be acknowledged.
Various sources of error are part of the assessment process
Error refers to factors other than what a test attempts to measure influencing test performance.
Error variance: the portion of a test score attributed to sources other than the trait being measured.
Both assessee and assessor characteristics contribute to error variance.
Potential sources of error should be examined.
Testing and assessment can be conducted fairly
Test developers strive for fairness in instruments if used according to test manual guidelines.
Problems may arise if tests are used with inappropriate populations.
Some issues may be more philosophical than psychometric in nature.
Testing and assessment benefit society
There is a need for quality tests across various life areas, promoting benefits to individuals and society.
Psychometric Considerations
What Constitutes a “Good Test”?
Reliability: Refers to the consistency of the measuring tool, indicating the precision with which a test measures and the extent of error presence in measurements.
Validity: Determines whether the test measures what it intends to measure.
Administration, scoring, and interpretation of tests should be straightforward for trained examiners.
A good test serves utility for individual test-takers or society.
Reliability and Validity Connection
Reliability alone does not ensure validity; both must be considered for effective test assessments.
Evaluation of Tests
Considerations for Using a Test
Reasons for instrument/method selection.
Exist published guidelines for the instrument used?
Assess reliability and validity of the test.
Cost-effectiveness evaluation.
Explore inferences made from scores and their generalizability.
Norms and Standardization
Norms in Testing and Assessment
Deriving meaning from test scores involves evaluating an individual score against a norm-referenced group of test-takers
Norms: Test performance data compiled from a standardized sample used for evaluation or interpretation of individual test scores.
A normative or standardization sample is the comparison group for individual test-takers.
Developing Norms through Standardization
Standardization: Process of administering a test to a representative sample to establish a normative sample.
Sampling: Test developers choose a certain population with shared observable characteristics.
Stratified Sampling: Inclusion of different subgroups (strata) from the population.
Stratified-Random Sampling: Ensures every member of the population has an equal chance of inclusion.
Other sampling methods:
Purposive Sample: Arbitrarily selects a sample believed to be representative.
Incidental/Convenience Sample: A convenient sample that may not represent the population.
Impact of Sampling Techniques
Generalizations must be cautious regarding convenience samples.
Consideration of exclusionary criteria's impact on results.
Standardizing Tests
Standard administration procedures, including instructions.
Recommended settings and required materials.
Data collection, analysis, and summation using descriptive statistics (e.g., measures of central tendency and variability).
Clear characterization of the standardization sample.
Types of Norms
Various types of norms include:
Percentile norms
Age norms
Grade norms
Cultural/racial norms
National norms
Local norms
Subgroup norms
Scoring Systems
Fixed Reference Group Scoring Systems
Utilizes the distribution of scores from previous administrations to calculate future test scores.
Example: The SAT uses this system.
Norm-Referenced vs. Criterion-Referenced Interpretation
Norm-referenced tests: Compare individual performance against a normative group.
Criterion-referenced tests: Evaluate whether test-takers meet predefined criteria (e.g., driving exam performance).
Practical Application Example
Construct of emotion regulation as operationalized in research (Ritschel et al., 2015).
Identification of populations with psychometric support for the Difficulties in Emotion Regulation Scale (DERS; Gratz & Roemer, 2004).
Standardization sample characteristics for DERS.
Analysis of psychometric properties of DERS across college subgroups.
Cultural Considerations in Testing
Responsible test users must research test norms for appropriateness with their targeted test populations.
Understanding the culture and context of the test-taker is crucial for accurate result interpretation.
Culturally Informed Assessment: Some ‘Do’s’ and ‘Don'ts’
Do: Be aware of cultural assumptions underlying tests.
Do Not: Assume tests impact all groups equally.
Do: Consult community members for assessment appropriateness.
Do Not: Assume agreement across all cultural groups for test items.
Do: Incorporate assessments that respect the worldview of specific cultural populations.
Do Not: Hold a “one-size-fits-all” view of assessment.
Do: Acknowledge equivalence issues (language and constructs) across cultures.
Do Not: Analyze data without considering cultural contexts.