Chapter 4: Psychological Testing, Assessment Assumptions, and Test Quality
Course Overview and Testing Schedule
- Calendar and Chapter Coverage:
- As of September 9th, Test #2 is scheduled for slightly more than a week out, covering Chapters 3 through 5.
- Chapter 4 addresses fundamental testing assumptions, criteria for evaluating tests, and norms used in test development.
- Chapter 5 addresses reliability across three dedicated sessions, followed by Chapter 6 on validity.
Fundamental Assumptions of Testing and Assessment
Assumption 1: Psychological Traits and States Exist:
- Personality Traits: Defined as enduring characteristics of an individual that remain relatively stable throughout life and across diverse situations.
- Examples of traits include shyness, self-esteem, social anxiety, depression, neuroticism, extroversion, openness to experience, agreeableness, conscientiousness, impulsivity, and anxiousness (incorporating the Big Five personality traits).
- States: Temporary emotional or psychological conditions that fluctuate depending on time, context, and environment.
- Emotional states exist on independent ranges rather than a simple bipolar continuum (e.g., happiness ranges from "not at all happy" to "very happy"; sadness ranges from "not at all sad" to "very sad"; the true opposite of love is indifference rather than hate).
- Inference Requirement: Psychological constructs cannot be directly physically measured (e.g., intelligence cannot be measured by opening a skull with a ruler). Constructs must be inferred from observable external behaviors, verbalizations, social interactions, and responses to standardized test items.
Assumption 2: Psychological Traits and States Can Be Quantified and Measured:
- General scientific principle: If a phenomenon exists, it can be measured and quantified.
- The operationalization of a construct into a test is a research-driven process, often resulting in multiple competing instruments for a single trait.
- Prominent measures developed to evaluate shyness include:
- McCroskey Shyness Scale
- Revised Cheek and Buss Shyness Scale
- Social Reticence Scale
- Social Avoidance and Distress Scale
- Interaction Anxiousness Scale
- Test selection depends on specific assessment goals: clinical vs. research applications, and self-report instruments vs. behavioral observation measures.
Assumption 3: Test-Related Behavior Predicts Non-Test-Related Behavior:
- Responses on an assessment instrument are assumed to forecast real-world behaviors in non-testing situations.
- Higher scores on occupational aptitude tests are assumed to predict superior job performance.
Assumption 4: All Tests Have Limits and Imperfections:
- Competent test users acknowledge that every assessment possesses inherent limitations, boundaries, and potential error.
- Medical analogy: An abnormal blood test result indicating a severe condition does not result in immediate radical treatment (e.g., immediate amputation); medical professionals order follow-up testing to confirm findings and eliminate false positives.
- Clinical diagnostic decisions are never based on a single test score, but rather on a comprehensive battery of multiple assessments combined together.
- Incompetent test users incorrectly assume that a test provides an infallible, exact score (a primary flaw seen in educational high-stakes testing).
Assumption 5: Unfair and Biased Assessment Procedures Can Be Identified and Reformed:
- The presence of bias in a test does not require discarding the assessment entirely; rather, test developers can isolate specific sources of bias and reform item content.
- Item reviewers evaluate questions to remove racial, sexual, or cultural bias and eliminate content that induces extraneous anxiety unrelated to the construct being measured.
- An example of extraneous anxiety: A test question presenting a scenario about caring for a terminally ill relative may trigger acute emotional distress in an examinee currently managing that real-life trauma, artificially depressing performance on unrelated test items.
Assumption 6: Testing and Assessment Offer Powerful Benefits to Society:
- Despite inherent imperfections, psychological assessment provides substantial societal value by quantifying human characteristics and predicting critical real-world outcomes.
Scale Analysis: Revised Cheek and Buss Shyness Scale
Structure and Items:
- The scale utilizes a 5-point Likert-type response format ranging from strongly disagree to strongly agree, where higher agreement reflects higher levels of shyness.
- The instrument consists of 13 standardized items, including:
- "I feel tense when I'm with people I don't know well."
- "I'm socially somewhat awkward when a group of people have trouble thinking the right things to talk about."
- "It's hard for me to act natural when I'm meeting new people."
- "I have trouble looking someone right in the eye."
- "I'm more shy with members of the opposite sex." (Noted as having a weaker factor loading, indicating it may require removal).
- "I feel nervous when speaking to someone in authority."
Scoring Parameters:
- Minimum possible score:
- Maximum possible score:
Behavioral Prediction Research:
- Empirical studies using the scale demonstrate that individuals scoring higher in shyness exhibit distinct behavioral patterns during group decision-making tasks compared to low scorers.
- High scorers make significantly different attributions regarding the underlying causes of their performance compared to low-shyness peers.
Test Imperfections and Sources of Error Variance
Definition of Measurement Error:
- Error does not refer to a mistake made by the test developer.
- Error variance is defined as any component of a test score attributable to sources other than the specific trait or ability being measured.
Assessee-Related Sources of Error:
- Recent emotional experiences immediately prior to testing (e.g., entering an assessment directly after receiving positive social feedback versus bombing a public speaking speech).
- Physical states such as acute illness or sleep deprivation.
- Example of fatigue impact: An examinee staying up until 4:30 AM to watch the longest-finishing tennis match in US Open history (which began around 11:00 PM and extended four hours past its scheduled recording block) will experience compromised test performance.
Assessor-Related Sources of Error:
- Demeanor and behavior of the administrator (e.g., a calm, reassuring approach versus a hostile, aggressive presentation).
- Failure to establish proper rapport during individually administered tests (e.g., Wechsler intelligence scales or Stanford-Binet Intelligence Scales).
Administration and Environmental Sources of Error:
- Deviations from standardized administration protocols (e.g., skipping explicit verbal instructions regarding standardized bubble sheet completion for ACT or SAT administration).
- Logistical or facility failures (e.g., broken electronic Scantron grading machines or missing pencil sharpeners).
- Time of day during administration (e.g., early morning testing at 8:00 AM versus 10:00 AM for non-morning individuals).
Test Construction Case Study: MCAT Behavioral Science Section
Contracted Development Parameters:
- Item writers were paid $500 per unit to construct a 500-word stimulus passage accompanied by 10 multiple-choice questions.
Content Examples and Bias Review:
- Behavioral science questions evaluated psychological concepts (e.g., an item describing a US patient hugging and saying "I love you" to a doctor after every routine visit, asking which norm or boundary this behavior violates).
- Standard biological questions evaluated intracellular locations (e.g., whether a process occurs in the cytoplasm, cell membrane area, or mitochondria).
- Items underwent rigorous multi-stage review panels specifically designed to eliminate racial, sexual, and socioeconomic bias.
- An item regarding stereotypes and their impact on medical decision-making underwent intense scrutiny due to sensitive content, but was retained because of its essential relevance to behavioral science in medicine.
Primary Criteria for Evaluating Test Quality
Reliability:
- Refers to the internal consistency and stability of measurement across time.
- Because personality traits are defined as stable attributes, a reliable test must yield consistent scores across repeated administrations.
- Test-Retest Reliability: Evaluated by administering the identical instrument to the same sample at two different time points (e.g., one month apart) and calculating the correlation coefficient between the two sets of scores.
Validity:
- Defined as the extent to which a test measures what it purports (claims) to measure.
- Behavioral Validation Examples for Shyness:
- Verbal Behavior: High-shyness individuals placed in a room with strangers speak significantly less than low-shyness individuals.
- Social Outcomes: High-shyness college freshmen assessed at the start of the academic year (e.g., at ECSU or UCF) report making fewer friends by the end of the year compared to non-shy peers.
Questions & Discussion
Sources of Assessee and Assessor Error Variance:
- Prompt: How do assessees and assessors specifically introduce error variance into test scores?
- Response: Assessees introduce error through transient state fluctuations such as acute physical illness, emotional trauma prior to testing, or severe fatigue. Assessors introduce error through hostile or variable administration styles, failure to establish rapport during individual assessments (such as the Wechsler or Stanford-Binet tests), or failing to deliver verbatim standardized instructions during group assessments.
Addressing Test Anxiety in Test Administration:
- Prompt: How should testing anxiety be managed to minimize error variance?
- Response: Testing anxiety is addressed across three levels:
- Environmental Controls: Utilizing quiet testing rooms and providing a calming settling-in period before test initiation.
- Test Construction Strategy: Ordering initial test items from easy to moderate difficulty (e.g., simple identification of basic roles like test taker, user, developer, and publisher) to build initial confidence, preventing immediate cognitive freeze compared to starting with highly complex items.
- Formal Accommodations: Providing individual accommodations such as distraction-reduced testing environments and extended time limits.
Research Impact of Academic Accommodations:
- Prompt: Do academic accommodations introduce confounding error variance into test data?
- Response: Yes. From a strict methodological research standpoint, modifying administration conditions (e.g., differing testing environments and extra time) introduces variance because standardized conditions are no longer identical across all participants. However, in applied educational settings, accommodations are legally required individual interventions designed to remove the specific barrier posed by a disability or severe anxiety.