Psychological Testing and Assessment Practice Flashcards

Application of Tests in Critical Decision-Making

  • Psychological tests are employed globally to address complex clinical, legal, and organizational questions that have significant impacts on individual lives.
  • Emergency medical scenarios: Determining diagnoses for patients exhibiting symptoms like delusions of being an archangel (e.g., Dustin) followed by inconsolable crying.
  • Legal and Forensic context: Assessing whether a defendant in an armed robbery case is competent to stand trial and assist in their own defense, regardless of a long history of academic and conduct problems.
  • Corporate and Organizational Management: Decisions regarding hiring entry-level employees, transferring personnel, identifying individuals for promotion, or terminating employment.
  • Educational Admissions and Scholarships: Evaluating credentials to grant entry into special academic programs or award financial aid.
  • Family Law and Custody: Resolving bitter divorce disputes by assessing allegations of negligence, abusive behavior, or criminal history (such as shoplifting) to determine the best interest of children.
  • The efficacy of these decisions rests on assessment professionals having high confidence in the tools they employ and defining what constitutes a "good test."

Foundational Assumptions of Psychological Testing and Assessment

  • Assumption 1: Psychological Traits and States Exist

    • Human behavior involves an interplay of order and chaos. Scholars use a technical vocabulary to describe stability and change.
    • Trait Definition: A trait is any distinguishable, relatively enduring way in which one individual varies from another (Guilford, 19591959, p. 66).
    • State Definition: Characteristics that distinguish one person from another but are relatively less enduring than traits.
    • Traits have been categorized into thousands of terms, covering intelligence, cognitive style, interests, sexual orientation, psychopathology, and personality (Allport & Odbert, 19361936).
    • Modern cultural evolution has introduced new trait terms such as androgynous (absence of primacy of male/female traits) and gender non-binary (individuals outside the masculine-feminine continuum).
    • Psychological traits exist as constructs: informed, scientific concepts developed to explain behavior. They are not physical entities (like brain circuits) but are inferred from overt behavior (observable actions or test responses).
    • Traits are "relatively enduring"; they are not manifested 100%100\% of the time. Stability is evidenced by high correlations between scores across the lifespan, despite evolution in personality (e.g., becoming more conscientious with age).
    • Trait manifestation is situation-dependent. For example, a parolee may be subdued with an officer but violent with family; a person may be dull to a spouse but charming to business associates.
    • Contextual interpretation is vital: Praying aloud is viewed as religious in a church but potentially deviant at a movie theater.
    • Traits are relative attributions; calling someone "shy" implies a comparison to a hypothetical average person or a specific reference group (e.g., comparing a 2222-year-old male exotic dancer to other male dancers vs. the general male population).
  • Assumption 2: Psychological Traits and States Can Be Quantified and Measured

    • E.L. Thorndike (19181918) asserted that whatever exists at all exists in some amount and can be known through quantity and quality.
    • Measuring a trait requires a clear operational definition. For instance, "aggression" may be defined as self-reported physical harm, observed pushing/hitting on a playground, or socially aggressive acts like slander.
    • Test developers sample behaviors from a domain (a universe of indicative behaviors) to create test items.
    • Item Weighting: Items are often assigned different values based on their comparative importance. A social judgment question (e.g., gun safety in the home) might earn more credit than a rote American history question (e.g., identifying the second president).
    • Cumulative Scoring: The most common model where the magnitude of a trait is assumed to correspond to the sum of keyed responses (e.g., correct = 11, incorrect = 00).
  • Assumption 3: Test-Related Behavior Predicts Non-Test-Related Behavior

    • The goal of tasks like grid-blackening or screen-tapping is not to predict future tapping behavior but to provide an indication of broader aspects of behavior (e.g., work performance or mental disorders).
    • Tests may use samples of behavior to make predictions about future performance or postdictions to understand a past state of mind, such as a criminal defendant's mental state during a crime.
  • Assumption 4: All Tests Have Limits and Imperfections

    • Competent test users must understand test development, appropriate administration, target populations, and inherent limitations. These requirements are emphasized in professional codes of ethics.
  • Assumption 5: Various Sources of Error Are Part of the Assessment Process

    • Error refers to factors other than what a test attempts to measure that influence performance.
    • Error Variance: The component of a test score attributable to sources other than the measured trait. Sources include:
      • Assessees (e.g., having the flu or being affected by situational conditions).
      • Assessors (e.g., varying levels of professionalism in following administration instructions).
      • Measuring instruments (e.g., some tests are inherently better at measuring a construct than others).
    • Random Error: Factors purely based on chance. For example, Rammstedt et al. (20152015) found people rate themselves as less disciplined on sunny days compared to rainy days because sunny days offer more leisure opportunities.
  • Assumption 6: Unfair and Biased Assessment Procedures Can Be Identified and Reformed

    • Test publishers aim for fairness by following guidelines in test manuals. Controversy often stems from using tests with populations they were not intended for or from political debates about selection and hiring (e.g., affirmative action).
  • Assumption 7: Testing and Assessment Offer Powerful Benefits to Society

    • A world without tests would risk meritocracy being replaced by nepotism.
    • Tests allow for screening of surgeons, pilots, and military recruits, and facilitate the diagnosis of educational difficulties or neuropsychological impairments.

Technical Standards for Testing: Reliability, Validity, and Utility

  • Reliability: Refers to the consistency and precision of a measuring tool. A perfectly reliable tool measures the same way every time.
    • Example: Scale A (accurate and consistent), Scale B (consistently inaccurate by 0.3lb0.3\,lb but reliable), and Scale C (randomly inconsistent and unreliable).
  • Validity: The extent to which a test actually measures what it purports to measure for a specific purpose.
    • Scrutiny focuses on whether items adequately sample the construct's range and how scores relate to other behaviors or opposite constructs (e.g., a high introversion score should relate inversely to extraversion).
  • Utility: A good test is useful, actionable, and cost-effective. During World Wars I and II, group intelligence tests were developed because individual Binet tests were not cost-effective for screening thousands of recruits.
  • Normative Data: A standard for comparing measurement results. A "good test" for comparative purposes must have adequate norms.

The Science of Sampling and Standardization

  • Standardization: The process of establishing replicable procedures for administration, scoring, and interpretation, often including normative data. A standardized test comes with a manual providing detailed guidelines so all users can replicate the administration.
  • Sampling: Selecting a portion of the targeted population to be representative of the whole.
    • Population: The complete set of individuals with at least one common characteristic (e.g., all marathon runners).
    • Stratified Sampling: Proportionately representing different subgroups (strata) such as race, gender, and socioeconomic status to prevent bias.
    • Stratified-Random Sampling: Every member of the defined population has an equal chance of being included in the strata.
    • Purposive Sampling: Arbitrarily selecting a sample believed to be representative (e.g., testing a product in Cleveland to predict national sales).
    • Incidental/Convenience Sampling: Using a sample that is easily available (e.g., introductory psychology students) rather than the most appropriate. Generalizations from these samples must be cautious.

Detailed Categorization of Norms

  • Percentile Norms: Raw data converted to percentiles. The xthx^{th} percentile is the score at or below which x%x\% of scores fall.
    • Note: Percentile (ranking relative to a group) is distinct from percentage correct (raw score calculation).
    • Limitation: Percentiles can exaggerate differences in the middle of a normal distribution and minimize differences at the extremes.
  • Age Norms (Age-Equivalent Scores): Average performance of samples at various ages. Often used in physical measurements (height) but controversial in psychology (e.g., "mental age" in the Stanford-Binet).
  • Grade Norms: Average performance of testtakers in a specific school grade (e.g., a score of 6.46.4 represents performance at the 4th4^{th} month of 6th6^{th} grade). These are referred to as developmental norms.
  • National Norms: Derived from a sample nationally representative of the population (considering age, gender, community type, etc.).
  • National Anchor Norms: Equivalency tables comparing scores on different tests (e.g., equating scores from the "Best Reading Test" and the "XYZ Reading Test") using the equipercentile method based on common samples (co-norming).
  • Subgroup Norms: Normative data segmented by specific criteria (e.g., handedness, socioeconomic level).
  • Local Norms: Developed by users for specific populations (e.g., a local company norming a national test on its specific applicant pool).
  • Fixed Reference Group Scoring Systems: Uses a distribution of scores from one group to calculate scores for all future administrations (e.g., the SAT). The SAT used the 19411941 sample (11,00011,000 people) until 19951995, when it shifted to a 19901990 reference group of 22 million testtakers. A score of 500500 corresponds to the mean of the reference group.

Comparative Models: Norm-Referenced vs. Criterion-Referenced Evaluation

  • Norm-Referenced: meaning is derived by comparing an individual's score to the performance of a group.
  • Criterion-Referenced: meaning is derived by evaluating a score against a set standard (criterion).
    • Examples: Reading level requirements for high school diplomas, road tests for driving licenses, or state licensing for psychologists.
    • A student earns a karate black belt by meeting specific proficiency and self-discipline standards, regardless of others' performance.
  • Key Differences:
    • Norm-referenced focus: Ranking relative to others; identifies brilliance/superiority at the extremes of the normal curve.
    • Criterion-referenced focus: What the testtaker can do; whether they have achieved "mastery." Used in computer-assisted education.
    • Critique: Strict criterion-referenced approaches may lose data on relative performance and are less useful for evaluating doctoral-level or highly original abilities.

Professional and Cultural Context

  • Case Study: The Chicago Bulls: During the 1990s1990s, the Bulls used personality testing (16PF16PF) and behavioral interviewing to select college players and free agents. They aimed to evaluate competencies like resilience and team orientation rather than psychopathology. They eventually developed a validated regression formula for prediction.
  • Culturally Informed Assessment Guidelines:
    • Do: Be aware of cultural assumptions; consult cultural community members about test appropriateness; use methods that complement the assessee’s worldview; score and interpret findings within the cultural context.
    • Do Not: Assume tests impact all groups the same way; take a ‘one-size-fits-all’ approach; select tools without regard for the assessee’s background; assume translated tests are automatically equivalent to the original.
    • Historical context (e.g., growing up before or after the advent of satellites) must be considered in generations-based evaluation.

Questions & Discussion

  • Question: What are the elements of a "good test"?
    • Response: Criteria include clear instructions for administration, scoring, and interpretation; economy in time and money; and psychometric soundness (reliability and validity).
  • Question: Why is it difficult to predict violence by means of a test?
    • Response: Violence is a complex behavior influenced by situational variables that may not be fully captured in a test environment.
  • Question: Why are certain groups excluded from the standardization sample of intelligence tests?
    • Response: Groups such as those with uncorrected visual/hearing loss, those not fluent in English, or people on medications that depress performance are excluded to ensured the sample reflects a specific typical population without confounding variables that might distort the norms.
  • Question: Does Ben’s Deli "Cold Cut Preference Test" (CCPT) qualify as a standardized test?
    • Response: While it has specified procedures (rules for asking "What would you like?"), most professionals reserve the term "standardized" for tests with replicable administration/scoring and established norms.
  • Question: Should different norms be used for different groups in hiring?
    • Response: While race norming was outlawed by the Civil Rights Act of 19911991, scholars continue to work on methods to ensure equitable hiring through adjusted item analysis or other equitable selection procedures.