Chapter 5 Notes: Identifying Good Measurement

Chapter Overview

  • Ways to measure variables
    • Conceptual vs operational variable
    • Self-report
    • Observational measures
    • Physiological measure
    • Categorical vs quantitative variables
    • Scales: ordinal, interval, ratio
  • Reliability of measurement: Are the scores consistent?
    • Test-retest reliability
    • Interrater reliability
    • Internal reliability
    • Use of scatterplots and the correlation coefficient, denoted as rr
  • Validity of measurement: Does it measure what it’s supposed to measure?
    • Face validity
    • Content validity
    • Criterion validity
    • Convergent validity
    • Discriminant validity
    • Reliability vs. validity
  • Review: interpreting construct validity evidence

Measurement Validity of Abstract Constructs

  • Validity is about whether the operationalization is appropriate
  • Reliability is necessary but not sufficient for validity
  • Two subjective ways to assess validity:
    • Face validity: Does the measure look like what you want to measure?
    • Example: Head circumference looks appropriate for hat size
    • Content validity: The measure contains all parts that the theory says it should contain
    • Example: For intelligence, measures should assess planning, problem solving, abstract thinking, learning from experience, etc.
  • Three empirical ways to assess validity (in addition to reliability):
    • Criterion validity: Your measure is correlated with a relevant behavioral outcome
    • Convergent validity: Your self-report measure is more strongly associated with self-report measures of similar constructs
    • Discriminant validity: Your self-report measure is less strongly associated with self-report measures of dissimilar constructs
  • Reliability is necessary but not sufficient for validity
  • Interrater reliability: Two coders’ ratings of a set of targets are consistent with each other
  • Reliability vs validity relationship illustrated (often depicted as a Venn-like relation):
    • A measure can be reliable but not valid
    • A measure cannot be valid if it is not reliable

Validity of Measurement: Does It Measure What It’s Intended to Measure?

  • Validity types summarized: Face validity, Content validity, Criterion validity, Convergent validity, Discriminant validity
  • Key idea: Do not confuse reliability with validity; both matter, but validity requires more robust evidence

Face Validity and Content Validity: Does It Look Like a Good Measure?

  • These are subjective assessments of validity
  • Face validity: It looks like what you want to measure
    • Example: A measure of head circumference looks appropriate for hat size, but not for intelligence
    • Assessed by experts
  • Content validity: The measure contains all parts that your theory says it should contain
    • Example: Intelligence measures should assess planning, problem solving, abstract thinking, learning from experience, etc.

Criterion Validity: Does It Correlate with Key Behaviors?

  • Criterion validity: The measure under consideration is associated with a concrete behavioral outcome
  • Particularly important for self-report measures

Criterion Validity (continued): Known-Groups Paradigm

  • An empirical approach to Criterion Validity: Known-groups paradigm
    • The measure should discriminate among known groups that differ on the construct
    • Example context: Cortisol and public speaking differences, Beck Depression Inventory (BDI)

Known-Groups Evidence for Criterion Validity

  • Example data illustrating known-groups differences (sample characteristics, N, M, SD, reference):
    • American college students: N = 244, M = 23.7, SD = 6.4; Pavot & Diener (1993)
    • French Canadian college students (male): N = 355, M = 23.8, SD = 6.1; Blais et al. (1989)
    • Korean university students: N = 413, M = 19.8, SD = 5.8; Suh (1993)
    • Printing trade workers: N = 304, M = 24.2, SD = 6.0; George (1991)
    • Veterans Affairs hospital inpatients: N = 52, M = 11.8, SD = 5.6; Frisch (1991)
    • Women in shelters for abused spouses: N = 70, M = 20.7, SD = 7.4; Fisher (1991)
    • Male prison inmates: N = 75, M = 12.3, SD = 7.0; Joy (1990)
  • Note: N = Number of people in group; M = Group mean on subjective well-being; SD = Group standard deviation
  • Source: Adapted from Pavot & Diener, 1993, Table 1

Convergent Validity and Discriminant Validity: Does the Pattern Make Sense?

  • Convergent validity: Similar constructs should correlate with each other
  • Discriminant validity: Dissimilar constructs should show weaker or no correlation
  • Example referenced: Center for Epidemiologic Studies Depression scale (CES-D)
  • Core idea: Evidence should show a coherent pattern where measures of similar constructs converge while distinct constructs diverge

The Relationship Between Reliability and Validity

  • Key statements:
    • The validity of a measure is not the same as its reliability
    • A measure can be less valid than it is reliable, but it cannot be more valid than it is reliable
  • Everyday analogy:
    • A friend who gives you a ride every day is reliable because they help you be on time and attend class
    • A truly valid friend will support you in multiple contexts beyond just punctuality

Quick Reference: Key Points to Remember

  • Reliability vs. validity are distinct but related concepts
  • Reliability is necessary for validity, but not sufficient
  • There are subjective (face/content validity) and empirical (criterion/convergent/discriminant validity; known-groups) approaches to assessing validity
  • Known-groups paradigm provides practical evidence for criterion validity by showing the measure differentiates between groups known to differ on the construct
  • Convergent and discriminant validity help establish the construct validity by showing expected patterns of correlations with similar and dissimilar constructs
  • When evaluating a measure, consider the entire validity portfolio: face, content, criterion, convergent, and discriminant validity, in relation to reliability