Chapter 5 Notes: Identifying Good Measurement
Chapter Overview
- Ways to measure variables
- Conceptual vs operational variable
- Self-report
- Observational measures
- Physiological measure
- Categorical vs quantitative variables
- Scales: ordinal, interval, ratio
- Reliability of measurement: Are the scores consistent?
- Test-retest reliability
- Interrater reliability
- Internal reliability
- Use of scatterplots and the correlation coefficient, denoted as r
- Validity of measurement: Does it measure what it’s supposed to measure?
- Face validity
- Content validity
- Criterion validity
- Convergent validity
- Discriminant validity
- Reliability vs. validity
- Review: interpreting construct validity evidence
Measurement Validity of Abstract Constructs
- Validity is about whether the operationalization is appropriate
- Reliability is necessary but not sufficient for validity
- Two subjective ways to assess validity:
- Face validity: Does the measure look like what you want to measure?
- Example: Head circumference looks appropriate for hat size
- Content validity: The measure contains all parts that the theory says it should contain
- Example: For intelligence, measures should assess planning, problem solving, abstract thinking, learning from experience, etc.
- Three empirical ways to assess validity (in addition to reliability):
- Criterion validity: Your measure is correlated with a relevant behavioral outcome
- Convergent validity: Your self-report measure is more strongly associated with self-report measures of similar constructs
- Discriminant validity: Your self-report measure is less strongly associated with self-report measures of dissimilar constructs
- Reliability is necessary but not sufficient for validity
- Interrater reliability: Two coders’ ratings of a set of targets are consistent with each other
- Reliability vs validity relationship illustrated (often depicted as a Venn-like relation):
- A measure can be reliable but not valid
- A measure cannot be valid if it is not reliable
Validity of Measurement: Does It Measure What It’s Intended to Measure?
- Validity types summarized: Face validity, Content validity, Criterion validity, Convergent validity, Discriminant validity
- Key idea: Do not confuse reliability with validity; both matter, but validity requires more robust evidence
Face Validity and Content Validity: Does It Look Like a Good Measure?
- These are subjective assessments of validity
- Face validity: It looks like what you want to measure
- Example: A measure of head circumference looks appropriate for hat size, but not for intelligence
- Assessed by experts
- Content validity: The measure contains all parts that your theory says it should contain
- Example: Intelligence measures should assess planning, problem solving, abstract thinking, learning from experience, etc.
Criterion Validity: Does It Correlate with Key Behaviors?
- Criterion validity: The measure under consideration is associated with a concrete behavioral outcome
- Particularly important for self-report measures
Criterion Validity (continued): Known-Groups Paradigm
- An empirical approach to Criterion Validity: Known-groups paradigm
- The measure should discriminate among known groups that differ on the construct
- Example context: Cortisol and public speaking differences, Beck Depression Inventory (BDI)
Known-Groups Evidence for Criterion Validity
- Example data illustrating known-groups differences (sample characteristics, N, M, SD, reference):
- American college students: N = 244, M = 23.7, SD = 6.4; Pavot & Diener (1993)
- French Canadian college students (male): N = 355, M = 23.8, SD = 6.1; Blais et al. (1989)
- Korean university students: N = 413, M = 19.8, SD = 5.8; Suh (1993)
- Printing trade workers: N = 304, M = 24.2, SD = 6.0; George (1991)
- Veterans Affairs hospital inpatients: N = 52, M = 11.8, SD = 5.6; Frisch (1991)
- Women in shelters for abused spouses: N = 70, M = 20.7, SD = 7.4; Fisher (1991)
- Male prison inmates: N = 75, M = 12.3, SD = 7.0; Joy (1990)
- Note: N = Number of people in group; M = Group mean on subjective well-being; SD = Group standard deviation
- Source: Adapted from Pavot & Diener, 1993, Table 1
Convergent Validity and Discriminant Validity: Does the Pattern Make Sense?
- Convergent validity: Similar constructs should correlate with each other
- Discriminant validity: Dissimilar constructs should show weaker or no correlation
- Example referenced: Center for Epidemiologic Studies Depression scale (CES-D)
- Core idea: Evidence should show a coherent pattern where measures of similar constructs converge while distinct constructs diverge
The Relationship Between Reliability and Validity
- Key statements:
- The validity of a measure is not the same as its reliability
- A measure can be less valid than it is reliable, but it cannot be more valid than it is reliable
- Everyday analogy:
- A friend who gives you a ride every day is reliable because they help you be on time and attend class
- A truly valid friend will support you in multiple contexts beyond just punctuality
Quick Reference: Key Points to Remember
- Reliability vs. validity are distinct but related concepts
- Reliability is necessary for validity, but not sufficient
- There are subjective (face/content validity) and empirical (criterion/convergent/discriminant validity; known-groups) approaches to assessing validity
- Known-groups paradigm provides practical evidence for criterion validity by showing the measure differentiates between groups known to differ on the construct
- Convergent and discriminant validity help establish the construct validity by showing expected patterns of correlations with similar and dissimilar constructs
- When evaluating a measure, consider the entire validity portfolio: face, content, criterion, convergent, and discriminant validity, in relation to reliability