CHAPTER 5 - Identifying Good Measurement

Operationalization and Conceptual Definitions

  • Construct Validity in Psychological Research

    • Construct validity is one of the four big validities in psychological research, applying to frequency, association, and causal claims.

    • It measures how well a study's variables are operationalized, measured, or manipulated.

    • Evaluating construct validity is essential when measuring abstract psychological phenomena such as motivation, emotion, thinking, reasoning, gratitude, extroversion, and happiness.

  • Conceptual vs. Operational Definitions

    • Conceptual Definition (Construct): The theoretical definition of a variable in question.

    • Operational Definition (Operationalization): The specific decision made by a researcher regarding how to measure or manipulate a conceptual variable in a given study.

  • Operationalizing "Happiness" (Subjective Well-Being)

    • Conceptualization: Ed Diener and colleagues explicitly defined happiness conceptually as "subjective well-being" (a person's well-being from their own perspective).

    • Operationalization via Diener's Subjective Well-Being Scale:

      • A 55-item self-report questionnaire where participants evaluate their life satisfaction using custom criteria.

      • Responses are made on a 77-point scale ranging from 11 ("strongly disagree") to 77 ("strongly agree").

      • The five items are:

        1. In most ways my life is close to my ideal.

        2. The conditions of my life are excellent.

        3. I am satisfied with my life.

        4. So far, I have gotten the important things I want in life.

        5. If I could live my life over, I would change almost nothing.

      • Scoring Structure:

        • Minimum score (unhappiest): 55 (1+1+1+1+1=51 + 1 + 1 + 1 + 1 = 5).

        • Maximum score (happiest): 3535 (7+7+7+7+7=357 + 7 + 7 + 7 + 7 = 35).

        • Neutral score: 2020 (4+4+4+4+4=204 + 4 + 4 + 4 + 4 = 20).

      • Empirical Findings: Diener and Diener (1996) reported that most people score above 2020: 63%63\% of high school and college students and 72%72\% of adults with disabilities scored above 2020.

    • Operationalization via the Ladder of Life (Cantril, 1965):

      • A single-item self-report measure asking participants to imagine a ladder with steps numbered from 00 at the bottom (worst possible life) to 1010 at the top (best possible life).

      • Participants indicate which step they personally stand on at the present time.

      • Utilized daily by the Gallup organization in the Gallup-Healthways Well-Being Index.

  • Operationalizing Other Psychological Constructs

    • Extroversion:

      • Self-report via the 88-item extroversion subscale of the Big Five Inventory (BFI; John et al., 2008), which includes items like "I see myself as someone who is talkative."

      • Self-report via the 22-item extroversion subscale of the Ten-Item Personality Inventory (TIPI; Gosling et al., 2003), which includes items like "I see myself as extroverted, enthusiastic."

      • Peer reports obtained by asking participants' friends to rate their extroversion.

    • Gratitude Toward Partner:

      • Self-report questionnaire asking agreement with statements such as "I appreciate my partner" or asking how frequently someone thanks their partner.

      • Observational behavioral coding by watching couples interact and counting instances of explicit verbal thanks.

    • Gender Identity:

      • Self-report question: "Which best applies to you?" with options: Man, Woman, Nonbinary, and Another identity.

      • Observational/archival measurement by recording the gender identity stated on an individual's social media profile.

    • Wealth:

      • Self-report question asking: "What is your total annual household income before taxes?"

      • Observational measure by coding the monetary value of a person's vehicle on a scale from 11 (older, lower-status vehicle) to 55 (new, higher-status vehicle in good condition) (Piff et al., 2012).

Three Common Types of Operational Measures

  • Self-Report Measures

    • Operationalize variables by recording participants' answers to questions about themselves in questionnaires or interviews.

    • Examples include Diener's 55-item Life Satisfaction Scale, Cantril's Ladder of Life, the Big Five Inventory (BFI), and Holmes & Rahe's (1967) stressful life events checklist (recording frequencies of events like marriage, divorce, or moving).

    • When researching children, self-reports are often replaced with parent reports or teacher reports (e.g., evaluating classroom behavior, known vocabulary words, or recent developmental events).

  • Observational Measures (Behavioral Measures)

    • Operationalize variables by recording observable behaviors or physical traces of behaviors.

    • Examples:

      • Counting the frequency of smiles to operationalize moment-to-moment happiness.

      • Standardized Intelligence (IQ) tests administered in person, where test administrators observe intelligent behaviors (e.g., solving puzzles or pattern detection).

      • Counting physical traces, such as the number of tooth marks left on a pencil to operationalize stress.

      • Checking public records for documentation of recent marriages, divorces, or relocations.

      • Coding vehicle market values to operationalize social wealth.

  • Physiological Measures

    • Operationalize variables by recording biological data using technical equipment.

    • Facial Electromyography (EMG): Electronically records minute electrical changes in facial muscle movements (e.g., eye and cheek muscle activation patterns associated with smiling) to assess happiness.

    • Functional Magnetic Resonance Imaging (fMRI): Tracks blood flow patterns across brain regions while participants rest or perform tasks (e.g., Vickery et al., 2011 tracked brain activity during a computer rock-paper-scissors game, finding heightened blood flow across specific regions during wins versus losses).

    • Used to study neural differences in extroversion vs. introversion and high vs. low IQ scores (Dubois et al., 2018; Hsu et al., 2018).

    • Salivary Cortisol: Measures levels of the hormone cortisol released in saliva to evaluate stress levels (Carlson, 2009).

    • Skin Conductance: Electronically measures activity in the sweat glands of the hands or feet, which increases under high stress.

    • Electroencephalography (EEG): Detects distinct electrical activity patterns across brain regions near the scalp.

  • Validation of Physiological Measures

    • Physiological measurements are not inherently superior or more objective than self-report or observational measures.

    • Biological data must still be validated against behavioral or self-report criteria. For instance, linking fMRI brain scans to intelligence required administering observational IQ tests (Dubois et al., 2018), and linking neural patterns to happiness requires simultaneous self-reported happiness evaluations.

Scales of Measurement

  • Categorical vs. Quantitative Variables

    • Every variable must possess at least 22 levels.

    • Categorical Variables (Nominal Variables): Levels represent qualitative categories. Numerals assigned to categories during data entry carry no quantitative meaning.

      • Examples: First language (Spanish, English, Chinese, Arabic), species (1=rhesus macaque1 = \text{rhesus macaque}, 2=chimpanzee2 = \text{chimpanzee}, 3=bonobo3 = \text{bonobo}), nationality, music preferences, cell phone model.

    • Quantitative Variables (Continuous Variables): Levels are coded using meaningful numbers where higher values represent greater quantities of the construct.

      • Examples: Height (170 cm170\,\text{cm}), weight (65 kg65\,\text{kg}), Diener's well-being score (55 to 3535), IQ score, salivary cortisol level, degree of brain activity.

  • Three Types of Quantitative Measurement Scales

    • Ordinal Scale: Numerals represent a ranked order, but intervals between subsequent ranks are not necessarily equal.

      • Examples: Top 1010 best-selling book rankings, order of finishers in a swimming or marathon race, rankings of 1010 television shows from most to least favorite.

    • Interval Scale: Numerals represent equal distances between levels, but there is no true zero (a score of 00 does not represent complete absence of the variable).

      • Examples: Temperature in degrees Celsius (0∘C0^\circ\text{C} does not mean "no temperature"), IQ test scores, shoe sizes, 11-to-77 Likert-type agreement scales.

      • Ratio statements (e.g., "twice as hot" or "three times as happy") cannot be made using interval scales.

    • Ratio Scale: Numerals represent equal distances between levels, and there is a true zero (a value of 00 means literally "none" or zero quantity of the variable).

      • Examples: Number of exam questions answered correctly (0=zero correct answers0 = \text{zero correct answers}), number of eyeblinks in a stressful situation (0=zero blinks0 = \text{zero blinks}), number of TV episodes watched, height in centimeters.

      • Ratio comparisons (e.g., "answering twice as many problems") are statistically valid on ratio scales.

Measurement Reliability and Statistical Evaluation

  • Core Concept of Reliability

    • Reliability refers strictly to the consistency or stability of a measure's scores.

    • Establishing reliability is an empirical question requiring data collection.

  • Scatterplots and Correlation Coefficients

    • Scatterplots: Graphically display consistency by plotting initial test scores on the x-axis and retest scores on the y-axis. The proximity of points to a straight sloping line demonstrates measurement consistency.

    • Correlation Coefficient (rr): A single statistic ranging from −1.0-1.0 to 1.01.0 that quantifies both the direction and strength of an association between two variables.

      • Slope Direction: Positive (r>0r > 0, upward left-to-right slope), negative (r<0r < 0, downward left-to-right slope), or zero (r=.00r = .00, or near zero like .02.02 or −.04-.04).

      • Strength: Indicated by proximity to 1.01.0 or −1.0-1.0. Strong associations feature data points tightly clustered along a line (r=.70r = .70, r=.80r = .80, or r=−.70r = -.70); weak associations feature widely scattered points (rr close to 00).

  • Three Types of Reliability

    • Test-Retest Reliability:

      • Consistency of scores every time a measure is administered.

      • Crucial for constructs theoretically expected to remain stable over long periods (e.g., personality traits, IQ).

      • Evaluated by administering the same measure to the same group at two time points (weeks or months apart) and calculating rr.

      • An r≥.50r \ge .50 indicates strong test-retest reliability for stable traits (e.g., BFI extroversion).

      • Low test-retest reliability is expected and appropriate for constructs that fluctuate over time (e.g., fluctuating allergy symptoms or semester stress levels).

    • Interrater Reliability:

      • Degree to which two or more independent raters give consistent estimates or observations of the same behavior.

      • Relevant for observational/behavioral measures.

      • Evaluated for quantitative variables using the correlation coefficient rr (aiming for r≥.70r \ge .70).

      • Evaluated for categorical variables using Cohen's kappa (aiming for kappa close to 1.01.0, ideally ≥.80\ge .80).

      • A negative correlation coefficient rr between observers signifies severe raters' disagreement and unacceptable interrater reliability.

    • Internal Reliability (Internal Consistency):

      • Consistency of participant responses across multiple differently worded items within a single measure designed to assess the same construct.

      • Evaluated first via Average Inter-Item Correlation (AIC): The mean of all possible pairwise Pearson correlation coefficients computed between all individual items on a scale. Acceptable AIC values range between .15.15 and .50.50 (Clark & Watson, 2019).

      • Evaluated second via Cronbach's Alpha (Coefficient Alpha, α\alpha): A correlation-based statistic mathematically combining the AIC and the total number of scale items. The acceptable threshold for self-report scales is α≥.80\alpha \ge .80 (Clark & Watson, 2019).

      • Items failing to correlate with others (such as inserting "I am fond of polka dots" into Diener's well-being scale) lower α\alpha and must be revised or deleted.

Measurement Validity

  • Distinction Between Reliability and Validity

    • Reliability: Concerns how well a measure correlates with itself (consistency).

    • Validity: Concerns whether an operationalization measures what it is theoretically intended to measure (appropriateness of conclusions).

    • A measure can be reliable without being valid (e.g., using physical height as a measure of extroversion, or a faulty bathroom scale that consistently reads 50 lbs50\,\text{lbs} light every time).

    • A measure cannot be more valid than it is reliable; reliability is necessary, but not sufficient, for validity.

  • Subjective Evidence for Measurement Validity

    • Face Validity: The subjective appearance of a measure as a plausible operationalization of the target construct (often verified by consulting domain experts).

    • Content Validity: The subjective assessment of whether a measure captures all component parts of a theoretically defined construct.

      • Example: Gottfredson's (1997) conceptual definition of intelligence includes seven distinct components: reason, plan, solve problems, think abstractly, comprehend complex ideas, learn quickly, and learn from experience. An IQ test must include items assessing all seven components to maintain content validity.

  • Empirical Evidence for Measurement Validity

    • Criterion Validity:

      • Evaluates whether a measure correlates with expected concrete behavioral outcomes.

      • Correlational Evidence: Determining if test scores predict specific performance (e.g., comparing Sales Aptitude Test A vs. Test B scores against actual dollar sales figures; Test A showing stronger correlation demonstrates superior criterion validity).

      • BFI Extroversion Criterion Evidence: Self-reported BFI extroversion scores correlated negatively with time spent alone (r=−.27r = -.27) and positively with time spent talking (r=.30r = .30) across 9696 participants (Mehl et al., 2006).

      • Gallup Ladder of Life Criterion Evidence: Scores correlate with concrete behavioral outcomes such as work absenteeism due to illness.

    • The Known-Groups Paradigm:

      • A method for establishing criterion validity where researchers test whether scores on a measure discriminate between two or more groups whose physical or behavioral status is already confirmed.

      • Salivary Cortisol: Confirmed higher salivary cortisol in people assigned to give a public speech compared to audience members (Carlson, 2009).

      • Beck Depression Inventory (BDI; Beck et al., 1961):

        • 2121-item self-report scale assessing major depressive symptoms on a 00-to-33 point scale per item (total score range: 00 to 6363).

        • Psychiatrist-diagnosed depressed patients scored significantly higher on the BDI (M≈30M \approx 30) compared to non-depressed individuals (M≈11M \approx 11).

        • BDI scores increased monotonically across severity categories established by clinical psychiatric interviews (none, mild, moderate, severe).

      • Diener's Subjective Well-Being Scale: Group averages discriminate between populations with predictable differences in life quality (Pavot & Diener, 1993):

        • American college students: N=244N = 244, M=23.7M = 23.7, SD=6.4SD = 6.4

        • French Canadian college students (male): N=355N = 355, M=23.8M = 23.8, SD=6.1SD = 6.1

        • Korean university students: N=413N = 413, M=19.8M = 19.8, SD=5.8SD = 5.8

        • Printing trade workers: N=304N = 304, M=24.2M = 24.2, SD=6.0SD = 6.0

        • Veterans Affairs hospital inpatients: N=52N = 52, M=11.8M = 11.8, SD=5.6SD = 5.6

        • Women in shelters for abused spouses: N=70N = 70, M=20.7M = 20.7, SD=7.4SD = 7.4

        • Male prison inmates: N=75N = 75, M=12.3M = 12.3, SD=7.0SD = 7.0

      • Ladder of Life: Demonstrates known-groups validity via documented score drops during the COVID-19 pandemic (Witters & Harter, 2020), economic recessions, and following the 20212021 Taliban takeover in Afghanistan (Ray, 2022).

    • Convergent Validity:

      • Demonstrated when a self-report measure correlates strongly with other self-report measures assessing the same or theoretically similar constructs.

      • Example (Segal et al., 2008): BDI scores among 376376 adults correlated strongly and positively with the Center for Epidemiologic Studies Depression scale (CES-D; r=.68r = .68) and strongly and negatively with psychological well-being (r=−.65r = -.65).

    • Discriminant Validity (Divergent Validity):

      • Demonstrated when a self-report measure does not correlate strongly with measures of theoretically distinct or dissimilar constructs ("near neighbors").

      • Example (Segal et al., 2008): BDI scores correlated weakly with perceived physical health problems (r=.16r = .16), demonstrating that the BDI assesses depression rather than general physical illness.

      • Ensures screening tools discriminate between related conditions (e.g., distinguishing autism from language delays, or learning disabilities from overall intelligence/IQ).

Interpreting and Evaluating Construct Validity Evidence

  • Evaluating Journal Articles

    • Reliability and validity evidence for operationalized variables is found in the Method section of empirical journal articles.

    • Common Methodological Limitation: Many published studies report only internal reliability (Cronbach's alpha, α\alpha) without presenting criterion, convergent, or discriminant validity evidence (Flake et al., 2017).

    • Researchers must provide holistic evidence supporting measurement validity; without empirical validity data (such as criterion prediction or known-groups performance), readers should remain skeptical of a study's overall construct validity.