CHAPTER 5 - Identifying Good Measurement
Operationalization and Conceptual Definitions
Construct Validity in Psychological Research
Construct validity is one of the four big validities in psychological research, applying to frequency, association, and causal claims.
It measures how well a study's variables are operationalized, measured, or manipulated.
Evaluating construct validity is essential when measuring abstract psychological phenomena such as motivation, emotion, thinking, reasoning, gratitude, extroversion, and happiness.
Conceptual vs. Operational Definitions
Conceptual Definition (Construct): The theoretical definition of a variable in question.
Operational Definition (Operationalization): The specific decision made by a researcher regarding how to measure or manipulate a conceptual variable in a given study.
Operationalizing "Happiness" (Subjective Well-Being)
Conceptualization: Ed Diener and colleagues explicitly defined happiness conceptually as "subjective well-being" (a person's well-being from their own perspective).
Operationalization via Diener's Subjective Well-Being Scale:
A -item self-report questionnaire where participants evaluate their life satisfaction using custom criteria.
Responses are made on a -point scale ranging from ("strongly disagree") to ("strongly agree").
The five items are:
In most ways my life is close to my ideal.
The conditions of my life are excellent.
I am satisfied with my life.
So far, I have gotten the important things I want in life.
If I could live my life over, I would change almost nothing.
Scoring Structure:
Minimum score (unhappiest): ().
Maximum score (happiest): ().
Neutral score: ().
Empirical Findings: Diener and Diener (1996) reported that most people score above : of high school and college students and of adults with disabilities scored above .
Operationalization via the Ladder of Life (Cantril, 1965):
A single-item self-report measure asking participants to imagine a ladder with steps numbered from at the bottom (worst possible life) to at the top (best possible life).
Participants indicate which step they personally stand on at the present time.
Utilized daily by the Gallup organization in the Gallup-Healthways Well-Being Index.
Operationalizing Other Psychological Constructs
Extroversion:
Self-report via the -item extroversion subscale of the Big Five Inventory (BFI; John et al., 2008), which includes items like "I see myself as someone who is talkative."
Self-report via the -item extroversion subscale of the Ten-Item Personality Inventory (TIPI; Gosling et al., 2003), which includes items like "I see myself as extroverted, enthusiastic."
Peer reports obtained by asking participants' friends to rate their extroversion.
Gratitude Toward Partner:
Self-report questionnaire asking agreement with statements such as "I appreciate my partner" or asking how frequently someone thanks their partner.
Observational behavioral coding by watching couples interact and counting instances of explicit verbal thanks.
Gender Identity:
Self-report question: "Which best applies to you?" with options: Man, Woman, Nonbinary, and Another identity.
Observational/archival measurement by recording the gender identity stated on an individual's social media profile.
Wealth:
Self-report question asking: "What is your total annual household income before taxes?"
Observational measure by coding the monetary value of a person's vehicle on a scale from (older, lower-status vehicle) to (new, higher-status vehicle in good condition) (Piff et al., 2012).
Three Common Types of Operational Measures
Self-Report Measures
Operationalize variables by recording participants' answers to questions about themselves in questionnaires or interviews.
Examples include Diener's -item Life Satisfaction Scale, Cantril's Ladder of Life, the Big Five Inventory (BFI), and Holmes & Rahe's (1967) stressful life events checklist (recording frequencies of events like marriage, divorce, or moving).
When researching children, self-reports are often replaced with parent reports or teacher reports (e.g., evaluating classroom behavior, known vocabulary words, or recent developmental events).
Observational Measures (Behavioral Measures)
Operationalize variables by recording observable behaviors or physical traces of behaviors.
Examples:
Counting the frequency of smiles to operationalize moment-to-moment happiness.
Standardized Intelligence (IQ) tests administered in person, where test administrators observe intelligent behaviors (e.g., solving puzzles or pattern detection).
Counting physical traces, such as the number of tooth marks left on a pencil to operationalize stress.
Checking public records for documentation of recent marriages, divorces, or relocations.
Coding vehicle market values to operationalize social wealth.
Physiological Measures
Operationalize variables by recording biological data using technical equipment.
Facial Electromyography (EMG): Electronically records minute electrical changes in facial muscle movements (e.g., eye and cheek muscle activation patterns associated with smiling) to assess happiness.
Functional Magnetic Resonance Imaging (fMRI): Tracks blood flow patterns across brain regions while participants rest or perform tasks (e.g., Vickery et al., 2011 tracked brain activity during a computer rock-paper-scissors game, finding heightened blood flow across specific regions during wins versus losses).
Used to study neural differences in extroversion vs. introversion and high vs. low IQ scores (Dubois et al., 2018; Hsu et al., 2018).
Salivary Cortisol: Measures levels of the hormone cortisol released in saliva to evaluate stress levels (Carlson, 2009).
Skin Conductance: Electronically measures activity in the sweat glands of the hands or feet, which increases under high stress.
Electroencephalography (EEG): Detects distinct electrical activity patterns across brain regions near the scalp.
Validation of Physiological Measures
Physiological measurements are not inherently superior or more objective than self-report or observational measures.
Biological data must still be validated against behavioral or self-report criteria. For instance, linking fMRI brain scans to intelligence required administering observational IQ tests (Dubois et al., 2018), and linking neural patterns to happiness requires simultaneous self-reported happiness evaluations.
Scales of Measurement
Categorical vs. Quantitative Variables
Every variable must possess at least levels.
Categorical Variables (Nominal Variables): Levels represent qualitative categories. Numerals assigned to categories during data entry carry no quantitative meaning.
Examples: First language (Spanish, English, Chinese, Arabic), species (, , ), nationality, music preferences, cell phone model.
Quantitative Variables (Continuous Variables): Levels are coded using meaningful numbers where higher values represent greater quantities of the construct.
Examples: Height (), weight (), Diener's well-being score ( to ), IQ score, salivary cortisol level, degree of brain activity.
Three Types of Quantitative Measurement Scales
Ordinal Scale: Numerals represent a ranked order, but intervals between subsequent ranks are not necessarily equal.
Examples: Top best-selling book rankings, order of finishers in a swimming or marathon race, rankings of television shows from most to least favorite.
Interval Scale: Numerals represent equal distances between levels, but there is no true zero (a score of does not represent complete absence of the variable).
Examples: Temperature in degrees Celsius ( does not mean "no temperature"), IQ test scores, shoe sizes, -to- Likert-type agreement scales.
Ratio statements (e.g., "twice as hot" or "three times as happy") cannot be made using interval scales.
Ratio Scale: Numerals represent equal distances between levels, and there is a true zero (a value of means literally "none" or zero quantity of the variable).
Examples: Number of exam questions answered correctly (), number of eyeblinks in a stressful situation (), number of TV episodes watched, height in centimeters.
Ratio comparisons (e.g., "answering twice as many problems") are statistically valid on ratio scales.
Measurement Reliability and Statistical Evaluation
Core Concept of Reliability
Reliability refers strictly to the consistency or stability of a measure's scores.
Establishing reliability is an empirical question requiring data collection.
Scatterplots and Correlation Coefficients
Scatterplots: Graphically display consistency by plotting initial test scores on the x-axis and retest scores on the y-axis. The proximity of points to a straight sloping line demonstrates measurement consistency.
Correlation Coefficient (): A single statistic ranging from to that quantifies both the direction and strength of an association between two variables.
Slope Direction: Positive (, upward left-to-right slope), negative (, downward left-to-right slope), or zero (, or near zero like or ).
Strength: Indicated by proximity to or . Strong associations feature data points tightly clustered along a line (, , or ); weak associations feature widely scattered points ( close to ).
Three Types of Reliability
Test-Retest Reliability:
Consistency of scores every time a measure is administered.
Crucial for constructs theoretically expected to remain stable over long periods (e.g., personality traits, IQ).
Evaluated by administering the same measure to the same group at two time points (weeks or months apart) and calculating .
An indicates strong test-retest reliability for stable traits (e.g., BFI extroversion).
Low test-retest reliability is expected and appropriate for constructs that fluctuate over time (e.g., fluctuating allergy symptoms or semester stress levels).
Interrater Reliability:
Degree to which two or more independent raters give consistent estimates or observations of the same behavior.
Relevant for observational/behavioral measures.
Evaluated for quantitative variables using the correlation coefficient (aiming for ).
Evaluated for categorical variables using Cohen's kappa (aiming for kappa close to , ideally ).
A negative correlation coefficient between observers signifies severe raters' disagreement and unacceptable interrater reliability.
Internal Reliability (Internal Consistency):
Consistency of participant responses across multiple differently worded items within a single measure designed to assess the same construct.
Evaluated first via Average Inter-Item Correlation (AIC): The mean of all possible pairwise Pearson correlation coefficients computed between all individual items on a scale. Acceptable AIC values range between and (Clark & Watson, 2019).
Evaluated second via Cronbach's Alpha (Coefficient Alpha, ): A correlation-based statistic mathematically combining the AIC and the total number of scale items. The acceptable threshold for self-report scales is (Clark & Watson, 2019).
Items failing to correlate with others (such as inserting "I am fond of polka dots" into Diener's well-being scale) lower and must be revised or deleted.
Measurement Validity
Distinction Between Reliability and Validity
Reliability: Concerns how well a measure correlates with itself (consistency).
Validity: Concerns whether an operationalization measures what it is theoretically intended to measure (appropriateness of conclusions).
A measure can be reliable without being valid (e.g., using physical height as a measure of extroversion, or a faulty bathroom scale that consistently reads light every time).
A measure cannot be more valid than it is reliable; reliability is necessary, but not sufficient, for validity.
Subjective Evidence for Measurement Validity
Face Validity: The subjective appearance of a measure as a plausible operationalization of the target construct (often verified by consulting domain experts).
Content Validity: The subjective assessment of whether a measure captures all component parts of a theoretically defined construct.
Example: Gottfredson's (1997) conceptual definition of intelligence includes seven distinct components: reason, plan, solve problems, think abstractly, comprehend complex ideas, learn quickly, and learn from experience. An IQ test must include items assessing all seven components to maintain content validity.
Empirical Evidence for Measurement Validity
Criterion Validity:
Evaluates whether a measure correlates with expected concrete behavioral outcomes.
Correlational Evidence: Determining if test scores predict specific performance (e.g., comparing Sales Aptitude Test A vs. Test B scores against actual dollar sales figures; Test A showing stronger correlation demonstrates superior criterion validity).
BFI Extroversion Criterion Evidence: Self-reported BFI extroversion scores correlated negatively with time spent alone () and positively with time spent talking () across participants (Mehl et al., 2006).
Gallup Ladder of Life Criterion Evidence: Scores correlate with concrete behavioral outcomes such as work absenteeism due to illness.
The Known-Groups Paradigm:
A method for establishing criterion validity where researchers test whether scores on a measure discriminate between two or more groups whose physical or behavioral status is already confirmed.
Salivary Cortisol: Confirmed higher salivary cortisol in people assigned to give a public speech compared to audience members (Carlson, 2009).
Beck Depression Inventory (BDI; Beck et al., 1961):
-item self-report scale assessing major depressive symptoms on a -to- point scale per item (total score range: to ).
Psychiatrist-diagnosed depressed patients scored significantly higher on the BDI () compared to non-depressed individuals ().
BDI scores increased monotonically across severity categories established by clinical psychiatric interviews (none, mild, moderate, severe).
Diener's Subjective Well-Being Scale: Group averages discriminate between populations with predictable differences in life quality (Pavot & Diener, 1993):
American college students: , ,
French Canadian college students (male): , ,
Korean university students: , ,
Printing trade workers: , ,
Veterans Affairs hospital inpatients: , ,
Women in shelters for abused spouses: , ,
Male prison inmates: , ,
Ladder of Life: Demonstrates known-groups validity via documented score drops during the COVID-19 pandemic (Witters & Harter, 2020), economic recessions, and following the Taliban takeover in Afghanistan (Ray, 2022).
Convergent Validity:
Demonstrated when a self-report measure correlates strongly with other self-report measures assessing the same or theoretically similar constructs.
Example (Segal et al., 2008): BDI scores among adults correlated strongly and positively with the Center for Epidemiologic Studies Depression scale (CES-D; ) and strongly and negatively with psychological well-being ().
Discriminant Validity (Divergent Validity):
Demonstrated when a self-report measure does not correlate strongly with measures of theoretically distinct or dissimilar constructs ("near neighbors").
Example (Segal et al., 2008): BDI scores correlated weakly with perceived physical health problems (), demonstrating that the BDI assesses depression rather than general physical illness.
Ensures screening tools discriminate between related conditions (e.g., distinguishing autism from language delays, or learning disabilities from overall intelligence/IQ).
Interpreting and Evaluating Construct Validity Evidence
Evaluating Journal Articles
Reliability and validity evidence for operationalized variables is found in the Method section of empirical journal articles.
Common Methodological Limitation: Many published studies report only internal reliability (Cronbach's alpha, ) without presenting criterion, convergent, or discriminant validity evidence (Flake et al., 2017).
Researchers must provide holistic evidence supporting measurement validity; without empirical validity data (such as criterion prediction or known-groups performance), readers should remain skeptical of a study's overall construct validity.