9/8 reading
Context and purpose of measurement
The U.S. War on Poverty (1964) spurred a large expansion of social programs (job training, Head Start, Medicare/Medicaid, housing, food stamps).
The problem: unclear how to measure poverty and know if policy is succeeding; measurement quality affects trustworthiness of research and program evaluation.
Measurement is the process of systematically observing a world feature and recording it as numbers or categories; it turns observations into analyzable data.
Distinction: observation (qualitative) vs measurement (quantitative/structured recording).
Measurement is central to public service delivery, program evaluation, and organizational management.
What is measurement? core idea and simple examples
Definition: systematic observation and recording of a feature, resulting in a number or category.
Simple example: counting park passersby with a pad and tallying by categories (children vs. adults) and time intervals to see patterns.
Real-world scale example: Census Bureau’s decadal population count requires years of planning, large workforces, and huge expenditure.
Measurement vs qualitative observation: core concepts (conceptualization, validity, reliability) apply to qualitative observation as well, especially in content analysis and coding.
Performance measurement: used for administrative purposes, leadership strategy, public accountability; focuses on activities, outputs, and outcomes of programs or organizations.
Basic model and road map: a basic measurement model (Allen & Yen, 1979) lays out the key elements and what the chapter will cover.
The Basic Measurement Model (conceptual overview)
Construct (trait): the concept we want to measure; requires clear conceptualization.
Indicator (empirical measure): the observable used to measure the construct; the task of operationalization.
Real-world measures rarely capture the construct completely; measurement error is inevitable.
Validity: the extent to which the observed measure corresponds to the intended construct (construct → observed measure).
Reliability: the extent to which the measure is free from random noise (noise → observed measure).
Central claim: measurement is a balance of conceptualization, operationalization, validity, and reliability.
Conceptualization: defining what we want to measure
Clear, precise definition of the construct is the first step; many concepts are straightforward (e.g., count of park entrants) whereas others are not (e.g., poverty).
Poverty conceptualization debate: absolute vs. relative definitions; how much is enough and what is enough a function of social context and values.
Dialogue example about poverty illustrates conceptual challenges:
Is poverty absolute (minimum income for survival) or relative (below what others have)?
Should necessities be defined by past norms (e.g., indoor plumbing in 18th century) or current standards?
Value judgments and politics influence conceptualization; this complicates international comparisons (poverty, unemployment, crime, education).
Constructs often originate from policy initiatives or organizational aims (e.g., War on Poverty, customer service) and then require refinement for measurement.
Box 4.1 Is Poverty the Same Thing the World Over? contrasts Orshansky’s absolute poverty concept with European relative measures and UN multidimensional approaches; emphasizes that constructs can be policy-driven and theory-informed.
Theoretical inputs: constructs come from theory and models; theory helps identify what to measure; logic models (application of theory to program design/evaluation) also define constructs to measure.
Logic models narrative typically includes conceptualizations that must be refined into actual measures.
Conceptualization: sources and dimensions
Constructs can arise from policy language, legislation, mission statements, or management goals.
Box 4.1 emphasizes that poverty is politically charged; different regions use different conceptualizations with different thresholds.
The broader point: measurement requires translating a concept into observable indicators, a process shaped by theory, policy, and practical constraints.
Latent vs manifest constructs; dimensions
Manifest constructs: directly observable traits (e.g., height, weight).
Latent constructs: not directly observable (e.g., self-esteem, mathematical knowledge).
Determining manifest vs latent status is a key part of conceptualization.
Psychometrics focuses on measuring latent traits using composite indicators (e.g., multiple questionnaire items).
Dimensions (domains) of a construct: health as a multidimensional concept (e.g., physical functioning, pain, emotional well-being, etc.).
SF-36 health survey: eight dimensions:
1) Physical functioning
2) Role limitations due to physical health
3) Role limitations due to emotional problems
4) Energy/fatigue
5) Emotional well-being
6) Social functioning
7) Pain
8) General healthIntelligence often viewed with multiple dimensions (e.g., Stanford–Binet five dimensions: fluid reasoning, knowledge, quantitative reasoning, visual-spatial processing, working memory).
Dimensions can be turned into distinct measures or combined into a single composite measure.
Controversies often center on emphasis or exclusion of certain dimensions.
Operationalization follows conceptualization (define how to measure each dimension).
Operationalization: turning concept into measurement
After defining conceptually, researchers specify how to measure it (operationalization).
Example: poverty measurement operationalization by Mollie Orshansky.
Orshansky conceptualized poverty as absolute and defined “less than the minimum income required to get by.”
The Agriculture Department used the Economy Food Plan as a baseline for nutrition costs; observed that families spent about one third of income on food and two thirds on other necessities, thus defining threshold as three times the basic food expenditure (adjusted for family size; updated with inflation each year).
Box 4.2 outlines the operational definition of U.S. poverty:
Baseline threshold: Economy Food Plan cost × 3; adjusted for family size
Update thresholds to present-day dollars via CPI
Use CPS data on incomes and family size to determine poverty status
Instruments: tools to measure constructs (e.g., breathalyzer, sphygmomanometer, thermometer). Questionnaires and coding sheets are also instruments.
Protocols and personnel: standardized procedures (e.g., standing height measurements) and trained interviewers; some fields require licensed professionals.
Proxies and indicators: proxies substitute for hard-to-measure constructs; proxies can mislead, but can be useful when data are limited.
Proxies vs indicators:
Proxy: substitute measure when direct measurement is unavailable
Indicator: a measure that captures a latent construct and is used within a multi-item approach
Proxies can apply to respondents answering for others (proxy respondents) and proxy reporting.
Indicators: questionnaire items that reflect latent constructs; multiple indicators used to build a scale.
Instruments, protocols, and proxies together form the operationalization of a construct.
Instruments and related operationalization details
Instruments include questionnaires, coding forms, data-extraction protocols; they help operationalize measures.
Protocols ensure consistency (e.g., how interviewers should administer measures).
Research personnel: training and supervision affect measurement quality; certain fields require specialized professionals.
Proxy respondents: common in CPS (one respondent per household reports on others’ work and income), can introduce error.
Indicators and composite measures: often used for latent constructs; combine indicators to form a scale that reduces random error.
Composite measures: scales and indexes
Composite measures combine multiple items to measure a latent construct; indicators are the items.
Benefits of multi-item scales:
Wider content coverage
Better capture of the construct’s range and intensity
Reduction of random measurement error through aggregation
Example: Rosenberg self-esteem scale (10 items); items assess agreement with statements; some items reverse-coded.
The measurement model (conceptual) shows arrows from the latent self-esteem construct to indicators, with measurement error (noise) affecting indicators.
Indicators reflect a mix of the underlying latent construct and error; multi-item scales separate construct from error using statistical techniques (e.g., confirmatory factor analysis, structural equation modeling).
Simple approach: sum items to form a scale; random errors tend to cancel out.
Response formats: items use various formats; common is Likert scale (e.g., strongly agree to strongly disagree):
Two-response formats: Agree/Disagree
Five-point or seven-point scales: strongly agree to strongly disagree, etc.
Box 4.3 discusses Likert scales and common confusions about scale terminology (Likert scale vs. individual items vs. composites).
Scales vs indexes:
Scale: multi-item measure with highly correlated items reflecting a latent trait; items are intercorrelated (often validated as a single construct).
Index: composite measure where items are not necessarily highly correlated (e.g., Consumer Price Index).
Some scholars distinguish scales (interrelated items) from indexes (aggregates that may include unrelated items).
A scale can also be a measured construct if items are ordered by intensity (e.g., Bogardus social distance scale).
Box 4.4 discusses item difficulty and Item Response Theory (IRT) as a method to relate item difficulty to latent traits; computer-adaptive testing (CAT) uses IRT to tailor item difficulty to the test-taker.
IRT elevates measurement precision by using item difficulty to gauge the underlying trait (e.g., math ability in test items).
Cat values: CAT starts with a mid-difficulty item; subsequent items adapt to the respondent’s ability; more efficient and precise.
Validity: does the measure capture the intended construct?
Validity is whether a measure truly represents the intended construct and captures variation in it.
Validity can be challenging to establish; Box 4.5 outlines multiple forms and nuances; not all researchers use terms consistently.
Types of validity (overview):
Face validity: does the measure look like it measures the intended construct? (subjective and intuitive; can be misleading, e.g., lie detector translates nervousness to lying)
Content validity: does the measure cover all dimensions of the construct? Does it capture the full range of variation? (important for pain vs. health; if a pain item misses other health dimensions, content validity may be low)
Criterion-related validity: empirical strategies to validate measures, including:
Concurrent validity: measure aligns with an existing contemporaneous measure of the same construct
Predictive validity: measure predicts future related behavior
Convergent validity: measure correlates with related constructs as theory would predict (e.g., health measure correlates with age in a theoretically expected way)
Discriminant validity: measure shows low correlation with unrelated constructs (e.g., health measure vs. ideology)
Nomological validity: validity within a broader theoretical network; correlations are consistent with theory.
Box 4.5 presents a list of validity forms; Box 4.6 provides an example validity study (self-reported height/weight compared to measured values).
Example: European Social Survey health measure should correlate with age (convergent validity) and show weak correlations with unrelated traits like ideology (discriminant validity); a correlation table (e.g., r = -0.36 with age) can demonstrate convergent validity.
Validity depends on purpose: a measure valid for broad population happiness may not be valid for clinical diagnosis (Beck Depression Inventory is more detailed for clinical use).
Box 4.6 provides a formal validity study example; Box 4.10–4.11 provide questions and tips for assessing validity in practice.
Reliability and measurement error
Reliability: consistency of a measure; reliability is about random error (noise) and the repeatability of measurements.
Measurement error components (classical view): Xi = Ti + Bi + Ni, where:
Xi = observed value
Ti = true value (the true score)
Bi = systematic bias (bias)
Ni = random noise (random error)
Observation quality: aim for Bi ≈ 0 and Ni ≈ 0; real measurements rarely achieve this perfectly.
Reliability vs validity: a measure can be reliable but not valid (consistent but measuring the wrong thing); a measure can be valid but not reliable (on average hitting the target but with wide dispersion).
Why reliability matters depends on use: averages, relationships, classifications, and tracking changes over time each have different reliability requirements.
Box 4.7 and 4.8 introduce bias vs random error; discuss sources of bias (question wording, social desirability), and random error (calibration, data entry, respondent variability). Classical Test Theory (CTT) formalizes the decomposition of observed scores into true score, bias, and noise.
Reliability methods (how to assess consistency)
Test-retest reliability: administer the same test twice to the same group; correlate results. Useful for stable traits but may be affected by learning or time-related changes.
Interrater reliability: multiple raters/observers score the same subject; assess consistency among raters.
Internal consistency (split-half reliability): for scales, split items into two halves and correlate; higher correlation indicates better internal consistency; Cronbach’s alpha summarizes average split-half reliability across all possible splits; alpha ~ 0.70 often considered acceptable, but depends on use (higher stakes require higher reliability).
Parallel forms reliability: use different but equivalent forms of a test to assess consistency across versions; important when tests must change over time (e.g., yearly standardized tests).
Reliability is necessary but not sufficient for validity; a reliable measure can still be biased or fail to capture the intended construct.
Figure 4.4 and related text illustrate reliability concepts (good vs. poor reliability); Figure 4.5 shows how increasing random error affects averages and confidence intervals; Figure 4.6 shows how reliability impacts relationships.
Reliability in qualitative research: validity and reliability concepts apply; intercoder reliability; code-recode reliability; qualitative validity concerns (does interpretation capture participants’ experiences consistently?).
Validity vs reliability: a practical contrast
A measure can be valid but unreliable, or reliable but invalid, or both, or neither (bull’s-eye analogy in Figure 4.7):
Reliable but not valid: shots clustered but off-target
Valid but not reliable: centered on target on average but widely dispersed
Both reliable and valid: clustered tightly around target
Neither: dispersed and off-target
Implications for measurement in practice:
For job performance measures, self-reports may be reliable (consistent) but not valid (biased upward);
Supervisor assessments may be more valid but face reliability concerns (inter-rater differences).
In qualitative research, validity and reliability translate to credible, trustworthy interpretations; intercoder reliability and code validity are key concerns.
The chapter emphasizes that validity and reliability are context-dependent and different measures may be valid for different purposes.
Levels of measurement, units of analysis, and data types
Levels of measurement (two broad types):
Quantitative variables: numbers that refer to actual quantities (e.g., age, income, hours worked, weight). Unit of measurement matters (e.g., dollars, kilograms).
Categorical variables: numbers refer to categories; can be nominal or ordinal.
Important distinctions:
Level of measurement (nominal, ordinal, interval, ratio): affects allowable statistics.
Unit of measurement: the unit (e.g., dollars, kilograms) that defines the quantitative variable.
Unit of analysis: the object described by the measure (people, households, neighborhoods, organizations).
Box 4.9 clarifies unit vs level vs unit of analysis distinctions.
Examples:
Household income coded into 12 categories in ESS; although labeled in euros, this is a categorical variable (not a precise quantity) unless midpoints are used.
Income could be treated as a quantitative variable if midpoints are assigned to categories (midpoint approximation) or if a multi-item scale sums to a continuous score.
Turning categorical variables into quantitative measures
Dummy variables (indicator variables): two-value categories (0/1) used to represent presence/absence (e.g., Employed: 0=no, 1=yes).
Using dummy variables for multi-category nominal variables: create a separate dummy per category (e.g., White, Black, Hispanic, Asian, Other).
Midpoint approximation: for ordinal measures with categories, use midpoints of ranges to approximate a quantitative score (e.g., income categories become approximate euros).
Multi-item scales: add up ordinal indicators to form a composite score; often treated as quantitative for analysis.
Endpoint scales and thermometers: 1–7 or 1–10 scales with endpoints anchored; some argue equal-interval interpretation across the scale; feeling thermometers (0–100) used for attitudes toward groups or leaders.
Figure 4.8 example: feeling thermometer illustrating usage of scales to quantify attitudes.
Levels of measurement and unit of analysis in practice
Unit of analysis vs level of measurement interact: a variable can be categorical at the individual level but become quantitative when aggregated to a geographic level (e.g., poverty rate by census tract).
Poverty example: individual poverty is a binary (categorical) variable; poverty rate by tract or county is a continuous quantitative measure.
The measurement in the real world: trade-offs and choices
Measurement is rarely perfect; practitioners balance validity, reliability, cost, and feasibility.
Costs and practicality:
Objective measures (clinical exams, bank records, hair/hair analyses) can be more valid but expensive.
Longer questionnaires increase reliability but raise respondent burden and reduce response rates; shorter measures reduce burden but may sacrifice reliability.
Reliability tends to improve with more indicators; multi-item scales generally provide more reliable measurements, but longer instruments increase cost and respondent burden.
Validity-reliability trade-off: lengthy, nuanced assessments (e.g., essays) may be valid in capturing complex constructs but less reliable due to interrater variability and scoring concerns; shorter tests improve reliability but may oversimplify constructs.
Established measures often preferred for reliability/validity; inventing new measures risks lower comparability across time and studies.
High-stakes measurement can induce behavior changes (gaming) that threaten validity; examples include test prep and coaching; public sector examples include fraud in performance-based pay schemes.
Multi-dimensional measures ( dashboards ) vs single headline indicators: EU and UN often use multiple dimensions; some argue for a single summary indicator for policy clarity; both approaches have trade-offs regarding comprehensiveness and comparability.
Measurement aggregation can obscure important differences across dimensions; a dashboard approach may be preferable for a fuller picture; aggregation weights are inherently arbitrary.
The chapter concludes with a call to thoughtful measurement: define concepts clearly, justify instrumentation and protocols, assess validity and reliability, and ask critical questions about measures.
Critical questions and practical guidance (Box summaries)
Box 4.10: Critical questions to ask about measurement
What is the purpose and origin of the measure?
What is the conceptual definition and its dimensions?
How is the measure operationalized (instruments, personnel, protocols)? Is it a single indicator or a multi-item scale? Is it a proxy or proxy reporting?
How valid is the measure (face, content, criterion-related; specific forms like concurrent, predictive, convergent, discriminant, nomological)?
How reliable is the measure (evidence and strength of reliability tests)?
What is the level of measurement?
Box 4.11: Tips on doing your own research: measurement planning steps (develop conceptual definitions, search for established measures, plan operationalization, decide on single item vs multi-item scales, consider proxies, test validity/reliability, review existing literature).
Key terms (glossary-style references)
Bias, measurement bias, random measurement error (noise)
Construct, latent vs manifest construct
Conceptualization, operationalization
Instrument, protocol, proxy, proxy respondent, proxy reporting
Indicator, composite measure, scale, index
Validity (face, content, criterion-related, convergent, discriminant, nomological)
Reliability (test-retest, interrater, split-half, internal consistency, Cronbach’s alpha, parallel forms)
Item Response Theory (IRT), computer-adaptive testing (CAT)
Levels of measurement (nominal, ordinal, interval, ratio)
Unit of measurement, unit of analysis
Dimensionality and multi-item scales (SF-36 dimensions; Rosenberg self-esteem scale)
Qualitative validity and reliability concepts (code validity, intercoder reliability)
Technical concepts: constructs, indicators, proxies, and the measurement model
Connections to theory, policy, and practice
Measurement translates abstract policy concepts (like poverty) into observable data, enabling evaluation, accountability, and resource allocation.
The debate over poverty measures illustrates how theory, politics, and data availability shape measurement choices and policy implications.
Logic models and theory-driven measurement connect theoretical propositions to empirical tests via carefully defined constructs and indicators.
The balance between validity and reliability, and the choice between single-item measures vs. multi-item scales, reflect practical trade-offs in policy research and program evaluation.
The chapter emphasizes that measurements are always context-dependent: a measure can be valid for one purpose and not for another; the same measure can have different validity in different settings or times.
Real-world relevance and ethical considerations
Measurement choices affect policy conclusions, program funding, and public accountability.
Cost, respondent burden, and ethical concerns (privacy, deception) influence instrument design and data collection.
The potential for gaming high-stakes measurements requires thoughtful design to preserve validity (e.g., avoiding perverse incentives, multiple measures to capture dimensions).
Researchers should use established, validated measures when possible to ensure comparability and reliability across time and settings, while remaining mindful of potential biases and measurement drift over time.
Summary takeaways
Measurement is a structured process that involves building a bridge from abstract concepts to observable data, with emphasis on conceptualization, operationalization, validity, and reliability.
Concepts can be manifest or latent, multi-dimensional, and sometimes require composite indicators or scales to capture them adequately.
Validity and reliability are not absolute properties of a measure but properties of a measure in a given context and use; they must be studied and reported.
Real-world measurement involves trade-offs among costs, precision, burden, and ethical considerations; multi-dimensional measures and dashboards can provide richer information, while simple measures facilitate comparability and communication.
Critical questions and careful planning are essential for credible measurement in research and policy work.