9.10 reading
Generalizability and External Validity
Generalizability: the extent to which findings from a study can be projected to a larger population, time period, or different contexts.
External validity is another term for generalizability (Shadish, Cook, & Campbell, 2002).
In practice, researchers care about what the study implies for the broader world, not just the specific sample.
Katrina example: CBS poll of 725 adults was used to infer thinking of a few hundred million people; the value lies in broader implications, not the exact individuals polled.
Population, Sampling Frame, and Generalizability
Population of interest: the entire group the study aims to learn about (e.g., all U.S. adults in the Katrina poll).
Sampling frame: a concrete list or operational representation from which the sample is drawn (e.g., voter lists, phone numbers, organizational rosters).
Parameters vs statistics: the study aims to learn about population characteristics (parameters); the sample yields statistics that estimate these parameters.
The closer the sample’s results are to true population parameters, the more generalizable the results.
Broader population, geography, time, and groups increase generalizability; studying many places and times tends to improve external validity.
Random (probability) sampling tends to be more generalizable than nonrandom sampling; small random samples often beat large nonprobability samples for generalizability.
Examples: Katrina poll (random sample) vs Red Cross shelter study (convenience/nonrandom sampling) – both informative but with different generalizability limits.
Question: which features of a sample affect generalizability? This is a prelude to later sections.
Are Experiments More Generalizable?
Many biological, psychological, or economic processes are fairly universal, making findings generalizable even from small, unrepresentative samples (especially in controlled experiments).
Experiments are used to determine causal relationships ("what if" questions).
Examples:
Drug/medical trials often use clinical volunteers (generalizability can be limited).
Psychological experiments often use undergraduates and still yield generalizable laws of perception, cognition, and behavior.
Experimental economists study altruism and risk aversion with small samples (e.g., ultimatum game, prisoners’ dilemma).
Caveat: generalizability is not guaranteed; experiments can be criticized for homogeneous or idiosyncratic samples and limited external validity.
Replication and Meta-Analysis
Replication: repeating a study with different samples, places, times, or designs to test robustness and generalizability.
Replication enhances generalizability of findings from small or nonrandom samples.
Meta-analysis: pooling multiple studies to produce a larger, more generalizable estimate of a treatment effect or relationship.
Formal definitions: meta-analysis combines separate effects into a single, generalizable estimate.
Examples: air pollution and daily mortality across many cities; second-generation antipsychotics efficacy across 124 experiments (Davis, Chen, & Glick, 2003).
Applications: health, education, social work, criminal justice, job training, etc.
Relationships and Generalizability (Health and Happiness in Moldova)
Relationships among variables (not just descriptive percentages) tend to generalize better.
World Values Survey data: Moldova (small Eastern European country) has means and correlations similar to global patterns in health and happiness, despite Moldova’s low GDP per capita.
Global correlation (health vs happiness): ≈
All nations:
Moldova (n ≈ 974):
Nigerians (n ≈ 2,021):
Implication: even small country data can yield generalizable insights about broader relationships; experiments designed to test causal theories can exhibit good generalizability when focusing on relationships.
Generalizability of Qualitative Studies
Qualitative research often uses nonprobability, small samples and is not generalizable in the statistical sense.
However, qualitative work can generate generalizable theories about how and why things happen (hows and whys).
Key idea: deep, longitudinal observation can reveal universal features of a setting beyond surface-level specifics, given time, effort, openness, and researcher judgment.
Basic Sampling Concepts
Sampling: selecting people or elements from a population for inclusion in a study due to limited resources/time.
Population, sample, and inference:
Population: the entire group of interest.
Sample: a subset actually studied.
Inference: drawing conclusions about the population from the sample.
Population vs sampling frame vs units:
Population: who we want to learn about.
Sampling frame: operational representation of the population (the list or method used to select the sample).
The closer the frame fits the population, the better the potential inference.
Census: data on the entire population; sampling involves selecting a subset when a census is impractical.
Examples: air quality in a city (airshed); reimbursements in an organization; exits polls outside polling stations.
Steps in sampling (summary):
Define the population of interest.
Identify a sampling frame representing that population.
Select a subset (randomly, ideally) from the frame.
Contact sampled units and request participation.
Record responses/observations.
Summarize findings and infer about the population.
How Large Does My Sample Need to Be?
Precision matters: larger samples reduce random error (increase precision).
Important considerations:
If you plan to analyze subgroups (e.g., men vs women), the subgroup sample size matters for precision.
For a given precision, sample size does not depend on population size; larger populations don’t require larger samples to achieve the same precision.
Quick guidelines introduced (to be formalized later):
Larger desired precision → larger sample size.
Subgroup analyses require larger overall samples to ensure adequate subgroup sizes.
The census (complete population) is only practical for small populations or readily available frames; otherwise sampling is preferable.
Practical note: final achievable sample size often constrained by time/resources; researchers must trade off precision for feasibility.
Problems and Biases in Sampling
Two major difficulties in achieving random/representative samples:
Coverage: whether the sampling frame adequately covers the population of interest.
Nonresponse: whether selected units actually respond/participate.
Sampling bias: systematic differences between the sample and population caused by problems in the sampling process.
Distinction: sampling bias (systematic) vs precision (random error).
Coverage bias vs nonresponse bias:
Coverage bias arises when the sampling frame misses parts of the population or includes ineligible units.
Nonresponse bias arises when those who respond differ on the outcome of interest from those who do not.
Coverage problems in telephone surveys: unlisted numbers; random digit dialing (RDD) helps mitigate but not eliminate coverage bias (cell-phone-only households, younger individuals).
Nonresponse issues: response rate = contact rate × cooperation rate.
Example: 50% contact rate × 70% cooperation rate = 35% response rate.
Real-world: many surveys have response rates below 50% (Pew, 2012a, as low as 9% in some cases).
Nonresponse bias models (propensity to respond P, outcome Y, other variables X and Z):
Reverse cause model: Y causes P (e.g., recycling behavior influences willingness to respond).
Common cause model: a variable Z drives both P and Y (e.g., residency status affecting both survey response and club preferences).
Separate causes model: P and Y are influenced by different factors; nonresponse may be essentially random and ignorable.
Coverage bias assessment steps (Box 5.1): define target population, identify sampling frame, assess systematic differences, determine relation of coverage to outcomes, predict bias direction.
Nonresponse bias assessment steps (Box 5.1): identify nonresponse, assess differences between responders and nonresponders, determine relation of response propensity to outcome, predict bias direction.
Ethics of nonresponse: voluntary participation is essential; coercion is unethical; IRBs oversee recruitment ethics; deception to reduce nonresponse is unethical.
Nonresponse and ethics: even with nonresponse challenges, respect for participants remains paramount; researchers should be conscientious and transparent.
Sampling bias vs generalizability: sampling bias is a subset issue; generalizability depends on broader considerations like population definition and replication across studies.
Nonprobability sampling: often used in practice for practical reasons; includes voluntary, convenience, snowball, and purposive sampling.
Box 5.2: Sampling Bias definitions (sampling frame coverage, nonresponse, voluntary and other biases).
Box 5.5–5.6: Critical questions for evaluating sampling in studies and tips for conducting sampling in practice.
Ethics of Nonresponse and Nonprobability Sampling
Ethics of recruitment: informed consent, voluntary participation, avoidance of coercion.
Nonresponse is an acknowledged cost of respecting participants' rights; sometimes reduction strategies involve incentives, multiple contact modes, and follow-ups.
Transparency about response rates and potential biases is essential for proper interpretation.
Nonprobability Sampling
Why used: cost, practicality, access, or particular research aims (e.g., exploratory or qualitative work).
Types:
Voluntary sampling: participants respond to an explicit call for volunteers (e.g., Craigslist ads). May suffer from volunteer bias.
Convenience sampling: recruit those readily available (e.g., Red Cross shelter refugees; university subject pools).
Snowball/Respondent-driven sampling: initial respondents refer others; useful for hard-to-reach populations.
Purposive sampling: select individuals with specific characteristics to study particular phenomena or theoretical categories.
Internet sampling contrasts:
Open web polls: voluntary and self-selected; often biased by self-selection.
Internet access panels: managed panels with weighting; can approximate probability samples but still vulnerable to biases.
Weighting (post hoc) can adjust for known differences, but cannot fix unobserved biases.
Practical questions: open web polls tend to overrepresent highly motivated groups (e.g., gun-rights advocates in gun-control polls); internet panels may be less biased when well-designed and properly weighted.
Purposive qualitative sampling emphasizes theory-building and causal reasoning over population representativeness; sequential design may build toward generalizable causal theories.
Random (Probability) Sampling vs Randomized Experiments
Random sampling (probability sampling): select elements from a population to generalize to that population; observational in nature.
Randomized experiments: assign participants to treatments to test causal effects; goals are internal validity and causation, not representativeness of a population.
In practice, many studies use random samples for observational studies or nonrandom samples for experiments; rare to have both random sampling and random assignment in the same study.
Simple Random Sampling: Concept and Calculation
Simple random sampling: every element has an equal chance of selection.
Example: unemployment rate estimation via simple random sample of n = 400 from the labor force.
Suppose the sample has p̂ = 22 unemployed out of 400 → p̂ = 0.055 (5.5%).
Sampling variability: any other sample of the same size would yield a different p̂ due to chance.
Sampling distribution: the distribution of p̂ across many samples; center tends toward the population parameter P; shape approximates a normal distribution for large enough samples.
Standard error (SE): the spread of the sampling distribution; for a proportion:
The SE depends on the population variance and the sample size.
Confidence intervals (CI) and margins of error:
A common 95% CI uses roughly two standard errors:
Example: if p̂ = 0.055 and SE ≈ 0.011, then CI ≈ 0.055 ± 0.022 → (0.033, 0.077).
The concept of a sampling distribution helps explain why CI/widens or tightens with sample size.
What Is the True Sample Size? Observed vs True Sample
True sample: the actual randomly selected units from the population (the theoretical basis for inference).
Observed sample: the units actually contacted and who participated; may be smaller than the true sample due to nonresponse.
Illustrative example: a study of employer disseminated quality info cited 1,365 employees interviewed (60% of the true sample of 2,275 initially selected).
Nonresponse can substantially reduce effective sample size and bias results; even good response rates may not reflect the true population if nonresponse is systematic.
A useful caution: reported statistics in publications often refer to the observed sample; the true sampling frame and response rate are essential for proper interpretation.
Exit Polls, Coverage, and Framing Effects
Exit polls sample voters leaving polling places; coverage may miss absentee or early voters, altering the frame.
Coverage issues can lead to systematic biases in outcomes (e.g., misestimating political shares due to who is captured by the frame).
Open Web Polls vs Internet Panels: Bias and Weighting
Open web polls: high volunteer bias; people with strong opinions are more likely to respond.
Internet access panels: panelists recruited to participate in surveys; weighting can help adjust for demographic differences, but bias may persist depending on who joins the panel and why.
Mixed evidence on bias: some internet panels yield results similar to traditional probability samples; others show persistent bias.
The 2012 U.S. presidential election highlighted variability among polls depending on method; some online methods performed well, others did not.
Weighting: adjusting the sample to reflect known population characteristics (e.g., gender, region); but cannot fix unobserved biases.
Purposive Sampling and Qualitative Research (Table/Grid Example)
Purposive sampling grid (Table 5.3): selecting a range of cases to capture theoretical diversity; used in case studies and comparative research.
Qualitative sampling aims for depth and causal understanding rather than statistical representativeness; generalizability comes from theoretical saturation and logical reasoning, not randomization.
Random (Probability) Sampling: Forms and Practicalities
Forms of random sampling: systematic, stratified, cluster, multistage, PPS (probability proportional to size), random digit dialing (RDD).
Simple random sampling as the foundation; complex methods adapt to real-world constraints.
Complex sampling often requires design effects and weighting to produce valid standard errors and confidence intervals.
The Contribution of Random Sampling
Random sampling underpins major government surveys (e.g., CPS, NCVS, NHIS, NAEP) and social surveys (e.g., GSS, ANES, European Social Survey, World Values Survey).
Distinction from randomized experiments:
Random sampling aims to estimate population parameters;
Randomized experiments aim to test causal effects between treatments and outcomes; participants are often volunteers.
Simple Random Sampling Details and Example: Unemployment Rate
Simple random sampling: equal probability of selection for each unit.
Example calculation: estimating unemployment rate with n = 400, p̂ = 0.055.
Sampling variability and the concept of a sampling distribution justify E = margin of error and CI.
Confidence Intervals and Margins of Error (Box/Concepts)
Margin of error (MOE) and CI are based on sampling variability and the assumed sampling distribution.
A typical 95% CI uses roughly ±2 × SE, assuming normal approximation.
Important caveats:
MOE is approximate and depends on assumptions (e.g., p = 0.5 for the most conservative SE).
Subgroup CIs are larger when subgroup sizes are smaller.
MOE does not account for non-sampling errors (measurement error, question wording, processing errors).
Interpretation example: a poll with p̂ = 0.67 and MOE = ±3.1 percentage points means 95% CI for the population proportion lies between 63.9% and 70.1%.
Sample Size and the Precision of Government Statistics
Increasing sample size tightens CIs and reduces MOE; larger samples improve precision for policy decisions but have diminishing returns (MOE decreases roughly with the square root of sample size).
Example: n = 400 versus n = 10,000; larger samples yield much narrower CIs (e.g., unemployment rate CI from roughly 5.1% to 5.9% for n = 10,000 in a similar scenario).
CPS (60,000 households) demonstrates the scale needed for national statistics; state/local precision is lower due to smaller sub-sample sizes.
How to Determine an Appropriate Sample Size (Practical Rules)
Guidelines for simple random samples:
Decide desired precision (MOE) and confidence level.
Use the conservative formula for proportions (p = 0.5 for maximum variability):
For 95% CI (Z ≈ 2) and E = 0.03, n ≈ 1111 (round up for safety).
If subgroups are planned (e.g., Hispanics ~17% of the population), adjust n to ensure sufficient subgroup precision:
Required total n ≈ (subgroup n) / 0.17 = 1,111 / 0.17 ≈ 6,535.
Population size (N) does not affect required n for a given precision in large populations; large populations require the same n as smaller populations.
If anticipated response rate is less than 100%, inflate the initial sample to achieve the desired final number of completed interviews.
In practice, sometimes a census (surveying all units in a frame) is preferable when frames are readily available and low-cost (e.g., organizational web surveys with email lists).
Problems and Biases in Sampling (Detailed)
Coverage bias: when the sampling frame misses parts of the population or includes ineligible units, leading to biased estimates.
Telephone surveys face coverage problems due to unlisted numbers and cell-phone-only households.
Exit polls: coverage limitations may exclude absentee voters; coverage must be assessed for potential bias.
Nonresponse bias: when nonresponders differ systematically from responders on outcomes of interest.
Response rate = contact rate × cooperation rate.
Low response rates increase risk of bias; disclosure of actual response rates is critical for assessment.
Nonresponse bias directions:
Reverse cause model: Y affects P (response propensity);
Common cause model: Z affects both P and Y (confounding by a third factor);
Separate causes model: P and Y influenced by different factors; nonresponse may be ignorable.
Coverage vs nonresponse bias: not all coverage problems cause bias; the relationship between coverage, response propensity, and outcome determines bias.
Open web polls vs internet panels: open polls are highly biased; panels with weighting can improve accuracy but depend on proper adjustment and non-observed biases.
Volunteer bias (nonresponse in voluntary samples) occurs when volunteers differ systematically from the target population on the study’s outcomes.
Ethics of sampling: respect for voluntary participation; use of institutional review boards (IRBs); deceptive practices to increase response rates are unethical.
Nonprobability Sampling: Details, Pros, and Cons
Voluntary sampling: exposed to volunteer bias; may attract participants with strong opinions or particular traits.
Convenience sampling: easy access but may suffer from coverage bias; useful for exploratory purposes.
Snowball sampling: useful for hard-to-reach populations; relies on referrals; may introduce bias if initial seeds are not representative.
Quota sampling: nonprobability method that fills quotas to mirror known population shares; can improve representativeness but risk of unobserved bias.
Internet sampling:
Open web polls: high bias due to self-selection; results may not generalize.
Internet access panels: subject to bias but can be weighted; may approximate probability sampling if designed carefully.
Open vs weighted internet samples: weights can adjust some known biases, but unobserved variables may still bias results.
Purposive Sampling and Qualitative Research (Deep Dive)
Purposive sampling: select individuals with specific characteristics or insights to develop theories or understand causal mechanisms.
Sequential, theory-driven design: researchers build generalizable causal theories across connected studies rather than aiming for population representativeness.
Qualitative sampling emphasizes depth, context, and the development of causal explanations, not statistical generalizability.
Potential biases in qualitative work: nonresponse or nonparticipation can still affect the breadth and depth of theory; researchers should consider who is missing and whether those differences matter for findings.
Random (Probability) Sampling: Methods and Theory
Random sampling uses chance to select population elements, enabling statistical inference about population characteristics.
Distinction from randomized experiments: the latter assign treatments; the former select samples to infer population parameters.
Practical reality: random sampling is often combined with nonrandom elements due to constraints; still, random sampling remains the best method for generalizability when feasible.
The Simplicity and Power of Simple Random Sampling
Simple random sampling provides the foundation for statistical theory and inference.
Other sampling methods (systematic, stratified, cluster, multi-stage, PPS, RDD) are variations designed for practicality but rely on the same underlying logic.
Illustrative unemployment example shows how a simple random sample can generate useful population-level estimates, while recognizing sampling variability.
The Observed Sample vs the True Sample; Reporting and Transparency
In published work, researchers may report results derived from the observed sample, with less emphasis on the true random sample size or the response rate.
This distinction matters because a low response rate can substantially bias results even if the observed sample looks representative.
Examples show how response rates and frame selection influence interpretation of results in real studies.
Systematic Sampling Methods: Systematic, Stratified, Multistage, and PPS
Systematic sampling: select every k-th element from a list, starting at a random point; often used in exit polling and field surveys.
Stratified sampling: divide population into strata (e.g., regions), sample within each stratum; improves precision when strata are homogeneous and distinct.
Disproportionate (oversampling) stratification: oversample smaller subgroups to ensure precise estimates for those groups; requires weighting to adjust analysis for population representation.
Poststratification weighting: adjust final weights after data collection to align sample distributions with known population distributions (e.g., W = P/p for female vs male samples).
Sampling with probabilities proportional to size (PPS): select units with probability proportional to their size; common in surveys where units vary in size (e.g., firms in national economic surveys).
Multistage and cluster sampling: sample geographic areas, then clusters within areas, then households, then individuals; common in large national surveys (CPS, GSS, ANES, Eurobarometer).
Design effects and intraclass correlation (rho): clustering increases SE; design effect (def) quantifies the loss of precision due to complex sampling; effective sample size n_eff = n / def.
Complex survey corrections: use design effects or specialized software commands (Stata, SAS) to adjust SEs and CIs; essential for large public-use surveys.
Sampling with the Internet and Open Web Polls
Open web polls: strong self-selection bias; results may reflect passion rather than representative opinions.
Internet access panels: active recruitment and weighting can offer more representative results, but biases persist depending on the panel’s composition and weighting strategy.
Conflicting evidence on bias: some internet panels perform comparably to random samples; others show substantial bias.
Real-world electoral forecasts show that some online methods performed well in certain elections; the field remains debated.
Exit Polls, Coverage, and Framing: Case Studies
Exit polls illustrate multi-stage sampling: sampling stations, then selecting individuals (e.g., every 20th voter) to interview.
Coverage bias concerns arise when some voters (absentee, early voters) are not included in the sampling frame.
The choice of frame can influence measured outcomes (e.g., voting shares) due to who is included or excluded.
Box 5.1: Steps in Assessing Coverage and Nonresponse Bias
Define carefully the target population.
For coverage bias:
Identify who is in the frame but not in the target population, and who is in the target population but not in the frame.
Assess systematic differences between those in the frame and those in the target population.
Determine whether the propensity to be covered relates to the study variable.
Use this relationship to predict the direction of bias.
For nonresponse bias:
Identify who was not contacted or refused.
Assess systematic differences between responders and nonresponders.
Determine whether propensity to respond relates to the study variable.
Use this to predict the direction of bias.
Box 5.2: Sampling Bias Terms
Sampling bias: systematic difference between an estimate from a sample and the population truth due to the sampling design.
Subtypes include:
Sample selection bias
Coverage bias
Nonresponse bias
Voluntary response bias (volunteer bias)
Box 5.3: Steps in Assessing Volunteer Bias
Define the target population.
Identify who volunteers and how they volunteer.
Assess systematic differences between volunteers and nonvolunteers.
Determine whether the propensity to volunteer relates to the study outcome and predict the direction of bias.
Box 5.4: Confidence Intervals and Precision Relationships
Relationship summary: Smaller SE → Smaller MOE → Narrower CI → More precision; larger SE → Larger MOE → Wider CI → Less precision.
Box 5.5–5.6: Practical Questions and Tips
Critical questions to ask about sampling when reading studies (generalizability, population definition, frame, response rates, weighting, design effects).
Tips for doing your own sampling: define population, select frame, consider census vs sample, probability vs nonprobability, proper random sampling methods, weighting, anticipate biases, and use design effects when necessary.
Box 5.6: Tips on Doing Your Own Research: Sampling
Define population precisely; know population size.
Identify an appropriate sampling frame.
Decide between census vs sample; consider practicality.
Choose probability sampling when possible; otherwise justify nonprobability choices.
If using disproportionate or cluster sampling, plan for weights and design effects.
Estimate needed sample size and account for subgroup analyses.
Anticipate bias reduction strategies (coverage/nonresponse).
Chapter Resources and Exercises (Overview)
Key terms: Census, Cluster sampling, Confidence interval, External validity, Generalizability, Meta-analysis, Multistage sampling, Nonresponse bias, Poststratification weighting, Probability sampling, Random digit dialing, Sampling frame, Sampling bias, Simple random sampling, Stratified sampling, Systematic sampling, Weighting, etc.
Exercises (selected): evaluating bias in alumni salary survey; nonresponse and coverage in school communications; patient satisfaction sampling; evaluating parenting style studies; Internet sampling biases; credibility of online polls; writing path diagrams for nonresponse bias.
Advanced discussions include Iraqi mortality estimates, cluster sampling challenges, ethical considerations, and politics of research in controversial topics.
Katrina and Public Health Sampling in Context (Illustrative Examples)
Katrina poll (CBS, 2005): 725 adults; 77% felt federal response inadequate; 80% believed response was not as fast as it could have been.
Shelter study (Mills, Edmondson, & Park, 2007): 132 shelter residents; evacuation took 4 days on average; 63% injured; 81% separated from family; 63% directly exposed to corpses.
Takeaway: sampling can yield important nationwide insights even from small, nonrepresentative samples, but bias and limitations must be acknowledged; generalizability is a central concern.
Practical Takeaways for Exam Preparation
Always define population, sampling frame, and sample: know what the study can legitimately say about.
Distinguish between sampling bias (coverage, nonresponse, voluntary response) and precision (sampling error).
Understand the difference between descriptive findings (percentages, means) and relationships (correlations, regressions) in terms of generalizability.
Be comfortable with core formulas:
Response rate:
Standard error (proportion):
Confidence interval:
Simple sample size for a proportion (conservative): with p = 0.5 for maximum variability and Z ≈ 1.96 for 95% CI.
Recognize when weighting, stratification, or multi-stage designs are essential and how they affect standard errors (design effect, effective sample size).
Be able to critique a study’s generalizability by examining the sampling frame, response rate, and potential biases, and to discuss ethical considerations around recruitment and consent.
Short Answer
The chapter emphasizes that generalizability is a nuanced, context-dependent goal. Random sampling generally supports external validity, but replication and meta-analysis are often necessary to establish robust, broad-based conclusions. Nonprobability samples can still contribute valuable causal or theoretical insights, especially in qualitative and exploratory research, provided researchers are transparent about limitations and biases and use appropriate methods (e.g., weighting, design corrections) where feasible.