Notes on Geddes (1990): Selection Bias in Comparative Politics

Abstract

  • Demonstrates how selecting cases for study based on outcomes on the dependent variable (DV) biases conclusions in comparative politics.
  • Outlines the logic of explanation and shows violation when only cases that achieved the outcome are studied.
  • Examines three influential studies in comparative politics, comparing original conclusions (case- DV–driven samples) with tests using samples not correlated with the DV.
  • In each instance, conclusions based on the uncorrelated samples differ from the original conclusions.
  • Argues that comparative politics has norms about appropriate research strategy and evidence; a durable convention is selecting cases by DV, which risks bias.

The Nature of the Problem

  • The problem arises from the logic of explanation: to explain why some countries achieve a certain outcome (e.g., rapid growth), researchers search for antecedent factors X through Z that A and B share but C through G do not.
  • If one studies only cases A and B, one learns only what they have in common; without also studying C–G, one cannot know whether X–Z are truly causal antecedents.
  • A sample consisting only of high-outcome cases may fail to reveal that other cases also possess X–Z, making the hypothesis dubious.
  • Graphical intuition (Fig. 1 and Fig. 2 in the article):
    • Fig. 1 shows an assumed positive relationship where high-DV cases share a high level of X.
    • Fig. 2 shows that even with a positive relationship in the full population, a sample restricted to high-DV cases could display no apparent relationship.
  • Key inference issues:
    • In the statistical literature, the second kind of faulty inference (selection on the DV leading to an apparent lack of relationship when one exists) is central (Achen 1986; King 1989).
    • In nonquantitative work, the first kind (inferring a simple causal relationship from a restricted set) is common.
  • Practical illustration: selecting NICs (newly industrializing countries) with rapid growth and repression of labor leads to the claim that repression contributes to growth, but such a sample distorts the true relationship.

A Straightforward Case of Selection on the Dependent Variable

  • Researchers often select a few successful NICs (e.g., Taiwan, South Korea, Singapore, Brazil, Mexico) to argue that repression or cooptation of labor contributed to growth.
  • Different scholars (e.g., Johnson 1987; O'Donnell 1973; Cardoso 1973; Deyo 1984, 1987) have advanced various mechanisms linking labor repression to growth.
  • Geddes cautions that this selection on DV does not by itself prove causation; other countries with comparable repression have not prospered, so the claim requires testing on an uncorrelated sample.
  • Key steps to properly test the hypothesis:
    • Identify the universe of cases where the hypothesis should apply (all developing countries in this instance).
    • Develop or obtain measures for the variables (growth rate and labor repression).
    • Select a sample from the universe with criteria uncorrelated with the DV.
    • If the universe is too large, use a random sample; but randomization does not guarantee absence of correlation if the universe itself is conditioned by prior outcomes.
  • Sample and measurement specifics used in the test:
    • Universe: all developing countries with World Bank data, excluding high-income oil exporters, Communist governments, civil-war-dominated periods (> one-third of the study period), and extremely small countries (<1,000,000 inhabitants).
    • DV: growth rate (1960–1982) using World Bank (1984) data for GNP per capita; this window focuses on pre-debt-crisis dynamics.
    • Labor repression measure: a scale from 1 to 4, based on Country Reports on Human Rights Practices (U.S. State Department, 1981 and 1989). Countries scored on:
    • S = 1: unions are free to organize and leaders are chosen by unions; strikes are legal and frequent; unions participate in politics.
    • S = 2: unions free to organize with some government controls; strikes legal but regulated or infrequent; violence may limit rights but not abolish organizing.
    • S = 3: unions constrained by government/dominant party links; strikes legal in some cases but heavily regulated; violence against workers may be moderate.
    • S = 4: unions illegal or fully controlled; strikes severely constrained or absent; violence against workers severe.
    • Countries assigned scores; opinions on borderline placements acknowledged but argued to be as precise as possible given data limitations.
  • The test against the labor repression hypothesis yields three key illustrations:
    • Fig. 3 (NICs sample): A scatter of five NICs shows repression is moderately high across these cases; no obvious simple relationship between repression and growth emerges in this tiny, purposively chosen subset.
    • Fig. 4 (East Asia focus): When focusing on East Asian NICs (Singapore, South Korea, Taiwan, Indonesia, etc.), a seemingly positive relationship appears with a regression R' = 0.56; a difference-of-means test between high vs. low repression categories is significant at p = 0.02, suggesting a tancy that repression might be associated with growth in this subset.
    • Fig. 5 (Third World broader sample): A broader Third World sample shows little to no simple relationship; the regression slope is slightly negative and R^2 = 0.07, implying virtually no explanatory power of labor repression for growth in this larger, more representative sample.
  • Interpretation and caveats:
    • The simple bivariate test cannot disconfirm the hypothesis; the apparent relation in Fig. 4 could be an artifact of sample selection (East Asia bias) rather than a generalizable causal link.
    • The apparent relationship in Fig. 3 and Fig. 4 arises from selecting cases that are already above the dotted line in the hypothetical figures, which biases inference.
    • If the sample is broadened or additional controls are added, the apparent link weakens or disappears (Fig. 5).
    • The sample of East Asian NICs is correlated with geographic region, which is itself correlated with growth; this is another form of selection on the DV (geography-as-proxy problem).
    • The conclusion is not that labor repression never contributes to growth, but that the simple, direct, bivariate relation inferred from highly selected samples does not generalize to a representative universe of cases.
  • Takeaway: If analysts had tested a more representative sample and included appropriate controls, they would likely reach different conclusions about the relationship between labor repression and growth.

The Two Tasks Crucial to Testing Any Hypothesis

  • Task 1: Identify the universe of cases to which the hypothesis should apply.
  • Task 2: Develop valid measures for the variables.
  • If the universe is too large, random sampling is recommended to avoid DV-correlated selection, but randomization does not guarantee an absence of correlation.
  • If the universe is such that extremely successful cases have weeded out failures, even random samples may effectively be DV-selected; this is a caveat about the limits of any sampling approach.
  • The author’s operationalization for the labor repression variable is argued to be as precise as the available descriptions allow, and the testing approach is designed to illustrate methodological points rather than to definitively prove or disprove the labor-repression hypothesis.

Selection Bias in Path-Dependent Arguments (Skocpol’s Case)

  • Theda Skocpol’s States and Social Revolutions (1979) blends selection on the DV with a path-dependent argument: she selects France, Russia, and China (three revolutions) and contrasts them with cases where revolutions did not occur at strategic points.
  • Central chain (Fig. 7): External military threats cause state officials to initiate reforms; if the dominant class has an independent economic base and political power, it can oppose reforms; peasants’ autonomy or solidarity can translate into revolution if the elite splits.
  • Skocpol’s comparative look includes contrasting cases (Prussia, Japan, Britain, Germany) to test whether similar structural conditions lead to revolutions; in some cases, elites do split but revolutions do not occur due to weak peasant mobilization or other factors.
  • Main takeaway: The cross-case comparison strengthens an argument but is weaker as a test of the mechanism if cases are selected by DV. A broader, uncorrelated sample would provide a stronger test, though perfect testing is often impractical.
  • Geddes’ assessment: While Skocpol’s cases enhance persuasiveness, a rigorous test would require examining all nations with the structural features Skocpol identifies (village autonomy, dominant-class independence). Spainish-American countries are offered as a nonrandom set that can be used for testing; they share the necessary structural traits but are not randomly selected.
  • Findings (Fig. 8–Fig. 10): With selected Spanish-American cases that fit Skocpol’s structure, several revolutions occurred and several cases with high external threat did not lead to revolution; one Bolivia case fits the theory. The presence of multiple counterexamples suggests that Skocpol’s linkage from threat to revolution is not universal.
  • Caveats: Measurement of “threat,” “village autonomy,” and “dominant-class independence” is difficult; the test is not definitive, but it challenges the universality of Skocpol’s claim. Operationalizations may place Nicaragua and Mexico into different cells than Skocpol would, and broader testing might alter conclusions.
  • Overall implication: Path-dependent arguments based on DV-selected cases require broader testing with cases not chosen for their DV position to avoid biased conclusions.

Selection of the Endpoints of a Time-Series Analysis

  • Time-series endpoint selection can bias conclusions when testing hypotheses using time-bound data.
  • Prebisch’s secular decline in terms of trade was influential but later studies showed it depended on the period chosen; different endpoints yield different conclusions.
  • The core problem: If a researcher selects an endpoint because the DV attains an extreme value at that end, the analysis is effectively selecting on the DV.
  • An illustration (Hirschman): Inflation and political learning in Chile.
    • Hirschman argues inflation can temporally ease conflicts and be tempered as groups learn; Chile’s inflation data (1930–1961; table 2 in the text) shows that the stabilization in 1960–1961 was followed by later inflationary episodes, suggesting that the early data points do not necessarily demonstrate lasting political learning or stabilization.
    • Figure 11 (Chile inflation 1930–72) shows that even though stabilization occurred in the late 1950s/early 1960s, two data points (around 1960–1961) were below the trend and might be misinterpreted as evidence of long-term learning; with more data, the inference might weaken.
  • Lesson: Conclusions that hinge on a few time points or endpoints can be fragile; more data over time may weaken or reverse apparent relationships.
  • Numerical example from Hirschman’s Chilean inflation data (1930–1961 in Table 2): yearly inflation rates range from negative rates to double-digit positives (e.g., -5% in 1930, 84% in 1955, 71% in 1954), illustrating volatility and the danger of drawing causal conclusions from a few endpoints.

Conclusion and Implications for Research Practice

  • Choosing cases for study on the basis of their DV scores biases conclusions; apparent causal variables may simply be common to the DV-selected cases and not causally generalizable.
  • Relationships observed in small, DV-selected samples may disappear or reverse when the sample includes a broader, more representative set of cases.
  • Time-series endpoint choices can heavily influence conclusions; results may be contingent on the time window and endpoint selection.
  • Geddes emphasizes that case studies selected on the DV are valuable for developing hypotheses and theories about mechanisms, but they are not adequate for rigorous theory testing or for accumulating generalizable knowledge.
  • To build robust theoretical knowledge in comparative politics, researchers should:
    • Use samples not determined by the DV when testing causal claims.
    • Employ representative or random sampling where possible, or justify the sampling frame and selection criteria transparently.
    • Incorporate multiple tests and controls (multivariate analyses, geographic controls, time-series robustness checks).
    • Combine qualitative case studies (for mechanism-building) with broader, uncorrelated tests (for theory testing).
  • The article acknowledges the continued value of DV-selected case studies for exploring mechanisms, anomalies, and plausible causal variables, but calls for more rigorous standards of evidence when testing theories.

Key Concepts and Definitions

  • Selection bias: Bias arising when case selection is correlated with the DV or with factors that affect the DV, leading to spurious or misleading conclusions.
  • Dependent variable (DV): The outcome variable the researcher aims to explain (e.g., growth rate, revolution occurrence).
  • Independent variable (X): A potential antecedent factor thought to influence the DV (e.g., labor repression).
  • Universe of cases: The complete set of cases to which a hypothesis should apply (e.g., all developing countries).
  • Uncorrelated sample: A sample whose case selection is not correlated with the DV (or the key explanatory variable), enabling valid inferences about the population.
  • Endpoints in time-series: The start and end points chosen to analyze a time series; selecting endpoints that correspond to extreme values can bias conclusions.
  • Path-dependent argument: An argument in which outcomes depend on the sequence of prior events; testing requires careful case selection and consideration of multiple pathways.

Notable Numerical References and Formulas

  • Labor repression scoring: S ∈ {1, 2, 3, 4} with specific criteria for unions and strikes; higher values indicate tighter repression.
  • Growth data window: Growth rate measured using GNP per capita for 1960–1982 (World Bank 1984 data).
  • Regression and association metrics:
    • Fig. 4: R' = 0.56 for the regression relating growth to labor repression in the East Asia NICs sample.
    • Fig. 5: Overall regression in a larger Third World sample with slope ≈ negative and R2=0.07R^2 = 0.07, indicating little explanatory power from repression alone.
  • Regional growth comparisons (Table 1): Average growth rates 1960–82 and 1965–86 by region:
    • East Asia: 1960–82 = 5.2; 1965–86 = 5.1
    • South Asia: 1960–82 = 1.4; 1965–86 = 1.5
    • Africa: 1960–82 = 1.0; 1965–86 = 0.5
    • Latin America: 1960–82 = 2.2; 1965–86 = 1.2
    • Middle East & North Africa: 1960–82 = 4.7; 1965–86 = 3.6
  • Endpoints/time-series example (PréBisch/Hirschman) discussed via Chile inflation data (1930–1961 table of rates includes: 1930 = -5%, 1955 = 84%, 1954 = 71%, etc.), illustrating sensitivity to endpoint choice.

References (selected items cited in the discussion)

  • Achen, Christopher. 1986. The Statistical Analysis of Quasi-Experimental Data. University of California Press.
  • Achen, Christopher, and Duncan Snidal. 1989. Rational Choice and Comparative Case Studies. World Politics 41:143–169.
  • Skocpol, Theda. 1979. States and Social Revolutions: A Comparative Analysis of France, Russia, and China. Cambridge University Press.
  • Hirschman, Albert O. 1973. Journeys Toward Progress: Studies of Economic Policy Making in Latin America. Norton.
  • Prébisch, Raúl. 1950. The Economic Development of Latin America and Its Principal Problems. United Nations.
  • Deyo, Frederic. 1984, 1987; Johnson, Chalmers. 1987; O'Donnell, Guillermo. 1973; Cardoso, Fernando Henrique. 1973; Koo, Hägen. 1987.
  • World Bank. 1984. World Development Report 1984; World Bank. 1988. World Development Report 1988.
  • U.S. Department of State. 1981, 1989. Country Reports on Human Rights Practices.
  • Ramos, Joseph. 1986. Neo-Conservative Economics in the Southern Cone of Latin America, 1973–1983.
  • Valenzuela, Arturo. 1978. The Breakdown of Democratic Regimes: Chile.

Connections to Practical Research Practice

  • Use of uncorrelated samples is essential for testing causal claims; avoid relying solely on DV-selected case studies when building general theories.
  • Case studies remain valuable for hypothesis generation and mechanism exploration; they should be complemented by broader tests and robust sampling methods.
  • Researchers should be explicit about sampling frames, measurement validity, and potential sources of bias, including geography, time period, and regime type.
  • When time-series endpoints are involved, researchers should test robustness across multiple end points and consider structural breaks, regime shifts, and period-specific effects.

Summary Takeaway

  • Selection on the dependent variable can bias both qualitative and quantitative analyses, leading to misleading conclusions about causal relationships.
  • A combination of mechanism-focused DV-selected case studies and broader, uncorrelated-sample tests provides a more reliable path to accumulating theoretical knowledge in comparative politics.
  • Rigorous sampling, clear variable definitions, and sensitivity analyses across samples and time periods are essential for credible theory testing.