1/58
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Quantitative methods (three uses)
Causal inference, prediction, and measurement — the three core purposes political scientists use quantitative methods for.
Research question
The specific question a study is trying to answer; always identify this first when reading a paper (e.g., "Does social pressure affect turnout?").
Causal relationship (X → Y)
A cause-and-effect connection where X (the treatment) affects Y (the outcome). The arrow's direction shows X influences Y, not the reverse.
Treatment variable (X)
The variable representing the "cause" being studied — e.g., whether an individual receives a social pressure message. Often binary: Xi = 1 if treated, 0 if not.
Outcome variable (Y)
The variable representing the effect being measured — e.g., Votedi = 1 if voter i voted, 0 if not. Can be binary or non-binary.
Treatment condition
The condition where the treatment is present for an individual (Xi = 1).
Control condition
The condition where the treatment is absent for an individual (Xi = 0).
Individual causal effect
The change in an individual's outcome (Y) caused by receiving the treatment: ΔYi = Yi(Xi=1) − Yi(Xi=0). Cannot be directly observed for any one person.
Potential outcome
The outcome an individual WOULD have under a given treatment condition — e.g., Yi(Xi=1) is what would happen to person i if treated, Yi(Xi=0) if not.
Factual outcome
The potential outcome actually observed — the outcome under the condition the individual actually received in reality.
Counterfactual outcome
The potential outcome that was NOT observed — what would have happened under the condition the individual did not receive. Can never be directly observed.
Fundamental problem of causal inference
The core problem that we can never observe both the factual and counterfactual outcome for the same individual at the same time — so individual causal effects can never be directly measured.
Average causal effect (average treatment effect)
The average of individual causal effects across a group — the average change in Y caused by moving from control to treatment across all individuals in that group.
Confounding variable
A pre-treatment characteristic (e.g., age, income, education) that differs systematically between treatment and control groups and could explain differences in outcomes instead of the treatment itself.
Comparable groups
Treatment and control groups that have similar pre-treatment characteristics (observed and/or unobserved) that might affect the outcome. When groups are comparable, one group's factual outcome approximates the other's counterfactual.
Randomized experiment
A study design in which treatment assignment is determined through a random process (e.g., a coin flip), rather than by individual choice — considered the "gold standard" for causal inference.
Random assignment
Randomly determining WHO WITHIN THE STUDY receives treatment vs. control. Different from random sampling — this is about assignment to conditions, not selection into the study.
Random sampling
Randomly choosing a SUBSET OF THE POPULATION to include in a study. Different from random assignment — this is about who's in the study, not who gets the treatment.
Observational study
A study design where treatment assignment is NOT controlled by the researcher — it results from individual choices or naturally occurring events, making confounding a bigger risk.
Why random assignment works
With a large enough sample, chance is the only systematic difference between treatment and control groups, so they end up similar on average across ALL pre-treatment characteristics — including unobserved ones like political interest.
Sample size caveat for randomization
Random assignment only produces comparable groups "on average" if n is large enough — it's difficult or impossible to create truly comparable groups from a very small number of observations (e.g., n=10).
Difference-in-means estimator
A formula for estimating the average causal effect: the average outcome in the treatment group minus the average outcome in the control group. Only valid when treatment and control groups are comparable.
When difference-in-means is NOT valid
When treatment and control groups are NOT comparable — e.g., in most observational studies, where confounding variables may differ systematically between groups.
Internal validity
The degree of confidence that a study's results reflect a TRUE cause-and-effect relationship, rather than being driven by outside/confounding factors. Asks: "Do I trust this causal claim within this study?"
External validity
The extent to which a study's results can be GENERALIZED to other people, settings, times, and real-world situations. Asks: "Does this finding hold up elsewhere?"
Pre-treatment characteristics
Traits that exist BEFORE treatment is assigned and might independently affect the outcome (e.g., age, income, education level). Must be similar across treatment/control groups for valid causal comparison.
Treatment variable notation
Xi = 1 if individual i receives the treatment; Xi = 0 if individual i does not receive the treatment (control).
Outcome variable notation (voting example)
Votedi = 1 if registered voter i voted; Votedi = 0 if registered voter i did not vote.
Potential outcome notation
Yi(Xi=1) = the potential outcome for individual i under the treatment condition. Yi(Xi=0) = the potential outcome for individual i under the control condition.
Individual causal effect formula
ΔYi = Yi(Xi=1) − Yi(Xi=0). The change in individual i's outcome caused by moving from control to treatment. Cannot be computed in practice since only one of the two terms is ever observed.
Individual causal effect, voting example
ΔVotedi = Votedi(pressurei=1) − Votedi(pressurei=0). The effect of the social pressure message on individual i's probability of voting.
Average causal effect formula (conceptual)
Average causal effect = average of all the individual causal effects (ΔYi) across a group. Since individual ΔYi can't be observed, this is instead approximated using group-level comparisons.
Difference-in-means estimator formula
Estimated average causal effect = (average Y in treatment group) − (average Y in control group). Only a valid estimate of the average causal effect when treatment and control groups are comparable.
What makes the difference-in-means formula "valid"
It equals the true average causal effect only when treatment/control groups have similar pre-treatment characteristics (i.e., under random assignment with adequate sample size) — otherwise confounding can bias the estimate.
General causal notation X → Y
X is always the treatment (cause) and Y is always the outcome (effect) — which real-world variables map onto X and Y depends on the research question and the nature of the relationship being studied.
Comparable groups (comparability)
Treatment and control groups are "comparable" when they have similar pre-treatment characteristics — both observed (age, income, education) and unobserved (e.g., political interest) — that might otherwise affect the outcome. When groups are comparable, the only systematic difference between them is the treatment itself, so one group's factual outcome is a good approximation of the other group's counterfactual outcome, making the difference-in-means estimator a valid estimate of the average causal effect. Random assignment (with a large enough n) is what produces comparability, since chance becomes the only source of difference between groups.
Direction (of a relationship)
Whether two variables move together or oppositely, shown by a + or − sign. Positive direction: both variables increase together (e.g., education ↑, income ↑). Negative/inverse direction: one increases as the other decreases (e.g., poverty rate ↑, graduation rate ↓).
Magnitude (of an effect)
The strength or size of an effect/difference/relationship — "how big is the impact?" Whether a given magnitude is "large" or "small" depends on subject-matter context (e.g., a 1.8 percentage point turnout increase could be huge or trivial depending on the race).
Unit (of measurement)
The concrete real-world scale a statistic is measured in (dollars, years, percentage points, score points). Essential for interpreting a coefficient correctly — "income increases by 5 per year of education" is meaningless without knowing if that's $5, $5,000, or 5%.
Two greatest challenges in political data analysis
(1) The fundamental problem of causal inference (never observing the counterfactual), and (2) confounding variables.
Observational data
Data collected about naturally occurring events, where treatment assignment is outside the researcher's control — unlike a randomized experiment, we cannot assume treatment and control groups are comparable.
Confounding variable (Z)
A variable that affects BOTH the probability of receiving treatment X AND the outcome Y. Confounders make treatment and control groups non-comparable, obscuring the true causal relationship between X and Y. Denoted Z.
Why confounders are a problem
They make treatment/control groups non-comparable, so the difference-in-means estimator can no longer be trusted. In the presence of confounders, correlation does NOT necessarily imply causation — a third variable (the confounder) may be driving the apparent relationship between X and Y.
Correlation vs. causation (with confounders)
High correlation between two variables doesn't mean one causes the other — a confounder (third variable) could be independently affecting both, making them move together or oppositely even with no direct causal link.
Why random assignment eliminates confounders
Random treatment assignment breaks the link between any potential confounder (Z) and the treatment (X) — nothing related to the outcome is also related to the likelihood of receiving treatment, so groups stay comparable on average (both observed and unobserved characteristics).
Statistically controlling for confounders
When using observational data (no random assignment), researchers must identify and measure relevant confounding variables and statistically control for them, to make treatment/control groups more comparable within the analysis.
Population vs. sample
The population is the entire group we ultimately want to understand; the sample is the specific subset of data we actually have access to and analyze.
Broad class of treatments vs. specific version
We're usually interested in the effect of a general treatment category (e.g., "aspirin"), but a study typically only tests one specific version of it (e.g., one brand of aspirin) — raising questions about how well that version represents the broader class.
Internal validity (expanded)
The extent to which a study's causal findings are valid WITHIN its specific context — i.e., is the estimated average causal effect valid for the sample/treatment analyzed, and can it be interpreted causally? Depends on (1) a research design that can credibly estimate causal effects, and (2) justifiable identifying assumptions. Core issue: avoiding/ruling out confounding variables.
Internal validity: strong vs. weak
Strong when treatment and control groups can be considered comparable within the analysis (after research design and controls are taken into account). Weak otherwise.
External validity (expanded)
The extent to which a study's causal findings can be generalized BEYOND its specific context — i.e., is the estimated average causal effect valid for a broader population and treatment class? Depends on BOTH (1) whether the sample is representative of the target population, and (2) whether the specific treatment studied represents the broader class of treatments.
External validity: strong vs. weak
Strong when the sample is representative of the population you want to generalize to AND the specific treatment studied is similar in nature to the broader treatment class you want to generalize to.
Why sample representativeness matters for external validity
An average causal effect is the mean effect for a SPECIFIC group. If the population you're generalizing to differs meaningfully from your sample (e.g., single-party sample vs. mixed-party population), the true effect in that new population may differ from your estimate.
Field experiments
A type of randomized experiment where treatment is randomly assigned in real-world (rather than lab) settings — often used to improve external validity while keeping the internal validity benefits of randomization.
Randomized experiments: validity tradeoff
Tend to have STRONG internal validity (random assignment eliminates confounders, in expectation) but relatively WEAK external validity (participants often rely on volunteers and may be unrepresentative; treatment may be unrealistic).
Observational studies: validity tradeoff
Tend to have STRONG external validity (samples and treatments are often realistic/representative since you observe the real population and real-world treatment) but relatively WEAK internal validity (unobserved confounders may remain uncontrolled).
Ideal research design
Combines random sampling (→ strong external validity, assuming a realistic treatment) with random treatment assignment (→ strong internal validity, by eliminating confounders). Rarely fully achievable in practice, but is the design to aim for.
Does studying the whole population fix external validity?
Not entirely — it eliminates sampling bias for that group, but you can still face external validity problems generalizing findings to other times, places, or theoretical populations.