Analysis of Political Data

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/58

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 3:38 PM on 9/11/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

59 Terms

1
New cards

Quantitative methods (three uses)

Causal inference, prediction, and measurement — the three core purposes political scientists use quantitative methods for.

2
New cards

Research question

The specific question a study is trying to answer; always identify this first when reading a paper (e.g., "Does social pressure affect turnout?").

3
New cards

Causal relationship (X → Y)

A cause-and-effect connection where X (the treatment) affects Y (the outcome). The arrow's direction shows X influences Y, not the reverse.

4
New cards

Treatment variable (X)

The variable representing the "cause" being studied — e.g., whether an individual receives a social pressure message. Often binary: Xi = 1 if treated, 0 if not.

5
New cards

Outcome variable (Y)

The variable representing the effect being measured — e.g., Votedi = 1 if voter i voted, 0 if not. Can be binary or non-binary.

6
New cards

Treatment condition

The condition where the treatment is present for an individual (Xi = 1).

7
New cards

Control condition

The condition where the treatment is absent for an individual (Xi = 0).

8
New cards

Individual causal effect

The change in an individual's outcome (Y) caused by receiving the treatment: ΔYi = Yi(Xi=1) − Yi(Xi=0). Cannot be directly observed for any one person.

9
New cards

Potential outcome

The outcome an individual WOULD have under a given treatment condition — e.g., Yi(Xi=1) is what would happen to person i if treated, Yi(Xi=0) if not.

10
New cards

Factual outcome

The potential outcome actually observed — the outcome under the condition the individual actually received in reality.

11
New cards

Counterfactual outcome

The potential outcome that was NOT observed — what would have happened under the condition the individual did not receive. Can never be directly observed.

12
New cards

Fundamental problem of causal inference

The core problem that we can never observe both the factual and counterfactual outcome for the same individual at the same time — so individual causal effects can never be directly measured.

13
New cards

Average causal effect (average treatment effect)

The average of individual causal effects across a group — the average change in Y caused by moving from control to treatment across all individuals in that group.

14
New cards

Confounding variable

A pre-treatment characteristic (e.g., age, income, education) that differs systematically between treatment and control groups and could explain differences in outcomes instead of the treatment itself.

15
New cards

Comparable groups

Treatment and control groups that have similar pre-treatment characteristics (observed and/or unobserved) that might affect the outcome. When groups are comparable, one group's factual outcome approximates the other's counterfactual.

16
New cards

Randomized experiment

A study design in which treatment assignment is determined through a random process (e.g., a coin flip), rather than by individual choice — considered the "gold standard" for causal inference.

17
New cards

Random assignment

Randomly determining WHO WITHIN THE STUDY receives treatment vs. control. Different from random sampling — this is about assignment to conditions, not selection into the study.

18
New cards

Random sampling

Randomly choosing a SUBSET OF THE POPULATION to include in a study. Different from random assignment — this is about who's in the study, not who gets the treatment.

19
New cards

Observational study

A study design where treatment assignment is NOT controlled by the researcher — it results from individual choices or naturally occurring events, making confounding a bigger risk.

20
New cards

Why random assignment works

With a large enough sample, chance is the only systematic difference between treatment and control groups, so they end up similar on average across ALL pre-treatment characteristics — including unobserved ones like political interest.

21
New cards

Sample size caveat for randomization

Random assignment only produces comparable groups "on average" if n is large enough — it's difficult or impossible to create truly comparable groups from a very small number of observations (e.g., n=10).

22
New cards

Difference-in-means estimator

A formula for estimating the average causal effect: the average outcome in the treatment group minus the average outcome in the control group. Only valid when treatment and control groups are comparable.

23
New cards

When difference-in-means is NOT valid

When treatment and control groups are NOT comparable — e.g., in most observational studies, where confounding variables may differ systematically between groups.

24
New cards

Internal validity

The degree of confidence that a study's results reflect a TRUE cause-and-effect relationship, rather than being driven by outside/confounding factors. Asks: "Do I trust this causal claim within this study?"

25
New cards

External validity

The extent to which a study's results can be GENERALIZED to other people, settings, times, and real-world situations. Asks: "Does this finding hold up elsewhere?"

26
New cards

Pre-treatment characteristics

Traits that exist BEFORE treatment is assigned and might independently affect the outcome (e.g., age, income, education level). Must be similar across treatment/control groups for valid causal comparison.

27
New cards

Treatment variable notation

Xi = 1 if individual i receives the treatment; Xi = 0 if individual i does not receive the treatment (control).

28
New cards

Outcome variable notation (voting example)

Votedi = 1 if registered voter i voted; Votedi = 0 if registered voter i did not vote.

29
New cards

Potential outcome notation

Yi(Xi=1) = the potential outcome for individual i under the treatment condition. Yi(Xi=0) = the potential outcome for individual i under the control condition.

30
New cards

Individual causal effect formula

ΔYi = Yi(Xi=1) − Yi(Xi=0). The change in individual i's outcome caused by moving from control to treatment. Cannot be computed in practice since only one of the two terms is ever observed.

31
New cards

Individual causal effect, voting example

ΔVotedi = Votedi(pressurei=1) − Votedi(pressurei=0). The effect of the social pressure message on individual i's probability of voting.

32
New cards

Average causal effect formula (conceptual)

Average causal effect = average of all the individual causal effects (ΔYi) across a group. Since individual ΔYi can't be observed, this is instead approximated using group-level comparisons.

33
New cards

Difference-in-means estimator formula

Estimated average causal effect = (average Y in treatment group) − (average Y in control group). Only a valid estimate of the average causal effect when treatment and control groups are comparable.

34
New cards

What makes the difference-in-means formula "valid"

It equals the true average causal effect only when treatment/control groups have similar pre-treatment characteristics (i.e., under random assignment with adequate sample size) — otherwise confounding can bias the estimate.

35
New cards

General causal notation X → Y

X is always the treatment (cause) and Y is always the outcome (effect) — which real-world variables map onto X and Y depends on the research question and the nature of the relationship being studied.

36
New cards

Comparable groups (comparability)

Treatment and control groups are "comparable" when they have similar pre-treatment characteristics — both observed (age, income, education) and unobserved (e.g., political interest) — that might otherwise affect the outcome. When groups are comparable, the only systematic difference between them is the treatment itself, so one group's factual outcome is a good approximation of the other group's counterfactual outcome, making the difference-in-means estimator a valid estimate of the average causal effect. Random assignment (with a large enough n) is what produces comparability, since chance becomes the only source of difference between groups.

37
New cards

Direction (of a relationship)

Whether two variables move together or oppositely, shown by a + or − sign. Positive direction: both variables increase together (e.g., education ↑, income ↑). Negative/inverse direction: one increases as the other decreases (e.g., poverty rate ↑, graduation rate ↓).

38
New cards

Magnitude (of an effect)

The strength or size of an effect/difference/relationship — "how big is the impact?" Whether a given magnitude is "large" or "small" depends on subject-matter context (e.g., a 1.8 percentage point turnout increase could be huge or trivial depending on the race).

39
New cards

Unit (of measurement)

The concrete real-world scale a statistic is measured in (dollars, years, percentage points, score points). Essential for interpreting a coefficient correctly — "income increases by 5 per year of education" is meaningless without knowing if that's $5, $5,000, or 5%.

40
New cards

Two greatest challenges in political data analysis

(1) The fundamental problem of causal inference (never observing the counterfactual), and (2) confounding variables.

41
New cards

Observational data

Data collected about naturally occurring events, where treatment assignment is outside the researcher's control — unlike a randomized experiment, we cannot assume treatment and control groups are comparable.

42
New cards

Confounding variable (Z)

A variable that affects BOTH the probability of receiving treatment X AND the outcome Y. Confounders make treatment and control groups non-comparable, obscuring the true causal relationship between X and Y. Denoted Z.

43
New cards

Why confounders are a problem

They make treatment/control groups non-comparable, so the difference-in-means estimator can no longer be trusted. In the presence of confounders, correlation does NOT necessarily imply causation — a third variable (the confounder) may be driving the apparent relationship between X and Y.

44
New cards

Correlation vs. causation (with confounders)

High correlation between two variables doesn't mean one causes the other — a confounder (third variable) could be independently affecting both, making them move together or oppositely even with no direct causal link.

45
New cards

Why random assignment eliminates confounders

Random treatment assignment breaks the link between any potential confounder (Z) and the treatment (X) — nothing related to the outcome is also related to the likelihood of receiving treatment, so groups stay comparable on average (both observed and unobserved characteristics).

46
New cards

Statistically controlling for confounders

When using observational data (no random assignment), researchers must identify and measure relevant confounding variables and statistically control for them, to make treatment/control groups more comparable within the analysis.

47
New cards

Population vs. sample

The population is the entire group we ultimately want to understand; the sample is the specific subset of data we actually have access to and analyze.

48
New cards

Broad class of treatments vs. specific version

We're usually interested in the effect of a general treatment category (e.g., "aspirin"), but a study typically only tests one specific version of it (e.g., one brand of aspirin) — raising questions about how well that version represents the broader class.

49
New cards

Internal validity (expanded)

The extent to which a study's causal findings are valid WITHIN its specific context — i.e., is the estimated average causal effect valid for the sample/treatment analyzed, and can it be interpreted causally? Depends on (1) a research design that can credibly estimate causal effects, and (2) justifiable identifying assumptions. Core issue: avoiding/ruling out confounding variables.

50
New cards

Internal validity: strong vs. weak

Strong when treatment and control groups can be considered comparable within the analysis (after research design and controls are taken into account). Weak otherwise.

51
New cards

External validity (expanded)

The extent to which a study's causal findings can be generalized BEYOND its specific context — i.e., is the estimated average causal effect valid for a broader population and treatment class? Depends on BOTH (1) whether the sample is representative of the target population, and (2) whether the specific treatment studied represents the broader class of treatments.

52
New cards

External validity: strong vs. weak

Strong when the sample is representative of the population you want to generalize to AND the specific treatment studied is similar in nature to the broader treatment class you want to generalize to.

53
New cards

Why sample representativeness matters for external validity

An average causal effect is the mean effect for a SPECIFIC group. If the population you're generalizing to differs meaningfully from your sample (e.g., single-party sample vs. mixed-party population), the true effect in that new population may differ from your estimate.

54
New cards

Field experiments

A type of randomized experiment where treatment is randomly assigned in real-world (rather than lab) settings — often used to improve external validity while keeping the internal validity benefits of randomization.

55
New cards

Randomized experiments: validity tradeoff

Tend to have STRONG internal validity (random assignment eliminates confounders, in expectation) but relatively WEAK external validity (participants often rely on volunteers and may be unrepresentative; treatment may be unrealistic).

56
New cards

Observational studies: validity tradeoff

Tend to have STRONG external validity (samples and treatments are often realistic/representative since you observe the real population and real-world treatment) but relatively WEAK internal validity (unobserved confounders may remain uncontrolled).

57
New cards

Ideal research design

Combines random sampling (→ strong external validity, assuming a realistic treatment) with random treatment assignment (→ strong internal validity, by eliminating confounders). Rarely fully achievable in practice, but is the design to aim for.

58
New cards

Does studying the whole population fix external validity?

Not entirely — it eliminates sampling bias for that group, but you can still face external validity problems generalizing findings to other times, places, or theoretical populations.

59
New cards
Percentage vs. percentage point (magnitude framing)
A percentage is a level (turnout = 42%); a percentage point (pp) is the absolute difference between two percentages — what difference-in-means produces (42% − 40% = 2 pp). pp and % diverge with small baselines: turnout 40%→45% is 5 pp but a 12.5% relative rise; a minor candidate 1%→3% is only 2 pp but a 100% relative rise (tripling). Always name the unit before judging an effect's size.