Quantitative Research 101
Qualitative vs Quantitative Approaches
Quantitative Analysis: studying social phenomena using scientific ways
Rationalism (reasoning through logic) and empiricism (using data, observation, which can be a qualitative or quantitative approach to resonate with what’s true). Sometimes the rationalism is false, which makes the conclusion false as well. In empiricism, we can overgeneralize because we rely on a few generalizations, a minimal number of examples compared to the whole community and jump to conclusions.
Social science is a science because we use a scientific approach (data and research), unlike philosophy.
Scientific research: an inductive, more qualitative approach, reasoning (from specific or basic observation to studying a worldwide and bigger general topic to theory), a deductive, more quantitative approach (from general explanation theory, hypothesis, gathering data, then to a particular observation to a conclusion), so deductive, big to small
Theory: general explanation
When gathering evidence, you may use a qualitative or quantitative approach depending on the nature of your study. For quantitative research, we use data, measuring variables to find a pattern and then a conclusion that explains the connection/relationship between the data and the theory. Qualitative research involves examining complexities, meanings, reactions (feelings and perceptions), emotions, and lived experiences, which explains the relationship between these qualitative factors and the theory.
As social scientists, you may decide to take on a quantitative approach or a qualitative approach, depending on what you are studying.
The choice between qualitative and quantitative methods is driven by research questions, what you want to measure, and how you want to interpret findings.
You may gather evidence using either or both approaches, and you may follow up with qualitative methods after an initial quantitative analysis (or vice versa).
Conceptualization and Operationalization
Conceptualization: defining what you want to study at a high level (the concept).
Operationalization: defining how you will measure the concept in real terms (the measurement).
Example: inflation
Conceptualization: what inflation is in theory.
Operationalization: how you measure inflation (e.g., price level indices over periods).
In practice, researchers may study relationships by collecting data and interpreting results, using either qualitative, quantitative, or mixed methods.
Data Types and Data Collection Methods
Qualitative data (nonnumeric): text, audio, images, video, etc.
Quantitative data (numeric): counts, measurements, statistics, survey results, etc.
Qualitative data collection methods mentioned:
In-depth questionnaires
Focus groups
Observations of behavior in settings (e.g., Rastafarian camp, parties at night)
Open-ended interviews
Quantitative data collection methods mentioned:
Surveys with measurable quantities
Testing hypotheses and establishing causal relationships
Studying large populations to enable generalization
Nonnumeric qualitative evidence is interpreted through systematic coding to identify patterns and categories.
Qualitative Data: Methods and Examples (from the transcript)
Example study: qualitative exploration of penile size and perceptions of women’s sexual satisfaction
Used focus groups with older and younger women, separated and mixed groups
Used pictures to elicit reactions (e.g., responses to different sizes)
Observed reactions to images, capturing lived experiences and perceptions
Observational qualitative work:
Observed cultural practices in a Rastafarian camp
Attended parties at night to observe behaviors and reactions
Coding qualitative data:
Different words or phrases are coded to identify similar meanings and patterns
Purpose of qualitative approach examples: to assess lived experiences, perceptions, and meanings behind behaviors
Qualitative analysis is appropriate when studying how people feel, perceive, or experience phenomena, or when the research question focuses on subjective experiences rather than numerical measurement
Quantitative Data: Methods and Examples (from the transcript)
Quantitative analysis is used when the goal is to study relationships, test hypotheses, and generalize findings to a larger population
Example mentions:
Relationship between number of hours students practice QA per week and overall performance
Studying the impact of a policy change across time periods (before vs after)
Analyzing demographic factors (age, gender, income) and their effect on voting behavior in a local election
Large-scale studies using measurable quantities to determine relationships or causality
Quantitative approach helps with:
Testing hypotheses and establishing causal relationships
Studying large populations for generalization
Economic analysis and market research
Analyzing public opinion
A common caution: when describing social phenomena with a quantitative lens, one must ensure appropriate sampling, measurement, and control for confounding factors to avoid spurious conclusions
Example of a qualitative vs quantitative decision based on study goals:
If you want to know how people experience and perceive a policy change, qualitative methods may be preferred
If you want to determine if there is a systemic relationship between demographics and voting behavior, a quantitative approach may be appropriate
Coding and Analyzing Qualitative Data
After collecting qualitative data (texts, transcripts, recordings, images), you code the data to create categories and themes
Coding helps transform rich qualitative material into analyzable units while preserving meaning
The process involves looking for similarities and patterns across different respondents, groups, or contexts
Historical Emergence and Theoretical Foundations
Rise of quantitative analysis in the social sciences:
Late 19th to early 20th century shift toward rigorous scientific methods in sociology, economics, psychology, etc.
The goal was to measure social phenomena objectively, like natural scientists measure physical phenomena
The use of variables and quantification allows objective measurement of concepts like inflation or deviance
Borrowed and adapted measurement scales:
Researchers reuse or adapt scales to measure constructs (e.g., depression scales) from prior studies
Scales can be manipulated or adjusted for specific research purposes
Emile Durkheim as a key proponent of positivism in sociology:
Studied suicide to show social factors influence seemingly individual phenomena
Pushed for quantitative, sociological explanations based on social factors rather than purely individual psychology
The advent of computers transformed data processing and analysis:
Early data processing enabled larger datasets and more complex analyses
This laid the groundwork for modern statistics, econometrics, and big data approaches
Connection to econometrics:
Emerged prominently around World War II
Econometrics combines statistics with economic theory to measure relationships between variables
Pearson contributed a foundational method for measuring correlation between variables
Caution about interpretation with new technologies:
The transcript notes that data collection, big data, and AI bring opportunities but require careful interpretation and ethical considerations
Tools, Data, and Technology: Modern Context
Ubiquity of data collection devices (e.g., wearables, smartphones) tracking activity and behavior
Data-driven feedback in everyday life (e.g., screen time, activity suggestions) can influence behavior and self-perception
Advancements in data science (big data, AI) expand what can be measured and analyzed, but also raise ethical considerations around privacy, bias, and interpretation
The reliance on these tools often blurs the line between qualitative insights and quantitative measurements, leading to an integrated approach in research that combines both methods.
Connections to Foundational Principles and Prior Coursework
Alignment with the scientific method: formulation of hypotheses, measurement, data collection, analysis, and interpretation
The shift toward quantification supports reliability, validity, and the ability to generalize findings from samples to populations
Foundational figures (e.g., Durkheim) illustrate the transition from individualistic explanations to social-structure explanations
Scales, measurements, and the use of established instruments reflect cross-study comparability and cumulative knowledge building
Ethical, Philosophical, and Practical Implications
Positivism emphasizes objective measurement and observable phenomena, but researchers must remain aware of potential bias and measurement error
Large-scale data collection raises concerns about privacy, surveillance, consent, and data governance
The use of AI and big data requires careful interpretation to avoid overgeneralization, misinterpretation, or reinforcing existing biases
Practical considerations:
Choosing the right method for the research question
Ensuring robust data collection, coding, and analysis procedures
Balancing depth (qualitative richness) with breadth (quantitative generalizability)
Exercise and Application Notes (from transcript)
When facing unfamiliar or incomplete readings, focus on the behavior and variables involved; identify factors such as age, gender, and other demographics
Open-ended interviews may be used to explore how different groups experience a phenomenon
Hypotheses mentioned: exploring whether relationships exist between satisfaction and productivity; whether demographic factors influence voting behavior
Remember: in large studies with measurable quantities, you aim to identify relationships and potentially causal links, while controlling for confounding variables
Quick Takeaways
Qualitative methods illuminate lived experiences, perceptions, and meanings; data are nonnumeric and analyzed through coding and thematic extraction
Quantitative methods quantify relationships, test hypotheses, and enable generalization to larger populations; data are numeric and analyzed with statistical models
Conceptualization defines what is studied; operationalization defines how it is measured
Computers, econometrics, and big data have expanded social science capabilities but require careful, ethical interpretation
Foundational theories (e.g., Durkheim’s sociology, positivism) underpin the rationale for systematic measurement in the social sciences
Measurement, Quantitative Methods, and the Research Process
As content advances, more advanced aspects will be introduced (data analysis, measurement, etc.).
Data sources will include: Statistics Canada and other statistical institutes; data sets available for unemployment, inflation, etc.
Software and tools introduced: MATLAB, R, Excel, Python; these are used for data analysis.
Before computers, data collection (e.g., thousands of questionnaires) required manual measurement and analysis; now we have computational tools to analyze data.
The course will cover measurement as a key concept, with emphasis on numerical representation of social phenomena.
Example: measuring student success across SAGEF programs in Quebec; what constitutes and quantifies “success” (e.g., average scores, completion rates).
Discussion of data sources and data availability as a foundation for analysis.
Emphasis on real-world relevance: measuring unemployment, inflation, and other macro indicators.
The course will touch on ethical, practical, and transparency considerations in research (e.g., sharing data and replication).
During the pandemic, policy actions (e.g., central banks lowering interest rates) were guided by prior research to stimulate consumption and investment; short-term effects included housing market activity and potential long-run inflation.
Ten-minute break noted during lecture; students should rest and return.
Measurement and Quantitative Methods: Core Concepts
Quantitative methods allow precise measurement and quantification of social variables; phenomena can be expressed in numbers or values.
Example: measure student success across SAGEF programs; possible metrics include average scores at different stages of the program.
Poverty measurement concepts:
Absolute poverty: absence of basic necessities; defined by a threshold for income or consumption that cannot cover essentials.
Relative poverty: assessed in relation to others in the economy; focuses on inequality or lack of resources compared to the population.
When measuring poverty, researchers must specify whether they use absolute or relative definitions and how income/consumption thresholds are determined.
Population vs. Sample:
Population: the entire group of interest to study.
Sample: a representative subset drawn from the population.
Example: studying SageF students in Quebec; population might be all SageF students in Quebec; a sample might be students from several institutions.
Unit of analysis vs. Individual:
An individual is a member of the study (e.g., a person, a family, a country).
The variable is a characteristic of the individuals or units (e.g., age, income, education level).
Constant vs. Variable:
A variable can take different values across individuals (e.g., age, income).
A constant does not vary across individuals in a given study (e.g., the institution being studied if focusing on one school).
Car example to illustrate units and variables:
Individuals: cars in a dataset (Make and Model, etc.).
Variables: type of body (station wagon, midsize, large), transmission type (manual, automatic), number of cylinders, city MPG, highway MPG, annual fuel cost.
The vehicle make/model may be considered a constant within a fixed study sample; MPG, fuel cost, etc., are variables that differ across cars.
Quantitative versus Qualitative variables:
Qualitative (categorical) variables (e.g., gender) can be coded numerically for analysis (e.g., 1 = male, 2 = female); numerical coding does not imply a rank or value judgment.
Within a quantitative framework, these coded values can be used in various analyses (e.g., chi-square tests) to examine relationships with other variables.
Measurement vs. operationalization:
Conceptualization: defining a concept using other concepts or existing definitions, often drawn from prior research.
Operationalization: specifying the procedures or measurements used to quantify a variable (the exact steps, instruments, or data sources).
Inflation and price-level concepts:
Conceptual definition: inflation is the change in the average price level of goods and services in an economy over a period of time.
Operationalization: use the Consumer Price Index (CPI) with a basket of goods and services tracked over time; inflation rate is the percentage change in CPI.
Formula for inflation rate (example):
ext{Inflation rate} = rac{Pt - P{t-1}}{P_{t-1}} imes 100 ext{ exttt{A0%}}.For a related macro measure, GDP growth can be measured similarly:
ext{GDP Growth}t = rac{GDPt - GDP{t-1}}{GDP{t-1}} imes 100 ext{ exttt{A0%}}.
Conceptualization and definitions in practice:
Often you derive definitions from prior studies (e.g., inflation definitions from existing literature) and cite sources (e.g., John et al.).
When studying poverty, specify whether you adopt absolute or relative definitions and provide clear thresholds or comparison bases.
Data collection and sampling considerations:
With large-scale data, researchers use sampling to infer about the population; the larger the sample, the more reliable the inference.
Replicability requires that other researchers can reproduce results using the same methods and data sources; sharing data and code enhances transparency.
Large-scale surveys and data sources:
World Value Survey (WVS): data collected in over 100 countries on various variables.
CPS: U.S. Current Population Survey; used for time-series and cross-country comparisons.
Time-use or time-availability studies in various countries to analyze effects on labor supply and human capital.
Applications of quantitative methods:
Investigate causal relationships (e.g., hours spent studying and QA scores) using regression analysis.
Use correlational analysis and regression equations to predict one variable from another.
Comparative studies: cross-cultural and cross-temporal comparisons (e.g., healthcare costs and outcomes under single-payer vs. multi-payer systems across OECD countries).
Policy and decision making: evidence-based policy making relies on empirical evidence and statistical support.
Data sharing and transparency:
Modern research emphasizes posting data sources and measurement schemes so others can reproduce analyses.
There is concern about publishing only statistically significant results; transparency helps address bias and fosters replication.
Policy context example from economics:
During the COVID-19 pandemic, central banks lowered interest rates to boost consumption and investment; this policy response has complex short-term and long-term effects (e.g., housing market activity and inflation pressures).
The Research Process: From Topic to Theory to Data
Initial steps:
Identify a research problem or topic (e.g., social media usage among teenagers).
Review the literature to see what has been studied and what gaps exist (e.g., prior work on Instagram vs. TikTok usage).
Formulate a research question based on identified gaps.
Design the methodology (qualitative, quantitative, or mixed) and decide who will be studied and how data will be gathered (interviews, questionnaires, etc.).
Theory and methodology integration:
Read literature to uncover relevant theories and models that explain relationships you want to study.
Decide on data collection methods (surveys, interviews, field research, content analysis, existing data) and sampling strategy.
Collect data, analyze, and draw conclusions.
Communicate findings clearly and situate them within existing literature and theory.
The non-linear nature of research (Babi 2017):
Unlike a linear flow, the research process often cycles back and forth.
Example flow (not strictly linear): topic → literature → theory → conceptualization → operationalization → measurement → variables → sampling → data collection → data cleaning (handling missing data) → analysis → conclusions → policy implications, with feedback loops that may lead to revising earlier steps.
Conceptualization and operationalization in practice:
Conceptualization: define a variable in terms of other concepts or existing definitions (e.g., inflation as a price-level change).
Operationalization: specify how to measure that concept (e.g., CPI with a basket of goods; calculate inflation as the percentage change in CPI).
For poverty, choose between relative vs absolute definitions and specify measurement thresholds or criteria.
Defining constants and variables in practice:
Decide which factors are variables and which could be constants given the study scope (e.g., if focusing on a single university’s SageF program, the institution is a constant; if studying across multiple institutions, the university variable becomes a factor).
Conceptualization and measurement example recap:
Inflation: conceptual definition = average price level change; measurement = CPI with basket of goods; formula for inflation rate as shown above.
Poverty: conceptualization as absolute or relative; measurement method detailed based on chosen definition; example thresholds for household income or consumption.
Practical notes for conducting research:
Determine the definition of variables from prior research when possible.
Specify how you will measure variables to enable replication.
Be prepared to adjust sampling, measurement, or analysis plans based on data availability or unexpected findings.
Key takeaways for topic two (to be revisited in topic three):
Variables, measurements, research questions, and hypotheses form the core of empirical inquiry.
Conceptualization and operationalization are the bridge from abstract ideas to measurable constructs.
Understanding units of analysis, population vs. sample, and the role of constants vs. variables is essential for study design.
Key Definitions and Concepts (Summary)
Population: the entire group of interest to study.
Sample: a subset representing the population.
Unit of analysis: the entity being studied (e.g., individuals, households, countries).
Individual: a member of the study population (could be a person, family, country, etc.).
Variable: a characteristic that can take different values across individuals or units (e.g., age, income).
Constant: a characteristic that does not vary across individuals in a given study (e.g., the institution in a single-site study).
Qualitative (categorical) variable: describes categories (e.g., gender, race); can be coded numerically for analysis.
Quantitative variable: takes numerical values (e.g., age, income, MPG).
Conceptualization: defining a concept using other concepts or existing definitions.
Operationalization: specifying how to measure a concept (the procedures, instruments, data sources).
CPI (Consumer Price Index): used to measure inflation via a basket of goods and services tracked over time.
Inflation rate: ext{Inflation rate} = rac{Pt - P{t-1}}{P_{t-1}} imes 100 ext{ exttt{A0%}}.
GDP Growth rate: ext{GDP Growth}t = rac{GDPt - GDP{t-1}}{GDP{t-1}} imes 100 ext{ exttt{A0%}}.
Replicability: the ability of other researchers to reproduce results using the same methods and data; emphasizes transparency and data sharing.
World Value Survey (WVS): large-scale, cross-country survey used for comparative research.
CPS (Current Population Survey): U.S. survey used for labor and demographic analysis.
Cross-cultural and cross-temporal comparison: comparing variables across different cultures or time periods to draw inferences.
Evidence-based policy making: using empirical evidence and statistics to inform policy decisions.
Practical Implications and Real-World Relevance
Data availability and methodological choices influence what can be studied and how results are interpreted.
Transparency and replicability help ensure credibility of research, reduce bias, and enable verification by others.
Policy relevance arises when research informs government decisions, such as healthcare funding or education programs.
Ethical considerations include proper data sharing, avoiding selective reporting, and acknowledging limitations in data and methods.
The integration of theory, measurement, and data analysis enables robust understanding of social phenomena and supports informed decision-making.
Quick Reference Formulas
Inflation rate via CPI:
GDP growth rate:
Other measures (conceptual): population vs. sample; variable vs. constant (descriptions, not numeric formulas).
Study Tips Based on Lecture Content
Before collecting data, clearly define whether you are measuring absolute or relative concepts (e.g., poverty).
When working with large data, plan sampling to ensure representativeness of the population of interest.
Consider how you will document and share your data and methods to facilitate replication.
Use CPI-based inflation and GDP growth formulas to ground macroeconomic analyses in concrete calculations.
Remember the non-linear nature of research: be prepared to revise questions, definitions, or methods as you learn from data and literature.
Population, sample, and unit of analysis
Unit of analysis: individuals.
Population vs. sample
Population example: Canadians aged 18 to 25 (or 18 to 35) — the broader group you want to study.
Sample example: A subset of Canadians aged 18 to 35 that you actually study.
A population can be broader than the sample you study; you may move up or down the age range, etc.
Broad concepts:
Population: the entire group you’re interested in describing or making inferences about.
Sample: a subset of the population that you actually collect data from.
The goal is to use information from the sample to infer something about the population.
Conceptualization and operationalization
Conceptualization: defining a concept at a theoretical level (e.g., socioeconomic status).
Operationalization: specifying how a concept will be measured in practice.
Example: socioeconomic status may be operationalized using individual income.
Key idea: bridge between abstract concepts and concrete measurements used in data collection.
Classifications of variables; Population vs. parameter vs. statistic
Parameter vs. statistic:
Parameter: a numerical value that describes a characteristic of the population. Not typically known because you can’t measure every member of the population.
Statistic: a numerical value that describes a characteristic of the sample. Used to make inferences about the population parameter.
Examples:
Population parameter example: the average income of Canadians aged 18 to 35:
where is the population mean.Sample statistic example: the average income of 300 individuals aged 18 to 35 (from the sample):
where is the sample mean.
Commonly, researchers use the sample statistic to make inferences about the population parameter.
Example from a text/lecture:
Researchers surveyed 15,624 American high school students (grades 9–12) and found that 27.2% were in grade 9.
The percentage of all American high school students who are in grade 9 is 27.5%.
The 27.2% is a sample statistic (\hat{p}); the 27.5% is a population parameter (p).
The percentage of those surveyed who were in grade 9 and had carried a gun to school was 4.5% (\hat{p}_{\text{gun}}), a sample statistic.
Quick recap of notation:
Population proportion:
Sample proportion:
Population mean:
Sample mean:
From topic organization to hypothesis development
Formulating research questions as the first step in the research process.
Example topic: Breakfast consumption and academic performance.
Broad topic: breakfast consumption.
Research question example: What is the relationship between breakfast consumption and academic performance among high school students?
Variables identified: breakfast consumption (independent) and academic performance (dependent).
Operationalization example: academic performance measured by grades (e.g., GPA or letter grades).
Bad vs. good research questions:
Bad: How does social media affect people’s behavior? (too broad)
Good: What effect does daily use of Instagram have on attention span? (narrowed to one platform and a specific outcome)
Research questions can guide multiple sub-questions depending on scope.
Literature review informs hypotheses; hypotheses are testable statements derived from the research questions and prior work.
In practice, topics can be explored with convenience samples (e.g., a single high school) which limits generalizability but can still meet study objectives.
Two key types of hypotheses in hypothesis testing
Null hypothesis (H0): there is no effect or no relationship between the variables.
Alternative hypothesis (H1 or HA): there is an effect or a relationship between the variables.
Form and content:
If the topic is breakfast and academic performance measured by grades: H0 might state there is no relationship between breakfast consumption and academic performance; H1 would state there is a relationship.
Examples of hypothesis forms:
Relationship/association example: There is a relationship between breakfast consumption and academic performance among high school students.
No relationship example: There is no relationship between hours of sleep and alertness in class.
Difference example: There is no difference in mean height between men and women: vs. .
No effect example: The use of a calculator has no effect on student test scores.
Notational example (mean comparison):
Null:
Alternative:
When data are numeric, you may state hypotheses in parametric form; when data are categorical, you may use proportions and/or counts in tests like (\chi^2).
Directionality:
A directional (one-tailed) alternative might be: H1: \mu{\text{men}} < \mu_{\text{women}}.
A non-directional (two-tailed) alternative: .
How hypotheses relate to theory and prior work:
Hypotheses are informed by the literature and prior findings.
A study’s results can contribute to theory and inform future research.
Hypotheses influence the choice of analysis methods (e.g., t-tests, chi-square, correlation) depending on variable type and measurement scale.
Practical note: hypotheses should be testable and falsifiable; the research design should be capable of supporting or refuting them.
Observational vs experimental designs
Observational studies:
Researchers observe and measure variables without manipulating the participants' environment or behavior.
Methods include surveys, interviews, and naturalistic observation.
Pros: more naturalistic; less ethical/logistical burden in some cases.
Cons: susceptible to confounding or lurking variables since the researcher does not control the assignment to groups or conditions.
Experimental studies:
Researchers deliberately impose treatments or conditions on participants and observe outcomes.
Key features: random assignment to conditions/groups and control of extraneous variables.
Goal: establish causality (effect of X on Y) by reducing alternative explanations.
Example: testing whether aspirin reduces the risk of heart attack; use of placebo controls; random assignment to treatment vs. placebo.
Confounding and lurking variables:
In observational studies, unseen variables can influence both the predictor and outcome, leading to spurious associations.
In experiments, random assignment helps balance these confounders across groups, increasing internal validity.
Practical psychology example:
Lab experiments with humans (or animals) allow control over conditions to infer causality; observational methods cover surveys, interviews, or watching behavior in natural settings.
Quick practice prompts (to distinguish designs):
If researchers randomly select 1,000 teenagers for a survey, and then report that teens who eat with family at least five times a week have better grades, is this observational or experimental?
If researchers set breakfast conditions (some eat breakfast, some skip) and then compare grades, is this observational or experimental?
Important caveat about generalization:
Random sampling supports generalization to the population, but small samples (e.g., only 3 individuals) may give little information about the population.
Sampling, generalization, and measures
Random sampling and inference:
A random sample provides a basis for inferring properties about the population from which the sample was drawn.
The strength of generalization increases with larger, well-randomized samples.
Population vs. sample revisited:
Population: the whole group of interest.
Sample: the actual individuals studied.
Additional notes on measurement and analysis (connecting to future topics):
The type of data (categorical vs. numeric) and measurement level affects which statistical tests are appropriate (e.g., (\chi^2) for associations between categorical variables; t-tests or ANOVA for comparing means; correlation for linear relationships).
The choice of hypothesis test depends on the scale of measurement and the research question (difference, relationship, effect).
Summary of practical takeaways
Always start with a clear topic and narrow it to a specific research question with defined variables.
Distinguish population from sample; identify the corresponding parameter vs statistic.
Use conceptualization and operationalization to turn abstract ideas into measurable variables.
Formulate hypotheses (null and alternative) that are testable and falsifiable; consider directional vs non-directional options.
Choose study design (observational vs experimental) based on research goals, feasibility, and the possibility of inferring causality; be mindful of confounding factors in observational studies.
Realize that results contribute to theory but generalization depends on sampling and design quality.
Use the workbook and practice problems to reinforce these concepts and to become proficient at formulating proper hypotheses and selecting appropriate analyses.
Quick LaTeX reference for notes
Population mean:
Sample mean:
Population proportion: ; Sample proportion:
Hypothesis notations: ;
Example null for mean comparison:
Example alternative:
Directional example: H1: \mu{\text{men}} < \mu_{\text{women}}
Relationship/association (correlation) example: vs.
Common tests (based on data): chi-square for categorical data, t-test/ANOVA for mean comparisons, correlation for relationships
Regression analysis for predicting outcomes and examining relationships between variables.
Experimental Design and Random Assignment
Primary goal in experiments: determine if a treatment has an effect by comparing a treatment group to a baseline control group. The control group receives no treatment (or standard treatment) so that everything else is held constant across groups.
If the treated group outperforms the control group, the treatment is considered effective.
Random assignment (randomization)
Random assignment helps create groups that are roughly equivalent at the start of the experiment, reducing confounding (where other factors could explain observed differences).
How it works (example): you have 100 participants and want two groups of 50. Use a random number generator to assign 50 to Group A and the remaining 50 to Group B.
This process prevents self-selection into groups and helps ensure comparability.
Key example: long-hand note taking study (Oppenheimer, Mueller)
Compare handwritten notes (vs alternative note-taking methods) to assess effect on learning or recall.
Control for confounders: participants’ self-reported study hours can be biased; design ensures the main difference is the note-taking method, not other factors.
Important principle: well-designed experiments avoid confounding by ensuring groups are treated the same except for the treatment.
Pfizer COVID-19 vaccine randomized trial (illustrative, real-world example)
Population size: 43,548 volunteers.
Randomization: half were randomly assigned to receive two vaccine doses 21 days apart; the other half were randomly assigned to receive two saline placebo shots (placebo) to mimic vaccine administration.
Outcome: after several months, the vaccine was determined to be $95\%$ effective in preventing COVID-19.
Purpose of random assignment in this trial: ensure that any differences in infection rates are due to the vaccine, not other factors (e.g., health status, exposure, or behavior).
Interpretation: if the vaccine group shows substantially lower infection rates than the placebo group, the difference is attributable to the vaccine efficacy rather than other variables.
How random assignment could be implemented in practice: assign each volunteer an ID (e.g., 1 to 43,548), use a random number generator to select IDs for the vaccine group, and assign the rest to placebo.
Conceptual takeaway: randomization creates roughly equivalent groups at baseline, enabling causal inference about the treatment effect.
Observational studies vs experiments
Observational studies often suffer from confounding because the treatment is not randomly assigned.
Well-designed experiments mitigate confounding by ensuring that all aspects except the treatment are the same across groups.
In observational settings, researchers may rely on self-reported data or natural variation, but causal claims are weaker due to potential confounders.
Random assignment: deeper understanding
Random assignment is a statistical technique (often using random number generators) to allocate participants to groups.
Goal: produce roughly equivalent groups at the outset to isolate the effect of the treatment.
Simple illustrative example (numbers): with 100 individuals, randomly assign 50 to the treatment and 50 to the control using a RNG.
Sampling vs experimentation context (transition to sampling techniques)
Population vs sample: the population is the entire group of interest; a sample is a subset used to make inferences about the population.
Why sample? It’s often too costly or impossible to study everyone.
Random sampling allows inference about the population if the sample is representative and large enough.
Non-random sampling can introduce bias and limit the generalizability of results.
When random sampling is not possible or ethical
Some topics involve vulnerability or harm, making random assignment or random sampling inappropriate (e.g., rape victims or incest survivors).
In such cases, researchers rely on voluntary participation or convenience samples, which may bias results and limit population-level inferences.
Research ethics and seminars (context from the course)
There is mention of a mandatory research ethics seminar as part of the course.
Ethical considerations guide when randomization or certain sampling methods are appropriate or inappropriate.
Sampling strategies overview
Non-random sampling methods (potential bias):
Convenience sampling: select individuals who are easy to reach; may bias results if the sample is not representative.
Voluntary response (self-selected samples): people choose to participate; can bias estimates toward the views of those who strongly feel a certain way.
Systematic sampling: often treated as a probability method, but it can be non-random if the starting point is not random or if the list has hidden patterns.
Random sampling methods (probability sampling):
Simple random sampling: every individual in the population has an equal chance of being selected; every possible sample of a given size has equal probability.
Stratified random sampling: divide the population into strata (subgroups) based on a characteristic (e.g., gender, age, race), then perform random sampling within each stratum.
Proportional allocation: the sample from each stratum is proportional to its size in the population (e.g., if 50% male and 50% female in the population, aim for 50/50 in the sample).
Purpose: ensure representation of key subgroups and improve accuracy.
Systematic sampling: select a random start and then pick every kth element (k is the sampling interval, typically N/n).
Example method: listing the population in a natural order (e.g., alphabetical), choose a random start between 1 and k, then select the 1st, (1+k), (1+2k), etc., to obtain n samples.
Cluster sampling: groups (clusters) are sampled and then all members of selected clusters are sampled; useful when the population is spread out geographically.
Practical demonstrations and calculations
Simple random sampling activity (conceptual):
Draw a sample of size n from a population of size N using a random number generator to select IDs.
Compute the sample mean: where $X_i$ are the observed values in the sample.
Compare sample mean to population mean ; expect some variability due to sampling.
Sampling variability
Different random samples from the same population can yield different sample means.
This variability explains why larger samples tend to yield estimates closer to the population parameter.
Stratified sampling in practice
If the population has known proportions by strata (e.g., gender, race), the sample should reflect those proportions to avoid bias.
Example: if a population is 50% male and 50% female, stratified sampling with proportional allocation would aim for roughly equal representation in the sample.
Real-world applications and examples
National surveys and statistics agencies (e.g., Stats Canada, Bureau of Labor Statistics) use stratified and other probability sampling methods to obtain representative samples.
In healthcare, sampling plans for patient satisfaction, discharge studies, or other health metrics may use stratified or cluster designs to capture variation across hospitals or patient groups.
Key concepts to remember for exams
Population vs sample; sampling frame; sampling bias.
Random assignment (causal inference) vs random sampling (generalization).
Confounding and how randomization mitigates it.
Types of sampling methods and their biases: non-random (convenience, voluntary response) vs random (simple random, stratified, systematic, cluster).
Sampling variability and the idea that the sample mean is an estimator of the population mean: \text{efficacy} = 1 - \text{RR}k = N/n$$ for population size $N$ and sample size $n$.
Sampling variability: the natural variation in a statistic (e.g., the sample mean) from one random sample to another.
Proportional vs. stratified allocation: stratified sampling allocates samples to strata; proportional allocation mirrors population proportions.
Final notes for exam readiness
Be able to describe, in your own words, why random assignment helps establish causality.
Be able to differentiate between random sampling and random assignment and explain their respective purposes.
Be able to outline a basic sampling plan (which method you would use and why) given a hypothetical research question.
Be able to compute and interpret the sample mean and discuss how sampling variability affects estimates.
Be prepared to discuss ethical considerations that might prevent randomization or random sampling in sensitive topics.