Statistics: Collecting Data and Study Design
Chapter Objectives
Describe how the conclusions that can be drawn from a statistical study depend on the way in which data are collected.
Explain why random selection is an important component of a sampling plan and why random assignment is an important component in the design of a statistical experiment.
Critically evaluate the design of a statistical study.
Assess whether conclusions drawn from a study are appropriate, given a description of the statistical study.
Statistics and Variability
Statistical methods allow for the collection, description, analysis, and drawing of conclusions from data.
With any data collection, variability in the data is naturally expected.
Example: Monitoring Water Quality
As part of daily water quality monitoring efforts, an environmental control board selects water specimens from a particular well each day.
The concentration of contaminants in parts per million () is measured for each of the specimens, and the daily average of the measurements is calculated.
A histogram summarizing these average contamination values over days prior to any incident establishes baseline variability.

Chemical Spill Scenario:
A chemical spill occurs at a manufacturing plant located from the well.
It is unknown whether such a spill contaminates groundwater in the immediate area or whether a spill at this distance affects well water quality.
One month after the spill, water specimens are collected from the well, yielding a daily average contamination of .
Evaluating Spill Effects:
Before the spill, daily average contaminant concentrations naturally varied across days.
An average concentration of falls well within the normal pre-spill historical range shown on the histogram. Therefore, observing post-spill is not evidence that well contamination increased.
If the post-spill calculated average were , this value is much less common in the baseline distribution.
If the post-spill calculated average were , this value is entirely outside the typical pre-spill range, leading to the definitive conclusion that the well water contamination level increased.
Observational Studies vs. Experiments
Core Definitions
Population: The entire collection of individuals or objects about which information is desired.
Sample: A subset or portion of the population selected for study.
Observational Study: A study in which characteristics of a sample selected from one or more existing populations are observed. The objective is to use sample data to learn about the corresponding population. It is critical to obtain a sample that is representative of the population.
Experiment: A study in which the investigator observes how a response variable behaves under different experimental conditions (treatments). The investigator determines who is in each experimental group and sets the experimental conditions. It is critical to establish comparable experimental groups.
Example: Screen Time and Mental Health
Study Reference: "Is Screen Time Associated with Anxiety or Depression in Young People?" (BMC Public Health [2019]: 82).
Methodology: Researchers assessed associations between screen-based device usage at age (measured via questionnaire) and anxiety/depression symptoms at age (measured using a self-administered diagnostic instrument named the Revised Clinical Interview Schedule).
Result: Participants spending more time using computers were slightly more likely to experience symptoms of anxiety.
Classification: This is an observational study because researchers did not assign teenagers to specific usage groups or manipulate their daily computer usage.
Planning an Observational Study
Parameters vs. Statistics
Population Characteristic (Population Parameter): A numerical value that describes an entire population.
Statistic: A numerical value that describes a sample.
Representative Samples and Selection Methods
Generalizability depends entirely on obtaining a representative sample.
Convenience Sampling: Selecting an easily accessible group to form a sample. This approach is unreliable because convenience samples systematically fail to represent the broader population.
Simple Random Sample (SRS):
A simple random sample of size is selected such that every possible distinct sample of size has an equal chance of being selected.
represents the sample size (total number of individuals or objects sampled).
Sampling Frame: A comprehensive list of all objects or individuals within the target population.
Selection tools include random number generators and tables of random digits.
Sampling Procedures:
Sampling with Replacement: An individual or object, once selected, is returned to the population prior to the next selection. A single population member can appear multiple times in the sample.
Sampling without Replacement: An individual or object, once selected, is removed from the population for subsequent draws. The sample consists of unique individuals.
Sample Size vs. Population Size
Common Misconception: A sample must represent a large fraction/percentage of the population to be accurate.
Correct Principle: Provided the population is not extremely small, the risk of obtaining an unrepresentative sample depends on the absolute sample size (risk decreases as increases), NOT on the proportion of the population sampled. Random selection allows small fractions of a population to accurately reflect the entire population.
Overview of Sampling Methods
Simple Random Sampling: Every distinct sample of size has an equal probability of selection.
Stratified Sampling: Divides the population into non-overlapping subgroups (strata) based on relevant characteristics, then takes a separate random sample from each stratum.
Systematic Sampling: Divides an ordered population frame into intervals of size , randomly selects a starting point within the first individuals, and selects every individual thereafter.
Cluster Sampling: Divides the population into non-overlapping subgroups (clusters), randomly selects a subset of clusters, and includes ALL individuals within selected clusters in the final sample.
Types of Bias in Sampling
Bias: The systematic tendency for sample estimates to differ from the true population value.
Selection Bias: Introduced when the sampling frame or selection method systematically overrepresents or underrepresents specific segments of the population.
Nonresponse Bias: Occurs when data or responses are not collected from all individuals chosen for the sample.
Measurement (Response) Bias: Occurs when the data collection method or observation process systematically alters or distorts the true values.
Case Study: Fast-Food Caloric Intake
Study Reference: "What People Buy from Fast-Food Restaurants: Caloric Content and Menu Item Selection" (Obesity [2009]: 1369–1374).
Findings: The average lunch intake at New York City fast-food restaurants was .
Data Collection Design: Researchers randomly selected fast-food locations across NYC. At each location, adult customers were approached upon entering and asked to provide their receipt upon exiting alongside a brief survey.
Evaluation:
Population of Interest: Adults eating at fast-food restaurants in New York City.
Sampling Appropriateness: Randomly selecting locations ensured every fast-food outlet had an equal chance of inclusion, making it a reasonable approach given logistical constraints.
Representativeness: Patrons at randomly selected outlets can reasonably represent NYC fast-food diners overall.
Potential Biases:
Response Bias: Approaching customers before they ordered likely altered their purchasing habits.
Nonresponse Bias: Refusal by some customers to participate means non-respondents may differ systematically in their food choices from respondents.
Planning an Experiment
Experimental Terminology and Design Goals
Simple Comparative Experiment: Measures a response variable across multiple experimental conditions.
Experimental Conditions (Treatments): Specific environmental conditions or interventions applied to units.
Experimental Unit: The smallest unit to which a treatment is applied.
Experimental Design: The structured plan for an experiment, engineered to isolate variation in response caused by treatments from variation caused by background noise.
Confounding Variables
Confounding: Occurs when the effects of two variables on the response variable cannot be distinguished from one another.
Confounding Variable: A variable associated with both the explanatory variable and the response variable, offering an alternative explanation for observed differences in response.

Tortilla Chip Example:
Goal: Evaluate the effect of frying time (, , ) on chip moisture content.
Flawed Design: Batch 1 used exclusively for , Batch 2 for , and Batch 3 for .
Result: Differences in batch composition (moisture, oil, density) are confounded with frying time, making it impossible to attribute moisture differences solely to frying time.
Temperature and Exam Performance Example:
Explanatory Variable / Treatments: Different room temperatures.
Response Variable: Exam performance.
Experimental Units: Individual students or entire class sections.
Potential Confounding: If all low-temperature tests occur in early morning and high-temperature tests occur in late afternoon, time of day (and associated student fatigue) is confounded with temperature.
Key Design Strategies
Direct Control: Holding potential sources of variability constant at a fixed level across all experimental units to completely eliminate them as confounding factors.
Random Assignment: Randomly distributing experimental units across treatment groups to ensure remaining uncontrolled sources of variability affect groups only by chance. This creates equivalent groups where treatment is the only systematic difference.
Case Study: Online Review and Exam Performance
Scenario: first-year college students participate in an experiment testing whether an online review module improves exam scores.
Baseline Heterogeneity: Students differ significantly in academic achievement, as evidenced by their Math and Verbal SAT scores.

Role of Random Assignment:
Assigning students randomly to review vs. non-review groups ensures high-achieving and low-achieving students are balanced evenly across both groups across both Verbal and Math SAT dimensions.
Random assignment also balances unmeasured potential confounders such as student motivation.
Motivation Analysis:
In an observational study, motivation would be a confounding variable because highly motivated students self-select into completing the review and also possess superior study habits.
In a randomized experiment, random assignment breaks the association between motivation and treatment status, eliminating motivation as a confounding variable.
Methods for Executing Random Assignment
Physically drawing units or group tags from a container.
Utilizing a random number generator.
Using a physical chance mechanism like flipping a coin.
Control Features in Experiments
Control Group: An experimental group receiving no active treatment, serving as a baseline.
Placebo: A dummy treatment identical in appearance and delivery to the active treatment but devoid of active ingredients.
Single-Blind Experiment: An experiment in which subjects are unaware of which treatment condition they are receiving.
Double-Blind Experiment: An experiment in which neither the subjects nor the individuals evaluating/measuring the responses know which treatment each subject received.
Systematic Questions for Evaluating Experimental Quality
What specific question is the experiment designed to answer?
What are the experimental conditions (treatments)?
What is the response variable?
What are the experimental units, and how were they selected?
Does the design feature random assignment of units to treatments? If not, what confounding variables threaten internal validity?
Does the experiment feature a control group and/or a placebo group?
Does the experiment utilize single- or double-blinding where appropriate?
Determining Reasonable Conclusions
Generalization vs. Causation Guidelines
Generalizing Sample Results to the Population: Justified ONLY when Random Selection was used to obtain the sample from the target population.
Establishing Cause-and-Effect Relationships: Justified ONLY when Random Assignment was used to allocate experimental units to experimental conditions.
Case Study: Fresh Fruit and Diabetes
Article Headline: "Eating Fresh Fruit Lowers Risk of Type 2 Diabetes, Study Claims" (New York Post, June 4, 2021).
Study Scope: Observational tracking of Australians over years.
Findings: Individuals consuming at least servings of fruit daily exhibited lower odds of developing type 2 diabetes.
Evaluation of Claims:
It is invalid to claim cause-and-effect because participants were not randomly assigned to fruit consumption levels.
Potential Confounders: Fruit consumers may engage in more physical exercise, maintain healthier overall diets, or have higher income levels allowing for greater total healthcare investment. These variables offer plausible alternative explanations.
Summary of Common Errors in Statistical Design
Mistake 1: Inferring a cause-and-effect relationship from an observational study. (Avoid entirely due to uncontrolled confounding).
Mistake 2: Generalizing findings from an experiment that relies on volunteer subjects to the general population. (Only valid if it can be convincingly demonstrated that volunteers accurately reflect the target population).
Mistake 3: Generalizing conclusions from a poorly designed observational study or non-random sample to a population. (Generalization requires representative random sampling).
Practice and Conceptual Questions
Question 1: How could the sit-stand desk experiment (evaluating daily sitting time) be conducted as an observational study?
Answer: Observe daily sitting time among employees who independently chose to use sit-stand desks versus those who use standard desks, without intervening or assigning setups.
Question 2: How could the screen time observational study be conducted as an experiment?
Answer: Randomly assign participants to specific required daily screen usage times (e.g., , , ) and evaluate anxiety levels after a defined period.
Question 3: Would surveying students on one dorm floor be an appropriate method to estimate fall textbook spending?
Answer: No. Students living on a single dorm floor often share class year, academic track, or budget constraints, creating selection bias.
Question 4: When taking a stratified sample of students regarding remaining degree requirements, is it better to stratify by class standing (freshman, sophomore, junior, senior) or randomly assign students into four groups?
Answer: Stratify by class standing. Remaining degree requirements strongly depend on class standing (freshmen need many; seniors need few). Randomly dividing students into 4 groups would not guarantee balanced representation across class standings.
Question 5: Is assigning the first arriving students to laptops and the last arriving students to mobile phones a good experimental design for a reading comprehension study?
Answer: No. Arrival time correlates with conscientiousness, sleep schedules, or preparation time, creating confounding variables. Random assignment must be used instead.
Question 6: Should every experiment feature a placebo group and blinding?
Answer: No. Placebos and blinding apply primarily to human subjects where psychological expectations influence outcomes. Physical or non-human experiments (such as testing material strength or cooking parameters) do not require placebos or blinding.