Unit 1 Part 2: Exploring One-Variable Data and Collecting Data
Exploring One-Variable Data and Collecting Data


Population vs. Sample Inference: Because collecting data from an entire population (a census) is usually impractical, researchers use sample data to make statistical inferences about population parameters.
Activity: The Federalist Papers Authorship Study:
The Federalist Papers comprise essays published anonymously to support the ratification of the U.S. Constitution.


Arbitrary Selection: Selecting words arbitrarily from the passage resulted in word lengths of , yielding a sample mean length of:
* *Random Selection*: Using a random number generator to choose words from numbered words ( through ) yielded word lengths such as with mean length:
* 
* *True Population Parameter*: The true mean word length of the entire passage is:
Section 1F: Introduction to Data Collection
Statistical Inference: The process of using sample statistics to answer investigative questions about population parameters.
Functions of an Investigative Question:
Guides the data collection process.
Guides the choice of data analysis methods.
Indicates the types of conclusions that can be made.
Experiments:
The most reliable method to determine if changes in one variable cause changes in another variable.

Definition - Experiment: An experiment deliberately imposes treatments on experimental units to measure their responses.
Variables in Experiments:
Variables of interest can be either quantitative or categorical.

Definition - Response Variable: Measures an outcome of a statistical study.
Definition - Explanatory Variable: May help explain or predict changes in a response variable.
Terminology Note: Avoid using "independent variable" for explanatory variable and "dependent variable" for response variable, as independent and dependent have distinct mathematical meanings in statistics.
Factors, Levels, Treatments, and Units:

Definition - Factors & Levels: In an experiment, factors are explanatory variables that are manipulated and may cause a change in the response variable. The different values of a factor are called levels.

Definition - Treatment: A specific condition applied to the items or individuals in an experiment. If an experiment has one factor, the levels form the treatments. If there are multiple factors, a treatment is a combination of specific levels of these factors.
Definition - Experimental Unit & Subjects: An experimental unit is the item or individual to which a treatment is assigned. When experimental units are human beings, they are called subjects.
Example 1: Lobster Hatchery Experiment:
Context: To reverse declining lobster populations, lobsters are raised in hatcheries and released into the wild. Larger lobsters have higher survival rates.
Experimental Units: The lobsters.
Factors and Levels:
Amount of protein: levels ( and ).
Lipid-to-carbohydrate (L:CHO) ratio: levels (low, medium, high).
Total Treatments: distinct treatments:
protein / Low L:CHO ratio
protein / Medium L:CHO ratio
protein / High L:CHO ratio
protein / Low L:CHO ratio
protein / Medium L:CHO ratio
protein / High L:CHO ratio
Response Variable: Growth rate measured by carapace (upper shell section) length.
Observational Studies:
Used when conducting an experiment is unethical or impractical (e.g., forcing individuals to smoke to study lung cancer).

Definition - Observational Study: Observes items or individuals and measures variables of interest, but does not impose treatments.
Definition - Retrospective Study: Observational study where units are selected at a point in time and data are gathered from the past.
Definition - Prospective Study: Observational study where units are selected at a point in time and data are gathered both at that time and into the future.

Definition - Survey: An observational study in which data are collected from people by asking them a common set of questions.
Example 2: Coastal Visits and Overall Health:
Context: Study of people across countries. People who never visited the coast were times as likely to report bad health compared to those who visited the coast at least once a week.
Explanatory Variable: Whether a person visits the coast at least once per week (Categorical).
Response Variable: Whether a person claims to have bad health (Categorical).
Classification: Observational study, because researchers did not assign or force participants to visit or not visit the coast.
Extraneous and Confounding Variables:

Definition - Extraneous Variables: Variables other than the explanatory variable that may have an effect on the response variable in a statistical study.

Example Table - Vitamin D & Diabetes Risk:
Explanatory: Vitamin D concentration (Group 1: High, Group 2: Low).
Response: Diabetes status (Group 1: Less likely to have diabetes, Group 2: More likely to have diabetes).
Extraneous variables: Quality of diet (Better vs. Worse), Amount of exercise (More vs. Less), Amount of vitamin supplementation (More vs. Fewer).

Definition - Confounding Variable: An extraneous variable associated with both the explanatory variable and the response variable in a statistical study. Its presence makes it difficult to determine whether changes in the explanatory variable cause changes in the response variable.
Example 3: Confounding in Coastal Health Study:
Possible Confounding Variable: Quality of diet (seafood consumption).
Explanation: People who eat seafood often may visit coastal areas more frequently, and eating seafood regularly may also lead to better overall health. Because diet is associated with both coast visits and health outcomes, it is impossible to determine if coastal visits alone cause better health.
Generalizing Results of Statistical Studies:
Data collected from a random sample allows generalization (inference) to the larger population from which the sample was selected.
If observational or experimental units are not selected at random, inferences can only be made about items or individuals that are similar to those in the study.
Example 4 Contexts:
Context (a): four-year-old volunteer children assigned to fast-paced cartoons, educational cartoons, or art supplies. Executive function was measured. The largest population to generalize results to is all children similar to the volunteers in the experiment (since the sample was not chosen randomly from all children).
Context (b): Random sample of U.S. adults ( and older) surveyed on depression symptoms ( high school diploma or less vs. college degree). Results generalize to all U.S. adults age 18 and older (because random sampling was used).
Components of an Investigative Question:
Identify the variable(s) of interest.
Identify the type of inference, including parameter(s).
Indicate the scope of inference (population generalized to and possibility of cause-and-effect conclusions).
Example 5: "What are plausible values for the average hourly wage (in dollars) of all employed high school students in the United States?"
Variable: Hourly wage in dollars (Quantitative continuous).
Type of Inference: Estimate the average hourly wage of all employed U.S. high school students.
Type of Conclusion: Results can be generalized to the population of all employed U.S. high school students if randomly selected from that population.
Section 1G: Sampling Good and Bad

Definition - Random Sampling: Involves using a chance process to determine which members of a population are chosen for the sample.
Simple Random Samples (SRS):

Definition - Simple Random Sample (SRS): An SRS of size is a sample chosen in such a way that every group of items or individuals in the population has an equal chance to be selected as the sample.
Key Distinction: Giving every individual an equal chance of selection is necessary but not sufficient; every group of size must also be equally likely.

Definition - Sampling Without Replacement: An item or individual can be selected only once.
Definition - Sampling With Replacement: An item or individual can be selected more than once.

Procedure - Choosing an SRS with Technology:
Label: Give each member of the population a distinct integer label from to , where is the total population size.
Randomize: Use a random number generator to obtain different integers from to , ignoring repeated numbers.
Select: Choose the observational units corresponding to the randomly selected integers.
Example 1: Quidditch Team Drug Testing:
Context: Select players out of randomly for drug testing.
Procedure:
Label players alphabetically from to ().
Use a random number generator to pick unique integers from to (ignore numbers outside and repeated integers).
Test the players corresponding to those numbers (e.g., numbers and select Weasley and Lovegood).
Stratified Random Sampling:

Definition - Strata & Stratified Random Sample: Strata are non-overlapping groups of items or individuals in a population sharing characteristics associated with measured variables (homogeneous within). A stratified random sample is selected by taking an SRS from each stratum and combining them into one overall sample.
Effectiveness: Works best when units within each stratum are similar (homogeneous) and large differences exist between strata.
Advantage: Yields estimates with lower variability (more precise) than an SRS of the same total size.
Activity: Sampling Sunflowers:
Grid: field ( total squares), true mean healthy plants .

Comparison: SRS dotplot, stratified by rows dotplot, and stratified by columns dotplot all center roughly around . However, stratified sampling with rows as strata displays the narrowest spread (least variability) because irrigation ditches running along top/bottom created strong row-by-row differences.
Cluster Random Sampling:

Definition - Clusters & Cluster Random Sample: Clusters are groups located near each other. A cluster random sample is chosen by taking an SRS of entire clusters and including every member from selected clusters.
Ideal Structure: Units within a cluster should be diverse (heterogeneous), mirroring the full population on a mini-scale.
Purpose: Primarily practical (saves time and expense when populations are large and widely distributed).
Systematic Random Sampling:

Definition - Systematic Random Sample: Selected from an ordered arrangement of the population by randomly selecting one of the first items and picking every item thereafter.

Caution: If recurring hidden patterns match the interval , the sample will be non-representative.
Example 2: Hot Sauce Factory Pallet Quality Control:
Context: machines (A, B, C, D) produce pallets each. Each pallet holds boxes. Total boxes boxes. Sample target boxes.

(a) Stratified Random Sample: Stratify by machine (4 strata of boxes each). Label boxes to for Machine A, select distinct random integers. Repeat for Machines B, C, and D. Preferred over SRS because weight characteristics are similar within a machine group.
(b) Cluster Random Sample: Let each loading dock aisle ( aisles total) be a cluster of boxes (containing pallets from all machines). Randomly select unique aisles from , and sample all boxes in those aisles.
(c) Systematic Random Sample: Calculate interval . Number boxes to . Pick a random integer from to for the starting box, then select every box thereafter.
Poor Sampling Methods and Bias:

Definition - Bias: The design of a statistical study shows bias if the resulting sample statistic is systematically likely to overestimate or underestimate the population parameter due to flawed data collection.

Definition - Convenience Sample: Consists of members of the population that are easy to reach.

Definition - Voluntary Response Sample: Consists of people who self-select into the sample by responding to a general invitation.
Example 3: Gardener Tomato Yield Estimate:
Context: A gardener counts tomatoes across plants in the southernmost row closest to the house and calculates an overall garden estimate of tomatoes for total plants.
Sampling Method: Convenience sample.
Impact of Bias: Plants closest to the house receive more frequent watering, weeding, and superior sunlight on the south edge. The calculated estimate of tomatoes is an overestimate of the true garden total.
Flaws and Bias Types in Surveys:

Definition - Undercoverage: Occurs when some members of the population are less likely to be chosen or cannot be chosen for a sample (flaw in the sampling frame).

Definition - Nonresponse: Occurs when an individual selected for the sample cannot be contacted or refuses to participate.

Definition - Response Bias: Occurs when responses to a survey question consistently differ from the truth in the same way (caused by leading wording, interviewer characteristics, or lack of anonymity).
Example 4: Walmart Opening Survey:
Poll Question: "Are you against a Walmart opening in our township, which would increase traffic delays, noise, and pollution?"
Bias Explanation: Wording bias. The question uses leading, negative statements about traffic, noise, and pollution. The reported against the opening is a severe overestimate of the true proportion of opposed residents.
Section 1H: Principles of Experimental Design
4 Principles of a Well-Designed Experiment:
Comparison: Compares two or more treatments.
Random Assignment: Assigns treatments using a chance process to balance out extraneous variables.
Replication: Uses enough experimental units in each treatment group.
Direct Control: Keeps extraneous variables constant across all units to prevent confounding and reduce variation.
Control Groups, Placebos, and Blinding:

Definition - Control Group: Provides a baseline for comparing treatment effects. Can receive an inactive treatment, active treatment, or no treatment.
Definition - Placebo: A treatment with no active ingredient that is otherwise identical to other treatments.

Definition - Placebo Effect: Describes the fact that some subjects respond favorably to any treatment, even an inactive one.

Definition - Double-Blind: Neither the subjects nor those interacting with them and measuring the response variable know which treatment a subject is receiving.
Definition - Single-Blind: Either the subjects OR the people interacting with them and measuring response variables do not know which treatment a subject receives.
Example 1: Rwanda Literacy Program:
Context: Sectors assigned to Teacher Training (TT), Literacy Boost plus Teacher Training (LB+TT), or Control (neither).
Purpose of Control Group: Provides a baseline to evaluate literacy growth. Without it, researchers could not isolate treatment effects from external environmental changes.
Example 2: AzaSite Eyedrop Trial:
Context: AzaSite antibiotic drops compared against placebo drops for bacterial conjunctivitis.
Double-Blind Meaning: Patients and evaluating doctors both were unaware of drop assignments.
Importance: Prevents patient placebo expectations and eliminates subjective evaluation bias by examining physicians.
Random Assignment and Experimental Designs:

Definition - Random Assignment: Assigns experimental units to treatments using a chance process.

Definition - Completely Randomized Design: Experimental units are assigned to treatments completely at random.

Paper Slips Method vs. Coin Toss: Equal-sized treatment groups should be created using equal paper slips drawn without replacement from a container rather than coin flips, to avoid forcing unequal assignments.
Example 3: Rwanda Literacy Assignment:
Context: sectors assigned to TT, TT+LB, and Control ( sectors per group).
Assignment Method: Write "TT" on paper slips, "TT+LB" on slips, and "Control" on slips. Shuffle in a container, draw slip without replacement for each sector.
Purpose: Creates roughly equivalent baseline treatment groups to justify cause-and-effect conclusions.
Replication and Direct Control:

Definition - Direct Control: Keeping values of extraneous variables identical for all experimental units.

Definition - Replication: Assigning multiple experimental units to each treatment group.
Purposes of Direct Control: Prevents confounding and reduces response variable variation.
Example 4: Trikafta Cystic Fibrosis Trial:
Context: CF patients assigned to Trikafta daily or placebo for weeks. All participants took doses twice daily with fat-containing food and maintained existing CF therapies.
Directly Controlled Variables: Medication dosage (), duration ( weeks), twice-daily administration with fat-containing food.
Benefit: Reduces variation in lung function changes, making real treatment effects easier to isolate.
Randomized Block Designs:

Definition - Block & Blocking Variable: A block is a group of experimental units known before the experiment to be similar in a way expected to affect response. A variable used to form blocks is a blocking variable.

Definition - Randomized Block Design: Random assignment of units to treatments is carried out separately within each block.

Example 5: School Commute Route Experiment:
Context: Compare driving routes (A, B, C) across test days.
Blocking Variable: Time of day (Morning vs. Afternoon).
Design: Form two blocks ( mornings and afternoons). In the morning block, randomly assign mornings to Route A, to B, to C. Repeat assignment for the afternoon block.
Justification: Accounts for morning rush-hour traffic variability.
Matched Pairs Designs:

Definition - Matched Pairs Design: A randomized block design for comparing two treatments using blocks of size . Either two similar experimental units are paired and randomly assigned treatments, or each unit receives both treatments in a random order.
Example 6: Swim Drag Suits Experiment:
Context: swimmers testing regular vs. drag suits in a freestyle sprint.
Matched Pairs Setup: Each swimmer acts as their own block/pair, wearing both suits in a randomized order.
Procedure: Number swimmers to . Use a random number generator to pick unique integers. Those swimmers wear the drag suit first and regular suit second; the remaining wear the regular suit first and drag suit second. Record and compare times.
Section 1I: Inference for Experiments and Data Ethics
Statistical Significance:

Caffeine Experiment Data: volunteer students ( caffeine, no caffeine).
Caffeine group changes:
No caffeine group changes:
Observed difference in means:

Definition - Statistically Significant: When an observed difference in responses between groups in an experiment is so large that it is unlikely to be explained by chance variation in random assignment alone.

Definition - Randomization Distribution: The distribution of a statistic generated by repeatedly reassigning response values to treatment groups assuming no treatment effect.

Caffeine Simulation Result: In simulation trials, a difference of or greater occurred times (). Because , the difference is not statistically significant. It can plausibly be explained by chance variation in random assignment.
Example 7: Botox for Chronic Low Back Pain:
Context: patients ( Botox, saline placebo).

Results Table:
Botox: functional improvement, no improvement ().
Saline: functional improvement, no improvement ().
Difference in proportions: ( percentage points).

Simulation Analysis: In trials, a difference of or greater occurred times (). Because , the result is statistically significant.
The Scope of Inference:


Summary Decision Rules:
Random Selection of Units: Justifies inference about the broader population.
Random Assignment to Treatments: Justifies inference about cause and effect (if results are statistically significant).
Example 8: Botox Scope of Inference:
(a) Cause-and-Effect: Reasonable, because subjects were randomly assigned to treatments and the observed improvement difference was statistically significant.
(b) Population Generalization: Applies only to year old adults with chronic low back pain similar to study participants, because subjects were volunteers and not randomly selected from the general population.
Basic Data Ethics:

Institutional Review Board (IRB): All planned studies must be reviewed in advance by an IRB to protect participant safety and well-being.
Informed Consent: All participants must give informed consent prior to data collection.
Confidentiality: Individual participant data must be kept confidential; only aggregate group summaries may be released publicly.