Unit 1 Part 2: Exploring One-Variable Data and Collecting Data

Exploring One-Variable Data and Collecting Data

Calvin and Hobbes ComicUnit 1 Header
  • Population vs. Sample Inference: Because collecting data from an entire population (a census) is usually impractical, researchers use sample data to make statistical inferences about population parameters.

  • Activity: The Federalist Papers Authorship Study:

    • The Federalist Papers comprise 8585 essays published anonymously to support the ratification of the U.S. Constitution.


    Federalist Papers Activity Header

    Federalist Papers Passage
    • Arbitrary Selection: Selecting 55 words arbitrarily from the passage resulted in word lengths of 9,5,6,3,39, 5, 6, 3, 3, yielding a sample mean length of:

xˉ=9+5+6+3+35=5.2\bar{x} = \frac{9 + 5 + 6 + 3 + 3}{5} = 5.2

*   *Random Selection*: Using a random number generator to choose 55 words from numbered words (11 through 129129) yielded word lengths such as 4,3,6,3,104, 3, 6, 3, 10 with mean length:

xˉ=5.2\bar{x} = 5.2

*   ![Federalist Papers Annotated Random Selection](https://assets.knowt.com/pdf-flow-prod/011e89d1-9ca2-4f6b-8d38-e285f8f27cee-figures/83.png)
*   *True Population Parameter*: The true mean word length of the entire passage is:

μ=4.9\mu = 4.9

Section 1F: Introduction to Data Collection

  • Statistical Inference: The process of using sample statistics to answer investigative questions about population parameters.

  • Functions of an Investigative Question:

    1. Guides the data collection process.

    2. Guides the choice of data analysis methods.

    3. Indicates the types of conclusions that can be made.

  • Experiments:

    • The most reliable method to determine if changes in one variable cause changes in another variable.


    Experiment Definition
    • Definition - Experiment: An experiment deliberately imposes treatments on experimental units to measure their responses.

  • Variables in Experiments:

    • Variables of interest can be either quantitative or categorical.


    Response and Explanatory Variable Definition
    • Definition - Response Variable: Measures an outcome of a statistical study.

    • Definition - Explanatory Variable: May help explain or predict changes in a response variable.

    • Terminology Note: Avoid using "independent variable" for explanatory variable and "dependent variable" for response variable, as independent and dependent have distinct mathematical meanings in statistics.

  • Factors, Levels, Treatments, and Units:


    Factors and Levels Definition
    • Definition - Factors & Levels: In an experiment, factors are explanatory variables that are manipulated and may cause a change in the response variable. The different values of a factor are called levels.


    Treatment and Experimental Unit Definition
    • Definition - Treatment: A specific condition applied to the items or individuals in an experiment. If an experiment has one factor, the levels form the treatments. If there are multiple factors, a treatment is a combination of specific levels of these factors.

    • Definition - Experimental Unit & Subjects: An experimental unit is the item or individual to which a treatment is assigned. When experimental units are human beings, they are called subjects.

  • Example 1: Lobster Hatchery Experiment:

    • Context: To reverse declining lobster populations, lobsters are raised in hatcheries and released into the wild. Larger lobsters have higher survival rates.

    • Experimental Units: The lobsters.

    • Factors and Levels:

      1. Amount of protein: 22 levels (400 g/kg400\,g/kg and 500 g/kg500\,g/kg).

      2. Lipid-to-carbohydrate (L:CHO) ratio: 33 levels (low, medium, high).

    • Total Treatments: 2×3=62 \times 3 = 6 distinct treatments:

      1. 400 g/kg400\,g/kg protein / Low L:CHO ratio

      2. 400 g/kg400\,g/kg protein / Medium L:CHO ratio

      3. 400 g/kg400\,g/kg protein / High L:CHO ratio

      4. 500 g/kg500\,g/kg protein / Low L:CHO ratio

      5. 500 g/kg500\,g/kg protein / Medium L:CHO ratio

      6. 500 g/kg500\,g/kg protein / High L:CHO ratio

    • Response Variable: Growth rate measured by carapace (upper shell section) length.

  • Observational Studies:

    • Used when conducting an experiment is unethical or impractical (e.g., forcing individuals to smoke to study lung cancer).


    Observational Study Definition
    • Definition - Observational Study: Observes items or individuals and measures variables of interest, but does not impose treatments.

    • Definition - Retrospective Study: Observational study where units are selected at a point in time and data are gathered from the past.

    • Definition - Prospective Study: Observational study where units are selected at a point in time and data are gathered both at that time and into the future.


    Survey Definition
    • Definition - Survey: An observational study in which data are collected from people by asking them a common set of questions.

  • Example 2: Coastal Visits and Overall Health:

    • Context: Study of 14,70214,702 people across 1515 countries. People who never visited the coast were 2.132.13 times as likely to report bad health compared to those who visited the coast at least once a week.

    • Explanatory Variable: Whether a person visits the coast at least once per week (Categorical).

    • Response Variable: Whether a person claims to have bad health (Categorical).

    • Classification: Observational study, because researchers did not assign or force participants to visit or not visit the coast.

  • Extraneous and Confounding Variables:


    Extraneous Variables Definition
    • Definition - Extraneous Variables: Variables other than the explanatory variable that may have an effect on the response variable in a statistical study.


    Vitamin D and Diabetes Extraneous Variables Table
    • Example Table - Vitamin D & Diabetes Risk:

      • Explanatory: Vitamin D concentration (Group 1: High, Group 2: Low).

      • Response: Diabetes status (Group 1: Less likely to have diabetes, Group 2: More likely to have diabetes).

      • Extraneous variables: Quality of diet (Better vs. Worse), Amount of exercise (More vs. Less), Amount of vitamin supplementation (More vs. Fewer).


    Confounding Variable Definition
    • Definition - Confounding Variable: An extraneous variable associated with both the explanatory variable and the response variable in a statistical study. Its presence makes it difficult to determine whether changes in the explanatory variable cause changes in the response variable.

  • Example 3: Confounding in Coastal Health Study:

    • Possible Confounding Variable: Quality of diet (seafood consumption).

    • Explanation: People who eat seafood often may visit coastal areas more frequently, and eating seafood regularly may also lead to better overall health. Because diet is associated with both coast visits and health outcomes, it is impossible to determine if coastal visits alone cause better health.

  • Generalizing Results of Statistical Studies:

    • Data collected from a random sample allows generalization (inference) to the larger population from which the sample was selected.

    • If observational or experimental units are not selected at random, inferences can only be made about items or individuals that are similar to those in the study.

    • Example 4 Contexts:

      • Context (a): 6060 four-year-old volunteer children assigned to fast-paced cartoons, educational cartoons, or art supplies. Executive function was measured. The largest population to generalize results to is all children similar to the volunteers in the experiment (since the sample was not chosen randomly from all children).

      • Context (b): Random sample of U.S. adults (1818 and older) surveyed on depression symptoms (21.5%21.5\% high school diploma or less vs. 12.4%12.4\% college degree). Results generalize to all U.S. adults age 18 and older (because random sampling was used).

  • Components of an Investigative Question:

    1. Identify the variable(s) of interest.

    2. Identify the type of inference, including parameter(s).

    3. Indicate the scope of inference (population generalized to and possibility of cause-and-effect conclusions).

    • Example 5: "What are plausible values for the average hourly wage (in dollars) of all employed high school students in the United States?"

      1. Variable: Hourly wage in dollars (Quantitative continuous).

      2. Type of Inference: Estimate the average hourly wage of all employed U.S. high school students.

      3. Type of Conclusion: Results can be generalized to the population of all employed U.S. high school students if randomly selected from that population.

Section 1G: Sampling Good and Bad


Random Sampling Definition
  • Definition - Random Sampling: Involves using a chance process to determine which members of a population are chosen for the sample.

  • Simple Random Samples (SRS):


    Simple Random Sample Definition
    • Definition - Simple Random Sample (SRS): An SRS of size nn is a sample chosen in such a way that every group of nn items or individuals in the population has an equal chance to be selected as the sample.

    • Key Distinction: Giving every individual an equal chance of selection is necessary but not sufficient; every group of size nn must also be equally likely.


    Sampling With and Without Replacement Definition
    • Definition - Sampling Without Replacement: An item or individual can be selected only once.

    • Definition - Sampling With Replacement: An item or individual can be selected more than once.


    How to Choose an SRS with Technology
    • Procedure - Choosing an SRS with Technology:

      1. Label: Give each member of the population a distinct integer label from 11 to NN, where NN is the total population size.

      2. Randomize: Use a random number generator to obtain nn different integers from 11 to NN, ignoring repeated numbers.

      3. Select: Choose the observational units corresponding to the randomly selected integers.

  • Example 1: Quidditch Team Drug Testing:

    • Context: Select 22 players out of 1212 randomly for drug testing.

    • Procedure:

      1. Label players alphabetically from 11 to 1212 (1=Chang,2=Delacour,…,11=Weasley,12=Wood1=\text{Chang}, 2=\text{Delacour}, \dots, 11=\text{Weasley}, 12=\text{Wood}).

      2. Use a random number generator to pick 22 unique integers from 11 to 1212 (ignore numbers outside 1–121\text{--}12 and repeated integers).

      3. Test the players corresponding to those numbers (e.g., numbers 1111 and 88 select Weasley and Lovegood).

  • Stratified Random Sampling:


    Stratified Random Sample Definition
    • Definition - Strata & Stratified Random Sample: Strata are non-overlapping groups of items or individuals in a population sharing characteristics associated with measured variables (homogeneous within). A stratified random sample is selected by taking an SRS from each stratum and combining them into one overall sample.

    • Effectiveness: Works best when units within each stratum are similar (homogeneous) and large differences exist between strata.

    • Advantage: Yields estimates with lower variability (more precise) than an SRS of the same total size.

  • Activity: Sampling Sunflowers:

    • Grid: 10×1010 \times 10 field (100100 total squares), true mean healthy plants μ=102.46\mu = 102.46.


    Sampling Sunflowers Dotplots Comparison
    • Comparison: SRS dotplot, stratified by rows dotplot, and stratified by columns dotplot all center roughly around μ=102.46\mu = 102.46. However, stratified sampling with rows as strata displays the narrowest spread (least variability) because irrigation ditches running along top/bottom created strong row-by-row differences.

  • Cluster Random Sampling:


    Cluster Random Sample Definition
    • Definition - Clusters & Cluster Random Sample: Clusters are groups located near each other. A cluster random sample is chosen by taking an SRS of entire clusters and including every member from selected clusters.

    • Ideal Structure: Units within a cluster should be diverse (heterogeneous), mirroring the full population on a mini-scale.

    • Purpose: Primarily practical (saves time and expense when populations are large and widely distributed).

  • Systematic Random Sampling:


    Systematic Random Sample Definition
    • Definition - Systematic Random Sample: Selected from an ordered arrangement of the population by randomly selecting one of the first kk items and picking every kthk^{\text{th}} item thereafter.


    Systematic Sampling Voters Diagram
    • Caution: If recurring hidden patterns match the interval kk, the sample will be non-representative.

  • Example 2: Hot Sauce Factory Pallet Quality Control:

    • Context: 44 machines (A, B, C, D) produce 1010 pallets each. Each pallet holds 2525 boxes. Total boxes =4×10×25=1000= 4 \times 10 \times 25 = 1000 boxes. Sample target n=200n = 200 boxes.


    Factory Loading Dock Pallets Layout
    • (a) Stratified Random Sample: Stratify by machine (4 strata of 250250 boxes each). Label boxes 11 to 250250 for Machine A, select 5050 distinct random integers. Repeat for Machines B, C, and D. Preferred over SRS because weight characteristics are similar within a machine group.

    • (b) Cluster Random Sample: Let each loading dock aisle (1010 aisles total) be a cluster of 100100 boxes (containing pallets from all machines). Randomly select 22 unique aisles from 1–101\text{--}10, and sample all 200200 boxes in those 22 aisles.

    • (c) Systematic Random Sample: Calculate interval k=1000200=5k = \frac{1000}{200} = 5. Number boxes 11 to 10001000. Pick a random integer from 11 to 55 for the starting box, then select every 5th5^{\text{th}} box thereafter.

  • Poor Sampling Methods and Bias:


    Bias Definition
    • Definition - Bias: The design of a statistical study shows bias if the resulting sample statistic is systematically likely to overestimate or underestimate the population parameter due to flawed data collection.


    Convenience Sample Definition
    • Definition - Convenience Sample: Consists of members of the population that are easy to reach.


    Voluntary Response Sample Definition
    • Definition - Voluntary Response Sample: Consists of people who self-select into the sample by responding to a general invitation.

  • Example 3: Gardener Tomato Yield Estimate:

    • Context: A gardener counts 5353 tomatoes across 55 plants in the southernmost row closest to the house and calculates an overall garden estimate of 53×4=21253 \times 4 = 212 tomatoes for 2020 total plants.

    • Sampling Method: Convenience sample.

    • Impact of Bias: Plants closest to the house receive more frequent watering, weeding, and superior sunlight on the south edge. The calculated estimate of 212212 tomatoes is an overestimate of the true garden total.

  • Flaws and Bias Types in Surveys:


    Undercoverage Definition
    • Definition - Undercoverage: Occurs when some members of the population are less likely to be chosen or cannot be chosen for a sample (flaw in the sampling frame).


    Nonresponse Definition
    • Definition - Nonresponse: Occurs when an individual selected for the sample cannot be contacted or refuses to participate.


    Response Bias Definition
    • Definition - Response Bias: Occurs when responses to a survey question consistently differ from the truth in the same way (caused by leading wording, interviewer characteristics, or lack of anonymity).

  • Example 4: Walmart Opening Survey:

    • Poll Question: "Are you against a Walmart opening in our township, which would increase traffic delays, noise, and pollution?"

    • Bias Explanation: Wording bias. The question uses leading, negative statements about traffic, noise, and pollution. The 92%92\% reported against the opening is a severe overestimate of the true proportion of opposed residents.

Section 1H: Principles of Experimental Design

  • 4 Principles of a Well-Designed Experiment:

    1. Comparison: Compares two or more treatments.

    2. Random Assignment: Assigns treatments using a chance process to balance out extraneous variables.

    3. Replication: Uses enough experimental units in each treatment group.

    4. Direct Control: Keeps extraneous variables constant across all units to prevent confounding and reduce variation.

  • Control Groups, Placebos, and Blinding:


    Control Group and Placebo Definition
    • Definition - Control Group: Provides a baseline for comparing treatment effects. Can receive an inactive treatment, active treatment, or no treatment.

    • Definition - Placebo: A treatment with no active ingredient that is otherwise identical to other treatments.


    Placebo Effect Definition
    • Definition - Placebo Effect: Describes the fact that some subjects respond favorably to any treatment, even an inactive one.


    Double Blind and Single Blind Definition
    • Definition - Double-Blind: Neither the subjects nor those interacting with them and measuring the response variable know which treatment a subject is receiving.

    • Definition - Single-Blind: Either the subjects OR the people interacting with them and measuring response variables do not know which treatment a subject receives.

  • Example 1: Rwanda Literacy Program:

    • Context: Sectors assigned to Teacher Training (TT), Literacy Boost plus Teacher Training (LB+TT), or Control (neither).

    • Purpose of Control Group: Provides a baseline to evaluate literacy growth. Without it, researchers could not isolate treatment effects from external environmental changes.

  • Example 2: AzaSite Eyedrop Trial:

    • Context: AzaSite antibiotic drops compared against placebo drops for bacterial conjunctivitis.

    • Double-Blind Meaning: Patients and evaluating doctors both were unaware of drop assignments.

    • Importance: Prevents patient placebo expectations and eliminates subjective evaluation bias by examining physicians.

  • Random Assignment and Experimental Designs:


    Random Assignment Definition
    • Definition - Random Assignment: Assigns experimental units to treatments using a chance process.


    Completely Randomized Design Definition
    • Definition - Completely Randomized Design: Experimental units are assigned to treatments completely at random.


    Completely Randomized Design Caffeine Experiment Diagram
    • Paper Slips Method vs. Coin Toss: Equal-sized treatment groups should be created using equal paper slips drawn without replacement from a container rather than coin flips, to avoid forcing unequal assignments.

  • Example 3: Rwanda Literacy Assignment:

    • Context: 2121 sectors assigned to TT, TT+LB, and Control (77 sectors per group).

    • Assignment Method: Write "TT" on 77 paper slips, "TT+LB" on 77 slips, and "Control" on 77 slips. Shuffle in a container, draw 11 slip without replacement for each sector.

    • Purpose: Creates roughly equivalent baseline treatment groups to justify cause-and-effect conclusions.

  • Replication and Direct Control:


    Direct Control Definition
    • Definition - Direct Control: Keeping values of extraneous variables identical for all experimental units.


    Replication Definition
    • Definition - Replication: Assigning multiple experimental units to each treatment group.

    • Purposes of Direct Control: Prevents confounding and reduces response variable variation.

  • Example 4: Trikafta Cystic Fibrosis Trial:

    • Context: 403403 CF patients assigned to 375 mg375\,mg Trikafta daily or 375 mg375\,mg placebo for 2424 weeks. All participants took doses twice daily with fat-containing food and maintained existing CF therapies.

    • Directly Controlled Variables: Medication dosage (375 mg375\,mg), duration (2424 weeks), twice-daily administration with fat-containing food.

    • Benefit: Reduces variation in lung function changes, making real treatment effects easier to isolate.

  • Randomized Block Designs:


    Block and Blocking Variable Definition
    • Definition - Block & Blocking Variable: A block is a group of experimental units known before the experiment to be similar in a way expected to affect response. A variable used to form blocks is a blocking variable.


    Randomized Block Design Definition
    • Definition - Randomized Block Design: Random assignment of units to treatments is carried out separately within each block.


    Randomized Block Design Field Fertility Diagram
  • Example 5: School Commute Route Experiment:

    • Context: Compare 33 driving routes (A, B, C) across 1515 test days.

    • Blocking Variable: Time of day (Morning vs. Afternoon).

    • Design: Form two blocks (1515 mornings and 1515 afternoons). In the morning block, randomly assign 55 mornings to Route A, 55 to B, 55 to C. Repeat assignment for the afternoon block.

    • Justification: Accounts for morning rush-hour traffic variability.

  • Matched Pairs Designs:


    Matched Pairs Design Definition
    • Definition - Matched Pairs Design: A randomized block design for comparing two treatments using blocks of size 22. Either two similar experimental units are paired and randomly assigned treatments, or each unit receives both treatments in a random order.

  • Example 6: Swim Drag Suits Experiment:

    • Context: 2424 swimmers testing regular vs. drag suits in a 50-meter50\text{-meter} freestyle sprint.

    • Matched Pairs Setup: Each swimmer acts as their own block/pair, wearing both suits in a randomized order.

    • Procedure: Number swimmers 11 to 2424. Use a random number generator to pick 1212 unique integers. Those 1212 swimmers wear the drag suit first and regular suit second; the remaining 1212 wear the regular suit first and drag suit second. Record and compare times.

Section 1I: Inference for Experiments and Data Ethics

  • Statistical Significance:


    Pulse Rate Change Data and Dotplot
    • Caffeine Experiment Data: 2020 volunteer students (1010 caffeine, 1010 no caffeine).

      • Caffeine group changes: 8,3,5,1,4,0,6,1,4,0→Mean=3.28, 3, 5, 1, 4, 0, 6, 1, 4, 0 \rightarrow \text{Mean} = 3.2

      • No caffeine group changes: 3,−2,4,−1,5,5,1,2,−1,4→Mean=2.03, -2, 4, -1, 5, 5, 1, 2, -1, 4 \rightarrow \text{Mean} = 2.0

      • Observed difference in means: 3.2−2.0=1.23.2 - 2.0 = 1.2


    Statistically Significant Definition
    • Definition - Statistically Significant: When an observed difference in responses between groups in an experiment is so large that it is unlikely to be explained by chance variation in random assignment alone.


    Randomization Distribution Definition
    • Definition - Randomization Distribution: The distribution of a statistic generated by repeatedly reassigning response values to treatment groups assuming no treatment effect.


    Caffeine Simulation Randomization Distribution Dotplot
    • Caffeine Simulation Result: In 100100 simulation trials, a difference of 1.21.2 or greater occurred 1919 times (19%19\%). Because 19%≥5%19\% \ge 5\%, the difference is not statistically significant. It can plausibly be explained by chance variation in random assignment.

  • Example 7: Botox for Chronic Low Back Pain:

    • Context: 5050 patients (2525 Botox, 2525 saline placebo).


    Botox Experiment Results Contingency Table
    • Results Table:

      • Botox: 1717 functional improvement, 88 no improvement (17/25=0.6817/25 = 0.68).

      • Saline: 33 functional improvement, 2222 no improvement (3/25=0.123/25 = 0.12).

      • Difference in proportions: 0.68−0.12=0.560.68 - 0.12 = 0.56 (5656 percentage points).


    Botox Simulation Randomization Distribution Dotplot
    • Simulation Analysis: In 100100 trials, a difference of 0.560.56 or greater occurred 00 times (0%0\%). Because 0%<5%0\% < 5\%, the result is statistically significant.

  • The Scope of Inference:


    Scope of Inference Matrix

    The Scope of Inference Key Principles
    • Summary Decision Rules:

      • Random Selection of Units: Justifies inference about the broader population.

      • Random Assignment to Treatments: Justifies inference about cause and effect (if results are statistically significant).

  • Example 8: Botox Scope of Inference:

    • (a) Cause-and-Effect: Reasonable, because subjects were randomly assigned to treatments and the observed improvement difference was statistically significant.

    • (b) Population Generalization: Applies only to 18–5518\text{--}55 year old adults with chronic low back pain similar to study participants, because subjects were volunteers and not randomly selected from the general population.

  • Basic Data Ethics:


    Basic Data Ethics Summary
    • Institutional Review Board (IRB): All planned studies must be reviewed in advance by an IRB to protect participant safety and well-being.

    • Informed Consent: All participants must give informed consent prior to data collection.

    • Confidentiality: Individual participant data must be kept confidential; only aggregate group summaries may be released publicly.