Sampling, Bias, and Reproducibility Study Notes

Administrative and Course Policy Overview

  • Attendance and absence policy considerations:

    • Requirements regarding the number of sessions that can be excused without losing attendance points.

    • Protocols for submitting make-up tasks for missed sessions.

    • Notification requirements for excusing 2 days or submitting 2 make-up tasks.

Learning Objectives

  • Fundamental concepts covered in statistical sampling and analysis:

    • Principles of sampling methodologies and the structural role of randomness.

    • Quantitative comparison between the sample mean (xˉ\bar{x}) and the population mean (μ\mu).

    • Identification and categorization of systemic bias in sampling design.

    • Implementation of reproducible workflows using pseudo-random seeds.

    • Standards for transparently reporting underlying statistical assumptions.

Case Study: Analyzing News Claims and Population Inference

  • Headline Case Study: A poll reported by CBS News (updated August 20, 2025 at 10:38 AM EDT / CBS/AP) presented the headline: "Only 54% of U.S. adults say they drink alcohol, a record low. A new poll shows what's behind the decline."

CBS News article headline regarding U.S. adult alcohol consumption
  • Methodological Discussion: Critical evaluation of generalizations such as "U.S. Adults say…":

    • Extrapolating survey findings to the entire demographic of all U.S. adults requires evaluating whether the survey sample accurately reflects the population.

    • The phrasing "say they drink alcohol" introduces self-reporting considerations, where social desirability or memory recall can introduce measurement bias.

    • To draw valid conclusions about an overall population, the sample must be selected using a methodology that guarantees representative demographic coverage.

Population vs. Sample Definitions and Examples

  • Population: The complete, overarching group of individuals, items, or units about which information is sought.

  • Sample: A smaller, finite subset selected from the population that is directly observed, measured, and analyzed.

  • Concrete Examples:

    • Netflix Shows:

    • Population: Every show and movie available in the full Netflix catalog.

    • Sample: The specific three shows a student re-watches over the course of a semester.

    • Cat Videos:

    • Population: The total collection of all cat videos uploaded to YouTube.

    • Sample: The five specific cat videos watched at 2:00 AM instead of studying.

    • Coffee Cups:

    • Population: The total inventory of paper coffee cups in a store's supply.

    • Sample: The subset of black and brown paper coffee cups stacked on top of an espresso machine.

    • French Fries:

    • Population: The entire batch of French fries cooked in a restaurant fryer.

    • Sample: A single batch or serving collected on parchment paper for quality inspection or consumption.

Rationale for Sampling

  • Population Constraints: Full population enumeration (taking a census) is rarely feasible due to prohibitive costs, temporal constraints, or sheer physical scale.

  • Objective: Collect data from a representative sample to make statistical inferences about unobserved population parameters.

  • Comparative Example: Instead of interviewing a full population of 10,00010{,}000 university graduates, researchers collect data from a sample of 100100 graduated students.

Graduation ceremony representing a population of graduates

Simple Random Sampling (SRS)

  • Definition: A probability sampling technique wherein every single individual in the population possesses an equal probability of selection.

  • Conceptual Model: Equivalent to drawing names thoroughly mixed inside a hat.

  • Core Properties:

    • Free from researcher favoritism or systematic selection patterns.

    • Forms the mathematical foundation for classical inferential statistics and hypothesis testing.

The Role of Randomness in Statistics

  • Operational Definition: Introducing deliberate unpredictability into the selection process.

  • Key Functions:

    • Protects the selection process against conscious and unconscious human bias.

    • Ensures equitable probability of inclusion across all population units.

    • Provides the mathematical foundation necessary to apply probability theory to evaluate estimation errors and support scientific conclusions.

Computational Implementation of SRS using NumPy

  • Simple Random Sampling can be implemented in Python using the numpy library:

import numpy as np

# Define the population array
fruits = np.array(["apple", "banana", "cherry", "grape", "blueberry"])

# Initialize the pseudo-random number generator
rng = np.random.default_rng()

# Select a sample of size 3 without replacement
sample = rng.choice(fruits, size=3, replace=False)
print(sample)
# Output example: array(['blueberry', 'cherry', 'banana'], dtype='<U9')
  • Step-by-Step Code Analysis:

    1. import numpy as np: Loads the NumPy library for numerical operations.

    2. fruits = np.array(...): Defines a population array containing 5 distinct fruit elements.

    3. rng = np.random.default_rng(): Instantiates NumPy's default BitGenerator/Random Number Generator (PCG64).

    4. sample = rng.choice(fruits, size=3, replace=False): Calls the choice method to randomly extract n=3n = 3 elements without replacement (replace=False), ensuring no element can be drawn twice in a single sample.

    5. Output: Returns a array containing the randomly selected elements.

Population Mean vs. Sample Mean

  • Population Mean (μ\mu): The exact mean value of a variable computed across all NN units in the full population.

μ=1N∑i=1Nxi\mu = \frac{1}{N} \sum_{i=1}^{N} x_i

  • Sample Mean (xˉ\bar{x}): The arithmetic mean computed across nn observations contained within a specific sample.

xˉ=1n∑i=1nxi\bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i

  • Estimation Goal: Use the observed sample statistic (xˉ\bar{x}) to estimate the true, unobserved population parameter (μ\mu).

  • Sampling Variation: The inherent variability that causes sample means (xˉ\bar{x}) to differ from sample to sample when drawn from the same underlying population.

Estimating the Mean and Sample Size Effects

  • Numerical Demonstration:

    • Consider a population of N=100N = 100 exam scores with a true population mean of μ=75\mu = 75.

    • Taking independent samples of size n=10n = 10 yields varying estimates: xˉ1=73\bar{x}_1 = 73, xˉ2=77\bar{x}_2 = 77, and xˉ3=78\bar{x}_3 = 78.

    • Each distinct sample yields a slightly different sample mean due to natural sampling variability.

  • Effect of Sample Size on Estimation Precision:

    • Small Samples: Display greater spread, resulting in higher variance across estimates.

    • Large Samples: Point estimates cluster more tightly around the true population mean (μ\mu), increasing accuracy.

Target bullseye diagrams contrasting small spread-out sample estimates with large clustered sample estimates

Understanding Bias in Sampling and Measurement

  • Scientific Definition of Bias: A systematic flaw or tilt in design, collection, or measurement that causes results to be consistently incorrect in a specific direction.

  • Sampling Bias: Structural defects in sample selection that make the sample unrepresentative of the population.

  • Key Distinction from Random Variation:

    • Random Variation: Unbiased, scatter-based fluctuations that average out toward zero with repeated trials.

    • Bias: Directional distortion that persists regardless of repetition and does not cancel out by taking larger samples.

    • Impact: A biased sampling process yields systematically misleading conclusions.

  • Real-World Metaphors:

    • A bathroom scale that is miscalibrated and consistently measures 5 pounds5\text{ pounds} too heavy.

    • An instructor who calls exclusively on students seated in the front row of a lecture hall.

Specific Types of Sampling Bias

  • Convenience Sample: Selecting individuals who are most accessible or easiest to contact (e.g., surveying only students present in the library, while ignoring those in class, at home, or working).

  • Voluntary Response Bias: Allowing individuals to self-select into the survey group. This systematically overrepresents individuals with strong or extreme opinions.

  • Nonresponse Bias: Occurs when a significant portion of selected participants fail or decline to respond. If non-respondents systematically differ from respondents, the resulting sample becomes skewed.

  • Consequence: Each form of bias impairs the ability of the sample mean (xˉ\bar{x}) to reliably estimate the population mean (μ\mu).

Core Advantages of Random Sampling

  • Eliminates systematic selection bias by offering every individual equal selection probability.

  • Produces structurally representative samples that mirror population characteristics.

  • Validates using sample statistics to estimate population parameters.

  • Serves as an indispensable foundation for trustworthy empirical research and data science.

Reproducibility and Pseudo-Random Seeds

  • Concept of Reproducibility: Replicating identical experimental or computational results across separate executions of an analysis.

  • Problem: Standard random sampling yields unpredictable, varying outputs each time code is executed.

  • Solution: Setting a pseudo-random seed to fix the algorithm's starting state and generate an identical pseudo-random sequence.

  • Python/NumPy Implementation:

import numpy as np

# Define population array
fruits = np.array(["apple", "banana", "cherry", "grape", "blueberry"])

# Instantiate generator with a fixed seed for exact reproducibility
rng = np.random.default_rng(seed=42)

# Execute sample selection
sample = rng.choice(fruits, size=3, replace=False)
print(sample)
# Output array is locked and deterministic across execution runs
  • Application Utility: Essential for computational research, instructional demonstrations, peer verification, and collaborative code debugging.

Cultural Context of 'seed=42'

  • Cultural Origin: Derived from Douglas Adams' science fiction novel The Hitchhiker's Guide to the Galaxy.

  • Story Context: A giant supercomputer spends millions of years calculating the "Ultimate Question of Life, the Universe, and Everything," eventually concluding that the answer is 42.

  • Function in Computer Science: Serves as a standard, humorous convention for pseudo-random seed initialization.

  • Equivalence: Using seed=42 is functionally identical to using seed=123 or seed=2025; any specified integer guarantees deterministic reproducibility.

Reporting Statistical Assumptions and Parameters

  • Maintaining transparency and credibility in data analysis requires disclosing the following structural details:

    • Target Population: Explicit description of the entire population under investigation.

    • Sampling Strategy: Exact method utilized to gather observations (e.g., Simple Random Sampling vs. Convenience Sampling).

    • Sample Size (nn): Total count of units measured within the sample.

    • Reproducibility Details: Exact seed values specified in computational algorithms.

Summary of Key Takeaways

  • The sample mean (xˉ\bar{x}) provides an empirical point estimate for the true population mean (μ\mu).

  • A sample reflects its target population effectively when derived through random, unbiased selection methods.

  • Bias systematically corrupts accuracy and cannot be remedied by increasing sample size alone.

  • Pseudo-random seeds lock stochastic algorithms to ensure exact computational reproducibility.

  • Transparently reporting sampling methodologies and assumptions is required to establish statistical credibility.