Sampling, Bias, and Reproducibility Study Notes
Administrative and Course Policy Overview
Attendance and absence policy considerations:
Requirements regarding the number of sessions that can be excused without losing attendance points.
Protocols for submitting make-up tasks for missed sessions.
Notification requirements for excusing 2 days or submitting 2 make-up tasks.
Learning Objectives
Fundamental concepts covered in statistical sampling and analysis:
Principles of sampling methodologies and the structural role of randomness.
Quantitative comparison between the sample mean () and the population mean ().
Identification and categorization of systemic bias in sampling design.
Implementation of reproducible workflows using pseudo-random seeds.
Standards for transparently reporting underlying statistical assumptions.
Case Study: Analyzing News Claims and Population Inference
Headline Case Study: A poll reported by CBS News (updated August 20, 2025 at 10:38 AM EDT / CBS/AP) presented the headline: "Only 54% of U.S. adults say they drink alcohol, a record low. A new poll shows what's behind the decline."

Methodological Discussion: Critical evaluation of generalizations such as "U.S. Adults say…":
Extrapolating survey findings to the entire demographic of all U.S. adults requires evaluating whether the survey sample accurately reflects the population.
The phrasing "say they drink alcohol" introduces self-reporting considerations, where social desirability or memory recall can introduce measurement bias.
To draw valid conclusions about an overall population, the sample must be selected using a methodology that guarantees representative demographic coverage.
Population vs. Sample Definitions and Examples
Population: The complete, overarching group of individuals, items, or units about which information is sought.
Sample: A smaller, finite subset selected from the population that is directly observed, measured, and analyzed.
Concrete Examples:
Netflix Shows:
Population: Every show and movie available in the full Netflix catalog.
Sample: The specific three shows a student re-watches over the course of a semester.
Cat Videos:
Population: The total collection of all cat videos uploaded to YouTube.
Sample: The five specific cat videos watched at 2:00 AM instead of studying.
Coffee Cups:
Population: The total inventory of paper coffee cups in a store's supply.
Sample: The subset of black and brown paper coffee cups stacked on top of an espresso machine.
French Fries:
Population: The entire batch of French fries cooked in a restaurant fryer.
Sample: A single batch or serving collected on parchment paper for quality inspection or consumption.
Rationale for Sampling
Population Constraints: Full population enumeration (taking a census) is rarely feasible due to prohibitive costs, temporal constraints, or sheer physical scale.
Objective: Collect data from a representative sample to make statistical inferences about unobserved population parameters.
Comparative Example: Instead of interviewing a full population of university graduates, researchers collect data from a sample of graduated students.

Simple Random Sampling (SRS)
Definition: A probability sampling technique wherein every single individual in the population possesses an equal probability of selection.
Conceptual Model: Equivalent to drawing names thoroughly mixed inside a hat.
Core Properties:
Free from researcher favoritism or systematic selection patterns.
Forms the mathematical foundation for classical inferential statistics and hypothesis testing.
The Role of Randomness in Statistics
Operational Definition: Introducing deliberate unpredictability into the selection process.
Key Functions:
Protects the selection process against conscious and unconscious human bias.
Ensures equitable probability of inclusion across all population units.
Provides the mathematical foundation necessary to apply probability theory to evaluate estimation errors and support scientific conclusions.
Computational Implementation of SRS using NumPy
Simple Random Sampling can be implemented in Python using the
numpylibrary:
import numpy as np
# Define the population array
fruits = np.array(["apple", "banana", "cherry", "grape", "blueberry"])
# Initialize the pseudo-random number generator
rng = np.random.default_rng()
# Select a sample of size 3 without replacement
sample = rng.choice(fruits, size=3, replace=False)
print(sample)
# Output example: array(['blueberry', 'cherry', 'banana'], dtype='<U9')
Step-by-Step Code Analysis:
import numpy as np: Loads the NumPy library for numerical operations.fruits = np.array(...): Defines a population array containing 5 distinct fruit elements.rng = np.random.default_rng(): Instantiates NumPy's default BitGenerator/Random Number Generator (PCG64).sample = rng.choice(fruits, size=3, replace=False): Calls thechoicemethod to randomly extract elements without replacement (replace=False), ensuring no element can be drawn twice in a single sample.Output: Returns a array containing the randomly selected elements.
Population Mean vs. Sample Mean
Population Mean (): The exact mean value of a variable computed across all units in the full population.
Sample Mean (): The arithmetic mean computed across observations contained within a specific sample.
Estimation Goal: Use the observed sample statistic () to estimate the true, unobserved population parameter ().
Sampling Variation: The inherent variability that causes sample means () to differ from sample to sample when drawn from the same underlying population.
Estimating the Mean and Sample Size Effects
Numerical Demonstration:
Consider a population of exam scores with a true population mean of .
Taking independent samples of size yields varying estimates: , , and .
Each distinct sample yields a slightly different sample mean due to natural sampling variability.
Effect of Sample Size on Estimation Precision:
Small Samples: Display greater spread, resulting in higher variance across estimates.
Large Samples: Point estimates cluster more tightly around the true population mean (), increasing accuracy.

Understanding Bias in Sampling and Measurement
Scientific Definition of Bias: A systematic flaw or tilt in design, collection, or measurement that causes results to be consistently incorrect in a specific direction.
Sampling Bias: Structural defects in sample selection that make the sample unrepresentative of the population.
Key Distinction from Random Variation:
Random Variation: Unbiased, scatter-based fluctuations that average out toward zero with repeated trials.
Bias: Directional distortion that persists regardless of repetition and does not cancel out by taking larger samples.
Impact: A biased sampling process yields systematically misleading conclusions.
Real-World Metaphors:
A bathroom scale that is miscalibrated and consistently measures too heavy.
An instructor who calls exclusively on students seated in the front row of a lecture hall.
Specific Types of Sampling Bias
Convenience Sample: Selecting individuals who are most accessible or easiest to contact (e.g., surveying only students present in the library, while ignoring those in class, at home, or working).
Voluntary Response Bias: Allowing individuals to self-select into the survey group. This systematically overrepresents individuals with strong or extreme opinions.
Nonresponse Bias: Occurs when a significant portion of selected participants fail or decline to respond. If non-respondents systematically differ from respondents, the resulting sample becomes skewed.
Consequence: Each form of bias impairs the ability of the sample mean () to reliably estimate the population mean ().
Core Advantages of Random Sampling
Eliminates systematic selection bias by offering every individual equal selection probability.
Produces structurally representative samples that mirror population characteristics.
Validates using sample statistics to estimate population parameters.
Serves as an indispensable foundation for trustworthy empirical research and data science.
Reproducibility and Pseudo-Random Seeds
Concept of Reproducibility: Replicating identical experimental or computational results across separate executions of an analysis.
Problem: Standard random sampling yields unpredictable, varying outputs each time code is executed.
Solution: Setting a pseudo-random seed to fix the algorithm's starting state and generate an identical pseudo-random sequence.
Python/NumPy Implementation:
import numpy as np
# Define population array
fruits = np.array(["apple", "banana", "cherry", "grape", "blueberry"])
# Instantiate generator with a fixed seed for exact reproducibility
rng = np.random.default_rng(seed=42)
# Execute sample selection
sample = rng.choice(fruits, size=3, replace=False)
print(sample)
# Output array is locked and deterministic across execution runs
Application Utility: Essential for computational research, instructional demonstrations, peer verification, and collaborative code debugging.
Cultural Context of 'seed=42'
Cultural Origin: Derived from Douglas Adams' science fiction novel The Hitchhiker's Guide to the Galaxy.
Story Context: A giant supercomputer spends millions of years calculating the "Ultimate Question of Life, the Universe, and Everything," eventually concluding that the answer is 42.
Function in Computer Science: Serves as a standard, humorous convention for pseudo-random seed initialization.
Equivalence: Using
seed=42is functionally identical to usingseed=123orseed=2025; any specified integer guarantees deterministic reproducibility.
Reporting Statistical Assumptions and Parameters
Maintaining transparency and credibility in data analysis requires disclosing the following structural details:
Target Population: Explicit description of the entire population under investigation.
Sampling Strategy: Exact method utilized to gather observations (e.g., Simple Random Sampling vs. Convenience Sampling).
Sample Size (): Total count of units measured within the sample.
Reproducibility Details: Exact seed values specified in computational algorithms.
Summary of Key Takeaways
The sample mean () provides an empirical point estimate for the true population mean ().
A sample reflects its target population effectively when derived through random, unbiased selection methods.
Bias systematically corrupts accuracy and cannot be remedied by increasing sample size alone.
Pseudo-random seeds lock stochastic algorithms to ensure exact computational reproducibility.
Transparently reporting sampling methodologies and assumptions is required to establish statistical credibility.