Comprehensive Introduction to Statistics: Variables, Data Collection, Sampling, and Bias

Fundamentals of Data and Variable Types

  • Variable Classifications:

    • Categorical Variable: A variable that places an individual into one of several groups, labels, or physical attributes. Categorical variables represent qualitative traits rather than numerical measurements.

      • Examples: Gender identification, hair color, religion, ethnicity, academic major, movie genre, content rating (e.g., PG, PG-13, R, NC-17).

      • Special Numeric Cases:

        • ZIP Code: Classified as categorical. Although expressed as a number, a ZIP code describes a physical geographic location (an attribute). Calculating a mathematical value such as an average ZIP code is logically meaningless.

        • Area Code: Classified as categorical. It serves as a geographic location identifier rather than a quantitative measure.

        • Release Year (in specific data contexts): Classified as categorical when used as a temporal label identifier (e.g., categorizing movies by whether they were released in a given year or not), because calculating an average release year like 2019.652019.65 provides no meaningful quantitative context.

    • Quantitative Variable: A variable that takes numerical values for which arithmetic operations such as addition, subtraction, or finding an average make sense.

      • Examples: Height, income, runtime/duration of a movie (e.g., 1 hour 20 minutes1\text{ hour } 20\text{ minutes}), distance from a business location, fuel economy, box office revenue.

Population, Censuses, and Samples

  • Population: The entire group of individuals or items about which statistical information is desired.

    • Examples: All voting-eligible United States citizens, all students enrolled at Georgia State University (GSU), all employees who smoke.

  • Census: An exhaustive collection of data gathered from every single individual within a population.

    • United States Census: Conducted every 1010 years, attempting to survey all United States residents above age 1818 regarding income, gender, and household demographics.

    • Practical Limitations: True censuses are extraordinarily difficult or impossible to execute for large populations due to non-response, untraceable individuals, time constraints, and immense costs. For instance, surveying a class of 115115 or 800800 students entirely requires significant effort; attempting it on a national scale always leaves gaps.

  • Sample: A subset of individuals selected from a target population, used to collect data in order to draw conclusions and estimate unknown parameters about the entire population.

    • Sample Size (nn): The total number of individuals or items selected and successfully measured in a study.

Data Collection Methods: Observational Studies vs. Experiments

  • Observational Study:

    • Definition: A study in which conditions or characteristics of individuals are observed and variables of interest are measured, but no treatment is imposed on the subjects.

    • Examples: Observing animal behavior on a safari; reviewing existing hospital medical records; recording student hair colors in a lecture hall.

    • Primary Purpose: To describe a specific group, situation, or characteristic.

    • Critical Limitation: Observational studies cannot establish cause-and-effect relationships (causation) due to unmeasured confounding factors.

  • Experiment:

    • Definition: A study in which researchers deliberately impose specific treatments on individuals to measure and observe the resulting responses.

    • Examples: Evaluating a new blood pressure medication by administering standard treatment to one group and the new drug to another; applying different fertilizers to plants to observe height growth; testing different doses of aspirin to evaluate heart attack risk reduction.

    • Primary Advantage: Properly designed experiments can establish cause-and-effect relationships (causation).

Case Studies and Checkpoint Applications

  • Dataset Individuals: Individuals in a dataset do not need to be human beings; they can be objects, animals, vehicles, archaeological artifacts, or financial accounts.

  • System Assessment Rules: When completing assessment checkpoints, viewing the solution directly locks out further attempts and awards zero points (00). To re-attempt questions without penalty, select options to continue working.

  • Car Dealership Dataset Analysis:

    • Individuals: The car buyers.

    • Variables:

      • ZIP Code: Categorical

      • Sex: Categorical

      • Distance from Dealer: Quantitative

      • Car Model: Categorical

      • Fuel Economy: Quantitative

      • Price: Quantitative

  • Archaeological Dig Artifact Analysis:

    • Context: Project staff analyze pottery shards, stone tools, and artifacts. The project director randomly inspects 2%2\% of the artifacts.

    • Population: All artifacts collected from the archaeological site.

    • Sample: The 2%2\% subset of artifacts selected and checked.

  • General Motors Smoking Cessation Study:

    • Context: General Motors sponsored a study involving 878878 volunteer smoking employees split into two equal groups of 439439. One group received up to 750750 dollars as a financial incentive to quit smoking for a year; the other group was simply encouraged to stop smoking.

    • Study Type: Experiment (treatments and financial incentives were imposed).

    • Categorical Variables: Smoking status after one year (quit vs. did not quit) and treatment type (financial incentive vs. encouragement).

    • Quantitative Variables: None present among measured attributes.

    • Population: All employees who smoke.

    • Sample: The 878878 volunteer employees who participated (439+439=878439 + 439 = 878).

    • Result: Participants receiving financial incentives were 33 times more likely to quit smoking after one year.

  • Newspaper Reader Survey:

    • Context: A publisher inserts survey forms into 1,0001{,}000 randomly chosen copies of a weekly newspaper. A total of 189189 surveys are completed and returned.

    • Population: All readers of the local weekly newspaper.

    • Sample: The 189189 readers who completed and returned the survey (since actual data exists exclusively for respondents).

  • Popular Movies Dataset (2019 Releases):

    • Individuals: The 1212 specific movies featured in the dataset.

    • Variables:

      • Year Released: Categorical

      • Rating (PG-13, R, etc.): Categorical

      • Runtime/Duration: Quantitative

      • Genre: Categorical

      • Box Office Revenue: Quantitative

    • Unmeasured Attributes: The total number of viewers per movie is absent from the dataset and must be ignored.

Sampling Methods and Sources of Bias

  • Bias: A systematic distortion in the design of a statistical study that makes it consistently overestimate or underestimate the true population parameter.

  • Flawed Sampling Techniques:

    • Voluntary Response Sampling (Volunteer Sampling):

      • Definition: Occurs when individuals self-select to participate in a study.

      • Examples: Customer receipt surveys offering free food (e.g., QR codes on Chili's or Applebee's receipts); tear-off flyers or QR code surveys posted on dorm walls.

      • Flaw: Produces extreme opinion bias. Individuals who choose to respond typically hold unusually strong negative or positive views (e.g., Yelp review bias). Results cannot be generalized to the broader population.

    • Convenience Sampling:

      • Definition: Occurs when researchers select individuals who are easiest to reach or happen to be in the right place at the right time.

      • Examples: Surveying students walking outside Central Dining Hall.

      • Flaw: The selected individuals are rarely representative of the overall target population.

  • Three Primary Categories of Statistical Bias:

    • Undercoverage:

      • Definition: Occurs when certain groups within a population are systematically left out of the process of choosing the sample.

      • Examples: Estimating student height using only the university basketball team; conducting landline-only phone surveys (excluding people without landlines); election polling in 2016 that sampled heavily from urban democratic centers (e.g., New York City, Atlanta, Boston, San Francisco, Miami) while underrepresenting rural populations.

      • High School Parking Lot Example: End-of-year surplus funds needed to be spent on campus projects. An administrator surveyed students stepping off school buses to evaluate demand for expanding student parking lots. Bus riders do not drive or park at school, leading to undercoverage of student drivers.

    • Nonresponse Bias:

      • Definition: Occurs when an individual chosen for a sample cannot be contacted or refuses to participate.

      • Examples: Direct mail surveys where the vast majority of recipients discard the materials.

    • Response Bias:

      • Definition: A systematic pattern of inaccurate or false responses provided by survey respondents.

      • Causes: Authoritative interviewers, sensitive or illegal topic matters, social desirability, or poorly worded questions.

      • 2002 Ohio High School Drug Study Case: A high school implemented a drug prevention program. To evaluate efficacy, a police officer (school resource officer) walked down hallways directly asking students, "Do you do drugs?" Every student answered "No." The school published the program as 100%100\% effective, failing to recognize severe response bias caused by authority pressure.

      • Wording Bias: Asking questions with loaded phrasing (e.g., "Do you support decreasing tuition at Georgia State even though 75 kids will die because of it?") forcibly alters responses.

Simple Random Sampling and Sampling Variability

  • Simple Random Sample (SRS):

    • Definition: A sampling design where every set of nn individuals in the population has an equal chance to be selected as the sample.

    • Procedure: Number every individual in a population of size NN (e.g., numbering individuals from 11 to 250250) and use a random process to select nn individuals (e.g., drawing 3030 numbers).

    • Significance: SRS eliminates selection bias and serves as the baseline sampling requirement for valid statistical inference.

  • Sampling Variability:

    • Definition: The natural variation observed among sample statistics when different random samples of the same size nn are drawn from the exact same population.

    • Concept: Taking repeated random samples of size nn will yield slightly different numerical estimates (e.g., sample mean height of 5 ft 10 in5\text{ ft } 10\text{ in} versus other sample averages) each time.

    • Core Purpose of Statistics: To develop methods that use a single random sample to draw valid inferences about an entire population, despite the unavoidable presence of sampling variability.