Introductory Statistics - Chapter 1: Sampling and Data
Definitions of Statistics, Probability, and Key Terms
Statistics Definition: Statistics is the science dealing with the collection, analysis, interpretation, and presentation of data.
Descriptive Statistics: The practice of organizing and summarizing data. Data can be summarized using visual graphs or numerical summaries, such as calculating an average.
Inferential Statistics: Formal mathematical methods used to draw conclusions from data after studying probability and probability distributions. Inferential statistics uses probability to measure confidence in the correctness of drawn conclusions.
Probability: A mathematical tool used to study randomness, dealing directly with the likelihood or chance of an event occurring. Probability calculations provide the foundation for statistical inference.
Key Statistical Terms:
Population: A comprehensive collection of persons, things, or objects under study.
Sample: A portion or subset selected from the larger population to be studied in order to gather information about the population.
Parameter: A numerical characteristic that represents a property of an entire population.
Statistic: A numerical characteristic calculated from sample data that estimates a population parameter.
Representative Sample: A sample that accurately reflects the key characteristics of the population from which it was drawn.
Variable: A characteristic of interest measured for each entity in a population, conventionally denoted by capital letters such as or .
Numerical Variable: Takes on numerical values with equal units, such as weight in pounds or time in hours.
Categorical Variable: Places an entity into a specific category, such as political affiliation, eye color, gender, or ethnicity.
Data: The actual observed values recorded for a variable, which may be numbers or text.
Datum: A single individual value within a dataset.
Sleep Duration Example & Analysis:
When recording average night sleep duration rounded to the nearest half-hour (), data points cluster around specific central values.
Replicating the sleep study across different student populations (e.g., an English class versus a Physics class) would likely yield different distributions due to variances in workload, scheduling, or student demographics.
Study Application: Cumulative GPA Analysis:
Context: A study conducted at a local college analyzes the cumulative GPA of graduating students.
Population: All students who graduated from the college last year.
Sample: A group of students who graduated from the college last year, randomly selected.
Parameter: The average cumulative GPA of students who graduated from the college last year.
Statistic: The average cumulative GPA of students in the sample who graduated from the college last year.
Variable: The cumulative GPA of one individual student who graduated from the college last year.
Data: The set of individual observed cumulative GPAs, such as , , , and .
Note: The phrase "all students who attended the college last year" describes a broader population than the target population of graduates.
Study Application: Extracurricular Activities Survey:
Context: Researching the mean number of extracurricular activities in which high school students participate, based on a random survey of high school students.
Population: All high school students.
Sample: The high school students randomly surveyed.
Parameter: The mean number of extracurricular activities for all high school students.
Statistic: The mean number of extracurricular activities calculated from the surveyed students.
Variable: The number of extracurricular activities in which a single high school student participates.
Data: The observed individual counts of activities, such as , , and .
Data, Sampling, and Variation in Data and Sampling
Types of Data:
Qualitative Data: Results from categorizing or describing attributes of a population. Examples include hair color, blood type, ethnic group, and car models.
Quantitative Data: Values that are always numerical, generated by counting or measuring population attributes.
Quantitative Discrete Data: Data resulting from counting, taking on distinct, specific numerical values (often phrases beginning with "the number of"). Examples include counting phone calls per day (, , , ).
Quantitative Continuous Data: Data resulting from precise physical measurement, capable of taking any real value within an interval. Examples include athlete heights (, , ).
Data Type Classification Examples:
Gym machine counts (, , , , ): Quantitative discrete.
Lawn area measurements (, , , , ): Quantitative continuous.
House exterior colors (white, yellow, white, red, white): Qualitative.
Number of pairs of shoes owned: Quantitative discrete.
Vehicle type driven: Qualitative.
Vacation destination choice: Qualitative.
Distance from home to grocery store: Quantitative continuous.
Number of classes taken per school year: Quantitative discrete.
Class tuition amounts: Quantitative continuous.
Calculator model used: Qualitative.
Movie ratings (e.g., G, PG, PG-13, R or star ratings): Qualitative / Ordinal.
Political party affiliation: Qualitative.
Weights of sumo wrestlers: Quantitative continuous.
Poker winnings in dollars: Quantitative continuous.
Number of correct quiz answers: Quantitative discrete.
Attitude ratings toward government: Qualitative.
IQ scores: Quantitative discrete (when recorded as integer scores).
Student academic classifications (Freshman, Sophomore, Junior, Senior) in a pie chart: Qualitative.
Standardized exam scores grouped into class boundaries ( to less than , to less than , etc.) in a histogram: Quantitative continuous.
Graphical Displays for Qualitative Data:
Pie Charts: Circular charts divided into wedges, where each wedge is proportional in size to the percentage of individuals in that category.
Bar Graphs: Displays where bar lengths (vertical or horizontal) are proportional to the count or percentage of individuals in each category.
Pareto Charts: Specialized bar graphs where category bars are sorted in descending order from largest to smallest size.
Random Sampling Methods:
Simple Random Sampling: Every sample of a given size has an equal probability of selection. Example: Computer-generated random numbers selecting student names from an alphabetical list of students.
Stratified Sampling: The population is partitioned into non-overlapping subgroups called strata, and a proportional simple random sample is taken from each stratum. Examples: Sampling by academic department; selecting students from each class level (Freshman, Sophomore, Junior, Senior).
Cluster Sampling: The population is partitioned into clusters; a subset of clusters is randomly selected, and every member of the chosen clusters is surveyed. Examples: Randomly selecting four academic departments and surveying all students in them; selecting two class years randomly out of four and surveying all students in those years.
Systematic Sampling: A starting point is chosen randomly, and every individual is selected from an ordered list. Examples: Surveying every person at a checkpoint; selecting every student from an alphabetical list starting from a randomly generated index.
Non-Random Sampling and Replacement Rules:
Convenience Sampling: Non-random sampling utilizing readily available results or subjects. Examples: Interviewing classmates in an algebra class; surveying the first students walking past a library. Highly prone to bias.
Sampling with Replacement: Selected individuals are returned to the population prior to subsequent selections, allowing an individual to be chosen multiple times.
Sampling without Replacement: Selected individuals are permanently removed from the selection pool for subsequent draws. When the sample size is small relative to the population, sampling without replacement mathematically approximates sampling with replacement.
Sampling Errors, Nonsampling Errors, and Bias:
Sampling Errors: Errors caused by the actual process of sampling, such as having a sample size that is too small.
Nonsampling Errors: Errors caused by factors unrelated to the sampling process, such as defective counting equipment or biased question phrasing.
Sampling Bias: Created when data collection systematically favors certain population members over others, rendering conclusions invalid.
Sampling Method Identification Exercises:
Selecting players (ages ), players (ages ), and players (ages ): Stratified sampling.
Interviewing all HR personnel across five high-tech companies: Cluster sampling.
Interviewing female and male high school teachers: Stratified sampling.
Interviewing every cancer patient from an hospital roster: Systematic sampling.
Using computer-generated random numbers to select students: Simple random sampling.
Interviewing algebra classmates regarding average jean ownership: Convenience sampling.
Representativeness Analysis of Specific Samples:
High school GPA study using university honor students: Non-representative and heavily biased upward.
Most popular cereal under age 10 surveyed outside a supermarket for 3 hours: Biased by store location, time of day, and demographic access.
U.S. adult average income surveyed via state-based cluster sampling of U.S. congressmen: Severely biased; congressmen earnings do not reflect general population income.
Public transit proportions evaluated by interviewing Central Park bench sitters in NYC: Biased toward park visitors and non-representative of general transit usage.
Average stay cost in Massachusetts hospitals surveyed via simple random sampling of 100 hospitals statewide: Representative.
Critical Evaluation Factors for Statistical Studies:
Non-representative / Biased Samples: Yield invalid, inaccurate results.
Self-Selected Samples: Call-in or voluntary response surveys reflect overrepresented extreme views.
Sample Size Issues: Small samples reduce reliability, though unavoidable in constraints like automobile crash testing or rare medical trials.
Undue Influence: Framing questions or environment to distort subject responses.
Non-response or Refusal: High non-response shifts representative traits toward respondents with strong opinions.
Causality Fallacy: Confusing correlation with causation; association between variables does not imply one causes the other.
Self-Funded or Self-Interest Studies: Studies financed by parties with outcome interests require rigorous scrutiny regarding impartiality.
Misleading Data Presentation: Scaled graph manipulation, missing baseline context, or selective reporting.
Confounding: Overlapping effects of multiple factors preventing isolation of individual variable impact.
Detailed Study Scenarios & Bias Evaluation:
ABC College Textbook Spending Study ( upperclassmen population):
Sample 1: organic chemistry students spending .
Sample 2: Every senior in P.E. classes ( total) spending .
Evaluation: Neither sample represents the student population. Chemistry students take high-cost STEM courses, biasing spending upward; P.E. seniors take low-cost courses, biasing spending downward.
Radio Station Listener Preference Study ( audience size):
Method: Convenience sample of attendees at a station concert event ( preferred talk shows, preferred music).
Evaluation: Biased. Concert attendees inherently skew toward music preferences, underrepresenting home or commuting talk-show listeners.
Frequency, Frequency Tables, and Levels of Measurement
Rounding Off Rules:
Carry final calculated statistics to one additional decimal place beyond the raw data.
Perform rounding only on the final answer; do not round intermediate steps.
Four Levels of Measurement:
Nominal Scale Level: Qualitative data consisting of categories, names, labels, colors, or yes/no responses. Data cannot be meaningfully ordered (e.g., placing pizza before sushi has no ordinal meaning).
Ordinal Scale Level: Categorical data that can be ordered or ranked, but differences between ranks cannot be calculated or measured meaningfully (e.g., top five national parks ranked to ).
Interval Scale Level: Ordered numerical data where differences between values are meaningful and measurable, but no absolute natural zero point exists ( does not denote complete absence). Example: Celsius () and Fahrenheit () temperature scales (, but does not represent absence of thermal energy; and exist).
Ratio Scale Level: Ordered numerical data with measurable differences AND a true natural zero starting point ( denotes complete absence of the property). Ratios between values are mathematically valid. Example: Exam scores ( to points), where a score of is four times greater than a score of .
Measurement Scale Examples:
High school soccer player athletic ability (superior, average, above average): Ordinal.
Baking temperatures (, , , , ): Interval.
Colors of crayons in a 24-count box: Nominal.
Social Security numbers: Nominal.
Incomes measured in dollars: Ratio.
Satisfaction survey (, , ): Ordinal.
Preferred TV shows (comedy, drama, sci-fi, sports, news): Nominal.
Time of day on an analog watch: Interval.
Distance in miles to the closest grocery store: Ratio.
Calendar dates (, , , , ): Interval.
Heights of adult women aged 21–65: Ratio.
Letter grades (A, B, C, D, F): Ordinal.
Frequency Definitions:
Frequency (): The number of times a specific value occurs within a dataset.
Relative Frequency (): The proportion or fraction of times a value occurs relative to total outcomes : .
Cumulative Relative Frequency (): The accumulation of consecutive relative frequencies up to the current row.
Frequency Table Construction 1: Student Work Hours ( students):
Raw Data:
Data Value : Frequency , Relative Frequency , Cumulative Relative Frequency
Data Value : Frequency , Relative Frequency , Cumulative Relative Frequency
Data Value : Frequency , Relative Frequency , Cumulative Relative Frequency
Data Value : Frequency , Relative Frequency , Cumulative Relative Frequency
Data Value : Frequency , Relative Frequency , Cumulative Relative Frequency
Data Value : Frequency , Relative Frequency , Cumulative Relative Frequency
Frequency Table Construction 2: Soccer Player Heights ( male semiprofessional players):
Interval : Frequency , Relative Frequency , Cumulative Relative Frequency
Interval : Frequency , Relative Frequency , Cumulative Relative Frequency
Interval : Frequency , Relative Frequency , Cumulative Relative Frequency
Interval : Frequency , Relative Frequency , Cumulative Relative Frequency
Interval : Frequency , Relative Frequency , Cumulative Relative Frequency
Interval : Frequency , Relative Frequency , Cumulative Relative Frequency
Interval : Frequency , Relative Frequency , Cumulative Relative Frequency
Interval : Frequency , Relative Frequency , Cumulative Relative Frequency
Percentage Analysis:
Percentage of heights less than : (corresponds directly to at upper boundary ).
Percentage of heights falling between and : (or ).
Experimental Design and Ethics
Experimental Design Concepts:
Explanatory Variable: The independent variable manipulated by researchers to determine its effect on another variable.
Response Variable: The dependent variable that measures the outcome of interest caused by changes in the explanatory variable.
Treatments: Specific values or conditions of the explanatory variable assigned to experimental subjects.
Experimental Unit: A single individual object or entity measured within the experiment.
Observational Studies vs. Experiments:
Observational Study: Researchers measure variables without manipulating treatments or directly controlling environments. Observational studies cannot prove causality due to lurking variables.
Lurking Variables: Unmeasured or uncontrolled additional variables that obscure cause-and-effect relationships.
Random Assignment: Assigning subjects to treatment conditions randomly to distribute lurking variables equally across all groups.
Control Mechanisms and Blinding:
Control Group: A baseline group receiving a dummy or placebo treatment to isolate active treatment effects from experimental influence.
Placebo: A neutral treatment (e.g., sugar pill) with no active ingredient, controlling for the power of suggestion.
Single-Blind Experiment: Participants do not know whether they are assigned to active treatments or placebos.
Double-Blind Experiment: Both the participants AND the researchers interacting directly with them are unaware of treatment assignments.
Experimental Study Analysis 1: Aspirin & Heart Attack Study:
Context: men aged to are randomly split into two groups (aspirin daily vs. placebo daily for 3 years) to test heart attack prevention.
Population: Men between the ages of and
Sample: The recruited men participating in the study
Experimental Units: Each individual man participating in the trial
Explanatory Variable: Regular daily intake of aspirin
Response Variable: Heart attack occurrence status
Treatments: Daily aspirin pill vs. daily placebo pill
Experimental Study Analysis 2: Smell, Taste & Maze Performance:
Context: Participants complete mazes three times wearing floral-scented masks and three times wearing unscented masks, assigned at random to mask order sequence. Time and scent perception ratings are recorded.
Explanatory Variable: Mask scent type (floral-scented vs. unscented)
Response Variable: Maze completion time and subject impression rating of scent (positive, negative, or neutral)
Treatments: Floral-scented mask trial condition vs. unscented mask trial condition
Lurking Variables: Maze practice effect (learning curve across trials), mental fatigue, individual baseline spatial ability
Blinding Feasibility: Full participant blinding is impossible if the floral scent is noticeable, but double-blind data recording and timing analysis can be maintained.
Experimental Study Analysis 3: Birth Order and Personality:
Cannot be conducted as a randomized experiment because researchers cannot randomly assign birth order to individuals.
Main Problem: Observational nature introduces uncontrolled confounding factors (family size, socioeconomic status, parental age differences) that prevent causal conclusions.
Experimental Study Design 4: Texting While Driving Simulation:
Explanatory Variable: Driving condition (distracted by texting vs. undistracted driving)
Response Variable: Driver reaction brake time (seconds elapsed when leading vehicle brakes)
Treatments: Driving while actively texting vs. driving without phone interaction
Participant Selection: Must consider age, driving experience, reaction speed, and phone usage habits
Design Considerations: Randomly assigning participants to a repeated-measures design (where each driver tests under both conditions in randomized order) controls for individual driving skill variability better than splitting into separate groups.
Lurking Variables: Vehicle simulator familiarity, phone interface familiarity, fatigue, ambient lighting
Blinding Application: Drivers cannot be blinded to texting, but timing evaluators can be blinded to driver treatment conditions during video measurement analysis.
Ethical Standards in Statistical Research:
Minimize risks to participants and ensure risks are reasonable.
Obtain formal informed consent from human participants.
Protect participant privacy and data confidentiality strictly.
Unethical Behaviors & Corrective Actions in Neighborhood Surveying:
Selecting a comfortable street block: Introduces convenience selection bias. Correction: Randomly select blocks across the entire target community.
Skipping unreached homes without returning: Introduces non-response bias. Correction: Record unreached addresses and make repeated call-backs at varied times.
Fabricating responses for missed homes: Constitutes fraudulent data fabrication. Correction: Exclude uncollected responses entirely and document the non-response rate.
Unethical Behaviors & Corrective Actions in Commercial Juice Testing:
Study commissioned by apple juice vendor: Financial conflict of interest. Correction: Utilize independent third-party research administration.
Restricted choice options (only apple and cranberry): Selection bias. Correction: Include comprehensive choices or open-ended preference options.
Unblinded taste testing: Allows brand recognition bias. Correction: Implement blind taste testing.
Misleading advertising claims: Advertising Brand X vs Brand Y as "Most teens like Brand X as much as or more than Brand Y" distorts data where had no preference. Correction: Accurately report full comparative statistical distributions.