Statistic Or Population- Chapter 1
Core Definitions: Population, Sample, Parameter, and Statistic
Population:
Definition: The entire set of individuals, objects, items, or events of interest that a researcher or statistician wants to study and draw conclusions about.
Scope: Includes every single unit that fits the defined criteria of the study group.
Sample:
Definition: A subset or portion of a population from which data is actively gathered.
Function: Serves as the observational basis when collecting data on the entire population is impossible, impractical, or unfeasible.
Parameter:
Definition: A descriptive, quantitative measure that characterizes an entire population.
Key Examples: Population mean (), population proportion (), population variance (), standard deviation (), or any other quantitative attribute of the population.
Unobservable Nature: In the context of inferential statistics, a parameter is a "true measure of interest" that exists in the unknown and cannot be directly observed directly when data is gathered only on a subset.
Statistic:
Definition: An equivalent descriptive measure computed directly from sample data.
Function: Serves as an observable proxy or estimate for the unknown population parameter.
Convergence in Inferential Statistics:
When a sample is drawn in a representative manner (e.g., randomly or through representative sampling designs), the sample statistic approaches the true population parameter as the sample size increases for specific measures like a mean or a proportion.
Descriptive Statistics versus Inferential Statistics
Descriptive Statistics:
Definition: The practice of gathering, organizing, summarizing, and presenting data collected on a specific group to describe or reach conclusions exclusively about that exact same group.
Applicable Contexts:
Complete Population Data: Used when a researcher, organization, or statistician possesses data on the entire population of interest (administrative records or a full census), allowing direct calculation of true parameters without uncertainty.
Sample Understanding: Used as an initial exploratory step on a sample dataset to calculate sample statistics and understand the dataset on hand prior to making broader inferences.
Detailed Applications and Examples:
Class Midterm Exam Percentiles:
Scenario: An instructor calculates the , , and percentiles of midterm exam grades for a class.
Classification: Descriptive statistics.
Reason: The instructor possesses full grade data for the entire target population (the entire class tabulation) and directly calculates the exact parameters.
Wisconsin Department of Health Services SNAP / FoodShare Study:
Scenario: The Department wants to know the proportion of enrolled SNAP (FoodShare) households in Wisconsin that have at least one child under the age of .
Population: All households currently on SNAP in Wisconsin.
Data Source: Complete administrative case records for all SNAP households in the state (a census).
Classification: Descriptive statistics, as the state agency directly computes the parameter from complete administrative population records.
Inferential Statistics:
Definition: The practice of using sample data gathered from a subset of a population to draw conclusions, estimations, or generalizations about the broader population from which the sample was drawn.
Thought Exercise: Addresses situations where observing the entire population is unfeasible, utilizing mathematical frameworks to determine how much can be inferred about the population and with what degree of confidence.
Detailed Applications and Examples:
American Housing Survey (AHS) Gas Appliance Study:
Goal: Estimate the proportion of occupied housing units in the United States that have a gas-powered stove or oven.
Population: All occupied housing units in the United States.
Sample: A nationally representative sample of approximately housing units surveyed in the American Housing Survey.
Parameter of Interest: The true proportion of all occupied US housing units containing a gas-powered stove or oven.
Statistic: The sample proportion of housing units in the AHS dataset that contain a gas-powered stove or oven.
Classification: Inferential statistics, because a sample of units is used to infer characteristics of the broader US housing population.
US Household Daily Energy Usage Study ():
Goal: Determine the average daily energy usage of US households in
Population: All households in the United States during the year
Parameter: The true population mean daily energy usage across all US households in
Sample: Households surveyed in the Residential Energy Consumption Survey.
Statistic: The sample mean daily energy usage calculated from the surveyed households.
Classification: Inferential statistics.
Skittle Color Proportion Study:
Goal: Determine the proportion of all manufactured Skittles that are red.
Population: All Skittles produced.
Parameter: The true proportion of red Skittles across the total population.
Sample: bags of Skittles purchased from Woodman's, emptied and counted.
Statistic: The sample proportion of red Skittles, calculated as:
- Classification: Inferential statistics.
- Blackberry Jam Sugar Content Quality Control:
- Goal: A quality control manager at a food processing plant wants to measure the standard deviation of the sugar content (mass or concentration) of jars of blackberry jam produced at a specific plant.
- Population: All jars of blackberry jam produced at this specific plant.
- Parameter: The true standard deviation of sugar content across all jam jars produced at the plant.
- Inferential Procedure: Sample random jars (or select every jar produced during a day), measure the sugar content of the sample, and calculate the sample standard deviation as a proxy for the population parameter.
- Descriptive Alternative: If the manager pulled *every single jar* produced during an entire day and tested the sugar content across the complete production run, computing the exact standard deviation for that day would be an application of descriptive statistics.
Classification of Variables and Data Types
Qualitative vs. Quantitative Categorization:
Categorical (Nominal) Data:
Definition: Variables that divide observations into distinct qualitative categories or groups that have no inherent numerical value or natural scalar ordering.
Example: Person's hair color (e.g., brown, blonde, black, red).
Ordinal Data:
Definition: Data that categorizes observations into ordered categories where relative order or ranking is meaningful, but the precise mathematical distances between categories are unknown, undefined, or unequal.
Example: Likert scale survey responses evaluating self-reported level of concern about the environment on a scale ranging from "not at all" to "extremely". While an answer of "extremely" indicates higher concern than "not at all", the exact quantitative difference in concern between scale points cannot be measured.
Numerical Data:
Definition: Quantitative variables taking on scalar numeric values.
Measurement Scales for Numerical Data:
Interval Level Measurement:
Definition: Numerical data where differences between values are meaningful, but there is no true or absolute zero point (zero does not denote the complete absence of the measured attribute).
Ratio Level Measurement:
Definition: Numerical data with a true, meaningful absolute zero point representing the total absence of the attribute. True mathematical ratios are valid (e.g., a value of is exactly twice as much as a value of
Discrete-Valued versus Continuous-Valued Data:
Discrete-Valued Data:
Definition: Data that can only take on specific, distinct values on the real number line within any given range. Discrete variables generally represent counts or frequencies that progress in set increments.
Identification Rule: Any variable structured as "the number of [items/events]" is a discrete count variable taking whole number values such as , , , etc.
Count/Frequency Examples:
Number of people attending an event.
Number of shirts sold in a day.
Number of likes on an Instagram post.
Number of children in a household (, , , ; an individual household cannot contain children, even if a sample average is
Number of customers arriving at a store during an hour.
Number of unemployed people in a county.
Non-Count Discrete Examples:
Credit score (defined by system rules to only take integer values).
SAT score (takes defined integer step values).
Shoe size (can take integer and half-integer values like , , , , but cannot take arbitrary values like
Continuous-Valued Data:
Definition: Data that can conceivably take on any real number value within a continuous range or interval on the number line. Continuous variables represent physical measurements rather than counted quantities.
Conceptual Foundation: Defined by what the variable conceivably could be realized as before observation rather than what is observed in a specific static instance.
Distinction Rule: Continuous data is measured; discrete data is counted.
Examples:
Amount of time required to reach a destination.
Temperature ( or
Birth weight.
Percent change in stock price.
Earthquake magnitude measured on the Richter scale (can take exact continuous decimal values like
Human height.
County unemployment rate (a proportion or percentage of unemployed individuals, e.g.,
Detailed Evaluation of Data Classification Examples
County Unemployment Rate:
Classification: Numerical, Ratio-level, Continuous-valued.
Explanation: Unemployment rate is a proportion or percentage calculated as a ratio. It can take any real number value in a continuous range (e.g., ). A rate of represents absolute zero unemployment, and a rate of is twice as high as
Contrast with Unemployed Count: The number of unemployed people in a county is numerical, ratio-level, but discrete-valued (a whole number count).
Person's Hair Color:
Classification: Categorical.
Explanation: Divides individuals into discrete qualitative groups without natural scalar ordering or corresponding numbers.
Number of Children in a Household:
Classification: Numerical, Ratio-level, Discrete-valued.
Explanation: Represents a count. A household can have , , or children, but cannot have children. It possesses a true zero point ( children).
Self-Reported Environmental Concern (Likert Scale):
Classification: Ordinal.
Explanation: Responses ranging from "not at all" to "extremely" provide ranking information, but distance between scale options cannot be quantitatively measured.
Questions & Discussion
Setup Interlude:
Prompt / Dialogue: Justin inquired whether time was needed to set up equipment so he could plug in his device, confirming he would be ready shortly.
Earthquake Richter Scale and Zero Point:
Dialogue Question: Can an earthquake size on the Richter scale be or , and does represent the complete absence of an earthquake (distinguishing interval vs. ratio scales)?
Discussion Response: The Richter scale is continuous-valued because exact measurements can yield detailed real number values like . Regarding measurement level, if a zero on the scale represents a true zero point (absence of seismic activity), it operates as a ratio scale.
Customer Arrival Frequencies in Retail:
Dialogue Question: If measuring how many customers buy items at certain periods of the day (e.g., peak rush hour vs. ), what type of variable is involved?
Discussion Response: The specific metric—the number of customers arriving per hour—is discrete-valued because it measures a whole-number count of human beings.
Unchanging Counts vs. Conceivable Realization:
Dialogue Question: If the same number of people enter an area continuously and the observed number does not change, does that change its classification from discrete to something else?
Discussion Response: No. Data classification depends on what the variable conceivably could take on prior to measurement across the number line, not what static values happen to be observed in a given sample. A count of people remains discrete because it can only conceivably take on whole number values (, , , etc.).
Height Measurement Explanation:
Dialogue Question: Is human height discrete or continuous, and what is the simple rule to remember it?
Discussion Response: Height is continuous because it is measured along a continuous spectrum, whereas discrete data represents objects or events that are counted.
Flexibility in Classification Explanations:
Note on Assessment: Some statistical boundary decisions between categories contain nuance. On evaluations, clear logic demonstrating understanding of core principles (such as explaining why a rate is continuous vs. why a count is discrete) is essential.