Section 1.1 Notes: Introduction to the Practice of Statistics
Define Statistics and Statistical Thinking
Statistics is the comprehensive science that involves the systematic processes of collecting, organizing, summarizing, and analyzing information (data). Its primary goal is to draw conclusions or answer specific research questions based on this data. Crucially, statistics also provides a crucial measure of confidence in any conclusions drawn, quantifying the uncertainty associated with inferences.
Data are the fundamental facts, observations, or propositions from which conclusions are drawn or decisions are made. They meticulously describe characteristics of individual entities and inherently vary across individuals (e.g., varying heights, diverse hair colors, different sleep durations, varied calorie consumption). A central objective of statistics is to meticulously describe and understand the inherent sources of variability observed both across different individuals and over periods of time.
Variability is an intrinsic and pervasive characteristic of the real world; individuals are rarely identical (e.g., people naturally differ in height, amount of sleep, calories consumed). Understanding and quantifying this variability is the core challenge statistics addresses.
Discussions related to statistics often appear in everyday life, reinforced by data encountered in newspapers, social media platforms, television reports, billboards, and personal tracking devices. Connecting statistical theory to these real-world observations helps solidify understanding.
Explain the Process of Statistics
In the vast majority of studies, it is practically unrealistic or impossible to access and study every single individual of interest, which constitutes the population.
Population: This refers to the entire group of individuals (persons, objects, or items) that a researcher specifically intends to study and for which conclusions are desired. For instance, all high school students in the U.S.
Individual: This is a single, specific person or object that is an elementary member of the defined population. For example, one particular high school student.
Sample: This is a carefully selected subset of the population that is actually examined and studied. Because it is often more feasible and economical, data is collected from a sample, and insights from this sample are then used to make statements about the larger population.
Descriptive Statistics: This branch of statistics is dedicated to organizing and summarizing the data collected. It involves calculating numerical summaries (like means, medians, modes, standard deviations) and creating visual summaries such as tables, charts, and graphs (e.g., histograms, bar charts, pie charts, scatter plots) to describe the data's main features. Descriptive statistics helps present data in a comprehensible format, often highlighting patterns, central tendencies, and variability within the sample.
Statistic: A numerical summary that is precisely calculated based on the observed data from a sample. For example, the mean height of students in a sample.
Inferential Statistics: This sophisticated branch comprises methods designed to extend or generalize results obtained from a sample to the broader population from which the sample was drawn. Crucially, inferential statistics also provides a rigorous framework to assess the reliability and level of confidence in these extended results, often through techniques like confidence intervals and hypothesis testing.
Parameter: A numerical summary that describes a characteristic of the entire population. This value is typically unknown and is estimated using sample statistics. For example, the true mean height of all students in a population.
Example distinguishing parameter vs statistic:
Suppose the population proportion of all students on a campus who currently have a job is known to be . This value is a parameter because it describes the entire population.
If a sample of students is randomly surveyed, and the sample proportion of students with a job is found to be . This value is a statistic because it is derived solely from the sample.
The Process of Statistics (4 detailed steps):
Identify the Research Objective: This initial step involves clearly defining what question(s) need to be answered and precisely determining the population of interest for the study. It also specifies the specific characteristics (variables) or relationships that will be investigated.
Collect the Data Needed to Answer the Question(s): This is a critical step. Often, due to the difficulty, cost, or sheer impossibility of studying the entire population, a carefully selected sample is used. The quality of data collection is paramount; if data are collected poorly, even the most advanced statistical analyses will lead to meaningless or misleading conclusions. Proper sampling methods, such as random sampling, are essential to ensure the sample is representative of the population and minimize bias.
Describe the Data: Once collected, the data must be organized and summarized. This involves using descriptive statistics (numerical and graphical summaries like means, medians, standard deviations, frequency distributions, and visual displays such as histograms, box plots, and bar charts). This step provides an initial overview of the data, helps identify patterns, outliers, and guides the selection of appropriate inferential methods for later stages.
Perform Inference: In this final stage, statistical techniques are applied to extend the results obtained from the sample back to the population. Sophisticated methods are used to draw conclusions about the population parameters based on the sample statistics. Crucially, a level of reliability (e.g., confidence level, p-value) is quantified and reported, indicating the certainty or uncertainty associated with these generalizations. This allows for evidence-based decision-making.
Practical Emphasis: The integrity and validity of any statistical conclusion fundamentally rely on correct data collection procedures. Poorly collected data (e.g., due to bias, improper sampling, or measurement errors) will inevitably lead to meaningless, inaccurate, or unreliable conclusions. Therefore, appropriate data collection is the foundational pillar for conducting valid and trustworthy statistical inference.
EXAMPLE: Illustrating the Process of Statistics
Study: Investigating the association between school start times and the duration of sleep among high school students.
Sample: A specific group of U.S. adolescents were randomly selected to participate in the study, ensuring a representative subset of the population.
Sample Characteristics: The collected data revealed a mean age of years and a standard deviation of year for the adolescents in the sample. These are descriptive statistics from the sample.
Findings (Descriptive Summaries of the Sample):
High schools with start times at 8:30 a.m. or later were associated with a sleep duration that was, on average, minutes longer.
Furthermore, the study found that each 1-hour delay in school start time was associated with approximately minutes longer sleep duration.
Conclusion (Inferential Statement): Based on the analysis of the sample data, it was concluded that high school start times at 8:30 a.m. or later are positively associated with longer sleep duration when compared to schools with earlier start times. This is an inference about the broader population of high school students.
Source: Nahmod et al. (2019). Later high school start times associated with longer actigraphic sleep duration in adolescents. Sleep, 42(2), zsy212.
Solution steps (aligned with the detailed Process of Statistics):
Objective: The explicit aim was to determine if an association existed between a school's start time and the sleep duration experienced by its high school students.
Data Collection: Data were meticulously collected from randomly selected adolescents using appropriate methods (e.g., actigraphic sleep monitoring to ensure accuracy). The mean age of years and standard deviation of year were descriptive characteristics of this collected sample.
Data Description: The raw data on start times and sleep durations were summarized. This involved calculating differences in average sleep durations (e.g., minutes and minutes per hour delay) and potentially creating visual aids like scatter plots or comparison bar charts to illustrate these patterns within the sample data.
Inference/Draw Conclusions: Statistical methods were applied to generalize these sample findings to the larger population of U.S. high school students, leading to the confident conclusion that later start times are indeed associated with longer sleep durations, with a reported level of statistical reliability.
Distinguish between Qualitative and Quantitative Variables
Variables are the specific characteristics or attributes of individuals within a population that are observed or measured. The defining feature of a variable is that its values vary across different individuals.
Qualitative (Categorical) Variables: These variables serve to classify individuals into categories or groups based on some attribute or characteristic. The values are typically labels or names and do not have a meaningful numerical order or value that can be used in arithmetic operations. Examples of qualitative variables include gender, hair color, type of car, or political affiliation.
Analysis: For qualitative variables, statistical analysis typically involves counting frequencies, calculating proportions or percentages for each category, and visualizing data using bar charts, pie charts, or Pareto charts.
Quantitative Variables: These variables provide numerical measurements or counts of individuals. The values are inherently numerical, and most importantly, arithmetic operations such as addition and subtraction (and consequently, calculating means or differences) yield results that are genuinely meaningful. Examples include height, age, income, or the number of children.
Analysis: For quantitative variables, descriptive statistics often involve calculating measures of central tendency (mean, median) and dispersion (range, standard deviation). Data can be visualized using histograms, box plots, dot plots, or time-series plots.
Key Idea: The fundamental reason for distinguishing between these variable types is that the variability of a variable is precisely what statistics aims to understand, measure, and quantify. The type of variable dictates the appropriate statistical methods for both description and inference.
Examples for practice (classify each as qualitative or quantitative):
(a) Education level (e.g., High School, Bachelor's, Master's) → Qualitative
(b) Today’s high temperature (in degrees Celsius) → Quantitative
(c) Daily intake of whole grains (grams per day) → Quantitative
(d) Number of vending machines at a school → Quantitative
(e) Whether a student is prepared for class (Yes/No) → Qualitative
(f) Number of days per week a student eats lunch → Quantitative
(g) Name of a university (e.g., Harvard, Stanford) → Qualitative
Distinguish between Discrete and Continuous Variables
This distinction applies specifically to quantitative variables.
Discrete Variable: A quantitative variable is discrete if it has a finite or a countable number of possible values. These values typically result from counting and usually consist of whole numbers (e.g., ). A discrete variable cannot take on every possible value within any given interval between two specific values. For instance, you can have 2 or 3 children, but not 2.5 children.
Examples: Number of bedrooms in a house, number of cars owned, number of defects in a product, number of students in a classroom.
Continuous Variable: A quantitative variable is continuous if it has an infinite number of possible values that are not countable. These values typically result from measurement and may take on any value within a given interval (e.g., all real numbers between 0 and 1). The precision of a continuous measurement is limited only by the measuring instrument.
Examples: Height of a person, weight of an object, temperature, time taken to complete a task, income (though often treated as discrete in practice due to rounding, it's inherently continuous).
Examples for practice (classify and, where appropriate, label discrete vs continuous):
(a) Gender → Qualitative (not quantitative, so not discrete/continuous)
(b) Income status (e.g., middle income, low income, high income) → Qualitative (ordinal)
(c) Income (exact dollar amount) → Quantitative (Continuous) (in raw form, it can take any value, though often presented discretely due to rounding)
(d) Grade earned in Algebra (as a percentage, e.g., ) → Quantitative (Continuous)
(e) Respondents’ agreement on a Likert scale (e.g., strongly agree, agree, neutral, disagree, strongly disagree) → Qualitative (Ordinal)
(f) Number of children in a classroom → Quantitative (Discrete)
Definitions of data types linked to variables (for clarity):
The collection of specific observations or measurements that a variable assumes for a sample or population is called data.
Qualitative data refers to the observed values corresponding