Introduction to Statistics: Variables, Parameters, Descriptive and Inferential Statistics, and Sampling Error
Fundamentals of Variables and Measurement
Definition of a Variable: A variable is any characteristic, condition, or attribute that can change or assume different values across individuals or environments.
Classification of Variables:
Individual Characteristics: Characteristics that naturally differ from one person to another. Examples include:
Weight
Gender identity
Personality traits
Motivation levels
Behavioral patterns
Environmental Conditions: External contextual factors that vary across experimental settings or time. Examples include:
Ambient temperature
Time of day
Physical dimensions or size of a testing room
Environmental Impact Example (Weather and Restaurant Reviews):
A study investigated whether external weather conditions influence customer online restaurant reviews.
Variables: Environmental weather conditions and customer restaurant review ratings.
Findings: A direct empirical relationship was demonstrated. Customer reviews were consistently worse during bad weather, specifically on days with extremely hot or extremely cold temperatures.
Measurement Terminology:
Datum (Score / Raw Score): The single measurement or numerical value obtained for a specific individual in a study.
Dataset (Data): The complete set or collection of measurements/scores gathered from all individuals participating in a research study.
Population Parameters vs. Sample Statistics
Populations and Samples:
Population: The entire, comprehensive group of individuals, objects, or items of interest in a research investigation (e.g., all registered voters).
Sample: A specific subset selected from a population, intended to represent that population in a study (e.g., a group of middle school children).
Populations and Samples of Scores: Because research involves measuring individuals to derive numerical values, every sample of individuals produces a corresponding sample of scores, and every population of individuals produces a corresponding population of scores.
Quantitative Descriptors:
Parameter: A numerical characteristic or summary value that describes an entire population (e.g., the average score of a population).
Statistic: A numerical characteristic or summary value that describes a sample (e.g., the average score of a sample).
Research Workflow:
Investigations systematically begin with a research question concerning an unknown population parameter.
Because populations are typically too large to measure every individual, researchers draw a representative sample.
Sample data are measured to calculate sample statistics.
Sample statistics serve as the practical foundation for answering questions and drawing inferences about population parameters.
Categories of Statistical Procedures
Descriptive Statistics:
Definition: Statistical procedures utilized to organize, simplify, and summarize raw data into a manageable, comprehensible form.
Common Techniques:
Organizing raw scores into structured visual distributions using tables or graphs.
Calculating single summary values, such as an average, to represent an entire dataset containing hundreds of individual scores.
Inferential Statistics:
Definition: Statistical methods that evaluate sample data to draw generalized conclusions, infer properties, and make statements about broader populations.
Sampling Error and Discrepancies
Concept of Sampling Error:
Definition: The naturally occurring, unsystematic discrepancy or difference between a sample statistic and its corresponding population parameter.
Mechanism: Samples provide limited information about a population. Because different samples consist of different individuals with unique attributes, sample statistics naturally vary from sample to sample and differ from the true population parameter.
Inferential Challenge: Sampling error represents the fundamental problem that all inferential statistical methods must address.
Demonstration Example (Student Age Study):
Population Parameters: A defined population of college students with a true population mean age of .
Sample 1: A random sample of students selected from the population yields a mean age statistic of .
Sample 2: A second random sample of students selected from the same population yields a mean age statistic of .
Analysis: Both sample means differ from one another ( vs. ) and diverge from the population parameter (). This discrepancy illustrates sampling error occurring across hundreds of possible random samples.
Real-World Applications of Sampling Error
Political Polling and Margin of Error:
Context: A political poll conducted on a sample of registered voters reports that Candidate Brown leads with of the vote, Candidate Jones holds approval, and of respondents are undecided.
Margin of Error: The reported polling results include a margin of error of (plus or minus four percentage points).
Interpretation: The reported sample percentages are generalized to the target population of all potential voters. The margin of error is the explicit quantification of sampling error.
Classroom Division Example:
Procedure: A statistics classroom is partitioned into two groups by drawing an imaginary line from front to back through the middle of the room.
Measurement: Computing the average age, height, or Grade Point Average () for each side of the room.
Outcome: The two groups will almost certainly yield different averages.
Interpretation: An observed difference (such as a higher average age on the right side of the room) does not reflect a systematic force causing older students to sit on the right. Instead, it is the result of random, chance factors.
Core Role of Inferential Statistics: Inferential statistics determine whether observed differences between samples are caused by random chance factors (sampling error) or reflect genuine, meaningful relationships existing within the overall population.