Statistical Studies and Graphical Methods for Describing Data Distributions
Statistical Studies: Observation and Experimentation
Population vs. Sample:
Population: The entire collection of individuals or objects that a researcher wants to study.
Sample: A subset or part of the population selected for investigation and analysis.
Observational Study vs. Experiment:
Observational Study:
A study in which characteristics of a sample selected from one or more existing populations are observed.
Primary Goal: To use data obtained from the sample to draw conclusions and learn about the corresponding population.
Key Requirement: It is essential to obtain a sample that is representative of the population.
Experiment:
A study in which the researcher observes how a response variable behaves under different experimental conditions.
The individual conducting the study actively determines who will be placed in each experimental group and assigned to specific conditions.
Collecting Data: Planning an Observational Study
Parameter vs. Statistic:
Parameter: A characteristic or numerical value that describes an entire population.
Statistic: A numerical value that describes a sample drawn from a population.
Selecting a Representative Sample:
A sample must be representative of the underlying population to yield valid findings.
Careful consideration must be given to the specific method by which the sample is selected.
Simple Random Sample (SRS):
Definition: A simple random sample of size is a sample selected from a population in such a manner that every possible sample of the same size has an equal probability of being chosen.
Notation: The lowercase letter is used to denote the sample size, representing the total number of individuals or objects contained in the sample.
Selecting a Simple Random Sample:
Sampling Frame: A comprehensive list of all objects or individuals in the population.
Selection Methods:
Using a random number generator.
Using a table of random digits.
Sampling Types:
Sampling with Replacement: A method in which an individual or object selected from the population is returned to the population prior to the next selection. Under this method, a particular individual may be selected more than once.
Sampling without Replacement: A method in which once an individual or object is selected, it is not returned to the population before subsequent selections. This method guarantees that the sample consists exclusively of distinct individuals.
Sample Size Misconception and Truth:
Common Misconception: Believing that if a sample size is small relative to the total population size, the sample is inherently inaccurate or uninformative regarding the population.
Correction: Unless the population itself is very small, risk depends primarily on the sample size itself (where the risk of error is smaller with larger sample sizes), rather than the proportion of the population sampled.
Role of Random Selection: Random selection provides confidence that the resulting sample will accurately reflect the characteristics of the population.
Common Sampling Methods:
Simple Random Sample: Every sample of size has an equal chance of being selected.
Stratified Sampling: The population is divided into non-overlapping subgroups (strata) and random samples are selected from each subgroup.
Systematic Sampling: Selecting every individual from a list after a random starting point.
Cluster Sampling: Dividing the population into clusters, selecting a random sample of clusters, and surveying all individuals within chosen clusters.
Avoiding Bias in Observational Studies:
Bias Definition: The tendency for samples to differ from the corresponding population in some systematic way.
Major Types of Sampling Bias:
Selection Bias: Occurs when the sampling method systematically overrepresents or underrepresents certain parts of the population of interest.
Non-response Bias: Occurs when responses or data are not obtained from all individuals selected into the sample.
Measurement Bias (Response Bias): Occurs when the measurement or observation method produces values that systematically differ from the true value in some direction.
Collecting Data: Planning an Experiment
Fundamental Concepts of Experimental Design:
Response Variable: A variable whose value is measured under different experimental conditions.
Experimental Unit: The smallest unit or subject to which a treatment is applied.
Designing an Experiment:
The design must ensure that there are no systematic sources of variation in the response variable other than the experimental conditions being investigated.
Confounding Variables:
Definition: Two variables are confounded when their effects on the response variable cannot be distinguished from one another.
Impact: Confounding variables provide alternate explanations for observed differences in the response variable across experimental groups.
Goal of Experimental Design: A well-designed experiment actively protects against potential confounding variables.

Experimental Design Strategies:
Direct Control: Eliminating sources of variability by holding potential confounding factors constant at a fixed level across all conditions.
Random Assignment: Ensuring that remaining uncontrolled sources of variability produce only chance-like differences between treatment groups.
Strategies for Implementing Random Assignment:
Physically drawing experimental units or labeled tags from a container.
Utilizing a random number generator.
Utilizing a physical random mechanism (e.g., flipping coins, rolling dice).
Other Design Considerations:
Control Group: A baseline group that receives no active treatment or receives a standard treatment for comparison.
Placebo: An inactive or dummy treatment that appears identical to the active treatment.
Single-Blind Experiment: An experiment in which either the subjects or the individuals measuring the response are unaware of treatment assignments.
Double-Blind Experiment: An experiment in which neither the subjects nor the individuals administering treatments and measuring responses know which treatment each subject receives.
Questions to Ask when Evaluating an Experiment:
What question is the experiment attempting to answer?
What are the experimental conditions (treatments)?
What is the response variable?
What are the experimental units?
Is there a control or placebo group included?
Does the experiment involve blinding (single-blind or double-blind)?
Importance of Random Selection and Random Assignment
Scope of Conclusions Matrix:
Type of Conclusion | Reasonable When |
|---|---|
Results observed in the sample can be generalized to the broader population | Random selection was used to obtain the sample |
Differences in the response variable are caused by experimental conditions (cause-and-effect relationship) | Random assignment of subjects to treatments was used |
Common Mistakes to Avoid in Statistical Studies
Critical Rules and Precautions:
Common Mistake | Rule / Justification |
|---|---|
Drawing a cause-and-effect conclusion from an observational study | Never do this! Observational studies cannot establish causation due to potential unmeasured confounding variables. |
Generalizing results of an experiment that uses volunteers as subjects | Generalization is permissible only if it can be convincingly argued that the volunteers are representative of the targeted population. |
Generalizing conclusions from a poorly designed observational study | Generalization is justified only when the sample is likely to be representative of the target population. |
Graphical Methods for Describing Data Distributions
Selecting an Appropriate Graphical Display:
Choice of display depends on three main criteria:
The number of variables in the dataset.
The type of data (categorical vs. numerical).
The primary purpose of the graphical display.
Definitions of Variables and Data Types:
Variable: Any characteristic whose value can differ from one individual or object to another.
Data: The set of observations gathered on one or more variables.
Univariate Data: Data resulting from observations on a single variable.
Bivariate Data: Data resulting from observations on two variables recorded for the same individual or object.
Multivariate Data: Data resulting from observations on two or more variables recorded for the same individual or object.
Categorization of Data Types:
Categorical Variable: A variable whose observations fall into qualitative categories.
Numerical Variable: A variable whose observations are numerical values.
Discrete Numerical Variable: A numerical variable whose possible values correspond to isolated points along a number line.
Continuous Numerical Variable: A numerical variable whose possible values form an entire continuous interval along a number line.
Purposes of Graphical Displays:
To determine whether a relationship exists between variables in a dataset.
To display data distributions effectively.
To compare different groups or populations.
To explore patterns and underlying structures.
Bar Charts and Comparative Bar Charts
Standard Bar Charts:
Summarize categorical data in a frequency table and use that information to construct a visual display.
Frequency Distribution: A table displaying possible categories alongside their associated counts.
Frequency: The total number of times a given category occurs in a dataset.
Relative Frequency Formula:
Relative Frequency Definition: The proportion or fraction of total observations belonging to a specific category.
Comparative Bar Charts:
Used to provide a direct visual comparison of categorical distributions across two or more groups.
Dotplots, Stem-and-Leaf Displays, and Histograms
Dotplots and Comparative Dotplots:
Dotplots: A straightforward visual representation of numerical data suitable for small to moderate dataset sizes. A dot is plotted above the corresponding numerical axis for each data point.
Comparative Dotplots: Formed by aligning two or more dotplots along the exact same numerical scale to compare distributions.
Stem-and-Leaf Displays:
Used for small to moderate numerical datasets.
Breaks each observation into two components:
Stem: The leading digit(s) representing the base value.
Leaf: The final trailing digit representing the fine detail.
Example: For the numerical value , the stem is and the leaf is ().
Outliers:
Observations that are unusually small or unusually large relative to the remainder of the dataset.
Histograms:
Preferred graphical method for large numerical datasets.
Discrete Histograms: Displays each distinct value along with its frequency or relative frequency.
Histograms for Comparing Groups:
Always utilize relative frequency on the vertical axis when comparing groups of varying sizes.
Maintain identical horizontal and vertical scales across graphs.
Histograms with Unequal Interval Widths:
Must utilize a density scale on the vertical axis, defined as:
Common Distribution Shapes:
Unimodal: Contains a single prominent peak.
Bimodal: Contains two distinct peaks.
Multimodal: Contains three or more distinct peaks.
Symmetric: The left and right halves of the distribution are mirror images.
Positively Skewed (Right-Skewed): The distribution extends much farther out to the right (positive direction).
Negatively Skewed (Left-Skewed): The distribution extends much farther out to the left (negative direction).
Displaying Bivariate Numerical Data: Scatterplots and Time Series Plots
Bivariate Numerical Data:
Measurements recorded on two numerical variables, traditionally designated as and , yielding pairs of numbers .
Scatterplots:
A graph in which each paired observation is plotted as a single point on a standard rectangular coordinate grid.
Time Series Plots:
Displays numerical data collected over uniform intervals of time to analyze trends, seasonal patterns, and variations over time.
Uses line segments connecting consecutive data points chronologically.
Graphical Displays in the Media and Evaluation Criteria
Pie Charts:
Graph depicting the distribution of a categorical variable.
The complete circle represents of the dataset.
Slices represent individual categories, with slice size proportional to category frequency or relative frequency.
Most effective when the categorical variable contains a small number of distinct categories.
Evaluating Media Graphics (Critical Questions):
What context is provided for the graphic?
Is the original data source clearly identified?
What distributions or relationships between variables are depicted?
What narrative or story does the visual display convey?
Is there room for visual or statistical improvement?
Graphical Pitfalls to Avoid:
Area Principle: Area must remain strictly proportional to the frequency, relative frequency, or quantity represented.
Axis Truncation: Exercise extreme caution with broken axes or vertical axes that do not begin at , as they misrepresent relative proportions.
Pattern Interpretation: Avoid making unwarranted causal claims based solely on patterns observed in scatterplots.