Comprehensive Study Notes: Skewed Distributions, Quartiles, and Summary Statistics
Measures of Center and Spread for Symmetric and Skewed Distributions
Symmetric Distributions:
Center: Described using the mean.
Spread: Described using the standard deviation.
Data Distribution: The empirical rule is applied to determine where the majority of the dataset exists.
Skewed Distributions:
Center: Described using the median.
Spread: Described using the Interquartile Range ().
Data Distribution: The empirical rule, mean, and standard deviation are not appropriate or useful for skewed distributions or skewed histograms.
Determining Distribution Shape by Comparing Mean and Median:
Rather than relying strictly on visual estimation of histograms, compare the numerical values of the mean and median.
Perfectly Symmetric: The mean is exactly equal to the median ().
Basically Symmetric: The mean is almost equal to the median ().
Skewed: The mean is several units away from the median ().
Skewed Left: The mean is significantly smaller than the median ().
Skewed Right: The mean is significantly larger than the median ().
Numerical Example of Skewness:
A histogram dataset has a mean of and a median of .
Because , the distribution is not symmetric.
Because the mean () is larger than the median (), the distribution is skewed right, meaning data points are more spread out to the right of the tallest peak than to the left.
Questions & Discussion: Z-Scores for US Birth Weights
Question: When evaluating a birth weight measurement relative to all US births within a symmetric distribution, which mean and standard deviation parameters should be selected?
Response:
Mean: Use the overall US birth mean of \, (or relative unit), rather than the mean specifically for infants born one month early.
Standard Deviation: Use the population standard deviation of .
Z-Score Application: The z-score formula calculates how many standard deviations a specific data point is away from the mean, allowing one to classify the observation as usual or unusual.
Calculating the Median and Summary Statistics
Procedure for Finding the Median by Hand:
Arrange the data values in order from smallest to largest.
Identify the direct middle of the ordered list.
If the dataset contains an odd number of observations (), the median is the single value located exactly in the center at position .
If the dataset contains an even number of observations (), the median is the arithmetic average of the two middle numbers.
Resistance Property:
The median is resistant to extreme values and outliers, whereas the mean is sensitive to outliers.
Class A Dataset Analysis:
Dataset size: data points.
Median: The ordered data point, which equals .
Mean: Approximately .
Comparison: The mean () and median () are not equal, indicating the distribution is skewed.
Class B Dataset Analysis:
Dataset size: data points.
Median: The average of the two middle observations ( and values), which equals
Mean: .
Comparison: The mean () and median () are not equal, indicating the distribution is skewed.
StatCrunch Software Workflow for Summary Statistics:
Open a blank spreadsheet or table.
Enter raw data into labeled columns (e.g.,
Class A,Class B). Data does not need to be sorted manually; the software sorts automatically.Navigate to Summary Stats Columns.
Hold the
Ctrlkey on the keyboard to select multiple columns simultaneously.Click Compute to generate summary statistics (mean, median, range, quartiles).
Assessing Skewness and Range in Datasets
Evaluating Degree of Skewness:
When comparing two distributions, evaluate which is more dramatically skewed or whether one is sufficiently close to symmetric to permit the use of symmetric rules.
New York Income Case Example:
Histograms representing income in New York (scaled in thousands of dollars per year) show dramatic right-skewness.
The mean and median lie in significantly different positions on the histogram.
Because the mean and median are vastly different, the visual skewness is confirmed numerically. The median must be used to represent a typical income value.
The Range:
Definition: The total distance spanned by the dataset.
Formula:
Class A Range Calculation: For Class A data with a maximum value of and a minimum value of :
Properties: Range defines a window units wide containing all data points. Range does not indicate whether a distribution is symmetric or skewed and is severely affected by outliers.
Quartiles and Interquartile Range (IQR)
Quartiles:
Quartiles partition an ordered dataset into four equal segments, each containing of the data.
Median (): Splits the total dataset in half ( below, above).
First Quartile (): The median of the lower half of the dataset (left side of the overall median).
Third Quartile (): The median of the upper half of the dataset (right side of the overall median).
Manual Quartile Step-by-Step Calculation (Class A Example):
Class A has data values with overall median .
Finding :
Look at the data points to the left of .
Average the middle two values of these observations to get Q_1 = 71$.\n * 7127127177$.
Finding :
Look at the data points to the right of (values: , , , ).
Average the middle two values ( and ) to get Q_3 = 91$.\n * 9127791291$.
Central Spread: Between and lies approximately (or slightly more for odd ) of the dataset. Values falling between and represent usual or typical values.
Class B Quartiles and IQR:
Interquartile Range Formula:
Class B IQR Calculation:
Generating IQR in StatCrunch:
Navigate to Summary Stats Columns.
Select the target column.
Under the Statistics selection box, scroll down and click
IQR.Hold the
Ctrlkey while clicking to keep default summary statistics selected alongsideIQR.Click Compute.
Comparing Center Spread: Empirical Rule versus IQR
Graphical Representation of Quartile Spacing:
Symmetric Distributions: Standard deviation intervals are evenly spaced on both sides of the mean.
Skewed Distributions: Distance between quartiles is uneven.
For example, the distance from the median to may be significantly narrower than the distance from the median to .
Quartile Proportions: Regardless of symmetry or skewness, the data is strictly divided into fixed intervals:
of data values
of data values between and Median
of data values between Median and
of data values
Comprehensive Sample Dataset Comparison:
Dataset size: observations.
Calculated summary metrics:
Methodology Comparison:
Empirical Rule Approach (Incorrect for skewed data): Uses standard deviation () to establish a spread window of roughly total units ( units above the mean and units below the mean).
IQR Approach (Correct for skewed data): Uses central interquartile range () to establish a central spread window of units around the median.