Numerical Methods for Describing Data Distributions
Selecting Numerical Summaries
- Measures of center describe typical values along a number line.
- Measures of variability describe how much individual values differ from one another.
- Distribution shape dictates the choice of numerical summaries:

- Symmetric distributions: Described using sample mean (xˉ) and sample standard deviation (s).
- Skewed distributions or distributions with outliers: Described using sample median and interquartile range (iqr).
Describing Symmetric Distributions
- Sample Mean (xˉ): The arithmetic average of all observations in a sample.
xˉ=n∑x
- Population Mean (μ): The arithmetic average of all values in an entire population.
- Deviations from the Mean: Difference between an observation and the sample mean.
(x−xˉ)
- Sample Variance (s2): Sum of squared deviations from the mean divided by n−1.
s2=n−1∑(x−xˉ)2
- Sample Standard Deviation (s): Square root of the sample variance.
s=s2=n−1∑(x−xˉ)2
- Population Variance (σ2) and Population Standard Deviation (σ): Parameters representing variability for an entire population.
Describing Skewed Distributions and Handling Outliers
- Sample Median: The middle value when data are listed in order from smallest to largest.
- If sample size n is odd: The single middle value.
- If sample size n is even: The average of the two middle values.
- Quartiles:
- Lower Quartile: Median of the lower half of the data set.
- Upper Quartile: Median of the upper half of the data set.
- If n is odd, the overall median is excluded from both halves when calculating quartiles.
- Interquartile Range (iqr): Measuring variability for middle 50% of data, resistant to outliers.
iqr=upper quartile−lower quartile
Boxplots and Five-Number Summaries
- Five-Number Summary: Minimum value, lower quartile, median, upper quartile, and maximum value.
- Outlier Rule (1.5 IQR Rule): An observation is officially an outlier if it is:
- Greater than upper quartile+1.5(iqr)
- Less than lower quartile−1.5(iqr)
- Modified Boxplots: Display individual outliers as plotted dots and extend whiskers to the nearest non-outlier values.
- Comparative Boxplots: Display multiple boxplots on a shared numerical scale to compare center, variability, and shape across groups.

Measures of Relative Standing: z-Scores and Percentiles
- z-Score: Represents the number of standard deviations a data value x lies from the mean xˉ.
z-score=standard deviationdata value−mean=sx−xˉ
- Standardizing: The process of subtracting the mean and dividing by the standard deviation.
- Empirical Rule: Applies to mound-shaped and approximately symmetric distributions:
- Approximately 68% of observations fall within 1 standard deviation (xˉ±1s).
- Approximately 95% of observations fall within 2 standard deviations (xˉ±2s).
- Approximately 99.7% of observations fall within 3 standard deviations (xˉ±3s).

- Percentiles: The r\text{th percentile} is the value such that r% of the data fall at or below it.
- Lower quartile = 25\text{th percentile}
- Median = 50\text{th percentile}
- Upper quartile = 75\text{th percentile}
Key Pitfalls and Common Mistakes
- Do not summarize numerically coded categorical variables (such as zip codes or class standings) using mean, median, standard deviation, or IQR.
- Never rely solely on measures of center; always combine them with measures of variability and distribution shape.
- Avoid using mean and standard deviation on heavily skewed distributions or data sets containing notable outliers.
- Do not confuse the variability of values along the x-axis with the variation in frequency bar heights.
- Do not apply the Empirical Rule to distributions that are skewed or non-mound-shaped.