Numerical Methods for Describing Data Distributions

Selecting Numerical Summaries

  • Measures of center describe typical values along a number line.
  • Measures of variability describe how much individual values differ from one another.
  • Distribution shape dictates the choice of numerical summaries:

Summary table for selecting center and variability measures

  • Symmetric distributions: Described using sample mean (xˉ\bar{x}) and sample standard deviation (ss).
  • Skewed distributions or distributions with outliers: Described using sample median and interquartile range (iqriqr).

Describing Symmetric Distributions

  • Sample Mean (xˉ\bar{x}): The arithmetic average of all observations in a sample. xˉ=∑xn\bar{x} = \frac{\sum x}{n}
  • Population Mean (μ\mu): The arithmetic average of all values in an entire population.
  • Deviations from the Mean: Difference between an observation and the sample mean. (x−xˉ)(x - \bar{x})
  • Sample Variance (s2s^2): Sum of squared deviations from the mean divided by n−1n - 1. s2=∑(x−xˉ)2n−1s^2 = \frac{\sum (x - \bar{x})^2}{n - 1}
  • Sample Standard Deviation (ss): Square root of the sample variance. s=s2=∑(x−xˉ)2n−1s = \sqrt{s^2} = \sqrt{\frac{\sum (x - \bar{x})^2}{n - 1}}
  • Population Variance (σ2\sigma^2) and Population Standard Deviation (σ\sigma): Parameters representing variability for an entire population.

Describing Skewed Distributions and Handling Outliers

  • Sample Median: The middle value when data are listed in order from smallest to largest.
    • If sample size nn is odd: The single middle value.
    • If sample size nn is even: The average of the two middle values.
  • Quartiles:
    • Lower Quartile: Median of the lower half of the data set.
    • Upper Quartile: Median of the upper half of the data set.
    • If nn is odd, the overall median is excluded from both halves when calculating quartiles.
  • Interquartile Range (iqriqr): Measuring variability for middle 50% of data, resistant to outliers. iqr=upper quartile−lower quartileiqr = \text{upper quartile} - \text{lower quartile}

Boxplots and Five-Number Summaries

  • Five-Number Summary: Minimum value, lower quartile, median, upper quartile, and maximum value.
  • Outlier Rule (1.5 IQR Rule): An observation is officially an outlier if it is:
    • Greater than upper quartile+1.5(iqr)\text{upper quartile} + 1.5(iqr)
    • Less than lower quartile−1.5(iqr)\text{lower quartile} - 1.5(iqr)
  • Modified Boxplots: Display individual outliers as plotted dots and extend whiskers to the nearest non-outlier values.
  • Comparative Boxplots: Display multiple boxplots on a shared numerical scale to compare center, variability, and shape across groups.

Comparative boxplots showing score improvement for two practice groups

Measures of Relative Standing: z-Scores and Percentiles

  • z-Score: Represents the number of standard deviations a data value xx lies from the mean xˉ\bar{x}. z-score=data value−meanstandard deviation=x−xˉsz\text{-score} = \frac{\text{data value} - \text{mean}}{\text{standard deviation}} = \frac{x - \bar{x}}{s}
  • Standardizing: The process of subtracting the mean and dividing by the standard deviation.
  • Empirical Rule: Applies to mound-shaped and approximately symmetric distributions:
    • Approximately 68%68\% of observations fall within 11 standard deviation (xˉ±1s\bar{x} \pm 1s).
    • Approximately 95%95\% of observations fall within 22 standard deviations (xˉ±2s\bar{x} \pm 2s).
    • Approximately 99.7%99.7\% of observations fall within 33 standard deviations (xˉ±3s\bar{x} \pm 3s).

Empirical Rule percentages for mound-shaped distributions

  • Percentiles: The rr\text{th percentile} is the value such that r%r\% of the data fall at or below it.
    • Lower quartile = 2525\text{th percentile}
    • Median = 5050\text{th percentile}
    • Upper quartile = 7575\text{th percentile}

Key Pitfalls and Common Mistakes

  • Do not summarize numerically coded categorical variables (such as zip codes or class standings) using mean, median, standard deviation, or IQR.
  • Never rely solely on measures of center; always combine them with measures of variability and distribution shape.
  • Avoid using mean and standard deviation on heavily skewed distributions or data sets containing notable outliers.
  • Do not confuse the variability of values along the x-axis with the variation in frequency bar heights.
  • Do not apply the Empirical Rule to distributions that are skewed or non-mound-shaped.